Sparse Crossbar Interconnect for Flexible Image Data Routing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data interconnects in computer vision and imaging systems for mobile devices lack flexibility and efficiency, particularly in supporting various interconnect configurations and high bandwidth-low latency data transfers, which are essential for real-time imaging and computer vision applications.
Innovation Solution
A multi-port memory architecture with a sparse crossbar interconnect that allows flexible routing between hardware accelerators and memory, utilizing a circular buffer and tile buffer configuration, along with head and tail pointer buses for synchronization, to enable efficient data transfer and support different use cases.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a fixed interconnect configuration is used, then device complexity is reduced, but adaptability to different use cases deteriorates
Solution Approach 1:
The interconnect configuration is made dynamic through runtime reconfigurability. The system can change the interconnect topology (mesh, ring, tree, star, fully-connected) based on the specific use case requirements, allowing the architecture to adapt between high bandwidth needs and low latency needs without being locked into a fixed configuration.
Solution Approach 2:
The interconnect architecture implements multi-functionality by supporting multiple topological configurations within a single unified structure. The same physical interconnect resources can be dynamically allocated to different logical topologies depending on the application requirements, making the system universally applicable to various computer vision and imaging workloads.
2Adaptability or versatility
If memory is allocated in fixed partitions, then device complexity is reduced, but adaptability to different hardware accelerator chains deteriorates
Solution Approach 1:
Memory allocation is made dynamic through runtime partitioning. The system can reconfigure memory boundaries and allocations based on the specific hardware accelerator chain being executed, allowing different accelerators to receive appropriate memory resources without being constrained by static partitions established at compile time.
Solution Approach 2:
The memory space is segmented into multiple partitions that can be dynamically assigned to different hardware accelerators. This segmentation allows flexible allocation where each accelerator receives the specific memory portion it needs for its operation, while the overall memory resource pool remains shared and reconfigurable.
3Productivity
If high bandwidth data transfer is prioritized, then productivity is improved, but latency increases
Solution Approach 1:
The interconnect dynamically adjusts its operational characteristics based on the workload requirements. For bandwidth-intensive operations, the system configures for high throughput modes, while for latency-sensitive operations, it optimizes for low-latency paths, allowing the system to achieve both high productivity and low latency depending on the specific task.
Solution Approach 2:
The interconnect parameters (such as buffer sizes, arbitration policies, and routing algorithms) are changed based on the operational mode. When high bandwidth is required, parameters are adjusted to maximize throughput, and when low latency is critical, parameters are modified to minimize transmission delays, enabling the system to optimize for different performance dimensions as needed.
Data Source
AI summary
An image and vision processing architecture included a plurality of image processing hardware accelerators each configured to perform a different one of a plurality of image processing operations on image data. A multi-port memory shared by the hardware accelerators stores the image data and is configurably coupled by a sparse crossbar interconnect to one or more of the hardware accelerators depending on a use case employed. The interconnect processes accesses of the image data by the hardware accelerators. Two or more of the hardware accelerators are chained to operate in sequence in a first order for a first use case, and at least one of the hardware accelerators is set to operate for a second use case. Portions of the memory are allocated to the hardware accelerators based on the use case employed, with an allocated portion of the memory configured as a circular buffer.


