Vector Processor Memory Interconnect Network Architecture
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing memory interconnect network architectures for parallel processing processors, such as digital signal processors, are inefficient due to the need for multiple switch-based interconnection networks, which result in high area, power, and communication costs, while algorithms like broadcast loads do not fully utilize the network capacity.
Innovation Solution
Modifying the memory interconnect network architecture by replacing one switch-based interconnection network with a non-switch-based interconnection network, such as a broadcast bus, to reduce area and power requirements while maintaining performance for digital signal processing applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If multiple switch-based interconnection networks are used to support parallel data transfers, then data throughput and processing performance are improved, but area, power consumption, and communication costs increase significantly
Solution Approach 1:
The interconnection network is segmented into two distinct networks: a switch-based interconnection network for transferring different data operands to different processing elements, and a non-switch-based interconnection network for transferring the same data operand to multiple processing elements. This segmentation allows each network to be optimized for its specific function, reducing overall complexity and area requirements
Solution Approach 2:
Instead of using switch-based networks for all data transfers (including broadcast operations), the patent inverts the approach by using a non-switch-based network specifically for broadcast operations where the same data is sent to multiple processing elements. This inversion eliminates the need for switches in the broadcast path, significantly reducing area and power consumption
2Adaptability or versatility
If switch-based interconnection networks are used for all data transfers, then flexible data routing is achieved, but power consumption and device complexity increase
Solution Approach 1:
Different parts of the interconnection system are given different qualities: the switch-based network provides flexible routing for point-to-point transfers, while the non-switch-based network provides efficient broadcast capability. Each network is optimized for its specific quality requirements, reducing overall power consumption while maintaining necessary flexibility
Solution Approach 2:
The patent applies switch-based routing only partially - specifically for operations requiring different data to be sent to different processing elements. For broadcast operations where the same data is sent to multiple elements, the simpler non-switch-based network is used, avoiding the excessive power consumption of switches when full routing flexibility is not needed
3Productivity
If broadcast loads are implemented using switch-based networks, then data can be distributed to multiple processing elements, but network capacity is not fully utilized and area requirements increase
Solution Approach 1:
The broadcast function is extracted from the switch-based interconnection network and implemented as a separate non-switch-based interconnection network. This extraction allows the switch-based network to focus on point-to-point transfers where routing flexibility is essential, while the broadcast network handles collective distribution operations with simpler, more area-efficient hardware
Data Source
AI summary
The present disclosure provides a memory interconnection architecture for a processor, such as a vector processor, that performs parallel operations. An example processor may include a compute array that includes processing elements; a memory that includes memory banks; and a memory interconnect network architecture that interconnects the compute array to the memory. In an example, the memory interconnect network architecture includes a switch-based interconnect network and a non-switch based interconnect network. The processor is configured to synchronously load a first data operand to each of the processing elements via the switch-based interconnect network and a second data operand to each of the processing elements via the non-switch-based interconnect network.


