Vector Processor Array Using Identical Circuit Slices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current vector processor architectures face challenges in efficiently executing complex data operations due to high interconnect costs and latency, as well as difficulties in splitting applications across multiple processors, leading to performance degradation in high-data-rate communications and digital signal processing tasks.
Innovation Solution
A vector processor architecture is implemented using an array of identical circuit blocks assembled by abutment, allowing for a large vector processor core to be segmented into multiple identical slices, each with shared functionality and a unique identifier for conditional logic circuit enablement, reducing interconnect complexity and enhancing layout efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a large vector processor core is implemented using traditional interconnect networks, then processing capability is improved, but interconnect cost and latency increase
Solution Approach 1:
The vector processor core is divided into multiple identical slices, each capable of independent operation. This segmentation eliminates the need for complex interconnect networks between processors, as each slice can process data locally. The slices are arranged in a regular grid pattern with simple nearest-neighbor connections, dramatically reducing interconnect complexity and latency while maintaining high processing capability through parallel operation of multiple slices.
2Productivity
If applications are split across multiple processors, then processing power is improved, but performance degrades due to interconnect latency
Solution Approach 1:
Multiple processor slices are merged into a single integrated core structure where data can flow between slices through simple, fast inter-slice connections. This merging allows applications to be distributed across slices while maintaining low-latency communication, effectively combining the processing power of multiple units without the performance penalty of traditional multi-processor interconnects.
3Adaptability or versatility
If irregular functionality is implemented in vector processor, then versatility is improved, but layout complexity increases
Solution Approach 1:
Each slice contains identical basic functionality, but local quality is achieved through selective enabling of specific logic circuits within each slice based on configuration values. This allows irregular and diverse functionality to be implemented across the array of slices without requiring complex custom layouts for each functional unit. The regular slice structure maintains layout simplicity while the configurable logic circuits provide versatility.
Solution Approach 2:
Each slice is designed as a universal building block that can perform multiple functions through configuration. The slices contain logic circuits that can be selectively enabled or disabled based on configuration values, allowing the same hardware structure to implement different functionalities. This universality provides versatility while maintaining the simplicity of a regular, repeating layout pattern.
Data Source
Figure 1~2
Figure 3
Figure 4
AI summary
Systems, methods, and apparatuses relating to vector processor architecture having an array of identical circuit blocks are described. In one embodiment, a processor includes a single centralized circuit comprising an instruction decoder and a controller; and a plurality of circuit slices that each comprise an arithmetic logic unit, a multiplier, a register file, a local memory, and a same plurality of logic circuits and a packed data datapath in between, wherein each circuit slice includes a physical port that provides a unique identification value that identifies a circuit slice from the other circuit slices, and the controller is to broadcast a same configuration value to the plurality of circuit slices to cause a first circuit slice to enable a first logic circuit and enable a second logic circuit of the first circuit slice based on its unique identification value and the configuration value, and cause a second circuit slice to enable a same, first logic circuit and disable a same, second logic circuit of the second circuit slice based on its unique identification value and the configuration value.