In-Flight Reordering Circuitry for Vector Processing Engines
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Wireless computing devices face power consumption challenges due to limited battery life and resource constraints, particularly in baseband processors where post-processing reordering of output vector data samples delays subsequent processing operations, leading to underutilization of computational components.
Innovation Solution
Incorporating reordering circuitry in data flow paths between execution units and vector data memory to perform in-flight reordering of output vector data, eliminating the need for additional post-processing steps and reducing data flow limitations, thus allowing subsequent vector processing to be limited only by computational resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If post-processing reordering is performed after storing output vector data in memory, then data can be reordered according to required patterns, but processing delays occur and computational components are underutilized
Solution Approach 1:
The reordering circuitry performs reordering operations on output vector data before the data is stored in vector data memory. This preliminary action eliminates the need for subsequent post-processing reordering steps, allowing computational components to continue processing without waiting for memory storage operations to complete, thereby reducing processing delays and improving utilization of computational resources
Solution Approach 2:
A reordering circuitry is introduced as an intermediary component between the execution units and vector data memory. This intermediary performs the reordering function in-flight during data transfer, acting as a buffer that allows execution units to operate independently of memory storage timing constraints, thus eliminating processing bottlenecks
2Adaptability or versatility
If additional post-processing steps are added for reordering, then data flow flexibility is improved, but device complexity increases
Solution Approach 1:
The reordering function is merged with the existing data storage operation. The reordering circuitry is integrated into the data flow path between execution units and vector data memory, combining the reordering operation with the memory write operation into a single unified process. This eliminates the need for separate post-processing reordering steps, maintaining data flow flexibility while avoiding additional processing stages
3Productivity
If in-flight reordering is performed in data flow paths, then processing efficiency is improved, but device complexity increases
Solution Approach 1:
A reordering circuitry is introduced as an intermediary component between the execution units and vector data memory. This intermediary performs the reordering function in-flight during data transfer, acting as a buffer that allows execution units to operate independently of memory storage timing constraints, thus eliminating processing bottlenecks
Solution Approach 2:
The reordering circuitry performs reordering operations on output vector data before the data is stored in vector data memory. This preliminary action eliminates the need for subsequent post-processing reordering steps, allowing computational components to continue processing without waiting for memory storage operations to complete, thereby reducing processing delays and improving utilization of computational resources
Data Source
AI summary
Vector processing engines (VPEs) employing reordering circuitry in data flow paths between execution units and vector data memory to provide in-flight reordering of output vector data stored to vector data memory are disclosed. Related vector processor systems and methods are also disclosed. Reordering circuitry is provided in data flow paths between execution units and vector data memory in the VPE. The reordering circuitry is configured to reorder output vector data sample sets from execution units as a result of performing vector processing operations in-flight while the output vector data sample sets are being provided over the data flow paths from the execution units to the vector data memory to be stored. In this manner, the output vector data sample sets are stored in the reordered format in the vector data memory without requiring additional post-processing steps, which may delay subsequent vector processing operations to be performed in the execution units.


