Streaming engine with multi dimensional circular addressing selectable at each dimension
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modern digital signal processors face challenges with increasing workloads, memory system latency, and memory bandwidth issues, particularly in real-time data processing, where predictable but non-sequential memory access is difficult to achieve.
Innovation Solution
A streaming engine with a fixed data stream sequence specified by a control register, featuring an address generator and a steam head register, supports nested loops with linear or circular addressing modes, and allows for efficient data streaming to functional units, reducing the need for complex memory access operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional memory access methods are used for real-time data processing, then the system can handle basic workloads, but memory bandwidth is insufficient and latency increases under increasing workloads
Solution Approach 1:
The patent segments the data stream processing into multiple independent nested loops, each with its own address generator and control parameters. This segmentation allows parallel processing of different data segments, increasing memory access efficiency and reducing latency by distributing the processing workload across multiple concurrent operations rather than sequential access.
Solution Approach 2:
The patent implements preliminary action by pre-configuring address generation parameters, loop counts, and data stream characteristics in control registers before processing begins. The streaming engine is pre-programmed with addressing modes and buffer sizes, enabling immediate high-speed data acquisition and processing without runtime configuration delays, thus reducing overall processing latency.
2Reliability
If complex memory access operations are implemented to achieve predictable non-sequential access patterns, then data processing accuracy improves, but device complexity increases
Solution Approach 1:
The patent implements a universal addressing mode control mechanism that can generate both linear and circular addressing sequences using the same hardware structure. The address generators are multi-functional, capable of producing sequential, circular, and random access patterns through programmable control parameters, eliminating the need for separate dedicated circuits for each access mode and reducing overall device complexity while maintaining reliable predictable access.
Solution Approach 2:
The patent uses parameter changes to switch between different addressing modes and access patterns. By modifying control register parameters such as loop count, buffer size, and addressing mode selection, the system can dynamically adapt to different data processing requirements without hardware reconfiguration, achieving reliable non-sequential access through software-controlled parameter adjustment rather than complex hardwired logic.
3Productivity
If more memory bandwidth is allocated to handle increasing workloads, then processing capacity increases, but the number of cache misses increases due to non-sequential access patterns
Solution Approach 1:
The patent employs a nested buffer structure where multiple levels of circular buffers are organized hierarchically. The nested doll principle is applied by placing smaller circular buffers within larger buffer structures, allowing data to be cached at multiple granularity levels. This nesting enables efficient utilization of memory bandwidth by pre-fetching and caching data in an organized manner, reducing cache misses while maintaining the ability to handle non-sequential access patterns required for increased processing capacity.
4Adaptability or versatility
If traditional CPU resources are used for data streaming and memory management, then flexibility is maintained, but CPU availability for other computations decreases
Solution Approach 1:
The patent implements self-service by enabling the streaming engine to autonomously manage its own data stream configuration, address generation, and buffer management through dedicated control registers and hardware logic. The engine can reconfigure addressing modes, adjust buffer sizes, and control nested loop operations without CPU intervention, maintaining full adaptability for different data processing scenarios while freeing CPU resources for other computational tasks.
Data Source
AI summary
A streaming engine employed in a digital data processor may specify a fixed read-only data stream defined by plural nested loops. An address generator produces address of data elements for the nested loops. A steam head register stores data elements next to be supplied to functional units for use as operands. A stream template register independently specifies a linear address or a circular address mode for each of the nested loops.


