Streaming engine with multi dimensional circular addressing selectable at each dimension

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern digital signal processors face challenges with increasing workloads, memory system latency, and memory bandwidth issues, particularly in real-time data processing, where predictable but non-sequential memory access is difficult to achieve.

Innovation Solution

A streaming engine with a fixed data stream sequence specified by a control register, featuring an address generator and a steam head register, supports nested loops with linear or circular addressing modes, and allows for efficient data streaming to functional units, reducing the need for complex memory access operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional memory access methods are used for real-time data processing, then the system can handle basic workloads, but memory bandwidth is insufficient and latency increases under increasing workloads

Engineering Contradiction:
Improvememory access efficiencyVSAvoidmemory system latency
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent segments the data stream processing into multiple independent nested loops, each with its own address generator and control parameters. This segmentation allows parallel processing of different data segments, increasing memory access efficiency and reducing latency by distributing the processing workload across multiple concurrent operations rather than sequential access.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements preliminary action by pre-configuring address generation parameters, loop counts, and data stream characteristics in control registers before processing begins. The streaming engine is pre-programmed with addressing modes and buffer sizes, enabling immediate high-speed data acquisition and processing without runtime configuration delays, thus reducing overall processing latency.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If complex memory access operations are implemented to achieve predictable non-sequential access patterns, then data processing accuracy improves, but device complexity increases

Engineering Contradiction:
Improvepredictable memory accessVSAvoidaddress generation complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a universal addressing mode control mechanism that can generate both linear and circular addressing sequences using the same hardware structure. The address generators are multi-functional, capable of producing sequential, circular, and random access patterns through programmable control parameters, eliminating the need for separate dedicated circuits for each access mode and reducing overall device complexity while maintaining reliable predictable access.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent uses parameter changes to switch between different addressing modes and access patterns. By modifying control register parameters such as loop count, buffer size, and addressing mode selection, the system can dynamically adapt to different data processing requirements without hardware reconfiguration, achieving reliable non-sequential access through software-controlled parameter adjustment rather than complex hardwired logic.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If more memory bandwidth is allocated to handle increasing workloads, then processing capacity increases, but the number of cache misses increases due to non-sequential access patterns

Engineering Contradiction:
Improveprocessing capacityVSAvoidcache hit rate
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent employs a nested buffer structure where multiple levels of circular buffers are organized hierarchically. The nested doll principle is applied by placing smaller circular buffers within larger buffer structures, allowing data to be cached at multiple granularity levels. This nesting enables efficient utilization of memory bandwidth by pre-fetching and caching data in an organized manner, reducing cache misses while maintaining the ability to handle non-sequential access patterns required for increased processing capacity.

Inventive Principle:
Principle #7Nested doll (Nesting)

4Adaptability or versatility

If traditional CPU resources are used for data streaming and memory management, then flexibility is maintained, but CPU availability for other computations decreases

Engineering Contradiction:
Improvedata stream configurationVSAvoidCPU availability
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent implements self-service by enabling the streaming engine to autonomously manage its own data stream configuration, address generation, and buffer management through dedicated control registers and hardware logic. The engine can reconfigure addressing modes, adjust buffer sizes, and control nested loop operations without CPU intervention, maintaining full adaptability for different data processing scenarios while freeing CPU resources for other computational tasks.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250217301A1Streaming engine with multi dimensional circular addressing selectable at each dimension
Publication Date: 2025.07.03 TEXAS INSTRUMENTS INC
  • US20250217301A1 patent drawing
  • US20250217301A1 patent drawing
  • US20250217301A1 patent drawing

AI summary

A streaming engine employed in a digital data processor may specify a fixed read-only data stream defined by plural nested loops. An address generator produces address of data elements for the nested loops. A steam head register stores data elements next to be supplied to functional units for use as operands. A stream template register independently specifies a linear address or a circular address mode for each of the nested loops.