Butterfly Network Load Return Alignment for Cache Stall Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern digital signal processors face challenges with increasing workloads, memory bandwidth limitations, and scheduling issues when operating on real-time data, particularly in video encoding applications, where predictable memory access is crucial but difficult to achieve within existing address generation and memory access resources.

Innovation Solution

A digital data processor with a streaming engine that recalls a predetermined sequence of data elements for processing, utilizing a multilayer butterfly network to transform and align data streams, and employing precalculated inputs and simple combinatorial logic to generate control signals, allowing for efficient data transformations while reducing complexity in multiplexor control circuits.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a multilayer butterfly network is used to transform and align data streams, then data transformation capability and bandwidth are improved, but device complexity increases due to multiple multiplexers and control circuits

Engineering Contradiction:
Improvedata transformation capabilityVSAvoidmultiplexor control circuit complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent pre-calculates control signals for the butterfly network multiplexers using simple combinatorial logic based on input and output port indices. Instead of using complex control circuits to dynamically determine multiplexer settings, the control signals are pre-determined by the formula: control_signal = (output_port_index - input_port_index) / 2. This preliminary calculation approach simplifies the control circuitry while maintaining full data transformation capability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the control approach from complex dynamic control to simple parameter-based control. By using the port indices as parameters in a straightforward mathematical relationship, the control signals become predictable and easily generated, reducing the complexity of control circuits while preserving the network's transformation capabilities.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If conventional address generation and memory access resources are used, then device complexity is kept low, but memory access efficiency and bandwidth are insufficient for real-time data processing

Engineering Contradiction:
Improvememory access efficiencyVSAvoidaddress generation resources
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the memory access function from the main processor by introducing a dedicated streaming engine. This separate component handles all memory access operations, allowing the processor to focus on data processing. The streaming engine includes dedicated address generation units and buffer memory, creating a specialized subsystem that improves memory access efficiency without increasing the complexity of the main processor.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a buffer memory as an intermediary between the streaming engine and the main memory system. This buffer acts as a mediator that decouples memory access operations from processing operations, allowing the processor to continue working while data is being loaded or stored. The buffer memory absorbs the complexity of address generation and timing coordination, improving overall system efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of time

If data is loaded from memory through conventional pipelines, then device complexity is minimized, but cache miss stalls and processing delays increase

Engineering Contradiction:
Improvecache miss stallsVSAvoidstreaming engine architecture
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The patent merges the data loading and processing functions into a unified streaming engine architecture. The streaming engine combines address generation, memory access, buffer management, and data formatting in a single integrated subsystem. This merging eliminates the need for separate memory access operations that would cause cache misses, as data is directly streamed from memory through the buffer to the processor in a continuous flow, reducing processing delays.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The streaming engine performs preliminary actions by pre-loading data into the buffer memory before the processor needs it. The address generation unit continuously generates addresses and loads data into the buffer in advance, so that when the processor requests data, it is already available. This preliminary data preparation eliminates cache miss stalls and ensures continuous processing.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10530397B2Butterfly network on load data return
Publication Date: 2020.01.07 TEXAS INSTRUMENTS INC
  • US10530397B2 patent drawing
  • US10530397B2 patent drawing
  • US10530397B2 patent drawing

AI summary

A method is shown that is operable to transform and align a plurality of fields from an input to an output data stream using a multilayer butterfly or inverse butterfly network. Many transformations are possible with such a network which may include separate control of each multiplexer. This invention supports a limited set of multiplexer control signals, which enables a similarly limited set of data transformations. This limited capability is offset by the reduced complexity of the multiplexor control circuits.