Tensor Processing Weight Loading for Shared AI and DSP Arithmetic

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current integrated circuit devices struggle to efficiently perform calculations for both machine learning and digital signal processing applications, as circuitry optimized for one domain is often not well-suited for the other, leading to suboptimal performance and resource utilization.

Innovation Solution

A digital signal processing (DSP) block is introduced on integrated circuit devices, such as FPGAs, that can perform multiply, accumulate, and addition operations, utilizing flexible FPGA architecture to adapt to emerging algorithms, support both fixed-point and floating-point operations, and implement virtual bandwidth expansion for efficient arithmetic processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If circuitry is optimized for machine learning applications, then machine learning performance is improved, but digital signal processing performance deteriorates

Engineering Contradiction:
Improvemachine learning performanceVSAvoiddigital signal processing capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The DSP block is designed with dual functionality to perform both machine learning operations (multiply-accumulate for neural network inference) and digital signal processing operations (filtering, convolution, spectral analysis). The same hardware resources including multipliers, accumulators, and data paths are configured to serve both domains, eliminating the need for separate dedicated circuitry while maintaining high performance in both applications.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The processing block implements dynamic reconfiguration capabilities where the same hardware resources can be dynamically allocated and configured for different operations. The block can switch between ML inference modes and DSP modes based on runtime requirements, with configurable parameters such as precision (fixed-point or floating-point), data flow patterns, and operational modes to adapt to varying computational demands of different algorithms.

Inventive Principle:
Principle #15Dynamics

2Productivity

If circuitry is optimized for digital signal processing applications, then digital signal processing performance is improved, but machine learning performance deteriorates

Engineering Contradiction:
Improvedigital signal processing performanceVSAvoidmachine learning capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The DSP block is designed with dual functionality to perform both machine learning operations (multiply-accumulate for neural network inference) and digital signal processing operations (filtering, convolution, spectral analysis). The same hardware resources including multipliers, accumulators, and data paths are configured to serve both domains, eliminating the need for separate dedicated circuitry while maintaining high performance in both applications.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The processing block implements dynamic reconfiguration capabilities where the same hardware resources can be dynamically allocated and configured for different operations. The block can switch between ML inference modes and DSP modes based on runtime requirements, with configurable parameters such as precision (fixed-point or floating-point), data flow patterns, and operational modes to adapt to varying computational demands of different algorithms.

Inventive Principle:
Principle #15Dynamics

3Productivity

If dedicated hardware is used for each application domain, then application-specific performance is improved, but device complexity and resource utilization deteriorate

Engineering Contradiction:
Improveapplication-specific performanceVSAvoidresource utilization
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The DSP block is designed with dual functionality to perform both machine learning operations (multiply-accumulate for neural network inference) and digital signal processing operations (filtering, convolution, spectral analysis). The same hardware resources including multipliers, accumulators, and data paths are configured to serve both domains, eliminating the need for separate dedicated circuitry while maintaining high performance in both applications.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent merges previously separate ML processing units and DSP processing units into a single integrated DSP block. This consolidation combines the functionally similar computational elements (multipliers, accumulators, data memory, control logic) into unified hardware structures that can be shared between ML and DSP workloads, reducing overall device complexity and improving resource utilization efficiency.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11656872B2Systems and methods for loading weights into a tensor processing block
Publication Date: 2023.05.23 ALTERA CORP
  • US11656872B2 patent drawing
  • US11656872B2 patent drawing
  • US11656872B2 patent drawing

AI summary

The present disclosure describes a digital signal processing (DSP) block that includes a plurality of columns of weight registers and a plurality of inputs configured to receive a first plurality of values and a second plurality of values. The first plurality of values is stored in the plurality of columns of weight registers after being received. In a first mode of operation, the first and second pluralities of values are received via a first portion of the plurality of inputs. In a second mode of operation, the first plurality of values is received via a second portion of the plurality of inputs, and the second plurality of values is received via the first portion of the plurality of inputs. Additionally, the DSP block includes a plurality of multipliers configured to simultaneously multiply each value of the first plurality of values by each value of the second plurality of values.