Tensor Processing Weight Loading for Shared AI and DSP Arithmetic
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current integrated circuit devices struggle to efficiently perform calculations for both machine learning and digital signal processing applications, as circuitry optimized for one domain is often not well-suited for the other, leading to suboptimal performance and resource utilization.
Innovation Solution
A digital signal processing (DSP) block is introduced on integrated circuit devices, such as FPGAs, that can perform multiply, accumulate, and addition operations, utilizing flexible FPGA architecture to adapt to emerging algorithms, support both fixed-point and floating-point operations, and implement virtual bandwidth expansion for efficient arithmetic processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If circuitry is optimized for machine learning applications, then machine learning performance is improved, but digital signal processing performance deteriorates
Solution Approach 1:
The DSP block is designed with dual functionality to perform both machine learning operations (multiply-accumulate for neural network inference) and digital signal processing operations (filtering, convolution, spectral analysis). The same hardware resources including multipliers, accumulators, and data paths are configured to serve both domains, eliminating the need for separate dedicated circuitry while maintaining high performance in both applications.
Solution Approach 2:
The processing block implements dynamic reconfiguration capabilities where the same hardware resources can be dynamically allocated and configured for different operations. The block can switch between ML inference modes and DSP modes based on runtime requirements, with configurable parameters such as precision (fixed-point or floating-point), data flow patterns, and operational modes to adapt to varying computational demands of different algorithms.
2Productivity
If circuitry is optimized for digital signal processing applications, then digital signal processing performance is improved, but machine learning performance deteriorates
Solution Approach 1:
The DSP block is designed with dual functionality to perform both machine learning operations (multiply-accumulate for neural network inference) and digital signal processing operations (filtering, convolution, spectral analysis). The same hardware resources including multipliers, accumulators, and data paths are configured to serve both domains, eliminating the need for separate dedicated circuitry while maintaining high performance in both applications.
Solution Approach 2:
The processing block implements dynamic reconfiguration capabilities where the same hardware resources can be dynamically allocated and configured for different operations. The block can switch between ML inference modes and DSP modes based on runtime requirements, with configurable parameters such as precision (fixed-point or floating-point), data flow patterns, and operational modes to adapt to varying computational demands of different algorithms.
3Productivity
If dedicated hardware is used for each application domain, then application-specific performance is improved, but device complexity and resource utilization deteriorate
Solution Approach 1:
The DSP block is designed with dual functionality to perform both machine learning operations (multiply-accumulate for neural network inference) and digital signal processing operations (filtering, convolution, spectral analysis). The same hardware resources including multipliers, accumulators, and data paths are configured to serve both domains, eliminating the need for separate dedicated circuitry while maintaining high performance in both applications.
Solution Approach 2:
The patent merges previously separate ML processing units and DSP processing units into a single integrated DSP block. This consolidation combines the functionally similar computational elements (multipliers, accumulators, data memory, control logic) into unified hardware structures that can be shared between ML and DSP workloads, reducing overall device complexity and improving resource utilization efficiency.
Data Source
AI summary
The present disclosure describes a digital signal processing (DSP) block that includes a plurality of columns of weight registers and a plurality of inputs configured to receive a first plurality of values and a second plurality of values. The first plurality of values is stored in the plurality of columns of weight registers after being received. In a first mode of operation, the first and second pluralities of values are received via a first portion of the plurality of inputs. In a second mode of operation, the first plurality of values is received via a second portion of the plurality of inputs, and the second plurality of values is received via the first portion of the plurality of inputs. Additionally, the DSP block includes a plurality of multipliers configured to simultaneously multiply each value of the first plurality of values by each value of the second plurality of values.


