DSP Block Sparsity Operations via Multi-Level Crossbar Routing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Programmable logic devices, such as FPGAs, inefficiently implement artificial intelligence architectures due to inefficient sparsity operations, which hinder their performance in machine-learning and AI applications.

Innovation Solution

A digital signal processing (DSP) block is integrated into these devices, capable of implementing sparsity modes using multiple sparsity ratios through minimal routing resources, by decomposing tensor columns into sub-columns and employing multi-level crossbar architectures and multiplexer patterns to select inputs based on sparsity modes, thereby enabling efficient routing and cascading of data across multiple DSP blocks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If sparsity operations are implemented in programmable logic devices, then multiplication operations can be performed more efficiently, but the device complexity increases due to routing resource requirements

Engineering Contradiction:
Improvemultiplication operation efficiencyVSAvoidrouting resources
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The tensor columns are decomposed into sub-columns, and the sparsity operations are divided into multiple stages using multi-level crossbar architectures. This segmentation allows the system to handle sparsity operations in manageable segments rather than requiring all routing resources simultaneously, thus improving multiplication efficiency while controlling device complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs multi-level crossbar architectures that add spatial dimensions to the routing structure. By organizing crossbars in multiple levels rather than a single plane, the system can route sparse data more efficiently through additional dimensional pathways, improving productivity without linearly increasing routing resource complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If multiple sparsity ratios are supported, then adaptability for different AI architectures is improved, but the device complexity increases due to additional routing configurations

Engineering Contradiction:
Improvesparsity ratio supportVSAvoidrouting configurations
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The DSP block is designed with universal multi-level crossbar architectures that can be configured to support multiple sparsity ratios (e.g., 2:4, 4:8, 8:16). The same physical routing infrastructure serves multiple sparsity modes through reconfigurable multiplexer patterns, enabling adaptability across different AI architectures without requiring separate dedicated routing for each sparsity ratio, thus avoiding proportional complexity increases.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The routing configurations are made dynamic and reconfigurable rather than fixed. The system can adapt its routing patterns to match different sparsity ratios as needed, allowing a single device to serve multiple AI architecture requirements. This dynamic reconfiguration capability provides versatility while keeping the base device complexity manageable through shared infrastructure.

Inventive Principle:
Principle #15Dynamics

3Productivity

If tensor columns are decomposed into sub-columns, then sparsity operations can be performed more efficiently, but the processing time increases due to additional routing stages

Engineering Contradiction:
Improvesparsity calculation efficiencyVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The multi-level crossbar architectures are designed to maintain continuous data flow through the decomposition stages. Rather than introducing significant idle time between sub-column processing stages, the system keeps data moving continuously through the routing pipeline, minimizing the time penalty associated with tensor column decomposition while still achieving efficient sparsity operations.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

Data is pre-positioned and pre-configured in the crossbar structures before the actual multiplication operations begin. The routing paths are established in advance, and input data is staged in appropriate locations, so that when processing begins, the decomposition and multiplication can proceed with minimal delay, reducing the overall processing time impact.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4155901A1Systems and methods for sparsity operations in a specialized processing block
Publication Date: 2023.03.29 ALTERA CORP
  • EP4155901A1 patent drawingFigure 1
  • EP4155901A1 patent drawingFigure 2
  • EP4155901A1 patent drawingFigure 3

AI summary

This disclosure is directed to a digital signal processing (DSP) block that includes multiple weight registers configurable to receive and store a first plurality of values, and multiple multipliers that are each configurable to receive a respective value of the first plurality of values. The DSP block further includes one or more inputs configurable to receive a second plurality of values, and a multiplexer network configurable to receive the second plurality of values and route each respective value of the second plurality of values to a multiplier of the multipliers. The multipliers are configurable to simultaneously multiply each value of the first plurality of values by a respective value of the second plurality of values to generate a plurality of products. Additionally, the DSP block includes adder circuitry configurable to generate a first sum and a second sum based on the plurality of products.