Reconfigurable Systolic Neural Network Engine for Training and Inference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional processors are inefficient for performing neural network computations due to high processing time and power consumption, and existing hardware accelerators like GPUs are costly and limited to inference tasks only.

Innovation Solution

A special-purpose hardware accelerator, the systolic neural network engine, uses a systolic array with data processing units connected in a local region to perform computations during both training and inference, supporting reconfiguration for various neural network architectures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional processors are used for neural network computations, then device complexity is low and versatility is high, but processing time is excessive and power consumption is high

Engineering Contradiction:
Improveprocessing speedVSAvoidpower consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent replaces conventional von Neumann architecture processors with a systolic array architecture where data flows continuously through processing elements in a pipelined manner, eliminating the need for repeated memory access cycles and enabling parallel computation of neural network operations

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The neural network computation is divided into multiple processing elements arranged in a systolic array, where each element performs a specific computation on a portion of the data, enabling parallel processing and reducing overall computation time

Inventive Principle:
Principle #1Segmentation

2Productivity

If GPUs are used as hardware accelerators, then processing speed improves, but device cost increases and adaptability is limited to inference tasks only

Engineering Contradiction:
Improveprocessing speedVSAvoidtask versatility
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The systolic array architecture is designed to be reconfigurable, allowing the same hardware to dynamically adapt between training and inference modes by changing data flow patterns and computation operations, providing both speed and versatility

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent designs a universal systolic array platform that can perform both forward propagation (inference) and backward propagation (training) operations, as well as support different neural network architectures, eliminating the need for separate hardware for different tasks

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If systolic array architecture is implemented, then processing efficiency improves, but device complexity increases

Engineering Contradiction:
Improvecomputation efficiencyVSAvoidhardware complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent combines multiple functions (data storage, computation, and data transfer) into integrated processing elements within the systolic array, reducing the need for separate components and interconnects while maintaining high computation efficiency

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11769042B2Reconfigurable systolic neural network engine
Publication Date: 2023.09.26 SANDISK TECHNOLOGIES LLC
  • US11769042B2 patent drawing
  • US11769042B2 patent drawing
  • US11769042B2 patent drawing

AI summary

Some embodiments include a special-purpose hardware accelerator that can perform specialized machine learning tasks during both training and inference stages. For example, this hardware accelerator uses a systolic array having a number of data processing units (“DPUs”) that are each connected to a small number of other DPUs in a local region. Data from the many nodes of a neural network is pulsed through these DPUs with associated tags that identify where such data was originated or processed, such that each DPU has knowledge of where incoming data originated and thus is able to compute the data as specified by the architecture of the neural network. These tags enable the systolic neural network engine to perform computations during backpropagation, such that the systolic neural network engine is able to support training.