Configurable Offloading System FPGA Neural Network Acceleration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional approaches to handling artificial intelligence (AI), machine learning (ML), and neural network operations using dedicated custom integrated circuits (ASICs) are expensive and lack flexibility, while programmable semiconductor devices (PSDs) like FPGAs offer flexibility but may not fully address the need for high-speed processing.

Innovation Solution

A configurable offloading system (COS) utilizing a field-programmable gate array (FPGA) with an input memory, processing unit, and accelerator for parallel processing of neural network operations, allowing for offloading of intensive compute tasks from the main processor to specialized accelerators, thereby enhancing system performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If dedicated custom integrated circuits (ASICs) are used to handle AI, ML, and neural network operations, then processing speed and efficiency are improved, but cost increases and flexibility is reduced

Engineering Contradiction:
Improveprocessing speedVSAvoidflexibility
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent employs FPGAs which allow dynamic reconfiguration of hardware logic to match different neural network architectures and operations. The configurable logic blocks can be programmed to implement specific AI/ML functions while maintaining high-speed parallel processing capabilities, thus achieving both speed and flexibility.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The FPGA-based accelerator is designed with universal interfaces and configurable logic that can handle multiple types of neural network operations (convolutional layers, fully connected layers, activation functions). This multi-functional design allows a single device to replace multiple dedicated ASICs while maintaining processing efficiency.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If dedicated custom integrated circuits (ASICs) are used to handle AI, ML, and neural network operations, then processing efficiency is improved, but device cost increases

Engineering Contradiction:
Improveprocessing efficiencyVSAvoiddevice cost
Core Design Contradiction:
ProductivityVSEase of manufacture

Solution Approach 1:

The patent utilizes FPGAs which, while not as cheap as ASICs, are significantly more cost-effective for low-volume and prototype production. The reprogrammable nature allows for rapid iteration and optimization without the high NRE costs of ASIC manufacturing, making it economically viable for smaller-scale deployments.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Adaptability or versatility

If programmable semiconductor devices (FPGAs) are used instead of ASICs, then flexibility is improved, but processing speed may be reduced

Engineering Contradiction:
ImproveflexibilityVSAvoidprocessing speed
Core Design Contradiction:
Adaptability or versatilityVSSpeed

Solution Approach 1:

The patent divides the FPGA into multiple configurable logic blocks (CLBs) that can operate in parallel to process different portions of neural network computations simultaneously. This segmentation enables the FPGA to achieve high throughput comparable to ASICs while maintaining reconfigurability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent utilizes the spatial dimension of the FPGA by implementing parallel processing architectures where multiple computation units operate concurrently. This dimensional approach to parallelism compensates for the inherent speed limitations of reconfigurable devices compared to fixed ASIC implementations.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20220092398A1Method and Apparatus for Offloading Tasks to Accelerator for Enhancing System Performance Using Configurable Devices
Publication Date: 2022.03.24 GOWIN SEMICON CORP LTD
  • US20220092398A1 patent drawing
  • US20220092398A1 patent drawing
  • US20220092398A1 patent drawing

AI summary

A method and/or apparatus using programmable device for parallel processing logic operations is disclosed. The apparatus, such as a semiconductor integrated circuit die, includes an input memory, a processing unit, and an accelerator. The input memory is used to buffer input signals from an external component. The processing unit, such as a microcontroller, retrieves the input signals from the input memory and generates pre-processed data in accordance with the input signals. The first configured circuit containing configurable logic blocks (“LBs”) of a field programmable logic array (“FPGA”), in one embodiment, is programmed as an accelerator to perform one or more neural networking functions. For example, the accelerator is able to process a set of convolutional operation in response to at least a portion of the pre-processed data offloaded from the processing unit for identifying a result or reference.