Neural Network Computing Array with Merging Unit for Power Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Neural network computing arrays with large input and output data bit widths exhibit high power consumption and require significant chip area, limiting flexibility and efficiency in image and voice processing applications.

Innovation Solution

A computing array with process element groups arranged in two-dimensional rows and columns, featuring input, fetch and decode, operation, and output subunits, along with a merging unit that combines computing results from multiple process elements and outputs them through data lines with varying bit widths, optimizing data transmission and reducing power consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If computing arrays with large input and output data bit width are used to achieve high flexibility in neural network applications, then adaptability is improved, but power consumption increases and chip area increases

Engineering Contradiction:
ImproveflexibilityVSAvoidpower consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The computing array is divided into multiple process element groups, where each group contains four process elements arranged in a 2x2 configuration. Each process element is further segmented into functional subunits (input subunit, fetch and decode subunit, operation subunit, output subunit). This segmentation allows the system to achieve high flexibility through modular configuration while reducing power consumption by activating only the necessary segments for each computation task, rather than requiring all elements to operate at full bit width simultaneously.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If computing arrays with large input and output data bit width are used to achieve high flexibility in neural network applications, then adaptability is improved, but chip area increases

Engineering Contradiction:
ImproveflexibilityVSAvoidchip area
Core Design Contradiction:
Adaptability or versatilityVSArea of stationary object

Solution Approach 1:

The merging unit combines the output data from four process elements within each process element group. By merging data from multiple process elements and outputting through shared data lines, the system achieves high adaptability for different neural network configurations while reducing the total chip area. The merging unit allows multiple process elements to share common output pathways, eliminating the need for separate dedicated output lines for each process element, thus reducing overall interconnect area.

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If data lines with high bit width are used for output transmission to maintain high precision computing results, then measurement precision is improved, but power consumption increases

Engineering Contradiction:
Improvecomputing precisionVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system dynamically adjusts the bit width of data lines based on the specific computation requirements. The merging unit can configure data lines to use low bit width for internal processing between process elements, and switch to high bit width only when high-precision output transmission is required. This dynamic adaptation allows the system to maintain computing precision when needed while minimizing power consumption during data transmission by using the minimum necessary bit width for each transmission task.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12013809B2Computing array and processor having the same
Publication Date: 2024.06.18 BEIJING TSINGMICRO INTELLIGENT TECH CO LTD
  • US12013809B2 patent drawing
  • US12013809B2 patent drawing

AI summary

A computing array includes a plurality of process element groups, and each of the plurality of the process element groups includes four process elements arranged in two rows and two columns and a merging unit. Each of the four process elements includes an input subunit; a fetch and decode subunit configured to obtain and compile the instruction to output a logic computing type; an operation subunit configured to obtain computing result data according to the logic computing type and the operation data; an output subunit configured to output the computing result data. The merging unit is connected to the output subunit of each of the four process elements, and configured to receive the computing result data output by the output subunit of each of the four process elements, merge the computing result data and output the merged computing result data.