Neural Network Computing Array with Merging Unit for Power Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Neural network computing arrays with large input and output data bit widths exhibit high power consumption and require significant chip area, limiting flexibility and efficiency in image and voice processing applications.
Innovation Solution
A computing array with process element groups arranged in two-dimensional rows and columns, featuring input, fetch and decode, operation, and output subunits, along with a merging unit that combines computing results from multiple process elements and outputs them through data lines with varying bit widths, optimizing data transmission and reducing power consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If computing arrays with large input and output data bit width are used to achieve high flexibility in neural network applications, then adaptability is improved, but power consumption increases and chip area increases
Solution Approach 1:
The computing array is divided into multiple process element groups, where each group contains four process elements arranged in a 2x2 configuration. Each process element is further segmented into functional subunits (input subunit, fetch and decode subunit, operation subunit, output subunit). This segmentation allows the system to achieve high flexibility through modular configuration while reducing power consumption by activating only the necessary segments for each computation task, rather than requiring all elements to operate at full bit width simultaneously.
2Adaptability or versatility
If computing arrays with large input and output data bit width are used to achieve high flexibility in neural network applications, then adaptability is improved, but chip area increases
Solution Approach 1:
The merging unit combines the output data from four process elements within each process element group. By merging data from multiple process elements and outputting through shared data lines, the system achieves high adaptability for different neural network configurations while reducing the total chip area. The merging unit allows multiple process elements to share common output pathways, eliminating the need for separate dedicated output lines for each process element, thus reducing overall interconnect area.
3Measurement precision
If data lines with high bit width are used for output transmission to maintain high precision computing results, then measurement precision is improved, but power consumption increases
Solution Approach 1:
The system dynamically adjusts the bit width of data lines based on the specific computation requirements. The merging unit can configure data lines to use low bit width for internal processing between process elements, and switch to high bit width only when high-precision output transmission is required. This dynamic adaptation allows the system to maintain computing precision when needed while minimizing power consumption during data transmission by using the minimum necessary bit width for each transmission task.
Data Source
AI summary
A computing array includes a plurality of process element groups, and each of the plurality of the process element groups includes four process elements arranged in two rows and two columns and a merging unit. Each of the four process elements includes an input subunit; a fetch and decode subunit configured to obtain and compile the instruction to output a logic computing type; an operation subunit configured to obtain computing result data according to the logic computing type and the operation data; an output subunit configured to output the computing result data. The merging unit is connected to the output subunit of each of the four process elements, and configured to receive the computing result data output by the output subunit of each of the four process elements, merge the computing result data and output the merged computing result data.

