Reconfigurable Neural Processing Unit for Fan-In Fan-Out Adaptability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional neural network processor systems with hardware-based neurocores suffer from inefficiencies and ineffective utilization due to fixed fan-in/fan-out specifications, leading to suboptimal performance in areas like power consumption and chip performance.

Innovation Solution

A neural network processor system with reconfigurable neural processing units, comprising a plurality of neural processing cores, a router network, and a host processing unit, where each neural processing core can receive and store partial sum configuration information to form combinable neurosynaptic column chains, supporting larger fan-in and fan-out requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional hardware-based neurocores with fixed fan-in/fan-out specifications are used, then the system provides stable and reliable neural network computation, but the system suffers from inefficient neurocore utilization and suboptimal performance in power consumption and chip performance

Engineering Contradiction:
Improvefan-in/fan-out flexibilityVSAvoidpower consumption efficiency
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent implements dynamic reconfiguration of neurocores by introducing control register blocks that can be programmed with partial sum configuration information. This allows the fan-in/fan-out specifications of neurocores to change dynamically based on the specific neural network workload, rather than being fixed in hardware. The host processing unit coordinates this reconfiguration to optimize power consumption and performance for different computation scenarios.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the operational parameters of neurocores by storing partial sum configuration information in control register blocks. This configuration information modifies how neurocores process neural packets, enabling them to adapt their fan-in/fan-out capabilities. The parameter changes are achieved through software-controlled register programming rather than hardware redesign, allowing flexible adaptation to different neural network requirements.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If conventional hardware-based neurocores with fixed fan-in/fan-out specifications are used, then the system maintains simple and fixed architecture, but the system experiences ineffective neurocore utilization and inferior chip performance

Engineering Contradiction:
Improvechip performanceVSAvoidsystem architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent makes neurocores universal by equipping them with control register blocks that can be programmed with different partial sum configuration information. This allows the same neurocore hardware to perform multiple functions with different fan-in/fan-out specifications depending on the configuration. The host processing unit manages this universality by coordinating reconfiguration across the neural processing unit to achieve optimal chip performance for various neural network workloads.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system performs preliminary configuration by storing partial sum configuration information in control register blocks before neural network computation begins. The host processing unit prepares the neurocore configurations in advance, allowing the neurocores to be optimally tuned for the specific workload before execution. This preliminary action enables high chip performance without requiring complex runtime reconfiguration mechanisms.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If conventional hardware-based neurocores with fixed fan-in/fan-out specifications are used, then the system ensures deterministic behavior, but the system achieves suboptimal area utilization and power efficiency

Engineering Contradiction:
Improveneurocore reconfigurabilityVSAvoidchip area utilization
Core Design Contradiction:
Adaptability or versatilityVSArea of stationary object

Solution Approach 1:

The patent segments the neurocore functionality by separating the fixed hardware computation units from the configurable control register blocks. This segmentation allows the majority of the neurocore area to remain dedicated to efficient hardware computation while only small control registers require additional area. The host processing unit coordinates this segmented architecture to achieve optimal area utilization by configuring only the necessary control parameters for each workload.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240296311A1Neural network processor system with reconfigurable neural processing unit, and method of operating and method of forming thereof
Publication Date: 2024.09.05 AGENCY FOR SCI TECH & RES
  • US20240296311A1 patent drawing
  • US20240296311A1 patent drawing
  • US20240296311A1 patent drawing

AI summary

There is provided a neural network processor system including: a neural processing unit including a plurality of neural processing cores; a router network including a plurality of routers communicatively coupled to the plurality of neural processing cores, respectively; and a host processing unit communicatively coupled to the neural processing unit based on the router network and configured to coordinate the neural processing unit for performing neural network computations. Each neural processing core includes: a control register block configured to receive and store partial sum configuration information from the host processing unit; and a partial sum interface communicatively coupled to the control register block and configured to transmit a first partial sum neural packet generated by the neural processing core to a first another neural processing core of the plurality of neural processing cores and/or receive a second partial sum neural packet generated by a second another neural processing core of the plurality of neural processing cores, based on the partial sum configuration information stored in the neural processing core. Furthermore, a first set of neural processing cores of the plurality of neural processing cores are combinable based on the partial sum configuration information respectively stored therein to form a first neurosynaptic column chain. There is also provided a corresponding method of operating and a corresponding method of forming the neural network processor system.