Neural Network Circuit Architecture with Segmented Broadcast Bus
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing neural network circuits face challenges in implementing a complete signal processing chain on silicon, including inefficiencies in energy consumption, adaptability, and flexibility due to specialized processors, limited expandability, and variable weight dynamics, which are not adequately addressed by previous solutions.
Innovation Solution
A circuit architecture with 32 neuro-blocs, each capable of implementing a set of neurons, featuring a segmented diffusion bus, interconnection line, and a Broadcast and Computation Unit (BCU) that enables parallel communication and variable precision, allowing for efficient implementation of various neural network topologies and dynamic weight management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If specialized processors are used for neural network treatments, then neural network functionality is improved, but device complexity and energy consumption increase
Solution Approach 1:
The processor is designed with a unified architecture that can perform both conventional signal processing operations and neural network computations using the same hardware resources. The neuro-blocs can be configured to implement different neural network topologies and the broadcast bus can support various communication patterns, allowing a single processor to replace multiple specialized processors while reducing overall system complexity
2Ease of manufacture
If fixed-precision architecture is used, then manufacturing simplicity is improved, but adaptability to different applications deteriorates
Solution Approach 1:
The architecture implements variable precision arithmetic where the number of bits used for weight and activation representations can be dynamically adjusted based on the specific application requirements. Each neuro-bloc can operate with different precision levels, allowing the system to optimize between accuracy and resource usage for different neural network configurations without requiring multiple fixed-precision hardware versions
3Adaptability or versatility
If network size is increased, then neural network capability is improved, but silicon area and energy consumption increase
Solution Approach 1:
The neural network processor is divided into multiple identical neuro-blocs that can be instantiated in parallel. Each neuro-bloc is a self-contained unit that can process a portion of the neural network computations. By segmenting the overall computation across multiple identical modules, the system can scale network capability while efficiently utilizing silicon area through regular, repeatable circuit patterns rather than requiring a monolithic design
4Productivity
If parallel communication is implemented, then processing speed is improved, but energy consumption and device complexity increase
Solution Approach 1:
Multiple neuro-blocs share a common broadcast bus for communication, merging the communication infrastructure into a single shared resource rather than providing dedicated point-to-point connections between all processing units. This reduces the total number of communication channels required, lowering both energy consumption and device complexity while maintaining parallel processing capability through efficient broadcast-based data distribution
Data Source
Figure 1
Figure 2~3
AI summary
The invention applies in particular in respect of implementation of neural networks on silicon for the processing of diverse signals, including multidimensional signals such as images for example. More generally, the invention allows the efficient realization on silicon of a complete signal processing chain by the neural network approach. The circuit comprises at least: - a series of neuro-blocks grouped together by branches, a branch being formed of a group of neuro-blocks (1) and of a broadcasting bus, the neuro-blocks being connected to said broadcasting bus (2); - a routing unit (3) linked to the branches broadcasting bus, performing at least the routing and the broadcasting of data to and from said branches; - a transformation module (6) linked to the routing unit (3) by an internal bus (21) and enabled to be linked at the input of said circuit (10) to an external data bus (7), the module (6) performing the transformation of the input data into serial coded data. All the internal processing to the said circuit is performed according to a serial communication protocol.