Neural Network Processor Inter-Device Connectivity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current neural network processing engines lack efficient inter-device connectivity and flexibility in implementing artificial neural networks (ANNs), leading to suboptimal computational unit density and high power consumption.
Innovation Solution
A neural network processing engine with a chip-to-chip interface that seamlessly spreads an ANN model across multiple devices, featuring a hierarchical architecture with self-contained computational units, lean control, and dynamic resource assignment, enabling efficient computation and reduced power consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If neural network processing engines use traditional architectures, then implementation flexibility is maintained, but computational unit density is low and power consumption is high
Solution Approach 1:
The system is divided into multiple independent NN processor devices, each containing computation circuits with computing elements and dedicated memory elements. This segmentation allows parallel processing across devices while maintaining high computational density within each unit, and enables selective activation of only necessary processing units to reduce power consumption.
Solution Approach 2:
The patent extends the computational architecture from a single-device dimension to a multi-device spatial dimension through device-to-device interface circuits. This dimensional expansion allows the neural network model to be distributed across multiple physical devices, increasing overall computational unit density while enabling dynamic resource assignment to optimize power consumption based on actual processing needs.
2Adaptability or versatility
If neural network models are implemented on single devices, then inter-device connectivity complexity is avoided, but flexibility in implementing various ANN models is reduced
Solution Approach 1:
The NN processor devices are designed with universal interfaces and standardized computation circuits that can handle various neural network architectures (CNN, RNN, Transformer, etc.). The device-to-device interface circuits provide standardized connectivity protocols, allowing the same hardware platform to adapt to different ANN models without requiring complex custom inter-device wiring for each model type.
Solution Approach 2:
The system employs dynamic resource assignment where computation circuits and memory elements can be dynamically allocated and configured based on the specific requirements of different neural network models. The inter-device connectivity is dynamically established through programmable interface circuits that can reconfigure data flow paths according to the computational needs of the implemented ANN model.
3Productivity
If computational units are densely packed, then computational unit density increases, but power consumption per unit may increase
Solution Approach 1:
Each computation circuit is self-contained with its own dedicated memory elements, allowing it to operate independently with minimal communication overhead to other units. This self-sufficient design reduces the energy required for inter-unit data transfer and coordination, enabling high computational density without proportional increases in power consumption per unit.
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
A novel and useful neural network (NN) processing core incorporating inter-device connectivity and adapted to implement artificial neural networks (ANNs). A chip-to-chip interface spreads a given ANN model across multiple devices in a seamless manner. The NN processor is constructed from self-contained computational units organized in a hierarchical architecture. The homogeneity enables simpler management and control of similar computational units, aggregated in multiple levels of hierarchy. Computational units are designed with minimal overhead as possible, where additional features and capabilities are aggregated at higher levels in the hierarchy. On-chip memory provides storage for content inherently required for basic operation at a particular hierarchy and is coupled with the computational resources in an optimal ratio. Lean control provides just enough signaling to manage only the operations required at a particular hierarchical level. Dynamic resource assignment agility is provided which can be adjusted as required depending on resource availability and capacity of the device.