Neural Network Data Moving Controller for Matrix Operation Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current AI semiconductor technologies face limitations in computation power, leading to inefficiencies in neural network processing, particularly with deep learning applications, which require significantly more performance than traditional CPU- and GPU-based architectures, resulting in high power consumption and scalability issues.
Innovation Solution
A neural network system with a data moving controller that reorders multidimensional matrix data for efficient processing, utilizing an internal memory and operator for multidimensional matrix multiplication, and a data moving controller to manage data exchange between external and internal memory, optimizing access addresses for improved performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If CPU- and GPU-based distributed computing techniques are used to improve AI computing performance, then computation capability is enhanced, but power consumption increases to several kilowatts or more
Solution Approach 1:
The system segments the computing architecture into distinct functional units: external memory for bulk storage, internal memory for active processing, and specialized operators for specific operations. This segmentation allows each component to be optimized independently, enabling high-performance AI computing with reduced power consumption by avoiding the need for a single high-power general-purpose processor.
Solution Approach 2:
The patent introduces an intermediary data movement and processing architecture between external and internal memory, using specialized operators to bridge the gap. This intermediary structure enables efficient data flow and computation without requiring excessive power consumption associated with traditional CPU-GPU distributed systems.
2Productivity
If semiconductor processes are continuously miniaturized to improve computation density, then processing capability increases, but physical limitations on process scaling are reached
Solution Approach 1:
Instead of continuing to miniaturize in the same dimensional direction, the patent transitions to a different architectural dimension by implementing a heterogeneous system with multiple memory hierarchies and specialized operators. This dimensional shift allows continued performance improvement without further miniaturization, overcoming physical scaling limitations.
3Productivity
If deep learning applications are implemented to improve AI performance, then processing capability increases 1000 times or more, but power consumption becomes prohibitively high for commercialization
Solution Approach 1:
The patent applies local quality by designing specialized operators optimized for specific deep learning operations (such as convolution, activation functions, and pooling) rather than using general-purpose processors. Each operator is locally optimized for its specific function, enabling high deep learning performance with significantly reduced power consumption suitable for commercial applications.
Data Source
AI summary
Provided is a neural network system for processing data transferred from an external memory. The neural network system includes an internal memory storing input data transferred from the external memory, an operator performing a multidimensional matrix operation by using the input data of the internal memory and transferring a result of the multidimensional array operation as output data to the internal memory, and a data moving controller controlling an exchange of the input data or the output data between the external memory and the internal memory. The data moving controller reorders a dimension order with respect to an access address of the external memory to generate an access address of the internal memory, for the multidimensional matrix operation.


