Neural Network Data Moving Controller for Matrix Operation Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current AI semiconductor technologies face limitations in computation power, leading to inefficiencies in neural network processing, particularly with deep learning applications, which require significantly more performance than traditional CPU- and GPU-based architectures, resulting in high power consumption and scalability issues.

Innovation Solution

A neural network system with a data moving controller that reorders multidimensional matrix data for efficient processing, utilizing an internal memory and operator for multidimensional matrix multiplication, and a data moving controller to manage data exchange between external and internal memory, optimizing access addresses for improved performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If CPU- and GPU-based distributed computing techniques are used to improve AI computing performance, then computation capability is enhanced, but power consumption increases to several kilowatts or more

Engineering Contradiction:
ImproveAI computing performanceVSAvoidpower consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system segments the computing architecture into distinct functional units: external memory for bulk storage, internal memory for active processing, and specialized operators for specific operations. This segmentation allows each component to be optimized independently, enabling high-performance AI computing with reduced power consumption by avoiding the need for a single high-power general-purpose processor.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary data movement and processing architecture between external and internal memory, using specialized operators to bridge the gap. This intermediary structure enables efficient data flow and computation without requiring excessive power consumption associated with traditional CPU-GPU distributed systems.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If semiconductor processes are continuously miniaturized to improve computation density, then processing capability increases, but physical limitations on process scaling are reached

Engineering Contradiction:
Improvecomputation densityVSAvoidprocess scaling limitations
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

Instead of continuing to miniaturize in the same dimensional direction, the patent transitions to a different architectural dimension by implementing a heterogeneous system with multiple memory hierarchies and specialized operators. This dimensional shift allows continued performance improvement without further miniaturization, overcoming physical scaling limitations.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If deep learning applications are implemented to improve AI performance, then processing capability increases 1000 times or more, but power consumption becomes prohibitively high for commercialization

Engineering Contradiction:
Improvedeep learning performanceVSAvoidpower consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent applies local quality by designing specialized operators optimized for specific deep learning operations (such as convolution, activation functions, and pooling) rather than using general-purpose processors. Each operator is locally optimized for its specific function, enabling high deep learning performance with significantly reduced power consumption suitable for commercial applications.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11068394B2Neural network system including data moving controller
Publication Date: 2021.07.20 ELECTRONICS & TELECOMM RES INST
  • US11068394B2 patent drawing
  • US11068394B2 patent drawing
  • US11068394B2 patent drawing

AI summary

Provided is a neural network system for processing data transferred from an external memory. The neural network system includes an internal memory storing input data transferred from the external memory, an operator performing a multidimensional matrix operation by using the input data of the internal memory and transferring a result of the multidimensional array operation as output data to the internal memory, and a data moving controller controlling an exchange of the input data or the output data between the external memory and the internal memory. The data moving controller reorders a dimension order with respect to an access address of the external memory to generate an access address of the internal memory, for the multidimensional matrix operation.