AI Processor Matrix Multiplication Memory Access Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large-scale neural network algorithms require significant calculation time and power for processing, particularly when performing matrix multiplication operations, leading to low operation efficiency due to frequent memory accesses and high bandwidth demands.

Innovation Solution

A processor with multiple processing elements arranged in a two-dimensional matrix, where each element includes registers, performs matrix multiplication by loading matrices into registers, aligning elements, and performing operations to reduce memory access frequency and improve efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If large-scale neural network models are processed with GPU and CPU, then recognition accuracy is improved, but calculation time increases and power consumption increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidcalculation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The processor is divided into multiple processing elements arranged in a two-dimensional matrix, where each element includes at least one register. This segmentation allows parallel processing of matrix multiplication operations, reducing calculation time while maintaining recognition accuracy for large-scale neural network models.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Processing elements are arranged in a two-dimensional matrix configuration rather than traditional one-dimensional or scalar arrangements. This dimensional change enables simultaneous multi-element operations, improving computational throughput and reducing calculation time for matrix operations in neural networks.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If large-scale neural network models are processed with GPU and CPU, then recognition accuracy is improved, but power consumption increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The processor is divided into multiple processing elements arranged in a two-dimensional matrix, where each element includes at least one register. This segmentation allows parallel processing of matrix multiplication operations, reducing calculation time while maintaining recognition accuracy for large-scale neural network models.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The processor architecture enables continuous matrix multiplication operations through pipeline processing and register-based data retention, reducing the need for repeated memory accesses and thereby lowering power consumption during neural network inference.

Inventive Principle:
Principle #20Continuity of useful action

3Productivity

If traditional processors are used for matrix multiplication, then operation can be performed, but operation efficiency is low due to frequent memory accesses and high bandwidth demands

Engineering Contradiction:
Improveoperation efficiencyVSAvoidmemory access time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent combines multiple matrix multiplication operations into a single processor unit with integrated processing elements and registers. This merging eliminates the need for frequent data transfer between separate memory and processing units, reducing memory access time and improving operation efficiency for matrix multiplication tasks.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

Processing elements are arranged in a two-dimensional matrix configuration rather than traditional one-dimensional or scalar arrangements. This dimensional change enables simultaneous multi-element operations, improving computational throughput and reducing calculation time for matrix operations in neural networks.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20230169144A1Operation method, processor, and related product
Publication Date: 2023.06.01 CAMBRICON (XIAN) SEMICON CO LTD
  • US20230169144A1 patent drawing
  • US20230169144A1 patent drawing
  • US20230169144A1 patent drawing

AI summary

The present disclosure relates to an operation method, a processor, and related products that improve operation efficiency during matrix multiplication. The products include a storage component, an interface apparatus, a control component, and the an artificial intelligence chip. The artificial intelligence chip is connected to the storage component, the control component, and the interface apparatus, respectively. The storage component stores data. The interface apparatus implements data transfer between the artificial intelligence chip and an external device. The control component monitors a state of the artificial intelligence chip. .