Computing core, accelerator, computing method and apparatus, device, non-volatile readable storage medium, and system

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing computing components perform matrix multiplication in graph neural networks through matrix partitioning, leading to excessive temporary storage of intermediate results outside the component, causing inefficient data movement and reduced computing speed.

Innovation Solution

A computing core with first and second cache components, a matrix multiplication computing component, and an output component, which loads and computes matrix multiplication results in parallel, utilizing memories with different bandwidths based on computation needs, and stores results efficiently to optimize memory usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If matrix multiplication is performed through matrix partitioning operation, then the computing operation can be completed, but a large quantity of intermediate results need to be temporarily stored outside the computing component, causing many additional data movement operations and reducing computing speed

Engineering Contradiction:
Improvecomputing speedVSAvoiddata movement time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent merges the storage function into the computing component by integrating cache memory directly within the computing core. This allows intermediate results to be stored inside the computing component rather than outside, eliminating the need for repeated data movement between storage and computing units, thereby resolving the contradiction between completing matrix multiplication and avoiding excessive data movement overhead.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary cache memory component that serves as a buffer between the computing unit and external memory. This cache stores intermediate results during matrix multiplication operations, acting as a mediator that prevents the need to repeatedly access external memory, thus reducing data movement time while maintaining computing productivity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If intermediate results are stored outside the computing component, then storage capacity is available, but additional data movement operations are required and computing resources are wasted

Engineering Contradiction:
Improvestorage capacityVSAvoidcomputing resource waste
Core Design Contradiction:
Quantity of substanceVSLoss of energy

Solution Approach 1:

The patent combines storage and computing functions into a single integrated computing component with embedded cache memory. This merger eliminates the separation between storage and computing units, allowing intermediate results to be retained within the computing component without requiring external storage, thereby preventing computing resource waste from repeated data movement operations.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements local storage capacity within the computing component through integrated cache memory. Instead of relying on distant external storage, the cache provides local storage capacity right at the computing unit, enabling fast access to intermediate results and eliminating the energy waste associated with transferring data across the system boundary.

Inventive Principle:
Principle #3Local quality

3Speed

If matrix multiplication results are stored in high bandwidth memory, then fast access is achieved for traversal computation, but memory cost and system complexity increase

Engineering Contradiction:
Improvedata access speedVSAvoidmemory system complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent applies local quality by providing high-speed cache memory specifically for storing intermediate results that require fast access, while using standard external memory for other data. This selective approach ensures fast data access speed for critical operations without requiring the entire memory system to be high-bandwidth, thereby controlling device complexity while maintaining necessary performance.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the memory system into multiple levels: a small, fast cache memory integrated within the computing component for intermediate results, and a larger, slower external memory for bulk data storage. This segmentation allows the system to achieve fast access speed where needed without incurring the high cost and complexity of making the entire memory system high-bandwidth.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20260111365A1Computing core, accelerator, computing method and apparatus, device, non-volatile readable storage medium, and system
Publication Date: 2026.04.23 INSPUR SUZHOU INTELLIGENT TECH CO LTD
  • US20260111365A1 patent drawing
  • US20260111365A1 patent drawing
  • US20260111365A1 patent drawing

AI summary

Disclosed in the present application are a computing core, an accelerator, a computing method and apparatus, a device, a non-volatile readable storage medium, and a system in the technical field of computers. According to the present application, parallel computing can be performed to obtain matrix multiplication results of N rows of data or N columns of data in a first matrix and N columns of data or N rows of data in a second matrix, so that N final matrix multiplication results may be obtained at one time. And the computing efficiency and speed are improved. The computing core does not need to temporarily store an intermediate result, and are source-on-chip is saved. According to the present application, after the matrix multiplication results are obtained, which memory the matrix multiplication results are stored into can be determined according to a participation manner in which the matrix multiplication results participate in each round of computing in a next matrix multiplication operation, and hence, a storage format of the matrix multiplication results in the memory is consistent with an output format of the matrix multiplication results when participating in computing. Data is conveniently read in sequence in a continuous computing process, and matrix transposition does not need to be carried out. Therefore, the time overhead of accessing a memory can be reduced, and the efficiency is improved.