Cascaded Chiplet Accelerators for On-Chip AI Model Execution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large-scale artificial intelligence models are difficult to fit entirely on a processing chip due to size limitations, leading to substantial data communication bandwidth consumption and processing latency, as well as high power consumption from frequent parameter fetching and result writing.

Innovation Solution

A processing unit comprising a plurality of chiplets that can be configured to execute layers or blocks of layers of computational models, with interfaces to cascade and transfer data efficiently between chiplets, reducing the need for extensive data movement by keeping model parameters on-chip.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If a monolithic integrated circuit chip is used to execute artificial intelligence models, then the processing unit can execute computations, but the physical size limits the capacity to fit large-scale AI model parameters on the chip

Engineering Contradiction:
Improvecapacity to fit AI model parametersVSAvoidphysical size of chip
Core Design Contradiction:
Quantity of substanceVSArea of stationary object

Solution Approach 1:

The patent divides the monolithic integrated circuit chip into multiple smaller chiplets, each capable of executing portions of AI model computations. This segmentation allows the system to handle large-scale AI models by distributing parameters across multiple chiplets, effectively increasing the total parameter capacity without requiring a single large chip.

Inventive Principle:
Principle #1Segmentation

2Productivity

If parameters for each layer are brought onto the chip and results are written back to memory, then the computation can be performed, but this consumes substantial data communication bandwidth and creates processing latency

Engineering Contradiction:
Improvecomputation executionVSAvoidprocessing latency
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent merges multiple chiplets into a unified processing system where intermediate results can be retained on-chip across chiplet boundaries. This allows continuous computation across layer boundaries without frequent writes to external memory, reducing data communication bandwidth consumption and processing latency.

Inventive Principle:
Principle #5Merging (Combining)

3Productivity

If parameters are frequently fetched from memory and results are written back, then the computation can proceed, but this consumes a substantial amount of power

Engineering Contradiction:
Improvecomputation executionVSAvoidpower consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent implements preliminary action by pre-loading AI model parameters into the chiplet memory structures before computation begins. This allows the computation to proceed using on-chip parameters without frequent fetches from external memory during execution, significantly reducing power consumption associated with data movement.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20260030425A1Chiplet based computational accelerators and configuration methods
Publication Date: 2026.01.29 MEMRYX INC
  • US20260030425A1 patent drawing
  • US20260030425A1 patent drawing
  • US20260030425A1 patent drawing

AI summary

A processing unit can include a plurality of chiplets coupled in a cascade topology by a plurality of interfaces. A set of the plurality of cascade coupled chiplets can be configured to execute a plurality of layers or blocks of layers of a computational model. The set of cascade coupled chiplets can also be configured with parameter data of corresponding ones of the plurality of layers or blocks of layers of the computational model.