Heterogeneous Chiplet mixed precision Transformer acceleration system based on MRAM-SRAM (Magnetic Random Access Memory-Static Random Access Memory)

By adopting MRAM-SRAM heterogeneous Chiplet technology in the Transformer acceleration system, combining the advantages of MRAM and SRAM-NPU, it solves the problem that existing accelerators are difficult to balance efficiency, energy consumption and accuracy, and achieves efficient and low-power Transformer computing acceleration.

CN120066787APending Publication Date: 2025-05-30CHENGDU ESWIN SYST IC CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510154393.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing accelerators have difficulty balancing computing efficiency, energy consumption and accuracy simultaneously, especially in the deployment of large-scale Transformer models.

Method used

Using a heterogeneous Chiplet hybrid precision Transformer acceleration system based on MRAM-SRAM, the Chiplet is calculated through MRAM, and the Chiplet is calculated through SRAM-NPU is calculated through precise calculation, and the Chiplet is coordinated by scheduling and managing Chiplet and data interaction.

Benefits of technology

The optimal combination of devices of different process nodes is realized, combining the advantages of high-density and low-power storage and computing of MRAM and the high-precision calculation of SRAM-NPU, which balances computing efficiency, energy consumption and accuracy, and supports efficient acceleration of large-scale Transformer models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066787A_ABST
    Figure CN120066787A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous Chiplet hybrid precision Transformer acceleration system based on an MRAM-SRAM (Magnetic Random Access Memory-Static Random Access Memory), which adopts a Chiplet block integration technology, gives full play to the modularization advantages of high yield and low cost of the Chiplet block integration technology, and internally integrates two types of heterogeneous calculation cores of MRAM calculation Chiplet and SRAM-NPU calculation Chiplet and two types of scheduling management Chiplet and data interaction Chiplet control units. In a Transform calculation process, an MRAM (Magnetic Random Access Memory) calculates a Chiplet to carry out approximate token relationship prediction in a self-attention mechanism, an SRAM-NPU (Synchronous Random Access Memory-Network Processing Unit) calculates the Chiplet to realize accurate attention calculation by utilizing a high-speed read-write characteristic of an SRAM (Static Random Access Memory) and a flexible calculation capability of the NPU, meanwhile, a scheduling management Chiplet coordinates calculation task allocation, and a data interaction Chiplet ensures efficient data flow; according to the acceleration system, optimal combination of different process node devices is realized through a Chiplet heterogeneous integration technology, and three dimensions of calculation efficiency, energy consumption and precision are balanced by combining the high-density low-power-consumption storage and calculation integrated advantage of an MRAM and high-precision calculation of an SRAM-NPU.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence computing acceleration, and particularly to a heterogeneous Chiplet hybrid-precision Transformer acceleration system based on MRAM-SRAM. Background Art

[0002] With the wide application of Transformer models in the fields of natural language processing, computer vision, etc., their huge computational volume and storage requirements pose severe challenges to hardware systems; traditional GPU acceleration solutions have problems such as high power consumption and low efficiency, and it is difficult to meet the growing demand for large-scale model deployment.

[0003] In terms of model optimization, mixed-precision computing provides a new design idea for hardware acceleration by maintaining high precision in critical computing paths and using low precision in non-critical paths. However, existing ASIC-based customized accelerators, although having excellent performance, lack flexibility; FPGA-based reconfigurable accelerators, although having good programmability, have low resource utilization, and CIM (Computing-in-Memory)-based accelerators, although reducing data transfer overhead, have difficult precision control. Although existing accelerators support mixed-precision computing to varying degrees, due to the characteristic that the storage system design is mainly based on a single type of memory, it is difficult to fully utilize the advantages of mixed-precision computing and it is also difficult to meet the requirements of high performance and low power consumption at the same time; in addition, the yield problem and cost pressure in the large-scale integrated circuit manufacturing process make it difficult for accelerators using traditional SoC design methods to be mass-produced on a large scale, and there are problems in being unable to balance efficiency, energy consumption, and precision at the same time. Summary of the Invention

[0004] The main object of the present invention is to provide a heterogeneous Chiplet hybrid-precision Transformer acceleration system based on MRAM-SRAM, aiming to solve the problem of being unable to balance efficiency, energy consumption, and precision at the same time.

[0005] To achieve the above object, the present invention proposes a heterogeneous Chiplet hybrid-precision Transformer acceleration system based on MRAM-SRAM. The heterogeneous Chiplet hybrid-precision Transformer acceleration system based on MRAM-SRAM includes an MRAM computing Chiplet, an SRAM-NPU computing Chiplet, a scheduling and management Chiplet, and a data interaction Chiplet that are integrated through a silicon interposer and interconnected using a micro-bump array; The MRAM computing Chiplet is used for approximate computing in the self-attention mechanism; The SRAM-NPU computing Chiplet is used for accurate computing; The scheduling management Chiplet is used for dynamic scheduling and allocation of computing tasks; The data interaction Chiplet is used for efficient management of on-chip storage resources and optimization of data access.

[0006] In one embodiment, the MRAM computing Chiplet includes a magnetic random access memory array, a decision circuit, an analog computing unit, a quantization unit, and a local control logic unit; The magnetic random access memory array adopts a magnetic tunnel structure with a perpendicular magnetization direction for non-volatile storage and analog computing; The analog computing unit is used for approximate computing; The quantization unit is used for analog-to-digital signal conversion; The local control logic unit is used for communication connection with the SRAM-NPU computing Chiplet, the scheduling management Chiplet, and the data interaction Chiplet, and coordinates and controls the quantization unit and the analog computing unit.

[0007] In one embodiment, the SRAM-NPU computing Chiplet includes a static random access memory, a neural network processor array, a mixed-precision computing unit, and an instruction controller; The static random access memory is used for caching high-speed data; The neural network processor array is used for matrix operations; The mixed-precision computing unit is used for operations with configured precision; The instruction controller is used for controlling instruction decoding and the computing process.

[0008] In one embodiment, the data interaction Chiplet includes a precision control engine and a data flow manager; The precision control engine is used for analysis and management of mixed-precision computing; The data flow manager is used for optimizing data transmission efficiency.

[0009] In one embodiment, the data flow manager includes a precision analyzer, a quantization controller, and a data reconstruction unit; The precision analyzer is used for task feature extraction and operation sensitivity analysis to identify the critical path in the computing process; The quantization controller is used for real-time monitoring of computing quality and dynamically adjusting quantization parameters to ensure precision balance between different layers of the model; The data reconstruction unit is used for performing precision conversion and data reorganization, and correcting errors through a built-in precision compensation module.

[0010] In one embodiment, the scheduling management Chiplet includes a task scheduler, a resource manager, and a performance monitoring unit; The task scheduler is used for dynamically allocating computing tasks; The resource manager is used for monitoring and allocating system resources; The performance monitoring unit is used for collecting system operation status information.

[0011] In one embodiment, the scheduling management Chiplet further includes a configuration management unit, and the configuration management unit is used for performing system parameter configuration, status maintenance, fault detection, and recovery.

[0012] The technical solution of the present invention adopts the Chiplet block integration technology, gives full play to its modular advantages of high yield and low cost, and internally integrates two types of MRAM computing Chiplets and SRAM-NPU computing Chiplet heterogeneous computing cores, as well as two types of scheduling management Chiplets and data interaction Chiplet control units. During the Transformer calculation process, the MRAM computing Chiplet performs approximate token relationship prediction in the self-attention mechanism, and the SRAM-NPU computing Chiplet uses the high-speed read / write characteristics of SRAM and the flexible computing ability of NPU to achieve accurate attention calculation. At the same time, the scheduling management Chiplet coordinates the computing task allocation, and the data interaction Chiplet ensures efficient data flow; this acceleration system realizes the optimal combination of devices with different process nodes through the Chiplet heterogeneous integration technology, combines the advantages of high-density and low-power memory-computation integration of MRAM and the high-precision computing of SRAM-NPU to balance the three dimensions of computing efficiency, energy consumption, and precision. Description of the Drawings

[0013] Figure 1 It is the overall block diagram of the heterogeneous Chiplet hybrid-precision Transformer acceleration system based on MRAM-SRAM of the present invention; Figure 2 It is the module schematic diagram of the MRAM computing Chiplet of the heterogeneous Chiplet hybrid-precision Transformer acceleration system based on MRAM-SRAM of the present invention; Figure 3 It is the module schematic diagram of the hybrid-precision calculation of the heterogeneous Chiplet hybrid-precision Transformer acceleration system based on MRAM-SRAM of the present invention; Figure 4 It is the flow schematic diagram of the hybrid-precision calculation of the data interaction Chiplet of the heterogeneous Chiplet hybrid-precision Transformer acceleration system based on MRAM-SRAM of the present invention; Figure 5 This is a schematic diagram of the module for computing precision adaptation of the heterogeneous Chiplet hybrid-precision Transformer acceleration system based on MRAM-SRAM of the present invention; Figure 6 This is a schematic diagram of the task scheduling and memory optimization process for scheduling and managing Chiplets in the heterogeneous Chiplet hybrid-precision Transformer acceleration system based on MRAM-SRAM of the present invention. Detailed implementation manners

[0014] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention generally described and illustrated in the figures herein can be arranged and designed in a variety of different configurations.

[0015] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0016] It should be noted that like reference numerals and letters denote like items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0017] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "upper", "lower", "inner", "outer", "left", "right", etc. are based on the orientation or positional relationships shown in the drawings, or the orientation or positional relationships in which the inventive product is customarily placed during use, or the orientation or positional relationships commonly understood by those skilled in the art. These are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as limiting the present invention.

[0018] In addition, the terms "first", "second", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.

[0019] In the description of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms such as "setting" and "connection" should be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0020] Although existing accelerators support mixed-precision computing to varying degrees, they are limited by the fact that the storage system design is mainly based on a single type of memory, making it difficult to fully utilize the advantages of mixed-precision computing and to meet the requirements of high performance and low power consumption at the same time. In addition, the yield issues and cost pressures in the large-scale integrated circuit manufacturing process make it difficult for accelerators using traditional SoC design methods to be mass-produced, and there is a problem of being unable to balance efficiency, energy consumption, and accuracy at the same time.

[0021] In order to solve the above problems, the present invention proposes a heterogeneous chiplet mixed precision Transformer acceleration system based on MRAM (Magnetoresistive Random Access Memory, non-volatile magnetic random access memory)-SRAM (Static Random-Access Memory, static random access memory). The specific implementation methods of the present invention are described in detail below with reference to the accompanying drawings.

[0022] like Figures 1-6 As shown in the figure, the MRAM-SRAM-based heterogeneous chiplet mixed-precision Transformer acceleration system includes an MRAM computing chiplet, an SRAM-NPU computing chiplet, a scheduling management chiplet, and a data interaction chiplet, which are integrated through a silicon interposer and interconnected using a micro-bump array; The MRAM computing chiplet is used to perform approximate calculations in the self-attention mechanism; The SRAM-NPU computing Chiplet is used to perform precise calculations; The scheduling management Chiple is used to dynamically schedule and allocate computing tasks; The data interaction chiplet is used to efficiently manage on-chip storage resources and optimize data access.

[0023] In one embodiment, the MRAM computing chiplet includes a magnetic random access memory array, a decision circuit, an analog computing unit, a quantization unit, and a local control logic unit; The magnetic random access memory array adopts a magnetic tunnel structure with a vertical magnetization direction for non-volatile storage and analog computing; The analog computing unit is used for approximate computing; The quantization unit is used for analog-to-digital signal conversion; The local control logic unit is used for communication connection with the SRAM-NPU computing Chiplet, the scheduling and management Chiplet, and the data interaction Chiplet, and coordinates and controls the quantization unit and the analog computing unit.

[0024] In one embodiment, the SRAM-NPU computing Chiplet includes a static random access memory, a neural network processor array, a mixed-precision computing unit, and an instruction controller; The static random access memory is used for caching high-speed data; The neural network processor array is used for matrix operations; The mixed-precision computing unit is used for operations with configured precision; The instruction controller is used for controlling instruction decoding and the computing process.

[0025] In this embodiment, the MRAM computing Chiplet, the SRAM-NPU (Neural Network Processing Unit) computing Chiplet, the scheduling and management Chiplet, and the data interaction Chiplet are physically connected through a silicon interposer, and a micro-bump array is used to achieve high-bandwidth inter-chip interconnection. Among them, the micro-bump array adopts a high-density interconnection structure to achieve high-bandwidth and low-latency data transmission between chips; data exchange between Chiplets is carried out through an inter-chip high-speed interconnection bus, which adopts differential signal transmission technology to support bidirectional data transmission and improves signal integrity; the system clock and control signals are distributed through a dedicated control bus to ensure synchronous operation between Chiplets.

[0026] As Figure 2As shown, the MRAM computing Chiplet includes a magnetic random access memory array, an analog computing unit, a quantization unit, and local control logic. Among them, the magnetic random access memory array is used to implement the integration of storage and computing. It adopts a magnetic tunnel structure with a vertical magnetization direction and has non-volatile storage characteristics. The magnetic random access memory array is divided into multiple sub-blocks, and each sub-block is equipped with an independent R&W read / write circuit and an SA (Sense Amplifier) to achieve efficient access and processing of parallel data; the analog computing unit includes a current comparator and an analog multiplier. The current comparator uses a high-speed comparator circuit to achieve high-precision comparison of current magnitudes, and the analog multiplier uses a current mirror network to perform multiplication operations, supporting multi-level configurable precision; the quantization unit consists of an adaptive quantization threshold adjustment circuit, a multi-level quantization precision configuration module, and an ADC / DAC conversion circuit. This quantization unit can dynamically adjust quantization parameters according to actual computing requirements to ensure the optimal balance between computing precision and efficiency; the local control logic is composed of a state machine and an address decoder, and is responsible for basic control functions such as access address generation and decoding, operation timing control, and working mode switching to ensure the orderly progress of the entire computing process.

[0027] In this embodiment, the SRAM-NPU computing Chiplet includes a static random access memory, a neural network processor array, a mixed-precision computing unit, and an instruction controller; among them, the static random access memory is divided into an instruction cache area and a data cache area, and adopts a dual-port structure to achieve parallel access; the neural network processor array is composed of a processing unit array, and each processing unit supports multiplication and addition operations and is equipped with a local register bank; the mixed-precision computing unit can support operations with multiple precisions such as floating-point and fixed-point, and includes a floating-point adder, a multiplier, and a fixed-point arithmetic unit; the instruction controller is used to parse and distribute optimized operation instructions to achieve instruction-level parallelism. In addition, the SRAM-NPU computing Chiplet also includes a vector operation unit for accelerating vectorized computing. The vector operation unit adopts a systolic array structure to improve data reuse efficiency.

[0028] The heterogeneous Chiplet hybrid-precision Transformer acceleration system based on MRAM-SRAM of the present invention uses Chiplet heterogeneous integration technology to achieve the optimal combination of devices with different process nodes, that is, it combines the advantages of high-density and low-power storage and computing integration of MRAM and the high-precision computing of SRAM-NPU to simultaneously balance the three dimensions of computing efficiency, energy consumption, and precision, and achieve large-scale mass production of the acceleration system.

[0029] In one embodiment, the data interaction Chiplet includes a precision control engine and a data flow manager; The precision control engine is used for the analysis and management of mixed-precision computing; The data flow manager is used to optimize data transmission efficiency.

[0030] In one embodiment, the data flow manager includes a precision analyzer, a quantization controller, and a data reconstruction unit; The precision analyzer is used to perform task feature extraction and operation sensitivity analysis, and identify the critical path in the calculation process; The quantization controller is used to monitor the calculation quality in real time and dynamically adjust the quantization parameters to ensure the precision balance between different layers of the model; The data reconstruction unit is used to perform precision conversion and data recombination, and perform error correction through a built-in precision compensation module.

[0031] As Figure 3 shown, the data interaction Chiplet can be used to implement a mixed-precision calculation method to process calculation tasks in an acceleration system. The method steps include: First, in the task decomposition stage, the input Transformer calculation task is parsed, the computable units that can be executed in parallel are identified, and a task dependency graph is established; Then, enter the precision analysis stage, the system evaluates the precision sensitivity of each calculation node, generates a precision requirement mapping table, and formulates a precision configuration plan. Specifically, in the task allocation stage, approximate calculation tasks are allocated to the MRAM calculation Chiplet, and accurate calculation tasks are allocated to the SRAM-NPU calculation Chiplet, and corresponding scheduling strategies are generated; In the execution control stage, parallel calculation is started, the execution status is monitored in real time, and the calculation precision is dynamically adjusted as needed; Finally, in the result fusion stage, the results of each calculation unit are collected, precision conversion and alignment are performed, and the final output result is synthesized.

[0032] As Figure 4As shown, the workflow of the data interaction Chiplet includes task decomposition, precision allocation, and result fusion. Specifically, the task decomposition module includes a computational graph analysis engine, a dependency resolver, and a task partitioning controller; the computational graph analysis engine is used to parse the computational graph structure of the Transformer model, identify computationally intensive nodes and memory access intensive nodes, and establish a data flow relationship graph between the nodes; the dependency resolver is used to analyze the data dependencies and control dependencies between tasks and generate a task execution sequence; the task partitioning controller allocates different operators to corresponding computing units according to their computational characteristics. Among them, the Key-Query matrix multiplication operation in self-attention calculation is partitioned into the MRAM computing unit, while accurate calculations such as weighted summation of weights are allocated to the SRAM-NPU computing unit. The precision allocation module includes a precision requirement analyzer, an energy consumption evaluation unit, and a dynamic precision regulator; the precision requirement analyzer is used to evaluate the sensitivity of each computing node to precision, including a sensitivity quantization module and a precision constraint generator; the energy consumption evaluation unit is used to estimate the energy consumption overhead under different precision configurations, including an energy consumption model library and a dynamic evaluation engine; the dynamic precision regulator includes a precision mapping table and a configuration update controller, which adjusts the computational precision configuration in real time based on precision requirements and energy consumption constraints. This precision allocation module can perform multiple precision configuration schemes, including low-precision approximation for attention score calculation and mixed-precision processing for feed-forward networks. The result fusion module includes a data alignment unit, a precision conversion processor, and a result verification engine; the data alignment unit is used to process the alignment operations of data with different precisions, including a data reordering module and an alignment controller; the precision conversion processor is responsible for the conversion between different precision formats, including a fixed-point to floating-point unit, a precision extension unit, and a rounding controller; the result verification engine is used to ensure the correctness of the calculation results, including an error analysis module and a quality evaluator. This result fusion module is also used to correct the cumulative errors caused by low-precision calculations to achieve an adaptive precision compensation mechanism.

[0033] It can be understood that in the task decomposition stage, the computational graph is first analyzed to identify the structure and dependencies of the computational tasks, and then through task partitioning and planning, the complex computational tasks are decomposed into smaller subtasks to facilitate parallel processing and optimize resource allocation; in the precision allocation stage, the precision requirement analysis is first carried out to determine the required computational precision according to the characteristics of the computational tasks, and then through dynamic precision adjustment, the precision is dynamically adjusted according to real-time feedback during the operation to balance the computational precision and system efficiency. In addition, the acceleration system is also provided with a quantization parameter configuration module for setting corresponding computational parameters; in the result fusion stage, data alignment is first carried out to ensure that the data is correctly aligned between different computational units, and then the results of each computational unit are fused through result fusion; finally, result verification is carried out to verify the correctness of the computational results, and after precision conversion and error qualification inspection if necessary, the final results are output. If the error exceeds the acceptable range, precision compensation will be carried out to improve the accuracy.

[0034] As Figure 5 shown, the computational precision adaptive method is used to dynamically adjust the computational precision during operation, which includes four stages: precision analysis, quantization control, data reconstruction, and dynamic adjustment. Specifically, in the precision analysis stage, the input task is feature-extracted by a precision analyzer to analyze the computational density distribution, identify the data access pattern, and complete the operation sensitivity evaluation at the same time; among them, the operation sensitivity evaluation includes the precision sensitivity analysis of different computational nodes to evaluate the impact of precision adjustment on energy consumption and the performance overhead brought by computational precision adjustment, and the analysis results will affect the decision-making results of the data flow manager; in the quantization control stage, real-time computational quality monitoring and dynamic quantization adjustment are carried out, and these factors are comprehensively considered by tracking real-time errors and monitoring performance indicators to determine the dynamic quantization adjustment strategy for each computational node to achieve precision balance and parameter self-adaptation; in the data reconstruction stage, the quantization control results are further adjusted, and the precision conversion of the data is completed through format conversion and data reorganization, and the data is corrected for bias and precision compensated to reduce errors; in the dynamic adjustment stage, the data flow manager updates the precision configuration according to the evaluation results, adjusts the quantization parameters, and activates the precision compensation mechanism if necessary. The compensation mechanism can optimize the system performance through software and hardware cooperation on the premise of ensuring the computational precision.

[0035] In one embodiment, the scheduling and management Chiplet includes a task scheduler, a resource manager, and a performance monitoring unit; The task scheduler is used for dynamically allocating computational tasks; The resource manager is used for monitoring and allocating system resources; The performance monitoring unit is used for collecting system operation status information.

[0036] In one embodiment, the scheduling and management Chiplet further includes a configuration management unit, which is used for system parameter configuration, status maintenance, fault detection and recovery.

[0037] As Figure 6 shown, in this embodiment, the task scheduling and memory optimization method is implemented by the scheduling and management Chiplet to improve the system execution efficiency. The method includes a dynamic task scheduling engine, a memory access optimizer, and a resource management controller; specifically, the dynamic task scheduling engine includes a feature extraction module, a policy generator, and an execution scheduler; among them, the feature extraction module is used to analyze the computational features and data dependency relationships of tasks, extract the resource requirement features of tasks, and evaluate the priorities of tasks; the policy generator is used to formulate scheduling policies based on task features, generate task allocation schemes, and optimize load balancing; the execution scheduler is used to implement specific scheduling decisions, coordinate the parallel execution of multiple tasks, and handle the synchronization between tasks. The memory access optimizer includes a data layout optimization unit, a memory access behavior analyzer, and a cache manager; among them, the data layout optimization unit is used to analyze data access patterns, optimize the storage location of data in memory, and reduce data movement overhead; the memory access behavior analyzer is used to monitor memory access characteristics, identify memory access hotspots and bottleneck locations, and provide memory access optimization suggestions; the cache manager is used to implement a multi-level cache policy, optimize the cache replacement algorithm, and improve the cache hit rate. The resource management controller includes a load monitoring unit, a task migration engine, and a resource allocator; among them, the load monitoring unit is used to monitor the system resource usage in real time, detect performance bottlenecks, and evaluate the system load level; the task migration engine is used to execute dynamic task migration, achieve load balancing, and handle fault recovery; the resource allocator is used to manage computing and storage resources, dynamically adjust resource allocation policies, and optimize resource utilization efficiency.

[0038] It can be understood that in the task scheduling link, the input task queue is first received, and the key attributes of the task are obtained through feature extraction. The system generates a scheduling policy based on this feature and evaluates the resource requirements. This process focuses on monitoring computational features, data dependencies, and resource requirements, etc.; the memory management link consists of three parallel optimization branches: data localization optimization, memory access behavior analysis, and cache management. The optimization results of these branches will be unified into the memory access optimization module, and the system memory access efficiency is improved through collaborative work. The system tracks the memory access status in real time through the load monitoring module; the resource control link implements a resource dynamic adjustment mechanism and executes specific scheduling decisions. The system adjusts the resource allocation policy in a timely manner through task status update and completion degree judgment. The task migration and resource reallocation modules ensure the flexibility and adaptability of resource control. The entire process forms a closed-loop optimization mechanism. When the task is not completed, the system will return to the resource evaluation link for a new round of optimization; when the task is completed, the entire execution process ends to ensure that the system can continuously optimize the task execution efficiency and resource utilization rate.

[0039] In summary, the heterogeneous Chiplet mixed-precision Transformer acceleration system based on MRAM-SRAM of the present invention adopts the Chiplet block integration technology, gives full play to its modular advantages of high yield and low cost, and internally integrates two types of MRAM computing Chiplets and SRAM-NPU computing Chiplet heterogeneous computing cores, as well as two types of scheduling and management Chiplets and data interaction Chiplet control units. During the Transformer calculation process, the MRAM computing Chiplet performs approximate token relationship prediction in the self-attention mechanism, and the SRAM-NPU computing Chiplet uses the high-speed read / write characteristics of SRAM and the flexible computing ability of NPU to achieve accurate attention calculation. At the same time, the scheduling and management Chiplet coordinates the calculation task allocation, and the data interaction Chiplet ensures efficient data flow. The acceleration system realizes the optimal combination of devices with different process nodes through the Chiplet heterogeneous integration technology, combines the advantages of high-density and low-power storage and computing integration of MRAM and the high-precision computing of SRAM-NPU to balance the three dimensions of computing efficiency, energy consumption, and precision.

[0040] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A heterogeneous chiplet mixed-precision Transformer acceleration system based on MRAM-SRAM, characterized in that: The MRAM-SRAM-based heterogeneous chiplet mixed-precision Transformer acceleration system includes an MRAM computing chiplet, an SRAM-NPU computing chiplet, a scheduling management chiplet, and a data interaction chiplet that are integrated through a silicon interposer and interconnected using a micro-bump array; The MRAM computing chiplet is used to perform approximate calculations in the self-attention mechanism; The SRAM-NPU computing Chiplet is used to perform precise calculations; The scheduling management Chiple is used to dynamically schedule and allocate computing tasks; The data interaction chiplet is used to efficiently manage on-chip storage resources and optimize data access.

2. The MRAM-SRAM-based heterogeneous chiplet mixed-precision Transformer acceleration system according to claim 1, characterized in that: The MRAM computing chiplet includes a magnetic random access memory array, a decision circuit, an analog computing unit, a quantization unit, and a local control logic unit; The magnetic random access memory array adopts a magnetic tunnel structure with a perpendicular magnetization direction and is used for non-volatile storage and analog computing; The analog computing unit is used to perform approximate computing; The quantization unit is used to perform analog-to-digital signal conversion; The local control logic unit is used to communicate with the SRAM-NPU computing chiplet, the scheduling management chiplet, and the data interaction chiplet, and coordinate and control the quantization unit and the analog computing unit.

3. The MRAM-SRAM-based heterogeneous chiplet mixed-precision Transformer acceleration system according to claim 1, characterized in that: The SRAM-NPU computing chiplet includes a static random access memory, a neural network processor array, a mixed precision computing unit and an instruction controller; The static random access memory is used to cache high-speed data; The neural network processor array is used to perform matrix operations; The mixed precision computing unit is used for configuring precision operations; The instruction controller is used to control instruction decoding and calculation process.

4. The MRAM-SRAM-based heterogeneous chiplet mixed-precision Transformer acceleration system according to claim 1, characterized in that: The data interaction chiplet includes a precision control engine and a data flow manager; The precision control engine is used to analyze and manage mixed precision calculations; The data flow manager is used to optimize data transmission efficiency.

5. The MRAM-SRAM-based heterogeneous chiplet mixed-precision Transformer acceleration system according to claim 4, characterized in that: The data flow manager includes a precision analyzer, a quantization controller, and a data reconstruction unit; The precision analyzer is used to perform task feature extraction and computational sensitivity analysis to identify the critical path in the computation process; The quantization controller is used to monitor the calculation quality in real time and dynamically adjust the quantization parameters to ensure the accuracy balance between different layers of the model; The data reconstruction unit is used to perform precision conversion and data reorganization, and perform error correction through a built-in precision compensation module.

6. The MRAM-SRAM-based heterogeneous chiplet mixed-precision Transformer acceleration system according to claim 1, characterized in that: The scheduling management chiplet includes a task scheduler, a resource manager and a performance monitoring unit; The task scheduler is used to dynamically allocate computing tasks; The resource manager is used to monitor and allocate system resources; The performance monitoring unit is used to collect system operation status information.

7. The MRAM-SRAM-based heterogeneous chiplet mixed-precision Transformer acceleration system according to claim 6, characterized in that: The scheduling management chiplet also includes a configuration management unit, which is used to perform system parameter configuration, state maintenance, fault detection and recovery.

Citation Information

Cited By

  • Attention mechanism calculation method and device, storage medium and product

    CN121052308A

  • An attention mechanism calculation method, device, storage medium and product

    CN121052308B