A low-privilege adaptive dequantization and on-chip fusion inference method and system

CN122840131APending Publication Date: 2026-09-29JIANGNAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611320985.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-28
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0015]为此,本发明实施例提供了一种低比特权重自适应反量化与片上融合推理方法及系统,用于解决现有技术中因缺乏低比特原生算术支持而导致低比特权重难以在异构众核平台上高效执行,且现有反量化方案无法根据不同低比特格式与硬件资源自适应调整、以及权重搬运与计算阶段流水不均衡的问题

Benefits of technology

第一,本发明通过格式感知的硬件反馈式剪枝搜索方法,根据低比特权重格式与目标硬件资源特征,自动确定最优反量化路径与并行粒度,解决了固定反量化配置难以适配不同格式和不同硬件的问题,降低了人工调优和实机测试开销。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840131A_ABST
    Figure CN122840131A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for low-bit-weight adaptive dequantization and on-chip fusion inference, relating to the field of large language model inference optimization technology. The method includes: acquiring hardware resource feature information and low-bit-weight format feature information; determining the initial dequantization execution tendency based on the format feature information, and prioritizing candidate dequantization configurations based on the hardware resource feature information; running the candidate configurations on the target hardware to obtain real-machine feedback, and determining the target dequantization configuration from it; moving the low-bit-weights from global storage to on-chip local storage according to the target configuration and dequantizing them into high-precision data; continuously feeding the dequantized high-precision data into the computing unit in segments for matrix multiplication and addition operations, without forming a complete high-precision weight tensor. This invention improves the efficiency and hardware adaptability of low-bit-weight inference through format-aware configuration search, on-chip table lookup and SIMD dual-path dequantization, and a two-stage, double-buffered pipeline.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large language model inference optimization technology, and in particular to a low-bit-weight adaptive dequantization and on-chip fusion inference method and system. Background Technology

[0002] With the widespread application of large language models in scenarios such as intelligent question answering, code generation, text understanding, complex reasoning, and intelligent agent systems, the scale of model parameters, context length, and service concurrency continue to grow. This places higher demands on accelerator card memory capacity, off-chip memory access bandwidth, inter-card communication capabilities, and multi-machine interconnection capabilities during the inference phase. Unlike the training phase, inference not only requires long-term resident model weights but also necessitates maintaining intermediate states such as key-value caches (KV caches) during the generation process. In scenarios with long contexts and high concurrency, KV caches grow significantly with context length, batch size, and number of layers, becoming a significant storage overhead in large model inference. Existing inference frameworks improve memory utilization through mechanisms such as paged attention and continuous batch processing. KV cache quantization is also used to reduce cache footprint to support longer contexts and higher throughput.

[0003] In practical deployments, model weights and KV cache jointly determine whether a model can be deployed on a single GPU. When the model is large or needs to support long contexts and high concurrency, a single GPU's memory is often insufficient to accommodate the complete model and runtime cache. The system must then employ a single-machine multi-GPU or multi-machine multi-GPU approach for model partitioning, tensor parallelism, pipelined parallelism, or request scheduling. In this case, inference performance is not only determined by the single GPU's computing power but is also affected by inter-GPU communication, host-to-device transmission, and multi-machine network interconnection. For some domestically produced intelligent accelerator cards, single-card half-precision computing power has gradually approached the level of advanced accelerators, but memory bandwidth, inter-GPU interconnection, and multi-machine scheduling capabilities may still become efficiency bottlenecks. Therefore, reducing model weights and runtime cache usage can not only reduce memory requirements but also potentially reduce cross-GPU and cross-machine data transmission and synchronization, amplifying the system-level benefits of quantization inference.

[0004] Low-bit quantization is an important technique for reducing the resource overhead of large language model inference. This type of method typically compresses model weights from high-precision formats such as half-precision floating-point numbers (FP16), BF16, or single-precision floating-point numbers (FP32) to 8-bit floating-point numbers (FP8), 8-bit integers (INT8), or other low-bit formats, thereby reducing storage space and data transfer volume. INT8 quantization has been applied in many types of inference systems, while FP8, due to its better balance between representation range, storage compression ratio, and model accuracy, is gradually becoming an important low-bit format for large model inference. Some advanced graphics processing units (GPUs) already provide hardware-level support for low-bit computing such as FP8. For example, the NVIDIA H100 supports FP8 precision through its fourth-generation Tensor Cores and Transformer Engine; the NVIDIA Transformer Engine also provides FP8-related capabilities for architectures such as Hopper, Ada, and Blackwell.

[0005] However, different low-bit data formats have significant differences in bit field structure and numerical recovery path. Different intelligent acceleration platforms also differ in Single Instruction Multiple Data (SIMD) vector width, number of on-chip memory banks, lookup table capability, register resources, and data transport mechanism. It is difficult to achieve better execution efficiency simultaneously on different low-bit formats and different hardware platforms with the same dequantization implementation path and fixed parallel configuration.

[0006] Currently, low-bit, high-language model inference mainly falls into the following categories: Firstly, there is the inference method based on hardware-native low-bit computing units, which accelerates the low-bit matrix computing units or dedicated instructions provided within the chip to enable formats such as INT8 and FP8 to directly participate in matrix multiplication. This method can reduce software dequantization overhead, but it relies on the hardware itself having native low-bit arithmetic capabilities. It is not suitable for domestic heterogeneous many-core platforms that only support finite-precision computing formats such as FP16 and lack native low-bit computing units.

[0007] Secondly, a software framework adaptation method based on the NVIDIA General-Purpose Graphics Processing Unit (GPGPU) ecosystem is employed. This method leverages the Compute Unified Device Architecture (CUDA) programming model, Single Instruction Multiple Thread (SIMT) execution mechanism, Tensor Core data pathways, and mature inference frameworks to support low-bit format inference such as INT8 and FP8 through software simulation. While related frameworks or operators reduce low-bit inference overhead through weight rearrangement, block scheduling, and dequantization fusion, their optimization paths are closely tied to the GPGPU-specific instruction set, making direct migration to non-GPGPU-based domestic heterogeneous many-core platforms difficult.

[0008] Third, low-bit inference adaptation methods based on domestic CUDA compatibility or domestic artificial intelligence (AI) chip ecosystems are adapted and packaged for domestic GPUs or CUDA-like runtime environments, but are usually tied to specific chip architectures and software stacks. For platforms such as the Shenwei intelligent accelerator card, which uses explicit direct memory access (DMA) for data transfer, on-chip local storage, and heterogeneous many-core organization, and whose main execution path is FP16 matrix computation, existing solutions are difficult to reuse directly. Moreover, different low-bit formats, different inverse quantization paths, and different parallel granularities result in significant differences in the utilization efficiency of on-chip memory, SIMD execution resources, and pipeline scheduling. Operator optimization is usually required for specific hardware.

[0009] In summary, existing technologies generally have the following shortcomings: First, low-bit computing capabilities rely on native hardware support. Existing solutions depend on FP8 and INT8 low-bit matrix computation units provided within the chip, allowing low-bit weights to directly participate in matrix multiplication and addition operations. This requires the hardware to have native computation capabilities for the corresponding low-bit format. For domestically produced heterogeneous many-core platforms that primarily perform half-precision matrix computations such as FP16 and lack native low-bit arithmetic and conversion units, such solutions are not directly applicable.

[0010] Second, the optimization methods for the GPGPU ecosystem are not well-coordinated with the domestic heterogeneous many-core architecture. Domestic heterogeneous many-core platforms, such as the Shenwei intelligent accelerator card, typically employ explicit DMA data transfer, on-chip local storage, and heterogeneous core collaborative execution. Their data transfer granularity, on-chip storage organization, computational unit calling methods, and parallel execution models differ from GPGPUs. Existing low-bit optimization methods for CUDA are difficult to directly transfer, and simple reuse often fails to leverage the advantages of on-chip storage and explicit data transfer in domestic heterogeneous many-core platforms.

[0011] Third, the dequantization calculation characteristics of different low-bit numerical formats differ significantly. FP8_E4M3 (an 8-bit floating-point format with a 4-bit exponent and a 3-bit mantissa), FP8_E5M2 (an 8-bit floating-point format with a 5-bit exponent and a 2-bit mantissa), and INT8 differ in the sign bit, exponent bit, mantissa bit, and scaling method. A unified element-by-element parsing method or fixed-bit transformation process is prone to redundant bit operations and additional numerical corrections, while designing a separate conversion implementation for each format results in operator fragmentation, increasing adaptation and maintenance costs.

[0012] Fourth, fixed dequantization paths and parallel granularities are difficult to adapt to different hardware resources. Different intelligent acceleration platforms differ in terms of SIMD vector width, on-chip memory volume, and register resources. The same dequantization algorithm exhibits varying performance characteristics under different hardware or parallel granularities. Existing low-bit inference implementations often rely on preset parameters or developer experience to determine execution configurations. After changing the hardware platform or low-bit format, extensive manual testing and operator tuning are required, resulting in high adaptation costs.

[0013] Fifth, low-bit weights are difficult to efficiently reside on-chip and integrate with the computation pipeline. For domestic heterogeneous many-core platforms that employ explicit DMA transport and on-chip local storage, low-bit weight transport, on-chip dequantization, and matrix computation each have different execution latencies and resource requirements. If only the traditional data transport-computation double-buffering approach is used, it is difficult to simultaneously coordinate the three execution stages of weight prefetching, dequantization data generation, and matrix computation consumption, which can easily lead to inter-stage waiting and pipeline imbalance, limiting the deployment effectiveness of the low-bit model.

[0014] Therefore, existing low-bit inference technologies have not fully solved the problem of how to determine efficient dequantization configuration based on low-bit format structure and target hardware execution resources on domestic heterogeneous many-core intelligent acceleration platforms that lack native low-bit arithmetic support, and how to coordinate the multi-stage pipelined execution of weight transfer, on-chip dequantization and matrix computation. Summary of the Invention

[0015] To address these issues, this invention provides a low-bit weight adaptive dequantization and on-chip fusion inference method and system, which solves the problems in the prior art where the lack of native low-bit arithmetic support makes it difficult to efficiently execute low-bit weights on heterogeneous many-core platforms, and where existing dequantization schemes cannot adaptively adjust according to different low-bit formats and hardware resources, as well as the uneven pipeline in the weight transfer and calculation stages.

[0016] To address the aforementioned technical problems, embodiments of the present invention provide a low-bit-weighted adaptive dequantization and on-chip fusion inference method, applied to a heterogeneous many-core intelligent acceleration platform. The method includes: The hardware resource feature information and the format feature information of the low bit weight to be processed of the target hardware platform are obtained. The hardware resource feature information includes the width of the single instruction multiple data vector and the number of on-chip memory. The format feature information includes the exponent bit width, the mantissa bit width and the exponent bias. The initial dequantization execution tendency is determined based on the format feature information, and multiple candidate dequantization configurations are prioritized based on the hardware resource feature information. Each candidate dequantization configuration includes at least a dequantization path type and a parallel granularity parameter. Run at least some of the candidate dequantization configurations on the target hardware platform, obtain real-machine execution feedback information including measured execution time, and determine the target dequantization configuration from the candidate dequantization configurations based on the real-machine execution feedback information; According to the target dequantization configuration, the low bit weights are moved from global storage to on-chip local storage, and in the on-chip local storage, the low bit weights are dequantized into a high-precision data format supported by the target hardware platform according to the target dequantization configuration. The high-precision data obtained by dequantization is continuously fed into the computing unit of the target hardware platform in segments to participate in matrix multiplication and addition operations, without forming a complete high-precision weight tensor or writing it back to the global storage.

[0017] Preferably, the target dequantization configuration includes a lookup table dequantization path; the step of dequantizing the low bit weights into a high-precision data format supported by the target hardware platform includes: A decoding lookup table between low-bit encoded values ​​and high-precision data format values ​​is generated in advance, and the decoding lookup table is stored in the on-chip local memory; Using the encoded value with the low bit weight as an index, the corresponding high-precision data format value is obtained from the decoding lookup table through on-chip lookup instructions; Based on the number of storage banks in the on-chip local storage, multiple access copies or logical mapping copies of the decoding lookup table are configured so that the access addresses of parallel queries are allocated to different storage banks to reduce storage bank access conflicts.

[0018] Preferably, the target dequantization configuration includes a single instruction multiple data stream parallel shift dequantization path; the single instruction multiple data stream parallel shift dequantization path supports at least the following three data formats: a first low-bit floating-point format, with an exponent width of 4 bits and a mantissa width of 3 bits; a second low-bit floating-point format, with an exponent width of 5 bits and a mantissa width of 2 bits; and an integer low-bit format. The step of dequantizing the low-bit weights into a high-precision data format supported by the target hardware platform includes: For the first low-bit floating-point format, the low-bit data is extended to the bit width of the high-precision data format through the sign extension operation, the extended data is shifted to achieve bit field alignment, redundant high bits are erased by the bit-AND operation mask constant, and the exponential bias compensation factor and the block-level scaling factor are merged into a single floating-point correction coefficient to complete the numerical recovery in a single floating-point multiplication. For the second low-bit floating-point format, the low-bit data is extended to the bit width of the high-precision data format through sign extension operation, and the bit field alignment of the exponent and mantissa bits is completed through shift operation, without performing an independent exponent bias correction operation. For the aforementioned low-bit integer format, the integer data is extended to the bit width of the high-precision data format through a sign extension operation. Using the preset offset interval in the half-precision floating-point numerical representation, the sign-extended integer data is added to the preset offset constant, and then subtracted from the preset offset constant to obtain the mapping result from integer to floating-point number.

[0019] Preferably, determining the initial dequantization execution tendency based on the format feature information includes: Obtain the format parameters of the low bit weight, which include the exponent bit width, mantissa bit width, exponent bias, and data type identifier; Calculate the format structure difference value between the low bit weight format and the high precision data format. The format structure difference value is determined by a weighted sum of the exponent bit width difference, mantissa bit width difference, exponent bias difference, and type difference identifier. When the format structure difference value is less than or equal to the first preset threshold, the initial dequantization execution tendency is determined to be a single instruction multiple data stream parallel shift dequantization path, and the initial parallel granularity is set to high parallel granularity. When the format structure difference value is greater than the first preset threshold and less than or equal to the second preset threshold, the initial dequantization execution tendency is determined to be a single instruction multiple data stream parallel shift dequantization path, and the initial parallel granularity is set to medium parallel granularity. When the difference value of the format structure is greater than the second preset threshold, the initial dequantization execution tendency is determined to be the parallel lookup table dequantization path; The first preset threshold and the second preset threshold are predetermined based on the bit field structure compatibility between the high-precision data format and the low bit weight format. The high parallel granularity, medium parallel granularity, and low parallel granularity are parallel path levels predetermined based on the single instruction multiple data stream vector width of the target hardware platform.

[0020] Preferably, determining the target dequantization configuration from the candidate dequantization configurations based on the actual execution feedback information includes: Obtain at least one feedback metric from the measured execution time, memory access conflict count, register occupancy rate, and pipeline pause cycle of the candidate dequantization configuration on the target hardware platform. Based on the measured execution time before and after the change in parallel granularity, calculate the marginal performance gain corresponding to a unit change in parallel granularity. When the marginal performance gain is less than the preset marginal gain threshold, it is determined that the current parallel granularity direction has entered the gain saturation region, and the candidate configuration with higher parallel granularity and resource consumption not lower than the current configuration is removed from the candidate dequantization configuration or its test priority is reduced. For any first candidate configuration and second candidate configuration in the candidate dequantization configuration, if the first candidate configuration is less than or equal to the corresponding index value of the second candidate configuration in all four indicators of measured execution time, memory access conflict number, register occupancy rate and pipeline pause cycle, and at least one indicator is less than the corresponding index value of the second candidate configuration, it is determined that the second candidate configuration is dominated by the first candidate configuration, and the second candidate configuration is deleted from the candidate dequantization configuration. After completing at least one round of search and pruning, the candidate dequantization configuration with the lowest measured execution time among the remaining candidate dequantization configurations is selected as the target dequantization configuration.

[0021] Preferably, the step of moving the low-bit weights from global storage to on-chip local storage, dequantizing the low-bit weights in the on-chip local storage according to the target dequantization configuration into a high-precision data format supported by the target hardware platform, and continuously sending the dequantized high-precision data into the computing unit of the target hardware platform in segments for matrix multiplication and addition operations includes: A first weight buffer and a second weight buffer are set in the on-chip local storage to form a first-level double buffer structure; the current low-bit weight sub-block is moved to the first weight buffer through direct memory access, and the next low-bit weight sub-block is moved to the second weight buffer asynchronously; and the first weight buffer and the second weight buffer are swapped after the current low-bit weight sub-block is processed. The current weighted sub-block is divided into multiple low-bit weighted segments. A first high-precision segment buffer and a second high-precision segment buffer are set in the on-chip local storage to form a second-level double-buffered structure. While the high-precision weighted segment in the current high-precision segment buffer is read by the computing unit and matrix multiplication and addition operations are performed, the dequantization unit dequantizes the next low-bit weighted segment into a high-precision data format and writes it into another high-precision segment buffer. After the current segment is consumed, the first high-precision segment buffer and the second high-precision segment buffer are swapped. The first-level double-buffered structure and the second-level double-buffered structure are executed concurrently in time, so that the low-bit-weighted direct memory access prefetch, on-chip dequantization data generation and matrix multiplication and addition data consumption form a three-stage continuous pipeline.

[0022] Preferably, the method further includes: Obtain the block-level quantization parameters of the low bit weights, which are stored in the global storage. Multiple low bit weight elements share the same block-level quantization parameters. Based on the linear storage index of the low bit weight in the global storage, the linear storage index of the corresponding block-level quantization parameter is determined by closed-loop mapping calculation. During the dequantization process, the corresponding block-level quantization parameter position is directly determined based on the linear storage index of the current low-bit weight element, and the block-level quantization parameter is embedded into the dequantization calculation path.

[0023] This invention also provides a low-bit-weighted adaptive dequantization and on-chip fusion inference system, applied to a heterogeneous many-core intelligent acceleration platform. The system includes: The information acquisition module is used to acquire hardware resource feature information and format feature information of low bit weight to be processed from the target hardware platform. The hardware resource feature information includes the width of the single instruction multiple data vector and the number of on-chip memory blocks. The format feature information includes the exponent bit width, the mantissa bit width and the exponent bias. A configuration search module is used to determine the initial dequantization execution tendency based on the format feature information, prioritize multiple candidate dequantization configurations based on the hardware resource feature information, each candidate dequantization configuration includes at least a dequantization path type and parallel granularity parameters, run at least some of the candidate dequantization configurations on the target hardware platform and obtain real-machine execution feedback information including measured execution time, and determine the target dequantization configuration from the candidate dequantization configurations based on the real-machine execution feedback information; The weight transfer module is used to transfer the low-bit weights from global storage to on-chip local storage according to the target dequantization configuration. The dequantization execution module is used to dequantize the low bit weight into a high-precision data format supported by the target hardware platform in the on-chip local storage according to the target dequantization configuration. The on-chip fusion computing module is used to continuously send the high-precision data obtained by dequantization into the computing unit of the target hardware platform in segments to participate in matrix multiplication and addition operations, without forming a complete high-precision weight tensor or writing it back to the global storage.

[0024] Preferably, the configuration search module includes: The format analysis unit is used to calculate the format structure difference value based on the exponent bit width, mantissa bit width, exponent bias and data type identifier of the low bit weight, and to determine the initial dequantization execution tendency and initial parallel granularity based on the format structure difference value. The hardware resource analysis unit is used to evaluate the resource matching degree and prioritize the candidate dequantization configurations based on the single instruction multiple data stream vector width, number of on-chip memory, number of register resources and on-chip memory capacity of the target hardware platform. The real-machine testing unit is used to run candidate dequantization configurations on the target hardware platform and collect at least one feedback indicator among the measured execution time, memory access conflict count, register occupancy rate, and pipeline pause cycle. The pruning decision unit is used to calculate the marginal performance gain based on the feedback index, prune the dominated candidate configurations using the gain saturation criterion and resource allocation relationship, and output the target inverse quantization configuration when the preset convergence condition is met.

[0025] Preferably, the system further includes a flow scheduling module, which is used for: A first weight buffer and a second weight buffer are configured in the on-chip local memory to form a first-level double buffer structure, and the direct memory access controller is controlled to asynchronously prefetch the next low-bit weight sub-block during the processing of the current weight sub-block. A first high-precision fragment buffer and a second high-precision fragment buffer are configured in the on-chip local storage to form a second-level double buffer structure, and the dequantization execution module is controlled to perform dequantization on the next low-bit weight fragment while the computing unit consumes the current high-precision weight fragment. The timing of switching between the first-level double-buffered structure and the second-level double-buffered structure is coordinated so that the three stages of direct memory access prefetching, dequantization execution, and matrix multiplication and addition operations overlap in time.

[0026] As can be seen from the above technical solutions, this invention application has the following beneficial effects: First, this invention uses a format-aware hardware feedback pruning search method to automatically determine the optimal dequantization path and parallel granularity based on the low bit weight format and target hardware resource characteristics. This solves the problem that fixed dequantization configurations are difficult to adapt to different formats and different hardware, and reduces the overhead of manual tuning and actual testing.

[0027] Second, this invention constructs differentiated dequantization paths for different low-bit formats such as FP8_E4M3, FP8_E5M2 and INT8—achieving low-overhead discrete decoding through on-chip parallel word lookup and simplified bit field reconstruction through SIMD parallel shifting—taking into account the conversion efficiency of multiple low-bit formats under a unified framework, and avoiding redundant bit operations and format adaptation fragmentation problems caused by a single conversion method.

[0028] Third, this invention uses a two-stage, double-buffered pipeline mechanism to make the three execution stages of low-bit-weighted DMA prefetching, on-chip dequantization, and matrix multiplication and addition overlap and run in time. The high-precision data obtained by dequantization directly participates in the calculation in the form of fragments without being written back to global storage, which effectively reduces the waiting time between stages and the overhead of intermediate data transfer, and improves the utilization rate of on-chip storage and computing resources. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly described below. Referring to the drawings will make the features and advantages of the present invention clearer. The drawings are illustrative and should not be construed as limiting the present invention in any way. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a flowchart of a low-bit-weight adaptive dequantization and on-chip fusion inference method provided by an embodiment of the present invention; Figure 2 This is a schematic diagram of the overall structure of the method of the present invention applied to a heterogeneous many-core intelligent acceleration platform; Figure 3 This is a schematic diagram of the inverse quantization structure of the lookup table in this embodiment of the invention; Figure 4 This is a schematic diagram of a single-instruction multiple-data (SIMD) parallel shift and dequantization structure in an embodiment of the present invention; Figure 5 This is a schematic diagram of an on-chip local storage two-stage double-buffered pipeline in an embodiment of the present invention; Figure 6 This is a block diagram of a low-bit-weighted adaptive dequantization and on-chip fusion inference system provided in an embodiment of the present invention. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0031] In this embodiment of the invention, the heterogeneous many-core intelligent acceleration platform refers to an intelligent acceleration platform that adopts a collaborative execution architecture of a management core and multiple computing cores, and performs data transfer between global storage and on-chip local storage through explicit direct memory access (DMA). The global storage refers to high-bandwidth memory (HBM) on the accelerator card or other forms of off-chip main storage; the on-chip local storage (LDM) refers to the on-chip high-speed temporary storage provided by the computing core. The high-precision data format is half-precision floating-point (FP16) format. The low-bit weight includes at least one of 8-bit floating-point FP8_E4M3 format, FP8_E5M2 format, and 8-bit integer INT8 format, but the invention is not limited to these and can be extended to other low-bit data formats in other embodiments.

[0032] In this invention, dequantization refers to the process of restoring or converting low-bit data format values ​​to high-precision data format values. The on-chip fusion inference refers to the process where, after low-bit weight dequantization is completed in on-chip local storage, the high-precision data obtained from dequantization is not written back to global storage, but is directly sent to the computation unit to participate in matrix multiplication and addition operations, thus realizing a continuous on-chip execution chain of data transfer—dequantization—computation.

[0033] Example 1: To address the challenges of efficient execution of low-bit weights on heterogeneous many-core platforms due to the lack of native low-bit arithmetic support, the inability of existing dequantization schemes to adaptively adjust to different low-bit formats and hardware resources, and the pipeline imbalance between weight transport and computation stages, this invention proposes a low-bit weight adaptive dequantization and on-chip fusion inference method. Figure 1 As shown, the method includes the following steps: S1: Obtain the hardware resource feature information of the target hardware platform and the format feature information of the low bit weight to be processed. The hardware resource feature information includes the width of the single instruction multiple data vector and the number of on-chip memory banks. The format feature information includes the exponent bit width, mantissa bit width and exponent bias. S2: Determine the initial dequantization execution tendency based on the format feature information, and prioritize multiple candidate dequantization configurations based on the hardware resource feature information. Each candidate dequantization configuration includes at least the dequantization path type and parallel granularity parameters. S3: Run at least some candidate dequantization configurations on the target hardware platform, obtain real-machine execution feedback information including actual execution time, and determine the target dequantization configuration from the candidate dequantization configurations based on the real-machine execution feedback information; S4: According to the target dequantization configuration, move the low bit weights from global storage to on-chip local storage, and dequantize the low bit weights into a high-precision data format supported by the target hardware platform in the on-chip local storage according to the target dequantization configuration; S5: The high-precision data obtained by dequantization is continuously fed into the computing unit of the target hardware platform in segments to participate in matrix multiplication and addition operations, without forming a complete high-precision weight tensor or writing it back to global storage.

[0034] As can be seen from the above technical solution, this invention proposes a low-bit-weight adaptive dequantization and on-chip fusion inference method. First, by acquiring the hardware resource feature information of the target hardware platform (including the width of the single instruction multiple data stream vector and the number of on-chip memory blocks) and the format feature information of the low-bit weight to be processed (including the exponent bit width, mantissa bit width, and exponent bias), a data foundation is provided for the adaptive selection of subsequent dequantization configurations, enabling targeted optimization based on different hardware resources and low-bit format features. Second, the initial dequantization execution tendency is determined based on the format feature information, and multiple candidate dequantization configurations are prioritized based on the hardware resource feature information. Each candidate dequantization configuration includes at least the dequantization path type and parallel granularity parameters, thereby introducing prior knowledge of the format structure and hardware resource constraints into the configuration search process, reducing the blindness of the search space and lowering the overhead of subsequent real-world tuning. Then, at least some candidate dequantization configurations are run on the target hardware platform to obtain real-world execution feedback information, including measured execution time. Based on the real-world execution feedback information, the candidate dequantization configurations are selected from the... In the dequantization configuration, the target dequantization configuration is determined. Through real-machine testing and pruning search mechanisms, the final determined dequantization configuration is made to fully adapt to the execution characteristics of the target hardware, avoiding the uncertainty brought about by manual experience selection. Next, according to the target dequantization configuration, the low-bit weights are moved from global storage to on-chip local storage. In the on-chip local storage, the low-bit weights are dequantized into a high-precision data format supported by the target hardware platform according to the target dequantization configuration. This allows the low-bit weights to participate in the calculation efficiently in software on hardware platforms that lack native low-bit arithmetic support. At the same time, the dequantization process is completed in on-chip local storage, reducing the global storage access overhead. Finally, the high-precision data obtained by dequantization is continuously sent to the computing unit of the target hardware platform in units of fragments to participate in matrix multiplication and addition operations. It does not form a complete high-precision weight tensor, nor is it written back to global storage. This avoids the storage and movement of intermediate high-precision weight data, reduces off-chip memory access bandwidth occupation, and makes the three stages of low-bit weight movement, on-chip dequantization, and matrix multiplication and addition form a continuous on-chip data stream. Through the close connection and synergy between the above steps, the present invention achieves the overall realization of adaptively determining efficient dequantization configuration based on different low-bit formats and hardware resources on a heterogeneous many-core intelligent acceleration platform, and reduces inference latency through multi-stage pipelined execution, thereby improving the utilization efficiency of on-chip storage and computing resources.

[0035] The overall structure of the present invention is as follows Figure 2As shown. Low-bit weights are stored in the accelerator card's global memory (HBM) in formats such as FP8_E4M3, FP8_E5M2, or INT8. During inference execution, the weight data is moved from the global memory to the on-chip local memory (LDM) of the slave core via DMA, according to the block granularity required for linear layer or matrix calculations. After the low-bit weights enter the LDM, they are not restored to complete FP16 weight tensors and written back to the global memory. Instead, they are directly dequantized on-chip according to the target dequantization configuration determined in step S3, and participate in subsequent matrix multiplication or linear layer multiplication-addition calculations in the form of registers or LDM temporary data. The whole forms a hierarchical execution topology of HBM low-bit weight storage—DMA block transport—LDM on-chip dequantization—matrix calculation unit consumption. During the target hardware initial adaptation, model deployment, or compilation stage, this invention performs hardware feedback configuration search (i.e., steps S2 and S3) for low-bit format, dequantization path, and target hardware resources through a configuration search module. The search results form a mapping relationship between hardware, low-bit format, and target dequantization configuration, and the corresponding configuration is directly called during the inference stage.

[0036] Further, in step S1, the hardware resource feature information of the target hardware platform and the format feature information of the low bit weight to be processed are obtained.

[0037] Specifically, before executing the inference task on the target heterogeneous many-core intelligent acceleration platform, this embodiment of the invention first acquires the hardware resource characteristic information of the target hardware platform. This hardware resource characteristic information includes, but is not limited to: Single Instruction Multiple Data (SIMD) vector width, number of Local Local Memory (LDM) banks, number of register resources, and on-chip storage capacity. Simultaneously, the format characteristic information of the low-bit weight to be processed is acquired. This format characteristic information includes the exponent width, mantissa width, and exponent bias of the low-bit data format. For example, for the FP8_E4M3 format, its exponent width is 4 bits and its mantissa width is 3 bits; for the FP8_E5M2 format, its exponent width is 5 bits and its mantissa width is 2 bits; for the INT8 format, its data type identifier is identified as an integer type.

[0038] In one specific embodiment of the present invention, the Shenwei intelligent accelerator card is used as the target hardware platform, with a SIMD vector width of 256 bits and 32 LDM memory banks. The low-bit weights to be processed are FP8_E4M3 format weights of the Qwen series model, with an exponent bit width of 4, a mantissa bit width of 3, and an exponent bias of 7.

[0039] Furthermore, in step S2, the initial dequantization execution tendency is determined based on the format feature information, and the priority of multiple candidate dequantization configurations is sorted based on the hardware resource feature information.

[0040] Specifically, after acquiring the format feature information, this invention determines the initial dequantization execution tendency based on the bit-domain structure relationship between the low-bit-weight format and the high-precision data format (i.e., FP16). Specifically, the format parameters for the low-bit weight are acquired, including the exponent bit width e, the mantissa bit width m, the exponent bias β, and the data type identifier z. The data type identifier z is used to distinguish between floating-point and integer formats: z is 0 when both the format to be dequantized and FP16 are floating-point formats, and z is 1 when the format to be dequantized is an integer format.

[0041] Calculate the structural differences between the low bit weight format and FP16. The calculation method is as follows: , in, , and These are the exponent width (5 bits), mantissa width (10 bits), and exponent bias (15 bits) of FP16, respectively. , , and A preset weighting coefficient is used to balance the contribution of differences in different format parameters to the structural difference value. In a specific embodiment of the present invention, , , and The values ​​are 1.0, 1.0, 0.5 and 2.0 respectively, but the present invention is not limited thereto. In other embodiments, the values ​​can be adjusted according to the specific hardware platform and the characteristics of the low bit format.

[0042] When format structure difference value Less than or equal to the first preset threshold This indicates that the current low-bit format is quite similar to the FP16 bit-field structure, and the direct bit-field conversion path is relatively short. Therefore, the initial dequantization execution tendency is determined to be the SIMD parallel shift dequantization path, and the initial parallel granularity is set to high parallel granularity. .when Greater than And less than or equal to the second preset threshold At that time, the initial dequantization execution tendency is still determined to be the SIMD parallel shift dequantization path, but the initial parallel granularity is set to medium parallel granularity. .when Greater than This indicates that direct bit-field conversion requires more correction steps. Therefore, the initial dequantization execution tendency is determined to be a parallel lookup table dequantization path, and the initial parallel granularity is set to low parallel granularity. The first preset threshold and the second preset threshold Based on the bit-domain structure compatibility between FP16 and low bit-weighted formats, the high parallel granularity, medium parallel granularity, and low parallel granularity are parallel path levels predetermined based on the SIMD vector width of the target hardware platform.

[0043] In one specific embodiment of the present invention, for the FP8_E4M3 format, its exponential bit width difference with FP16 is... Difference in mantissa digits Exponential bias Type identifier The format structure difference value was calculated. In and Therefore, the initial dequantization execution tends to be a SIMD parallel shift dequantization path, and the initial parallel granularity is set to medium parallel granularity.

[0044] After determining the initial execution tendency and initial parallel granularity, this invention further prioritizes multiple candidate dequantization configurations based on the target hardware platform's SIMD vector width, LDM memory count, register resource count, and on-chip storage capacity. Each candidate dequantization configuration includes at least the dequantization path type and parallel granularity parameters. Specifically, for the parallel lookup table dequantization path, its candidate priority is mainly determined based on the matching degree between the number of parallel lookup paths and the number of LDM memory counts, as well as memory access conflicts; for the SIMD parallel shift dequantization path, its candidate priority is mainly determined based on SIMD execution width utilization, register usage, and instruction dependencies. This invention prioritizes candidate configurations with a higher degree of matching with the target hardware's on-chip resources for subsequent real-world testing.

[0045] Further, in step S3, at least some candidate dequantization configurations are run on the target hardware platform to obtain real-machine execution feedback information including the measured execution time, and the target dequantization configuration is determined from the candidate dequantization configurations based on the real-machine execution feedback information.

[0046] In this invention, an inverse quantization configuration to be tested Defined as: , in, Indicates low-bit data format; Indicates the type of dequantization path (including parallel lookup dequantization path and SIMD parallel shift dequantization path). Indicates the number of parallel paths (i.e., the number of elements processed in a single parallel run). Indicates the number of copies of the on-chip decoder lookup table (applicable only to the parallel lookup table dequantization path); Indicates the on-chip memory mapping method; Indicates the number of elements processed in a single SIMD operation (applicable only to SIMD parallel shift dequantization paths); Indicates the scaling factor processing method; These represent the pipeline configuration parameters. Together, these parameters constitute a complete description of the candidate inverse quantization configurations, providing a clear search space for subsequent priority ranking and live search.

[0047] Specifically, after prioritizing the candidate dequantization configurations, this invention actually runs the dequantization test kernels of each candidate configuration on the target heterogeneous many-core intelligent accelerator and records feedback indicators such as execution time, memory access conflict count, register occupancy rate, and pipeline pause cycle. Based on this real-world feedback information, this invention determines the target dequantization configuration from the candidate dequantization configurations through a pruning search mechanism.

[0048] Specifically, assuming parallel granularity is adopted The actual execution time was When the parallel granularity is changed from Increase to ( When ), the marginal performance gain corresponding to a unit parallel granularity. Represented as: , Let the preset marginal revenue threshold be... .when Less than When this occurs, it indicates that the performance gains from further increasing the parallel granularity are already limited, and the current search direction has entered a saturation region. For candidate configurations with higher parallel granularity and register usage, LDM usage, or pipeline pressure not lower than the current configuration, they are removed from the candidate dequantization configurations or their test priority is reduced. Greater than or equal to If this indicates that the current direction still has significant performance benefits, we should continue to search towards a higher parallel granularity and further test local parameters such as the number of elements processed in a single SIMD operation, the number of lookup table replicas, or the memory mapping method around the current configuration.

[0049] In addition to marginal performance gains, this invention further performs resource allocation pruning based on actual hardware feedback. Let candidate configurations be considered. and The execution times are respectively and The storage access conflict indicators are as follows: and The register usage metrics are as follows: and The production line stoppage indicators are as follows: and When both conditions are met: ,and ,and ,and , Furthermore, if at least one of the above relationships is strictly less than, then the candidate configuration is determined. Superior to candidate configuration Candidate configurations It is marked as the dominated configuration and removed from the candidate dequantization configuration.

[0050] After completing at least one round of real-world testing and pruning, the candidate dequantization configuration with the lowest measured execution time among the remaining candidate dequantization configurations is selected as the target dequantization configuration. Let the first... The optimal execution cost obtained by round search is Improvement rate between two adjacent search rounds Represented as: , When K consecutive rounds satisfy Less than the preset convergence threshold At this point, stop the real-machine search corresponding to the current hardware and the current low-bit format, and output the current optimal configuration as the target dequantization configuration. : , in, Indicates the target dequantization path; Indicates the parallel granularity of the target; Indicates the number of copies of the target lookup table; Indicates the target bank mapping method; Indicates the number of elements processed in a single SIMD operation; Indicates how the target scaling factor is handled; Indicates the target pipeline configuration.

[0051] Ultimately, this invention establishes a target hardware identifier. Low-bit format Inverse quantization configuration with target Mapping relationship between them: This mapping relationship is stored in the dequantization configuration mapping table. During the inference execution phase, the corresponding target configuration is directly obtained based on the current hardware identifier and low bit weight format, without the need for repeated real-world tuning.

[0052] In a specific embodiment of the present invention, taking the FP8_E4M3 format weight on the Shenwei intelligent accelerator card as an example, after the above-mentioned real-machine search and pruning, the target dequantization is configured as a parallel lookup table dequantization path, with 16 parallel query paths and 4 copies of the lookup table.

[0053] Further, in step S4, according to the target dequantization configuration, the low bit weights are moved from global storage to on-chip local storage, and in the on-chip local storage, the low bit weights are dequantized into a high-precision data format supported by the target hardware platform according to the target dequantization configuration.

[0054] Specifically, after determining the target dequantization configuration, this invention begins the low-bit weight transfer and dequantization process during the actual inference process. First, based on the computational block granularity of the current linear layer or matrix multiplication-addition task, the range of low-bit weight sub-blocks to be transferred each time is determined. The low-bit weights are stored in the accelerator card's global memory (HBM) in formats such as FP8_E4M3, FP8_E5M2, or INT8. During inference execution, the low-bit weights are asynchronously transferred from the global memory to the slave core's on-chip local memory (LDM) according to the block granularity via DMA.

[0055] After the low-bit weights enter the LDM, they are no longer restored to the complete FP16 weight tensor and written back to global storage. Instead, dequantization is directly performed on-chip according to the target dequantization configuration determined in step S3. Based on the dequantization path type in the target dequantization configuration, this invention selects the corresponding method from the parallel lookup table dequantization path or the SIMD parallel shift dequantization path to perform the conversion.

[0056] For the inverse quantization path of the merge lookup table, such as Figure 3 As shown, this invention utilizes the limited 8-bit low-bit encoding space (only 256 possible encoding values) to pre-generate a lookup table (LUT) between the low-bit encoded values ​​and the FP16 values, and keeps this lookup table resident in the LDM. During inference execution, the encoded value with the low bit weight is used as an index to retrieve the corresponding FP16 value from the lookup table using on-chip parallel lookup instructions. This method transforms the traditional element-by-element bit-field parsing process into discrete table entry access, reducing complex bit operations and floating-point calculations.

[0057] Furthermore, to improve the efficiency of parallel lookup tables, this invention configures the decoding lookup table as multiple access copies or logically mapped copies based on the number of banks in the LDM. This distributes the access addresses of parallel queries to different banks, reducing bank access conflicts and improving on-chip lookup throughput. Specifically, when the LDM has N banks, the decoding lookup table is copied into N copies. Adjacent parallel query threads access copies on different banks, enabling multiple parallel queries to run simultaneously without bank conflicts.

[0058] In one specific embodiment of the present invention, the FP8_E4M3 format has 256 possible encoded values, and the decoding lookup table size is 256 × 2 bytes (FP16 occupies 2 bytes) = 512 bytes. Four copies are stored in the LDM, occupying approximately 2KB of storage space. When performing 16-way parallel queries, address rearrangement evenly distributes the 16 queries across 16 different memory banks, eliminating memory bank access conflicts and achieving a lookup throughput 16 times that of a single-way query.

[0059] In a specific embodiment of the present invention, the dequantization of FP8_E4M3 format to FP16 by lookup table can be represented as: FP16_Value = LUT[FP8_Code], where LUT is a pre-generated lookup table containing 256 FP16 values, and FP8_Code is the current low bit weight encoding value (an integer in the range of 0 to 255). Since the lookup table is stored in LDM and access conflicts are eliminated through memory copying, the latency of a single lookup operation is much lower than that of traditional bit field parsing methods.

[0060] For SIMD parallel shift inverse quantization paths, such as Figure 4 As shown, this invention rapidly maps low-bit weights to FP16 representation using operations such as sign extension, shifting, masking, integer addition, and floating-point scaling, based on the sign bit, exponent bit, mantissa bit, and scaling rules of different low-bit formats. This path supports at least three data formats: a first low-bit floating-point format (FP8_E4M3, exponent width 4 bits, mantissa width 3 bits), a second low-bit floating-point format (FP8_E5M2, exponent width 5 bits, mantissa width 2 bits), and an integer low-bit format (INT8).

[0061] For the FP8_E4M3 format, the conversion process is completed in three stages: sign extension, bit field alignment, and scaling correction. Let the FP8_E4M3 weighted byte sequence be... The length of the vector block is (i.e., the number of elements processed in a single SIMD operation), the alignment shift parameter is... The mask constant is scaling factor (Where weight_scale is the block-level quantization scaling factor,) To compensate for the exponential bias difference between FP8_E4M3 and FP16, the conversion process can be expressed as the following algorithm: Step 1: Initialize the circular index ; Step Two: When Perform the following operations: Step 3: From the input sequence Reading continuous FP8_E4M3 elements form a vector ; Step 4: [Regarding...] Perform symbol extension operation This expands 8-bit data to 16-bit data. Step 5: Perform a left shift operation on the expanded data. To achieve bit field alignment; Step 6: Perform a bitwise AND operation on the shifted data. Using mask constants Eliminate redundant high-order bits; Step 7: Multiply the processed data by the fusion scaling factor. ,Right now Simultaneously, it completes exponential bias compensation and block-level inverse quantization scaling; Step 8: Write the transformation result into the output sequence. The corresponding position ; Step 9: Circular Index Increase Continue processing the next vector block.

[0062] In the above process, the sequential execution of steps four through seven constitutes the complete conversion path from FP8_E4M3 to FP16. Specifically, sign extension preserves the sign of the original data; the left shift operation shifts the exponent and mantissa bits to their correct bit positions in the FP16 data format; the bitwise AND operation clears high-order redundancy generated during the extension process; and the fusion scaling factor c simultaneously compensates for the exponent bias difference between FP8_E4M3 and FP16. The exponential bias compensation and the block-level dequantization scaling factor (weight_scale) are combined into a single coefficient before step S1, eliminating the need for additional independent floating-point correction steps and reducing the number of floating-point operations.

[0063] For the FP8_E5M2 format, which has the same exponent bias as FP16 (both 15), and whose mantissa field can completely cover the mantissa precision of FP8_E5M2, a simplification has been made based on the FP8_E4M3 conversion method described above. Specifically, the sign-extended FP8_E5M2 data only requires a single left shift operation to align the exponent and mantissa bit fields, without additional exponent correction or floating-point scaling steps. This is the scaling factor in step seven. The value is the weight_scale itself (excluding) The compensation factor reduces the number of instructions and floating-point operations.

[0064] For the INT8 format, this invention utilizes a special offset interval in the FP16 numerical representation to achieve fast mapping. Specifically, let... Given a sign-extended INT8 value, the conversion from INT8 to FP16 can be represented as: , in, As a preset offset constant, Integer operations, This method represents floating-point operations. It completes exponent correction and offset processing with a single integer addition and achieves the final mapping with a single floating-point subtraction, avoiding element-by-element conditional checks and loop branches, and fully utilizing the parallel processing capabilities of SIMD. In this invention, because the Shenwei intelligent accelerator card has a wide SIMD vector width, it can process 16 or more INT8 elements in a single operation, further amplifying the benefits of the above conversion method.

[0065] Furthermore, during dequantization, this invention also achieves low-overhead access to block-level quantization parameters through a closed-index mapping method. Specifically, the block-level quantization parameters (scaling factors) with low bit weights are obtained. These block-level quantization parameters are stored in global storage, and multiple low-bit weight elements share the same block-level quantization parameter. Based on the linear storage index of the low bit weights in global storage, the linear storage index of the corresponding block-level quantization parameter is determined through closed-index mapping. During dequantization, the position of the corresponding block-level quantization parameter is directly determined based on the linear storage index of the current low-bit weight element, and the block-level quantization parameter is embedded into the dequantization calculation path.

[0066] Taking the FP8 block quantization of the Qwen series models as an example, the weight tensor It can be abstracted as The last dimension, 1024, is the alignment dimension for calculation. The weights are linearly expanded in row-major order in global storage, and the correspondence between their linear indices and tensor coordinates can be uniformly represented as: , in, , , These represent the index coordinates of the weight tensor in the three dimensions. Under FP8 block quantization settings, the weights are accompanied by a two-dimensional quantization parameter matrix. Each quantization parameter corresponds to a set of weight sub-blocks of fixed size (e.g., 4×4 blocks). Based on this block structure, the weights can be linearly indexed. A closed mapping relationship is directly established with the corresponding quantization parameter index. Assuming S expands linearly in row-major order, the one-dimensional index of the quantization parameter... It can be represented as: , in, `mod` represents the floor operation, and `mod` represents the modulo operation. This closed expression directly provides the mapping process from the linear index of the weights to the block-level dequantization parameter index, avoiding explicit unpacking and conditional judgment of the weight coordinates during the execution phase, and enabling access to the quantization scaling factor to be embedded in the dequantization process with constant overhead.

[0067] Furthermore, in step S5, the high-precision data obtained by dequantization is continuously sent to the computing unit of the target hardware platform in segments to participate in matrix multiplication and addition operations, without forming a complete high-precision weight tensor or writing it back to global storage.

[0068] Specifically, after the low-bit weights are dequantized into FP16 data, this invention does not form a complete FP16 weight tensor from the dequantization result, nor does it write it back to global storage. Instead, it continuously sends the data in segments to the matrix computation unit (such as the FP16 matrix multiplication and addition unit) of the target hardware platform to participate in multiplication and addition operations. That is, the temporary FP16 data obtained by dequantization directly participates in matrix multiplication or linear layer multiplication and addition calculations in registers or LDM temporary buffers, forming a continuous execution chain of "low-bit weight transport - on-chip dequantization - matrix multiplication and addition". In this way, the amount of low-bit weight transport between global storage and on-chip local storage, as well as the amount of data movement between on-chip local storage and registers, are reduced.

[0069] In a specific embodiment of the present invention, using the low-bit weight W8 in FP8_E4M3 format and the input activation matrix A as input, the dequantization path is obtained through an on-chip parallel lookup table. Convert to FP16 weighted fragment Then Directly input into the matrix multiply-add unit for execution. The entire process does not produce a complete [product / service]. Tensor storage The fragment is overwritten by the new dequantization result immediately after the multiplication and addition are completed.

[0070] It is worth noting that during the execution of steps S4 and S5, this invention employs a two-stage, double-buffered pipeline mechanism to coordinate the three execution stages: low-bit weight transfer, on-chip dequantization, and matrix multiplication and addition. Specifically, as follows... Figure 5 As shown, the dual-stage, dual-buffered flow mechanism of the present invention includes: Phase 1 – Low-Bit Weight Transfer with Double Buffer: A first weight buffer (bufW0) and a second weight buffer (bufW1) are set up in the on-chip local memory to form a first-level double-buffered structure. The current low-bit weight sub-block is transferred to bufW0 via DMA, while the next low-bit weight sub-block is transferred to bufW1 asynchronously. After the current weight sub-block is processed, bufW0 and bufW1 are swapped.

[0071] The second stage—on-chip dequantization and matrix multiplication-addition double buffering—further divides the current weighted sub-block into multiple low-bit weighted segments. A first high-precision segment buffer (bufD0) and a second high-precision segment buffer (bufD1) are set up in the on-chip local storage, forming a second-level double-buffered structure. While the high-precision weighted segment in the current high-precision segment buffer (e.g., bufD0) is read by the computation unit and matrix multiplication-addition operations are performed, the dequantization unit dequantizes the next low-bit weighted segment into a high-precision data format and writes it to another high-precision segment buffer (e.g., bufD1). After the current segment is consumed, bufD0 and bufD1 are swapped.

[0072] The first-level and second-level double-buffered structures are executed concurrently in time, enabling a three-stage continuous pipeline for low-bit-weight DMA prefetching, on-chip dequantization data generation, and matrix multiplication-addition data consumption. Specifically, while the first stage is moving the next weighted sub-block from the HBM to the LDM spare buffer, the second stage is simultaneously performing dequantization and matrix multiplication-addition within the current weighted sub-block; while the current high-precision segment is consumed by the computation unit in the second stage, the dequantization unit is simultaneously generating the next high-precision segment. Through the coordinated execution of the two stages and two levels of double buffering, the three stages overlap and run in parallel in time, reducing inter-stage waiting.

[0073] In a specific embodiment of the present invention, taking a linear layer weight matrix of 4096×4096 and a block size of 512×512 as an example, the first stage sets up two 512×512 weight buffers. The DMA controller moves the current block to bufW0 while asynchronously moving the next block to bufW1. In the second stage, each weight sub-block is divided into 32 weight segments (each segment 512×16). When a segment in bufD0 is consumed by the matrix multiply-add unit, the dequantization unit dequantizes the next segment and writes it to bufD1. Through the coordination of the two-stage pipeline, the DMA transfer utilization reaches over 95%, and the dequantization unit and the computation unit achieve near-complete overlapping execution.

[0074] In another preferred embodiment of the present invention, when the target dequantization is configured as a parallel lookup table dequantization path, the dequantization operation in step S4 specifically involves: using the low-bit weighted encoded value as an index, searching for the corresponding FP16 value in the resident decoder lookup table in the LDM. When the number of parallel lookup paths is 16, 16 low-bit encoded values ​​are simultaneously sent to the lookup table unit, and the on-chip lookup instructions perform parallel searches in 16 memory bank replicas. A single lookup operation can complete the format conversion of 16 elements. Since the decoder lookup table is replicated according to the number of LDM memory banks and the addresses are rearranged, the memory bank access conflict rate of 16 parallel queries is reduced to near zero, and the lookup throughput increases linearly with the number of parallel paths. Subsequently, the 16 FP16 values ​​obtained from the lookup table are directly sent to the matrix multiply-add unit in vector form to participate in the partial multiply-add calculation of the current segment.

[0075] In another preferred embodiment of the present invention, when the target dequantization is configured as a SIMD parallel shift dequantization path, the dequantization operation in step S4 is as follows: B low-bit weight elements are read from the LDM to form a vector, and the corresponding sign extension, shift, mask, and scaling operations are performed according to the type of the current low-bit format. Taking the FP8_E4M3 format as an example, if B=16 (the SIMD vector width is 256 bits, each FP8_E4M3 element occupies 8 bits, i.e., 16 elements are processed at a time), then a single SIMD operation can complete the sign extension, shift, mask, and scaling of 16 elements. Subsequently, the resulting 16 FP16 values ​​are directly fed into the matrix multiply-add unit in vector form. Since the SIMD width utilization reaches 100%, and the dequantization step is compressed into two core bit operations (sign extension + shift) and one floating-point multiplication, the single-element dequantization overhead is extremely low.

[0076] After step S5 is completed, the calculation result of the current weight sub-block (i.e., the output of the linear layer or matrix multiplication and addition) is written back to global storage or retained on-chip for subsequent calculations. Subsequently, the next weight sub-block is ready in the spare buffer, and steps S4 and S5 are repeated until all weight sub-blocks of the current layer are calculated, and the complete linear layer output result is obtained.

[0077] Thus, the low-bit weight adaptive dequantization and on-chip fusion inference method of Embodiment 1 of the present invention has been completed. Through the above steps S1 to S5, the present invention realizes the adaptive determination of dequantization configuration based on different low-bit format structures and target hardware resources on a heterogeneous many-core intelligent acceleration platform lacking low-bit native arithmetic support, and achieves three-stage continuous pipelined execution of low-bit weight transfer, on-chip dequantization, and matrix multiplication and addition through a two-stage double-buffered pipeline mechanism.

[0078] Example 2: Embodiment 2 of the present invention provides a low-bit-weighted adaptive dequantization and on-chip fusion inference system, which is used to implement the low-bit-weighted adaptive dequantization and on-chip fusion inference method of Embodiment 1 above. Figure 6 As shown, the system is applied to a heterogeneous many-core intelligent acceleration platform, and specifically includes: an information acquisition module, a configuration search module, a weight transfer module, an inverse quantization execution module, and an on-chip fusion computing module.

[0079] The information acquisition module is used to acquire hardware resource feature information of the target hardware platform and format feature information of the low bit weights to be processed. The hardware resource feature information includes the width of the Single Instruction Multiple Data Vector and the amount of on-chip memory, while the format feature information includes the exponent bit width, mantissa bit width, and exponent bias. In a specific embodiment of the present invention, the information acquisition module acquires the above information by reading the register configuration of the target hardware platform and the header metadata of the model weight file.

[0080] The configuration search module is used to determine the initial dequantization execution tendency based on format feature information, prioritize multiple candidate dequantization configurations based on hardware resource feature information, and each candidate dequantization configuration includes at least dequantization path type and parallel granularity parameters. At least some candidate dequantization configurations are run on the target hardware platform and real-machine execution feedback information including measured execution time is obtained. The target dequantization configuration is determined from the candidate dequantization configurations based on the real-machine execution feedback information.

[0081] The configuration search module further includes the following sub-units: The format analysis unit calculates the format structure difference value based on the low-bit-weighted exponent width, mantissa width, exponent bias, and data type identifier, and determines the initial dequantization execution tendency and initial parallel granularity based on the format structure difference value. Specifically, the format analysis unit calculates... and according to With preset threshold , The comparison results determine the initial execution tendency and the initial parallel granularity.

[0082] The hardware resource analysis unit evaluates the resource matching degree and prioritizes candidate dequantization configurations based on the target hardware platform's single instruction multiple data vector width, number of on-chip memory banks, number of register resources, and on-chip memory capacity. For parallel lookup dequantization paths, the hardware resource analysis unit evaluates the matching degree between the number of parallel lookup paths and the number of LDM memory banks, as well as memory bank access conflicts; for SIMD parallel dequantization paths, the hardware resource analysis unit evaluates SIMD execution width utilization, register usage, and instruction dependencies.

[0083] The live test unit is used to run candidate dequantization configurations on the target hardware platform and collect at least one feedback metric from the measured execution time, memory access conflict count, register occupancy rate, and pipeline pause cycles.

[0084] The pruning decision unit calculates marginal performance gains based on feedback indicators, prunes candidate configurations that are subject to resource allocation using the gain saturation criterion and resource allocation relationships, and outputs the target inverse-quantized configuration when a preset convergence condition is met. Specifically, the pruning decision unit calculates the target inverse-quantized configuration based on the formula... Calculate the marginal performance gain, and prune low-gain configurations with high parallel granularity when M is less than a preset threshold; simultaneously, prune dominated configurations based on resource allocation relationships (all four indicators are not inferior and at least one indicator is superior). The improvement rate over K consecutive rounds of search is then considered. Less than the convergence threshold When the time comes, stop the search and output the target configuration.

[0085] The weight transfer module is used to transfer low-bit weights from global storage to on-chip local storage according to the target dequantization configuration. Specifically, the weight transfer module determines the range of weight sub-blocks to be transferred each time based on the computation block granularity of the current linear layer or matrix multiplication-addition task, and performs data transfer through the DMA controller. In a preferred embodiment, the weight transfer module is also used to configure a first weight buffer and a second weight buffer in the on-chip local storage to form a first-level double-buffered structure, and control the DMA controller to asynchronously prefetch the next low-bit weight sub-block during the processing of the current weight sub-block.

[0086] The dequantization execution module is used to dequantize low-bit weights into a high-precision data format supported by the target hardware platform in on-chip local memory according to the target dequantization configuration. Specifically, when the target dequantization configuration is a parallel lookup table dequantization path, the dequantization execution module uses the encoded value of the low-bit weight as an index to obtain the corresponding FP16 value from the resident decoding lookup table in the LDM; when the target dequantization configuration is a SIMD parallel shift dequantization path, the dequantization execution module performs corresponding sign extension, shift, masking, and scaling operations according to the low-bit format type. In a preferred embodiment, the dequantization execution module is also used to configure a first high-precision segment buffer and a second high-precision segment buffer in on-chip local memory to form a second-level double-buffered structure, and to perform dequantization on the next low-bit weight segment while the computing unit consumes the current high-precision weight segment.

[0087] The on-chip fusion computing module continuously feeds the high-precision data obtained from dequantization into the computing unit of the target hardware platform in segments for matrix multiplication and addition operations, without forming a complete high-precision weight tensor or writing it back to global storage. The on-chip fusion computing module directly feeds the FP16 temporary data output by the dequantization execution module into the matrix multiplication and addition unit in segments. After the current segment is calculated, it is immediately overwritten by the dequantization result of the next segment, realizing on-chip continuous execution of data transfer, dequantization, and multiplication and addition.

[0088] In one specific embodiment of the present invention, the system further includes a pipeline scheduling module. The pipeline scheduling module coordinates the switching timing of the first-level double-buffered structure and the second-level double-buffered structure, ensuring that the three stages of direct memory access prefetching, dequantization execution, and matrix multiplication-addition overlap in time. Specifically, the pipeline scheduling module configures a first weight buffer and a second weight buffer in the on-chip local memory to form a first-level double-buffered structure, controlling the DMA controller to asynchronously prefetch the next low-bit weight sub-block during the processing of the current weight sub-block; it also configures a first high-precision fragment buffer and a second high-precision fragment buffer in the on-chip local memory to form a second-level double-buffered structure, controlling the dequantization execution module to perform dequantization on the next low-bit weight fragment during the consumption of the current high-precision weight fragment by the computing unit; and coordinates the switching timing of the first-level double-buffered structure and the second-level double-buffered structure, ensuring that the three stages of DMA prefetching, dequantization execution, and matrix multiplication-addition overlap in time.

[0089] In a specific embodiment of the present invention, using the Shenwei intelligent accelerator card as the deployment platform and the Qwen series model in FP8_E4M3 format as the deployment object, the system execution flow of Embodiment 2 of the present invention is as follows: The information acquisition module acquires the SIMD width (256 bits) and the number of LDM memory banks (32) of the Shenwei accelerator card, as well as the format parameters of FP8_E4M3 (exponent bit width 4, mantissa bit width 3, exponent bias 7); the configuration search module determines the initial tendency as a SIMD parallel shift dequantization path (medium parallel granularity) through format analysis, and after actual machine testing and trimming... After branch search, the target configuration was finally determined to be a parallel lookup table dequantization path (16-way parallel lookup, 4 copies of the lookup table). The weight transport module moved low-bit weight blocks from HBM to LDM, and achieved prefetching and processing overlap through the first-level double buffer. The dequantization execution module converted the FP8_E4M3 weights into FP16 fragments through a parallel lookup table. The on-chip fusion computing module directly sent the FP16 fragments into the matrix multiplication and addition unit. The pipeline scheduling module coordinated the switching timing of the two-level double buffer, so that the three stages of DMA prefetching, dequantization execution, and matrix multiplication and addition formed a continuous pipeline. No complete FP16 weight tensor was generated during the entire execution process, effectively reducing on-chip storage pressure and off-chip memory access overhead.

[0090] This embodiment provides a low-bit-weight adaptive dequantization and on-chip fusion inference system for implementing the aforementioned low-bit-weight adaptive dequantization and on-chip fusion inference method. Therefore, the specific implementation of the low-bit-weight adaptive dequantization and on-chip fusion inference system can be found in the previous embodiment section of the low-bit-weight adaptive dequantization and on-chip fusion inference method. Thus, the specific implementation can be referred to the description of the corresponding embodiments. To avoid redundancy, it will not be repeated here.

[0091] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0092] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0093] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0094] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A low-bit-weighted adaptive dequantization and on-chip fusion inference method, characterized in that, The method, applied to a heterogeneous many-core intelligent acceleration platform, includes: The hardware resource feature information and the format feature information of the low bit weight to be processed of the target hardware platform are obtained. The hardware resource feature information includes the width of the single instruction multiple data vector and the number of on-chip memory. The format feature information includes the exponent bit width, the mantissa bit width and the exponent bias. The initial dequantization execution tendency is determined based on the format feature information, and multiple candidate dequantization configurations are prioritized based on the hardware resource feature information. Each candidate dequantization configuration includes at least a dequantization path type and a parallel granularity parameter. Run at least some of the candidate dequantization configurations on the target hardware platform, obtain real-machine execution feedback information including measured execution time, and determine the target dequantization configuration from the candidate dequantization configurations based on the real-machine execution feedback information; According to the target dequantization configuration, the low bit weights are moved from global storage to on-chip local storage, and in the on-chip local storage, the low bit weights are dequantized into a high-precision data format supported by the target hardware platform according to the target dequantization configuration. The high-precision data obtained by dequantization is continuously fed into the computing unit of the target hardware platform in segments to participate in matrix multiplication and addition operations, without forming a complete high-precision weight tensor or writing it back to the global storage.

2. The low-bit-weighted adaptive dequantization and on-chip fusion inference method according to claim 1, characterized in that, The target dequantization configuration includes a lookup table dequantization path; the step of dequantizing the low-bit weights into a high-precision data format supported by the target hardware platform includes: A decoding lookup table between low-bit encoded values ​​and high-precision data format values ​​is generated in advance, and the decoding lookup table is stored in the on-chip local memory; Using the encoded value with the low bit weight as an index, the corresponding high-precision data format value is obtained from the decoding lookup table through on-chip lookup instructions; Based on the number of storage banks in the on-chip local storage, multiple access copies or logical mapping copies of the decoding lookup table are configured so that the access addresses of parallel queries are allocated to different storage banks to reduce storage bank access conflicts.

3. The low-bit-weighted adaptive dequantization and on-chip fusion inference method according to claim 1, characterized in that, The target dequantization configuration includes a single instruction multiple data stream parallel shift dequantization path; the single instruction multiple data stream parallel shift dequantization path supports at least the following three data formats: a first low-bit floating-point format, with an exponent width of 4 bits and a mantissa width of 3 bits; a second low-bit floating-point format, with an exponent width of 5 bits and a mantissa width of 2 bits; And the low-bit format of integers; The step of dequantizing the low-bit weights into a high-precision data format supported by the target hardware platform includes: For the first low-bit floating-point format, the low-bit data is extended to the bit width of the high-precision data format through the sign extension operation, the extended data is shifted to achieve bit field alignment, redundant high bits are erased by the bit-AND operation mask constant, and the exponential bias compensation factor and the block-level scaling factor are merged into a single floating-point correction coefficient to complete the numerical recovery in a single floating-point multiplication. For the second low-bit floating-point format, the low-bit data is extended to the bit width of the high-precision data format through sign extension operation, and the bit field alignment of the exponent and mantissa bits is completed through shift operation, without performing an independent exponent bias correction operation. For the aforementioned low-bit integer format, the integer data is extended to the bit width of the high-precision data format through a sign extension operation. Using the preset offset interval in the half-precision floating-point numerical representation, the sign-extended integer data is added to the preset offset constant, and then subtracted from the preset offset constant to obtain the mapping result from integer to floating-point number.

4. The low-bit-weighted adaptive dequantization and on-chip fusion inference method according to claim 1, characterized in that, The step of determining the initial dequantization execution tendency based on the format feature information includes: Obtain the format parameters of the low bit weight, which include the exponent bit width, mantissa bit width, exponent bias, and data type identifier; Calculate the format structure difference value between the low bit weight format and the high precision data format. The format structure difference value is determined by a weighted sum of the exponent bit width difference, mantissa bit width difference, exponent bias difference, and type difference identifier. When the format structure difference value is less than or equal to the first preset threshold, the initial dequantization execution tendency is determined to be a single instruction multiple data stream parallel shift dequantization path, and the initial parallel granularity is set to high parallel granularity. When the format structure difference value is greater than the first preset threshold and less than or equal to the second preset threshold, the initial dequantization execution tendency is determined to be a single instruction multiple data stream parallel shift dequantization path, and the initial parallel granularity is set to medium parallel granularity. When the difference value of the format structure is greater than the second preset threshold, the initial dequantization execution tendency is determined to be the parallel lookup table dequantization path; The first preset threshold and the second preset threshold are predetermined based on the bit field structure compatibility between the high-precision data format and the low bit weight format. The high parallel granularity, medium parallel granularity, and low parallel granularity are parallel path levels predetermined based on the single instruction multiple data stream vector width of the target hardware platform.

5. The low-bit-weighted adaptive dequantization and on-chip fusion inference method according to claim 1, characterized in that, The step of determining the target dequantization configuration from the candidate dequantization configurations based on the actual execution feedback information includes: Obtain at least one feedback metric from the measured execution time, memory access conflict count, register occupancy rate, and pipeline pause cycle of the candidate dequantization configuration on the target hardware platform. Based on the measured execution time before and after the change in parallel granularity, calculate the marginal performance gain corresponding to a unit change in parallel granularity. When the marginal performance gain is less than the preset marginal gain threshold, it is determined that the current parallel granularity direction has entered the gain saturation region, and the candidate configuration with higher parallel granularity and resource consumption not lower than the current configuration is removed from the candidate dequantization configuration or its test priority is reduced. For any first candidate configuration and second candidate configuration in the candidate dequantization configuration, if the first candidate configuration is less than or equal to the corresponding index value of the second candidate configuration in all four indicators of measured execution time, memory access conflict number, register occupancy rate and pipeline pause cycle, and at least one indicator is less than the corresponding index value of the second candidate configuration, it is determined that the second candidate configuration is dominated by the first candidate configuration, and the second candidate configuration is deleted from the candidate dequantization configuration. After completing at least one round of search and pruning, the candidate dequantization configuration with the lowest measured execution time among the remaining candidate dequantization configurations is selected as the target dequantization configuration.

6. The low-bit-weighted adaptive dequantization and on-chip fusion inference method according to claim 1, characterized in that, The step of moving the low-bit weights from global storage to on-chip local storage, dequantizing the low-bit weights in the on-chip local storage according to the target dequantization configuration into a high-precision data format supported by the target hardware platform, and continuously feeding the dequantized high-precision data into the computing unit of the target hardware platform in segments for matrix multiplication and addition operations includes: A first weight buffer and a second weight buffer are set in the on-chip local storage to form a first-level double buffer structure; the current low-bit weight sub-block is moved to the first weight buffer through direct memory access, and the next low-bit weight sub-block is moved to the second weight buffer asynchronously; and the first weight buffer and the second weight buffer are swapped after the current low-bit weight sub-block is processed. The current weighted sub-block is divided into multiple low-bit weighted segments. A first high-precision segment buffer and a second high-precision segment buffer are set in the on-chip local storage to form a second-level double-buffered structure. While the high-precision weighted segment in the current high-precision segment buffer is read by the computing unit and matrix multiplication and addition operations are performed, the dequantization unit dequantizes the next low-bit weighted segment into a high-precision data format and writes it into another high-precision segment buffer. After the current segment is consumed, the first high-precision segment buffer and the second high-precision segment buffer are swapped. The first-level double-buffered structure and the second-level double-buffered structure are executed concurrently in time, so that the low-bit-weighted direct memory access prefetch, on-chip dequantization data generation and matrix multiplication and addition data consumption form a three-stage continuous pipeline.

7. The low-bit-weighted adaptive dequantization and on-chip fusion inference method according to claim 1, characterized in that, The method further includes: Obtain the block-level quantization parameters of the low bit weights, which are stored in the global storage. Multiple low bit weight elements share the same block-level quantization parameters. Based on the linear storage index of the low bit weight in the global storage, the linear storage index of the corresponding block-level quantization parameter is determined by closed-loop mapping calculation. During the dequantization process, the corresponding block-level quantization parameter position is directly determined based on the linear storage index of the current low-bit weight element, and the block-level quantization parameter is embedded into the dequantization calculation path.

8. A low-bit-weighted adaptive dequantization and on-chip fusion inference system, characterized in that, The system, applied to a heterogeneous many-core intelligent acceleration platform, includes: The information acquisition module is used to acquire hardware resource feature information and format feature information of low bit weight to be processed from the target hardware platform. The hardware resource feature information includes the width of the single instruction multiple data vector and the number of on-chip memory blocks. The format feature information includes the exponent bit width, the mantissa bit width and the exponent bias. A configuration search module is used to determine the initial dequantization execution tendency based on the format feature information, prioritize multiple candidate dequantization configurations based on the hardware resource feature information, each candidate dequantization configuration includes at least a dequantization path type and parallel granularity parameters, run at least some of the candidate dequantization configurations on the target hardware platform and obtain real-machine execution feedback information including measured execution time, and determine the target dequantization configuration from the candidate dequantization configurations based on the real-machine execution feedback information; The weight transfer module is used to transfer the low-bit weights from global storage to on-chip local storage according to the target dequantization configuration. The dequantization execution module is used to dequantize the low bit weight into a high-precision data format supported by the target hardware platform in the on-chip local storage according to the target dequantization configuration. The on-chip fusion computing module is used to continuously send the high-precision data obtained by dequantization into the computing unit of the target hardware platform in segments to participate in matrix multiplication and addition operations, without forming a complete high-precision weight tensor or writing it back to the global storage.

9. The low-bit-weighted adaptive dequantization and on-chip fusion inference system according to claim 8, characterized in that, The configuration search module includes: The format analysis unit is used to calculate the format structure difference value based on the exponent bit width, mantissa bit width, exponent bias and data type identifier of the low bit weight, and to determine the initial dequantization execution tendency and initial parallel granularity based on the format structure difference value. The hardware resource analysis unit is used to evaluate the resource matching degree and prioritize the candidate dequantization configurations based on the single instruction multiple data stream vector width, number of on-chip memory, number of register resources and on-chip memory capacity of the target hardware platform. The real-machine testing unit is used to run candidate dequantization configurations on the target hardware platform and collect at least one feedback indicator among the measured execution time, memory access conflict count, register occupancy rate, and pipeline pause cycle. The pruning decision unit is used to calculate the marginal performance gain based on the feedback index, prune the dominated candidate configurations using the gain saturation criterion and resource allocation relationship, and output the target inverse quantization configuration when the preset convergence condition is met.

10. The low-bit-weighted adaptive dequantization and on-chip fusion inference system according to claim 8, characterized in that, The system also includes a flow scheduling module, which is used for: A first weight buffer and a second weight buffer are configured in the on-chip local memory to form a first-level double buffer structure, and the direct memory access controller is controlled to asynchronously prefetch the next low-bit weight sub-block during the processing of the current weight sub-block. A first high-precision fragment buffer and a second high-precision fragment buffer are configured in the on-chip local storage to form a second-level double buffer structure, and the dequantization execution module is controlled to perform dequantization on the next low-bit weight fragment while the computing unit consumes the current high-precision weight fragment. The timing of switching between the first-level double-buffered structure and the second-level double-buffered structure is coordinated so that the three stages of direct memory access prefetching, dequantization execution, and matrix multiplication and addition operations overlap in time.