Model reasoning acceleration method and device, electronic equipment and storage medium

By converting the weight data of large language models from BF16 format to BFP8 format and performing quantization compression, the problem of low model weight compression efficiency in existing technologies is solved, thereby improving model inference efficiency and maintaining performance.

CN120996176APending Publication Date: 2025-11-21YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510872335.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies for compressing model weights in large language models suffer from complex operations, low quantization efficiency, and impact on model accuracy and stability. How to efficiently compress model weights to improve inference efficiency while maintaining model performance has become an urgent technical problem to be solved.

Method used

The block floating-point representation method is adopted to convert the weight data from BF16 format to BFP8 format. Quantization compression is performed by sharing the exponent and storing the mantissa separately. Dequantization is performed during inference calculation to restore the precision, thereby achieving efficient conversion of weight data.

Benefits of technology

It effectively reduces data storage space and transmission overhead, significantly improves the inference efficiency of the model, and ensures that the model performance is not affected.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996176A_ABST
    Figure CN120996176A_ABST
Patent Text Reader

Abstract

The invention provides a model reasoning acceleration method and device, electronic equipment and a storage medium, and relates to the technical field of computers, and the method comprises the following steps: obtaining weight data of a target model; the data format of the weight data is a first floating-point number format; converting the weight data from the first floating-point number format to a second floating-point number format to obtain quantized weight data; the precision of the second floating-point number format is lower than that of the first floating-point number format; loading the quantized weight data, and converting the quantized weight data from the second floating-point number format to the first floating-point number format to obtain inverse quantized weight data; and executing the reasoning task of the target model based on the inversely quantized weight data. According to the method and the device provided by the invention, the reasoning efficiency of the model can be remarkably improved, and meanwhile, the model performance is not influenced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer technology, and in particular to a model inference acceleration method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the rapid development of deep learning technology, the scale of large language models (hereinafter referred to as "large models") is constantly expanding, and the parameter quantity and computational complexity thereof are growing exponentially. This has caused a sharp increase in the demand for storage bandwidth and computing resources for inference computation, putting enormous pressure on hardware resources.

[0003] In order to alleviate this pressure, the academic and industrial communities have widely adopted low-precision representation formats to compress model weights and activation values, aiming to improve the efficiency of large models during inference and reduce storage and transmission overheads. However, these methods have many problems in practical applications: complex operations, low quantization efficiency, and often a significant reduction in model inference efficiency after quantization, even affecting the accuracy and stability of the model.

[0004] Therefore, how to efficiently compress model weights to significantly improve the inference efficiency of the model while ensuring that the model performance is not affected has become a technical problem that the industry urgently needs to solve. SUMMARY

[0005] The present application provides a model inference acceleration method, device, electronic equipment and storage medium, which solves the technical problem of how to efficiently compress model weights to significantly improve the inference efficiency of the model while ensuring that the model performance is not affected.

[0006] The present application provides a model inference acceleration method, comprising: obtaining weight data of a target model; the data format of the weight data is a first floating-point number format; converting the weight data from the first floating-point number format to a second floating-point number format to obtain quantized weight data; the precision of the second floating-point number format is lower than that of the first floating-point number format; loading the quantized weight data and converting the quantized weight data from the second floating-point number format to the first floating-point number format to obtain dequantized weight data; based on the dequantized weight data, performing an inference task of the target model.

[0007] In some embodiments, the first floating-point number format is a BF16 format; and the second floating-point number format is a BFP8 format.

[0008] In some embodiments, the converting the weight data from the first floating-point number format to a second floating-point number format to obtain quantized weight data comprises: dividing the weight data to obtain a plurality of data blocks; determining an exponent part and a mantissa part of each data in a current data block in the first floating-point number format; determining a maximum value of the exponent part of each data in the current data block as a shared exponent, and determining a right shift number of the mantissa based on a difference between the exponent part of each data and the shared exponent; performing a right shift operation on the mantissa part of each data in the current data block based on the right shift number of the mantissa to obtain a right shift operation mantissa part of each data; obtaining an exponent part and a mantissa part of each data in the current data block in the second floating-point number format based on the shared exponent and the right shift operation mantissa part of each data; obtaining quantized weight data based on the exponent part and the mantissa part of each data in each data block in the second floating-point number format.

[0009] In some embodiments, the obtaining an exponent part and a mantissa part of each data in the current data block in the second floating-point number format based on the shared exponent and the right shift operation mantissa part of each data comprises: determining the shared exponent as the exponent part of each data in the second floating-point number format; determining a sign bit of each data based on the positive and negative of each data; determining a mantissa bit of each data based on the right shift operation mantissa part of each data; determining the mantissa part of each data in the second floating-point number format based on the sign bit and the mantissa bit of each data.

[0010] In some embodiments, the converting the quantized weight data from the second floating-point number format to the first floating-point number format to obtain dequantized weight data comprises: determining the positive and negative of the dequantized weight data based on the sign bit of each data; determining a value of the dequantized weight data based on the mantissa part of each data in the second floating-point number format and a power of two; a power index in the power of two is determined based on the shared exponent.

[0011] In some embodiments, the performing an inference task of the target model based on the dequantized weight data comprises: obtaining input data; the data format of the input data is the first floating-point number format; input the dequantized weight data and the input data into the target model to obtain an inference result output by the target model.

[0012] The application provides a model inference acceleration device, comprising: a weight data obtaining module configured to obtain weight data of a target model, wherein the weight data is in a first floating-point number format; a quantization module configured to convert the weight data from the first floating-point number format to a second floating-point number format to obtain quantized weight data, wherein the second floating-point number format has a lower precision than the first floating-point number format; a dequantization module configured to load the quantized weight data and convert the quantized weight data from the second floating-point number format to the first floating-point number format to obtain dequantized weight data; a calculation module configured to perform an inference task of the target model based on the dequantized weight data.

[0013] The application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the model inference acceleration method when executing the computer program.

[0014] The application provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the model inference acceleration method.

[0015] The application provides a computer program product comprising a computer program, wherein the computer program is executable by a processor to implement the model inference acceleration method.

[0016] The model inference acceleration method, device, electronic device, and storage medium provided by the application comprise the following steps: obtaining weight data of a target model; the weight data is in a first floating-point number format; converting the weight data from the first floating-point number format to a second floating-point number format to obtain quantized weight data; the second floating-point number format has a lower precision than the first floating-point number format; loading the quantized weight data and converting the quantized weight data from the second floating-point number format to the first floating-point number format to obtain dequantized weight data; performing an inference task of the target model based on the dequantized weight data; since the weight data is quantized and compressed, the storage space of the data is reduced, and the transmission cost of the data is reduced; the weight data is dequantized during inference calculation, the precision of the weight data is restored, the inference efficiency of the model is significantly improved, and the performance of the model is ensured. BRIEF DESCRIPTION OF DRAWINGS

[0017] The accompanying drawings, which are incorporated herein and constitute part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles of the application.

[0018] In order to make the technical solution of the present application or the prior art clearer, the drawings needed in the following embodiment or prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.

[0019] Figure 1 is one of the flowcharts of the model inference acceleration method provided by the present application.

[0020] Figure 2 is another flowchart of the model inference acceleration method provided by the present application.

[0021] Figure 3 is a structural schematic diagram of the model inference acceleration device provided by the present application.

[0022] Figure 4 is a structural schematic diagram of the electronic device provided by the present application. DETAILED DESCRIPTION

[0023] In order to make the technical solution of the present application or the prior art clearer, the drawings needed in the following embodiment or prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.

[0024] It should be noted that the terms "first", "second", and the like in the present application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units or modules does not have to be limited to those steps or units or modules clearly listed, but can include other steps or units or modules that are not clearly listed or inherent to these processes, methods, products or devices.

[0025] To solve the problems in the related art, the present application provides a model inference acceleration method based on block floating point, which can effectively reduce the precision requirement while maintaining the dynamic range by sharing the exponent part and storing the mantissa separately in the weight data. At the same time, this method does not require users to quantize the model, and the operation is more simple. By using the block floating point method to compress the large model weight, the bandwidth bottleneck problem in the large model inference process can be effectively alleviated, and the large model inference efficiency can be improved. The above method can be used in large model inference tasks, and has high practical value and innovation value.

[0026] Figure 1 is one of the flowcharts of the model inference acceleration method provided by the present application, as shown in Figure 1 The method comprises steps 110, 120, 130 and 140.

[0027] Step 110, obtaining the weight data of the target model; the data format of the weight data is a first floating point number format.

[0028] Specifically, the execution subject of the model inference acceleration method provided by the embodiments of the present application is a model inference acceleration device. The device can be realized by software, such as a model inference acceleration program running in a computer; or by hardware, such as a computer or server executing the model inference acceleration method.

[0029] The target model refers to a deep learning model that needs to perform an inference task, which can be various types of large language models, etc. The model structure of the target model is not specifically limited in the embodiments of the present application.

[0030] Weight data is a key parameter in a deep learning model, which determines how the model processes and converts input data to achieve specific tasks such as classification, regression or generation.

[0031] Weight data is usually stored and processed in floating point number format. Floating point number format is a binary format used to represent real numbers, which can represent a very large range and a very small value with limited bits.

[0032] In the embodiments of the present application, the data format of the weight data is a first floating point number format. The first floating point number format can be BF16 format (Brain Floating Point 16). BF16 format is a 16-bit floating point number representation method.

[0033] Step 120, converting the weight data from the first floating point number format to a second floating point number format to obtain quantized weight data; the precision of the second floating point number format is lower than that of the first floating point number format.

[0034] Specifically, in the embodiments of the present application, the second floating-point number format can be a BFP8 format (Block Floating Point 8). The BFP8 format is an 8-bit floating-point number representation method mainly used for low-precision calculation in deep learning and machine learning. The precision of the BFP8 format is lower than that of the BF16 format.

[0035] Quantization refers to the process of mapping continuous or high-precision numerical values into a finite, discrete set of numerical values. In deep learning, quantization is often used to reduce the storage space and computational resource consumption of a model while maintaining the performance of the model. Specifically, quantization can convert high-precision floating-point number formats into low-precision floating-point number formats.

[0036] The weight data can be converted from the first floating-point number format to the second floating-point number format to obtain quantized weight data.

[0037] The quantized weight data can be stored. The quantized weight data can reduce storage space, improve computational efficiency, and reduce the overhead of data transmission.

[0038] Step 130, load the quantized weight data and convert the quantized weight data from the second floating-point number format to the first floating-point number format to obtain dequantized weight data.

[0039] Specifically, dequantization is the inverse process of quantization, i.e., mapping the quantized discrete values back to the original continuous or high-precision numerical values. Dequantization is often used to restore the original data when high-precision calculation is required.

[0040] When inference calculation is required, the quantized weight data can be loaded from storage first. Then the quantized weight data is converted from the second floating-point number format to the first floating-point number format to obtain dequantized weight data. The precision of the dequantized weight data is restored.

[0041] Step 140, based on the dequantized weight data, perform the inference task of the target model.

[0042] Specifically, the dequantized weight data and the input data are input into the target model, and the target model performs matrix multiplication and other operations to output the inference result, completing the inference task of the target model.

[0043] The model inference acceleration method provided by the embodiment of the present application comprises the following steps: obtaining weight data of a target model; the data format of the weight data is a first floating-point number format; converting the weight data from the first floating-point number format into a second floating-point number format to obtain quantized weight data; the precision of the second floating-point number format is lower than that of the first floating-point number format; loading the quantized weight data and converting the quantized weight data from the second floating-point number format into the first floating-point number format to obtain dequantized weight data; performing an inference task of the target model based on the dequantized weight data; since the weight data is quantized and compressed, the storage space of the data is reduced, and the transmission cost of the data is reduced; the weight data is dequantized during inference calculation, the precision of the weight data is restored, the inference efficiency of the model can be significantly improved, and the model performance is not affected.

[0044] It should be noted that each embodiment of the present application can be freely combined, the order can be changed or executed alone, and does not need to rely on or depend on a fixed execution order.

[0045] In some embodiments, converting the weight data from the first floating-point number format into the second floating-point number format to obtain the quantized weight data comprises: dividing the weight data to obtain a plurality of data blocks; determining the exponent part and the mantissa part of each data in the current data block in the first floating-point number format; determining the maximum value of the exponent part of each data in the current data block as a shared exponent, and determining the right shift number of the mantissa based on the difference between the exponent part of each data and the shared exponent; performing a right shift operation on the mantissa part of each data in the current data block based on the right shift number of the mantissa to obtain the right shift operation mantissa part of each data; obtaining the exponent part and the mantissa part of each data in the current data block in the second floating-point number format based on the shared exponent and the right shift operation mantissa part of each data; obtaining the quantized weight data based on the exponent part and the mantissa part of each data in each data block in the second floating-point number format.

[0046] Specifically, the weight data is usually represented in the form of a tensor. A tensor is a multidimensional array, which can be represented as a one-dimensional (vector), two-dimensional (matrix) or higher-dimensional array.

[0047] The weight data can be divided to obtain a plurality of data blocks. When dividing the weight data, each tensor can be divided into a fixed size sub-tensor to realize the blocking of the weight data. A sub-tensor is a data block. Each data block contains a certain number of elements, for example, 16 or 32 elements. Larger blocks can improve the compression ratio, but can reduce the precision; smaller blocks can better maintain the precision, but the compression ratio is relatively low. The block size can be set as a hyperparameter to balance the compression ratio and the precision.

[0048] Taking the current data block as an example, the exponent part and the mantissa part of each data in the first floating-point number format can be determined first. For example, the current data block is a data block with a block size of 4, and the first floating-point number format is the BF16 format. The weight data in the data block are 1.5, 3.25, 6.5, and 1.875, respectively. The exponent part, the mantissa part, and the real expression process of each data in the BF16 format are shown in Table 1. In the table, bit represents a bit.

[0049] Table 1: Data block in BF16 format

[0050] For each data in the current data block, the exponent part of each data can be extracted. The maximum value (the maximum valid exponent value) among them is denoted as E_max and determined as the shared exponent of the current data block. For the remaining data x_i, the difference d_i between the exponent part of each data and the shared exponent E_max-E_i is calculated. Wherein, E_i is the exponent part of the data x_i. i is the serial number of the data. The difference d_i is determined as the number of right shift of the mantissa. For example, in the above data block, the exponent of the data 6.5 is the maximum, which is 129, and can be used as the shared exponent of the block. The difference exponent of the four numbers is 2, 1, 0, and 2, respectively.

[0051] In addition to the shared exponent, the BFP8 format (the second floating-point number format) includes 1 sign bit and 7 mantissa bits, and the first bit of the mantissa bit is a hidden bit (hidden_bit). If the number represented by the mantissa bit is greater than or equal to 1, the hidden_bit is 1, otherwise it is 0.

[0052] According to the number of right shift of the mantissa, the mantissa part of each data in the current data block is right shifted to obtain the right shifted mantissa part of each data. For example, for each data, the mantissa needs to be right shifted by d_i bits (i.e., divided by 2^d_i) to align to the same exponent, so that the mantissa can be shared. For the part of the data that cannot be represented due to insufficient bits after right shifting, a nearest even rounding method is used for rounding.

[0053] The shared exponent can be taken as an exponent part of each data in the current data block in the second floating-point number format, and the right shift operation mantissa part can be taken as a mantissa part of each data in the current data block in the second floating-point number format, so that the weight data is converted from the first floating-point number format to the second floating-point number format. The weight data in the second floating-point number format is quantized weight data.

[0054] The model inference acceleration method provided in the embodiments of the present application converts the weight data from the first floating-point number format to the second floating-point number format to obtain quantized weight data, and effectively reduces the precision requirement while maintaining the dynamic range by sharing the exponent part in each data block and storing the mantissa part separately.

[0055] In some embodiments, based on the shared exponent and the right shift operation mantissa part of each data, the exponent part and the mantissa part of each data in the current data block in the second floating-point number format are obtained, including: determining the exponent part of each data in the second floating-point number format as the shared exponent; determining the sign bit of each data based on the positive and negative of each data; determining the mantissa bit of each data based on the right shift operation mantissa part of each data; determining the mantissa part of each data in the second floating-point number format based on the sign bit and the mantissa bit of each data.

[0056] Specifically, the exponent part of each data in the second floating-point number format is the same, which is the shared exponent.

[0057] According to the shared exponent and the right shift operation mantissa part of each data, the exponent part and the mantissa part of each data in the current data block in the second floating-point number format are obtained, as shown in Table 2.

[0058] Table 2 Data after right shift of mantissa

[0059] The same processing is performed on each data in each data block to obtain quantized weight data.

[0060] The mantissa part of each data in the second floating-point number format includes two parts, which are the sign bit and the mantissa bit. The sign bit can be determined according to the positive and negative of each data. The mantissa bit can be determined according to the right shift operation mantissa part of each data.

[0061] The weight data is stored in units of data blocks. Each data block contains an 8-bit shared exponent part and a mantissa part of each data element, wherein each mantissa part is also 8 bits, including 1 sign bit and 7 mantissa bits.

[0062] The model inference acceleration method provided by the embodiment of the application converts the weight data from the first floating-point number format to the second floating-point number format, quantizes and compresses the weight data, which is beneficial to reduce the storage space of the data and reduce the transmission overhead of the data.

[0063] In some embodiments, the quantized weight data is converted from the second floating-point number format to the first floating-point number format to obtain the dequantized weight data, including: determining the positive and negative of the dequantized weight data based on the sign bit of each data; determining the value of the dequantized weight data based on the mantissa part and the power of two of each data in the second floating-point number format; the power index in the power of two is determined based on the shared exponent.

[0064] Specifically, when performing inference calculation, the quantized weight data needs to be loaded into the local cache or register of the artificial intelligence processor to reduce the memory bandwidth consumption and power consumption. The artificial intelligence chip can be a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), a domain specific architecture (DSA), etc.

[0065] When performing calculation, the quantized weight data can be converted from the second floating-point number format to the first floating-point number format to obtain the dequantized weight data, so as to realize the recovery of the precision.

[0066] The sign bit can be used to determine the positive and negative of the dequantized weight data, and the value of the dequantized weight data can be determined based on the mantissa part and the power of two of each data in the second floating-point number format, which is expressed by the formula as follows: .

[0067] wherein, is the dequantized weight data; is the sign bit; is the hidden bit in the mantissa bit; is the mantissa bit except the hidden bit; is the shared exponent.

[0068] The model inference acceleration method provided by the embodiment of the application converts the quantized weight data from the second floating-point number format to the first floating-point number format to obtain the dequantized weight data, realizes the recovery of the precision, can significantly improve the inference efficiency of the model, and ensures that the model performance is not affected.

[0069] In some embodiments, based on the dequantized weight data, an inference task of the target model is performed, including: obtaining input data; the data format of the input data is a first floating-point number format; inputting the dequantized weight data and the input data into the target model to obtain an inference result output by the target model.

[0070] Specifically, the input data is usually provided in a first floating-point number format. The data format of the dequantized weight data is provided in a first floating-point number format. The dequantized weight data and the input data can be input into the target model, the input data is processed by the target model, matrix multiplication, convolution, etc. operations are performed, the forward inference process of the model is completed, and the inference result is obtained.

[0071] The model inference acceleration method provided by the embodiment of the application can significantly improve the inference efficiency of the model while ensuring that the performance of the model is not affected.

[0072] Figure 2 is a second flowchart of the model inference acceleration method provided by the application, as shown in Figure 2 the method is applied to inference acceleration of a large model, including: Step 210, obtaining large model weight data. Since the large model training often adopts BF16 data precision, the data format of the weight data trained for inference is BF16 format.

[0073] Step 220, quantizing the weight data into BFP8 format.

[0074] The weight data is divided into data blocks, and each tensor in the large model is divided into a fixed size data block. For the elements in each data block, the exponent part in the BF16 format representation is extracted, and the maximum valid exponent value in the block is found as the shared exponent of the block. In addition to the shared exponent, BFP8 includes 1 sign bit and 7 mantissa bits, and the first bit of the mantissa bit is a hidden bit (hidden_bit). If the number represented by the mantissa bit is greater than or equal to 1, hidden_bit is 1, otherwise it is 0. Due to the sharing of the exponent, for each element, the mantissa needs to be right shifted by d_i bits (i.e. divided by 2^d_i) to align to the same exponent, so that the mantissa can be shared. For the part of the data that cannot be represented due to the lack of bits after right shifting, the nearest even rounding method is used for rounding. The weight data is stored in BFP8 format. According to the data block as the storage unit, each block contains an 8-bit shared exponent part and a mantissa part of each data element, and each mantissa part is also 8 bits, including 1 sign bit and 7 mantissa bits.

[0075] Step 230, load the weight data in BFP8 format to the computing unit.

[0076] Load the quantized BFP8 format weight data into the local cache or register of the acceleration computing unit to reduce memory bandwidth consumption and power consumption.

[0077] Step 240, dynamically restore the weight data in BFP8 format to weight data in BF16 format.

[0078] Step 250, perform inference calculation using BF16 format.

[0079] Perform matrix multiplication, convolution and other operations on the weight data in BF16 format after precision restoration and the input data in BF16 format, and complete the forward inference process of the model.

[0080] Step 260, the complete large model inference process can be used in common tasks of large models such as question and answer, query.

[0081] The model inference acceleration method provided by the embodiment of the application adopts block floating point to quantize and compress the weight of the large model, which can effectively alleviate the bandwidth bottleneck problem in the large model inference process and improve the large model inference efficiency. The above method can be used in the large model inference task, and has high practical value and innovation value.

[0082] The device provided by the embodiment of the application will be described below. The device described below can be correspondingly referred to the method described above.

[0083] Figure 3 is a structural diagram of the model inference acceleration device provided by the application, as Figure 3 shown, the device comprises: The acquisition module 310 is configured to acquire weight data of a target model, and the data format of the weight data is a first floating point format. The quantization module 320 is configured to convert the weight data from the first floating point format to a second floating point format to obtain quantized weight data, and the precision of the second floating point format is lower than that of the first floating point format. The dequantization module 330 is configured to load the quantized weight data and convert the quantized weight data from the second floating point format to the first floating point format to obtain dequantized weight data. The computing module 340 is configured to perform an inference task of the target model based on the dequantized weight data.

[0084] The model inference acceleration device provided by the embodiment of the present application obtains weight data of a target model; the data format of the weight data is a first floating-point number format; the weight data is converted from the first floating-point number format to a second floating-point number format to obtain quantized weight data; the precision of the second floating-point number format is lower than that of the first floating-point number format; the quantized weight data is loaded and converted from the second floating-point number format to the first floating-point number format to obtain dequantized weight data; a reasoning task of the target model is performed based on the dequantized weight data; since the weight data is quantized and compressed, the storage space of the data is reduced, and the transmission overhead of the data is reduced; the weight data is dequantized during reasoning calculation to restore the precision of the weight data, which can significantly improve the reasoning efficiency of the model while ensuring that the model performance is not affected.

[0085] Figure 4 is a structural schematic diagram of an electronic device provided by the present application, as Figure 4 shown, the electronic device can include a processor (Processor) 410, a communications interface (Communications Interface) 420, a memory (Memory) 430 and a communications bus (Communications Bus) 440, wherein the processor 410, the communications interface 420, the memory 430 complete the communication among each other through the communications bus 440. The processor 410 can call the logic command in the memory 430 to execute the method described in the above embodiment, for example: obtain weight data of a target model; the data format of the weight data is a first floating-point number format; the weight data is converted from the first floating-point number format to a second floating-point number format to obtain quantized weight data; the precision of the second floating-point number format is lower than that of the first floating-point number format; the quantized weight data is loaded and converted from the second floating-point number format to the first floating-point number format to obtain dequantized weight data; a reasoning task of the target model is performed based on the dequantized weight data.

[0086] In addition, the logic commands in the memory described above can be implemented in the form of a software function unit and sold or used as a separate product, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of commands to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0087] The processor in the electronic device provided by the embodiments of the present application can call the logic instructions in the memory to implement the above-mentioned method, and the specific implementation manners are consistent with the above-mentioned method implementation manners, and the same beneficial effects can be achieved, which will not be described here.

[0088] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method provided by each of the above-mentioned embodiments.

[0089] The specific implementation manners are consistent with the above-mentioned method implementation manners, and the same beneficial effects can be achieved, which will not be described here.

[0090] The embodiments of the present application provide a computer program product, which includes a computer program, and the computer program is executed by a processor to implement the above-mentioned method.

[0091] The device embodiments described above are only schematic, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present embodiment scheme according to actual needs. Those skilled in the art can understand and implement it without creative labor.

[0092] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in the various embodiments or some parts of the embodiments.

[0093] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A model inference acceleration method, characterized in that, The method comprises the following steps: obtaining weight data of a target model; the data format of the weight data is a first floating-point number format; converting the weight data from the first floating-point number format into a second floating-point number format to obtain quantized weight data; the precision of the second floating-point number format is lower than that of the first floating-point number format; loading the quantized weight data and converting the quantized weight data from the second floating-point number format into the first floating-point number format to obtain dequantized weight data; performing an inference task of the target model based on the dequantized weight data.

2. The model inference acceleration method of claim 1, wherein, The first floating-point number format is a BF16 format; and the second floating-point number format is a BFP8 format.

3. The model inference acceleration method of claim 1, wherein, The conversion of the weight data from the first floating-point number format into the second floating-point number format to obtain the quantized weight data comprises the following steps: dividing the weight data to obtain a plurality of data blocks; determining the exponent part and the mantissa part of each data in a current data block in the first floating-point number format; determining the maximum value of the exponent part of each data in the current data block as a shared exponent, and determining the right shift number of the mantissa part based on the difference between the exponent part of each data and the shared exponent; performing a right shift operation on the mantissa part of each data in the current data block based on the right shift number of the mantissa part to obtain the right shift operation mantissa part of each data; obtaining the exponent part and the mantissa part of each data in the current data block in the second floating-point number format based on the shared exponent and the right shift operation mantissa part of each data; obtaining the quantized weight data based on the exponent part and the mantissa part of each data in each data block in the second floating-point number format.

4. The model inference acceleration method of claim 3, wherein, The obtaining of the quantized weight data based on the shared exponent and the right shift operation mantissa part of each data to obtain the exponent part and the mantissa part of each data in the current data block in the second floating-point number format comprises the following steps: determining the exponent part of each data in the second floating-point number format as the shared exponent; determining the sign bit of each data based on the positive and negative of each data; determining the mantissa bit of each data based on the right shift operation mantissa part of each data; determining the mantissa part of each data in the second floating-point number format based on the sign bit and the mantissa bit of each data.

5. The model inference acceleration method of claim 4, wherein, The conversion of the quantized weight data from the second floating-point number format into the first floating-point number format to obtain the dequantized weight data comprises the following steps: determining the positive and negative of the dequantized weight data based on the sign bit of each data; determining the numerical value of the dequantized weight data based on the mantissa part of each data in the second floating-point number format and the power of two; the power index in the power of two is determined based on the shared exponent.

6. The model inference acceleration method of claim 1, wherein, The performing of the inference task of the target model based on the dequantized weight data comprises the following steps: obtaining input data; the data format of the input data is a first floating-point number format; inputting the dequantized weight data and the input data into the target model to obtain an inference result output by the target model.

7. A model inference acceleration device, comprising: The method comprises the following steps: an obtaining module, configured to obtain weight data of a target model; The data format of the weight data is a first floating-point number format; a quantization module, configured to convert the weight data from the first floating-point number format to a second floating-point number format to obtain quantized weight data; The second floating-point number format has lower precision than the first floating-point number format; a dequantization module, configured to load the quantized weight data and convert the quantized weight data from the second floating-point number format to the first floating-point number format to obtain dequantized weight data; a calculation module, configured to perform an inference task of the target model based on the dequantized weight data.

8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the model inference acceleration method in any one of claims 1 to 6. 9.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the model inference acceleration method in any one of claims 1 to 6.

10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the model inference acceleration method in any one of claims 1 to 6. The computer program is executed by the processor to implement the model inference acceleration method in any one of claims 1 to 6.