Inference acceleration method and device applied to edge device and electronic device

By dividing the weight matrix and activation matrix of the pre-trained model into sub-blocks and quantizing them, the problems of high storage occupancy and computational cost of large-scale models in edge devices are solved, and efficient deployment of the model on edge devices is achieved.

CN120633870AActive Publication Date: 2025-09-12HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511113382.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-09-12
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

When large-scale models are deployed on edge devices, they occupy a lot of storage and have high computing costs, making them unable to operate normally.

Method used

The weight matrix of the pre-trained model is divided into N weight sub-blocks, and the activation matrix is ​​divided into M activation sub-blocks for quantization. If at least two weight sub-blocks have the same quantization bit width, they are quantized as a whole to optimize storage and calculation.

Benefits of technology

It reduces the storage space occupied and computing cost of the model and improves the reasoning speed of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633870A_ABST
    Figure CN120633870A_ABST
Patent Text Reader

Abstract

The invention provides a reasoning acceleration method and device applied to edge equipment and electronic equipment. The method comprises the following steps: dividing a weight matrix of a pre-training model into N weight sub-blocks, dividing an activation matrix of the pre-training model into M activation sub-blocks, and quantifying the weight sub-blocks in the pre-training model and the activation sub-blocks corresponding to the weight sub-blocks to obtain a target model; and if the quantization bit widths of the at least two weight sub-blocks are the same, performing quantization processing on the at least two weight sub-blocks as a whole based on the weight value quantization super-parameters corresponding to the at least two weight sub-blocks and the activation value quantization super-parameters corresponding to the activation sub-blocks corresponding to the weight sub-blocks. Weight sub-blocks with the same quantization bit width and corresponding activation sub-blocks are integrally processed, when the sub-blocks are loaded, a memory access mode is converted from random hopping to sequential reading and writing, weight values and activation values are quantized at the same time, the reasoning speed of the model is increased, and the storage space occupied by the model is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of model training technology, and in particular to an inference acceleration method, device, and electronic device applied to edge devices. Background Art

[0002] In actual scenarios, some models, such as large language models, have been widely used in fields such as natural language processing and computer vision. However, these models are large in overall scale, occupy a lot of storage, and have high computing costs. This makes the overall inference speed of these models slow and cannot be properly deployed in some devices with limited performance, such as edge devices. Summary of the Invention

[0003] In view of this, the present application provides an inference acceleration method, device and electronic device applied to edge devices to reduce the storage space occupied by the model and improve the model inference speed.

[0004] The technical solutions provided in this application are as follows: According to an embodiment of the first aspect of the present application, a method for inference acceleration applied to an edge device is provided, the method comprising: Divide the weight matrix of the pre-trained model into N weight sub-blocks; where N is greater than 1, and each weight sub-block has a corresponding quantization bit width; The activation matrix of the pre-trained model is divided into M activation sub-blocks; where M is less than or equal to N; each activation sub-block is configured with a corresponding quantization bit width; any activation sub-block corresponds to at least one weight sub-block; The weight sub-block and the activation sub-block corresponding to the weight sub-block in the pre-trained model are quantized to obtain a target model; the target model is deployed on an edge device for accelerated inference of the task; wherein, when the weight sub-block and the activation sub-block corresponding to the weight sub-block are quantized, if the quantization bit widths of at least two weight sub-blocks are the same, then based on the weight value quantization hyperparameters corresponding to the at least two weight sub-blocks and the activation value quantization hyperparameters corresponding to the activation sub-blocks corresponding to each weight sub-block, the at least two weight sub-blocks are quantized as a whole.

[0005] Optionally, the quantization bit width corresponding to each weight sub-block is determined according to the following method: Determining a reference weight range corresponding to the weight sub-block according to the weight value in the weight sub-block; The quantization bit width corresponding to the reference weight range is determined as the quantization bit width corresponding to the weight sub-block.

[0006] Optionally, the N weight sub-blocks have the same size; dividing the activation matrix of the pre-trained model into M activation sub-blocks includes: Determining the number M by which the weight matrix is ​​divided in the input channel dimension according to the size of the weight sub-block; According to M, the activation matrix corresponding to the weight matrix is ​​divided into M activation sub-blocks in the input channel dimension.

[0007] Optionally, the weight value quantization hyperparameter corresponding to any weight sub-block is determined according to the following method: Determining a quantization interval of the weight value corresponding to the weight sub-block according to the weight value in the weight sub-block and the quantization bit width corresponding to the weight sub-block; The weight value quantization interval is used to determine the weight value quantization hyper-parameter corresponding to the weight sub-block.

[0008] Optionally, the activation value quantization hyperparameter corresponding to any activation sub-block is determined according to the following method: Determining an activation value quantization interval corresponding to the activation sub-block according to the activation value in the activation sub-block and the quantization bit width corresponding to the activation sub-block; The activation value quantization interval is used to determine the activation value quantization hyper-parameter corresponding to the activation sub-block.

[0009] Optionally, the pre-trained model includes at least one MHSA layer structure and at least one FFN layer structure, the MHSA layer structure and the FFN layer structure include multiple sorting operators, and the sorting operators are used to sort the weight sub-blocks based on the quantization bit width; wherein, the weight sub-blocks with the same quantization bit width are sorted together.

[0010] Optionally, the method further comprises: If it is found that the first sorting operator in any layer structure is used to sort the first weight matrix, and the second sorting operator is used to sort the calculation results of the first weight matrix, then the first sorting operator and the second sorting operator are merged into a whole to sort the first weight matrix and the calculation results of the first weight matrix.

[0011] According to an embodiment of the second aspect of the present application, there is provided an inference acceleration device applied to an edge device, the device comprising: A first division unit is used to divide the weight matrix of the pre-trained model into N weight sub-blocks; wherein N is greater than 1, and each weight sub-block has a corresponding quantization bit width; The second division unit is configured to divide the activation matrix of the pre-trained model into M activation sub-blocks; wherein M is less than or equal to N; each activation sub-block is configured with a corresponding quantization bit width; and any activation sub-block corresponds to at least one weight sub-block; A quantization processing unit is used to quantize the weight sub-block and the activation sub-block corresponding to the weight sub-block in the pre-trained model to obtain a target model; the target model is deployed on an edge device for accelerated inference of the task; wherein, when the weight sub-block and the activation sub-block corresponding to the weight sub-block are quantized, if the quantization bit widths of at least two weight sub-blocks are the same, then based on the weight value quantization hyperparameters corresponding to the at least two weight sub-blocks and the activation value quantization hyperparameters corresponding to the activation sub-blocks corresponding to each weight sub-block, the at least two weight sub-blocks are quantized as a whole.

[0012] Optionally, the quantization bit width corresponding to each weight sub-block is determined according to the following method: Determining a reference weight range corresponding to the weight sub-block according to the weight value in the weight sub-block; Determining the quantization bit width corresponding to the reference weight range as the quantization bit width corresponding to the weight sub-block; And / or, the N weight sub-blocks have the same size; dividing the activation matrix of the pre-trained model into M activation sub-blocks, including: Determining the number M by which the weight matrix is ​​divided in the input channel dimension according to the size of the weight sub-block; Dividing the activation matrix corresponding to the weight matrix into M activation sub-blocks in the input channel dimension according to M; And / or, the weight value quantization hyperparameter corresponding to any weight sub-block is determined according to the following method: Determining a quantization interval of the weight value corresponding to the weight sub-block according to the weight value in the weight sub-block and the quantization bit width corresponding to the weight sub-block; Determine the weight value quantization hyper-parameter corresponding to the weight sub-block by quantizing the weight value interval; And / or, the activation value quantization hyperparameter corresponding to any activation sub-block is determined according to the following method: Determining an activation value quantization interval corresponding to the activation sub-block according to the activation value in the activation sub-block and the quantization bit width corresponding to the activation sub-block; Determine the activation value quantization hyper-parameter corresponding to the activation sub-block by quantizing the activation value interval; And / or, the pre-trained model includes at least one MHSA layer structure and at least one FFN layer structure, the MHSA layer structure and the FFN layer structure include multiple sorting operators, and the sorting operators are used to sort the weight sub-blocks based on the quantization bit width; wherein, the weight sub-blocks with the same quantization bit width are sorted together; And / or, the quantization processing unit is further configured to: If it is found that the first sorting operator in any layer structure is used to sort the first weight matrix, and the second sorting operator is used to sort the calculation results of the first weight matrix, then the first sorting operator and the second sorting operator are merged into a whole to sort the first weight matrix and the calculation results of the first weight matrix.

[0013] According to an embodiment of the third aspect of the present application, an electronic device is provided, comprising: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the method described in the first aspect.

[0014] As can be seen from the above technical solution, the present application obtains a target model by dividing the weight matrix of the pre-trained model into N weight sub-blocks and the activation matrix of the pre-trained model into M activation sub-blocks to quantize the weight sub-blocks and the activation sub-blocks corresponding to the weight sub-blocks in the pre-trained model; if the quantization bit widths of at least two weight sub-blocks are the same, then based on the weight value quantization hyperparameters corresponding to the at least two weight sub-blocks and the activation value quantization hyperparameters corresponding to the activation sub-blocks corresponding to each weight sub-block, the at least two weight sub-blocks are quantized as a whole. Among them, the weight sub-blocks and the corresponding activation sub-blocks of the same quantization bit width are processed as a whole. When loading these sub-blocks, the memory access mode is transformed from random jumps to sequential reading and writing, and the weight values ​​and activation values ​​are quantized at the same time, which improves the inference speed of the model and reduces the storage space occupied by the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0016] Figure 1 A flowchart of the inference acceleration method applied to edge devices provided in an embodiment of the present application; Figure 2 A schematic diagram of weight sub-block division and sorting provided in an embodiment of the present application; Figure 3 Schematic diagram of the model structure and sorting operator provided in the embodiments of this application; Figure 4 A schematic diagram of the reasoning of an abnormal behavior monitoring system under the conventional method provided in an embodiment of the present application; Figure 5 This is a schematic diagram of the reasoning of the abnormal behavior monitoring system after quantization parameters provided in an embodiment of the present application; Figure 6 A structural diagram of an inference acceleration device applied to edge devices provided in an embodiment of the present application; Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0017] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application, and to make the above-mentioned purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application are further described in detail below with reference to the accompanying drawings.

[0018] In actual scenarios, some models, such as large language models, have been widely used in fields such as natural language processing and computer vision. However, these models are large in overall scale, occupy a lot of storage, and have high computing costs. This makes it impossible to deploy these models normally on some devices with limited performance, such as edge devices.

[0019] Based on this, the currently commonly used method is to compress the model before deployment. Compression techniques include sparsity, quantization, distillation, and low-rank decomposition technology.

[0020] The following mainly introduces the quantification method.

[0021] In related schemes, the quantization method for weight values ​​is usually per-channel quantization.

[0022] The weight matrix W can be expressed as Cout*Cin, where Cout represents the number of output channels and Cin represents the number of input channels. Perchannel quantization means that each output channel corresponds to a different quantization hyperparameter.

[0023] The quantization method for activation values ​​is usually per-tensor quantization.

[0024] Among them, the activation matrix X can be expressed as K*Cin, where K represents the number of samples participating in a single operation, Cin represents the number of input channels, and the entire activation matrix X usually has only one quantization hyperparameter.

[0025] However, existing quantization methods still suffer from high storage usage and computational costs. This is especially true when deploying pre-trained models to devices with limited performance, such as edge devices, where the performance of the edge devices themselves may not be able to support the deployment of pre-trained models.

[0026] Based on this, this application proposes an inference acceleration method applied to edge devices to reduce storage occupancy and reduce computing costs.

[0027] Please refer to Figure 1 , Figure 1 This is a flowchart of the inference acceleration method applied to edge devices provided in an embodiment of the present application.

[0028] As an embodiment, the method may be applied to a controller in an electronic device, such as a central processing unit (CPU), etc., and this application does not impose any limitation on this.

[0029] It should be noted that the pre-trained model may include multiple layers, and each layer may include one or more weight matrices, and each weight matrix has a corresponding activation matrix; the processing of the weight matrix and the activation matrix in this application may be the processing of any weight matrix in any layer and the activation matrix corresponding to the weight matrix.

[0030] like Figure 1 As shown, the method may include the following steps: Step 101: Divide the weight matrix of the pre-trained model into N weight sub-blocks.

[0031] Wherein, N is an integer greater than 1, and each weight sub-block is configured with a corresponding quantization bit width.

[0032] In this embodiment, the weight matrix of the pre-training model can be directly obtained after the pre-training is completed. The weight matrix can be a matrix in the format of Cout*Cin.

[0033] As an embodiment, the sizes of the N weight sub-blocks can be the same, and the process of dividing the weight matrix to obtain N weight sub-blocks can be based on the dimension of the output channel, that is, dividing from the Cout dimension without breaking up the arrangement of Cin, or dividing according to a specified size; similarly, the sizes of the N weight sub-blocks can also be different, and this application does not limit this.

[0034] In this embodiment, each weight sub-block is assigned a corresponding quantization bit width. Quantization bit width refers to the number of bits used to represent a value. For example, a quantization bit width of 8 bits means that each weight value in the weight sub-block is represented by an 8-bit binary number. The lower the quantization bit width, the higher the data compression rate, but the greater the risk of precision loss.

[0035] As an embodiment, the quantization bit width corresponding to each weight sub-block can be determined according to the following method: Determine a reference weight range corresponding to the weight sub-block according to the weight value in the weight sub-block; The quantization bit width corresponding to the reference weight range is determined as the quantization bit width corresponding to the weight sub-block.

[0036] In this embodiment, the correspondence between the weight range and the quantization bit width can be pre-set. Considering that the weight sub-block with a higher weight value is more important and the accuracy it needs to maintain is also higher, a high quantization bit width can be set for a high weight range so that the weight sub-block with a higher weight value can be configured with a larger quantization bit width. For example, if the weight range is 0 to 0.6, the corresponding quantization bit width is set to 4 bits, and if the weight range is 0.6 to 1, the corresponding quantization bit width is set to 8 bits. The weight range and the number of quantization bit widths here can be adjusted according to actual needs, and this application does not impose any restrictions on this.

[0037] In addition, the quantization bit width can be selected according to the size of the operation core supported in the edge device. For example, if the edge device includes three types of operation cores: 8*8, 8*4, and 4*4 (for example, the 8*8 operation core indicates that the operation core is used to operate on two matrices with a quantization bit width of 8 bits), then the weight sub-block can be configured with two quantization bit widths of 8 bits or 4 bits.

[0038] The specific method for determining the reference weight range corresponding to the weight sub-block according to the weight value in the weight sub-block may be: Determine the maximum value of the absolute values ​​of the weight values ​​in the weight sub-block, and determine the reference weight range to which the maximum value belongs; or determine the average value of the absolute values ​​of the weight values ​​in the weight sub-block, and determine the reference weight range to which the average value belongs. This application does not impose any restrictions on this.

[0039] At this point, the description of step 101 ends, and step 102 is executed next.

[0040] Step 102: Divide the activation matrix of the pre-trained model into M activation sub-blocks.

[0041] Among them, M is less than or equal to N; each activation sub-block is configured with a corresponding quantization bit width; any activation sub-block corresponds to at least one weight sub-block.

[0042] In this embodiment, the activation matrix refers to the activation matrix corresponding to the weight matrix in step 101, and the activation matrix is ​​obtained by inputting sample data into the training model.

[0043] As an embodiment, when the N weight sub-blocks have the same size, a specific method of dividing the activation matrix of the pre-trained model into M activation sub-blocks may include: According to the size of the weight sub-block, the number M of divisions of the weight matrix in the input channel dimension is determined; according to M, the activation matrix corresponding to the weight matrix is ​​divided into M activation sub-blocks in the input channel dimension.

[0044] Specifically, the number M by which the weight matrix is ​​divided in the input channel dimension can be determined based on the size of the weight sub-block. For example, if the size of the weight matrix is ​​30*20, that is, Cout=30, Cin=20, and the size of the weight sub-block is 3*2, it means that the weight matrix is ​​divided into 100 weight sub-blocks, and the number M by which the weight matrix is ​​divided in the input channel dimension is 20 / 2=10.

[0045] After determining the number of divisions M along the input channel dimension, the activation matrix can be further divided into M activation sub-blocks along the input channel dimension based on M. For example, if the size of the activation matrix is ​​2*20, the activation matrix can be divided into M blocks along the input channel dimension, that is, the activation matrix can be divided into 10 2*2 activation sub-blocks. Here, each 2*2 activation sub-block corresponds to 10 3*2 weight sub-blocks.

[0046] In this embodiment, each activation sub-block is also configured with a corresponding quantization bit width, and the quantization bit width corresponding to each activation sub-block is the same or different. The quantization bit width configured for each activation sub-block is the same or different from the quantization bit width configured for the weight sub-block corresponding to the activation sub-block. This application does not impose any restrictions on this.

[0047] At this point, the description of step 102 ends, and step 103 is executed next.

[0048] Step 103: quantize the weight sub-block in the pre-trained model and the activation sub-block corresponding to the weight sub-block to obtain a target model.

[0049] Specifically, when quantizing the weight sub-block and the activation sub-block corresponding to the weight sub-block, if the quantization bit width of at least two weight sub-blocks is the same, the at least two weight sub-blocks are quantized as a whole based on the weight value quantization hyperparameters corresponding to the at least two weight sub-blocks and the activation value quantization hyperparameters corresponding to the activation sub-blocks corresponding to each weight sub-block.

[0050] In this embodiment, the quantization bit width of the weight sub-block can be used as a reference to quantize the weight sub-blocks with the same quantization bit width and the activation sub-blocks corresponding to the weight sub-blocks as a whole to improve the quantization processing efficiency.

[0051] The weight quantization hyperparameters of each weight sub-block and the activation quantization hyperparameters of each activation sub-block are determined according to the following method: According to the weight value in the weight sub-block and the quantization bit width corresponding to the weight sub-block, the weight value quantization interval corresponding to the weight sub-block is determined; and the weight value quantization hyperparameter corresponding to the weight sub-block is determined by the weight value quantization interval.

[0052] According to the activation value in the activation sub-block and the quantization bit width corresponding to the activation sub-block, the activation value quantization interval corresponding to the activation sub-block is determined; and the activation value quantization hyperparameter corresponding to the activation sub-block is determined by the activation value quantization interval.

[0053] Specifically, in this embodiment, the weight value quantization hyperparameter corresponding to the weight sub-block and the activation value quantization hyperparameter corresponding to the activation sub-block can be determined according to the following formula:

[0054] Among them, α is the absolute maximum value of the statistical distribution (for the weight sub-block, it is the maximum absolute value of the weight value in the weight sub-block; for the activation sub-block, it is the maximum absolute value of the activation value in the activation sub-block), L is the quantization bit width, and s represents the quantization interval, which is the quantization hyperparameter.

[0055] It should be noted that, in this embodiment, the quantization hyperparameters are all obtained by static quantization, that is, the weight value quantization hyperparameters corresponding to each weight sub-block and the activation value quantization hyperparameters corresponding to each activation sub-block are determined by pre-acquired sample data, and then the above hyperparameters are stored for use in model inference, so as to improve the operating efficiency in actual deployment, accelerate inference, and reduce the consumption of hardware resources; rather than determining the weight value quantization hyperparameters corresponding to each weight sub-block and the activation value quantization hyperparameters corresponding to each activation sub-block according to the data to be tested during the actual inference process of the model.

[0056] As an embodiment, the specific method of quantizing the weight sub-blocks with the same quantization bit width and the activation sub-blocks corresponding to the weight sub-blocks as a whole can be: sorting the weight sub-blocks according to the quantization bit width, sorting the weight sub-blocks with the same quantization bit width together and storing them continuously.

[0057] In this embodiment, the pre-trained model includes at least one MHSA (Multi-Head Self-Attention) layer structure and at least one FFN (Feed-Forward Network) layer structure. The MHSA layer structure and the FFN layer structure include multiple sorting operators, which are used to sort weight sub-blocks based on quantization bit width; weight sub-blocks with the same quantization bit width are sorted together. The sorting operators can be used to implement the above-mentioned process of sorting each weight sub-block according to quantization bit width, sorting weight sub-blocks with the same quantization bit width together, and storing them continuously.

[0058] As an embodiment, if it is found that the first sorting operator in any layer structure is used to sort the first weight matrix, and the second sorting operator is used to sort the calculation results of the first weight matrix, the first sorting operator and the second sorting operator are merged into a whole to sort the first weight matrix and the calculation results of the first weight matrix.

[0059] The specific process of the merge sort operator will be described below through specific embodiments and will not be repeated here.

[0060] In this embodiment, the specific quantization method can be implemented according to the following formula:

[0061] Among them, x q Refers to the quantized value, x refers to the value before quantization (for the weight sub-block, x is the weight value; for the activation sub-block, x is the activation value), s refers to the quantization hyperparameter corresponding to the value before quantization (for the weight sub-block, s is the weight value quantization hyperparameter, for the activation sub-block, s is the activation value quantization hyperparameter), L refers to the quantization bit width, Express The result is rounded to the nearest integer. clip() represents a truncation function. For example, clip(a, b, c) means limiting a to the specified minimum value b and maximum value c. If a is less than the minimum value b, it will be truncated to the minimum value b. If a is greater than the maximum value c, it will be truncated to the maximum value c.

[0062] It should be noted that in the process of quantizing the weight sub-blocks with the same quantization bit width and the activation sub-blocks corresponding to the weight sub-blocks as a whole, although the quantization bit width of each weight sub-block is the same, the quantization bit width of the corresponding activation sub-block is not necessarily the same as the quantization bit width of the weight sub-block, and even the quantization bit width of each activation sub-block is not necessarily the same. During the processing, the activation sub-blocks corresponding to each weight sub-block can be quantized based on the highest quantization bit width in the activation sub-block.

[0063] For example, the weight sub-blocks with a quantization bit width of 4 bits and the corresponding activation sub-blocks are calculated as a whole, but the activation sub-blocks corresponding to some 4-bit weight sub-blocks have a quantization bit width of 8 bits, and the activation sub-blocks corresponding to some 4-bit weight sub-blocks have a quantization bit width of 4 bits. In this case, an 8*4 operation core can be used to calculate the weight sub-blocks and the corresponding activation sub-blocks, that is, the activation sub-blocks are based on the highest quantization bit width of 8 bits, and the weight sub-blocks use a quantization bit width of 4 bits.

[0064] It can be seen that the above method treats the weight sub-blocks with the same quantization bit width as a whole, stores the weight sub-blocks with the same quantization bit width physically and continuously, and stores the activation sub-blocks corresponding to the weight sub-blocks with the same quantization bit width physically and continuously. When loading these sub-blocks, the memory access mode is transformed from random jumps to sequential reading and writing, which greatly improves the loading rate.

[0065] After completing the quantization of the pre-trained model and obtaining the target model, the target model can be deployed on the edge device to accelerate the inference of the task through the target model.

[0066] It should be noted that when the edge device receives the sample to be detected, it can quantize the sample to be detected according to the activation value quantization hyperparameter of the activation matrix corresponding to the input layer of the pre-trained model, and obtain the detection result based on the quantized sample to be detected and the activation value quantization hyperparameters and weight value quantization hyperparameters, and dequantize the detection result according to the activation value quantization interval corresponding to the input layer of the pre-trained model to obtain the target prediction result.

[0067] This concludes the description of step 103.

[0068] This concludes Figure 1 Description of the inference acceleration method applied to edge devices.

[0069] The present application obtains a target model by dividing the weight matrix of the pre-trained model into N weight sub-blocks and the activation matrix of the pre-trained model into M activation sub-blocks to quantize the weight sub-blocks and the activation sub-blocks corresponding to the weight sub-blocks in the pre-trained model; if the quantization bit width of at least two weight sub-blocks is the same, then based on the weight value quantization hyperparameters corresponding to the at least two weight sub-blocks and the activation value quantization hyperparameters corresponding to the activation sub-blocks corresponding to each weight sub-block, the at least two weight sub-blocks are quantized as a whole. Different weight sub-blocks are set with different weight value quantization hyperparameters; different activation sub-blocks are set with different activation value quantization hyperparameters, and sub-blocks with the same quantization bit width are processed as a whole, while the weights and activations are quantized, thereby reducing the storage space occupied by the model and the computational cost.

[0070] Below through Figures 2 to 5 The inference acceleration method applied to edge devices proposed in this application is described in detail.

[0071] Please refer to Figure 2 , Figure 2 A schematic diagram of weight sub-block division and sorting provided in an embodiment of the present application.

[0072] like Figure 2As shown, X is a 2*6 activation matrix and W is a 4*6 weight matrix, where each 1*1 block in the weight matrix represents a weight sub-block, and different colors of the weight sub-blocks represent different quantization bit widths. Each column in the activation matrix, that is, a 2*1 block, is an activation sub-block, corresponding to all the weight sub-blocks included in the corresponding column in the weight matrix.

[0073] After completing the division of weight sub-blocks and activation sub-blocks, the weight sub-blocks can be sorted according to the quantization bit width through the sorting operator, and the weight sub-blocks with the same quantization bit width can be arranged together to obtain three groups of weight sub-blocks W1, W2 and W3. The activation sub-blocks corresponding to these three groups of weight sub-blocks are X1, X2 and X3 respectively.

[0074] When quantizing the activation matrix and the weight matrix, the weight sub-blocks with the same quantization bit width can be quantized as a whole according to the quantization bit width configured for the weight sub-blocks W1, W2 and W3.

[0075] Specifically, the inference quantization process of a large model can be expressed by the following formula:

[0076] in The weight value quantization hyperparameter representing the weight matrix; The activation value quantization hyperparameter representing the activation matrix, Represents the weight value in the weight matrix; Represents the activation value in the activation matrix.

[0077] In this embodiment, due to the introduction of hybrid quantization, the weight matrix is ​​split into multiple weight sub-blocks, and the activation matrix is ​​split into multiple activation sub-blocks. Each weight sub-block can be expressed with different expression precision, so the above formula can be transformed into:

[0078] in, Represents the quantized hyperparameter of the weight value corresponding to the i-th weight sub-block; Represents the activation value quantization hyperparameter of the activation sub-block corresponding to the i-th weight sub-block, Represents the weight value in the i-th weight sub-block; Represents the activation value in the activation sub-block corresponding to the i-th weight sub-block.

[0079] Since the mixed blocks are different and the precision cannot be aligned, it is often necessary to move blocks of the same precision together, that is, the weight sub-blocks of the same quantization bit width can be quantized as a whole to speed up the processing speed and improve the processing efficiency.

[0080] This concludes Figure 2Description.

[0081] Please refer to Figure 3 , Figure 3 Schematic diagram of the model structure and sorting operator provided in the embodiments of this application.

[0082] like Figure 3 As shown, the pre-trained model includes at least one MHSA layer structure and at least one FFN layer structure. Figure 3 The left side shows the schematic diagram of the MHSA structure, and the right side shows the schematic diagram of the FFN structure. Both the MHSA and FFN layers include multiple sorting operators (stars in the figure represent sorting operators), which are used to sort weight sub-blocks based on the quantization bit width. The specific structure of the MHSA and FFN layers is a common term in the field of deep learning, so the specific components of each structure will not be detailed here.

[0083] In the FFN structure, the lower sorting operator is used to sort the weight sub-blocks in the weight matrix at the up and gate locations (denoted as the first weight matrix), while the upper sorting operator is used to sort the calculation results of the first weight matrix. At this point, the two sorting operators in the FFN structure can be merged (i.e., the hollow five-pointed star and the solid five-pointed star in the FFN structure in the figure are merged) to further improve hardware efficiency. The principle of merging sorting operators in FFN is as follows: O

[0084]

[0085]

[0086]

[0087] Among them, W down Refers to the weight matrix of the down module in FFN, W up Refers to the weight matrix of the up module in FFN; I refers to the input feature vector (output from the MHSA layer structure); f refers to the activation function (the SiLU function in the figure is Sigmoid-weighted Linear Unit); y refers to the original output of FFN (unsorted); sort() is the sorting operator; O represents the final output after sorting; and Respectively represent the W up The sorted weights and the W down The weight after sorting.

[0088] It can be seen that through the above formula, the sorting operation is advanced to the weight matrix preprocessing stage, so that only the pre-sorted weights need to be used at runtime. and , that is, the final output O can be calculated and sorted without explicit calculation, but by first as well as Sorting to get the sorted weight and , and then perform FFN calculation to obtain it directly, that is, sorting the entire output is equivalent to sorting the weights and then performing FFN structure calculation to obtain the result.

[0089] This concludes Figure 3 Description.

[0090] In this embodiment, the actual application scenario can be any intelligent scenario, such as video surveillance (abnormal behavior detection), customer service dialogue, machine translation, automatic video generation and other task scenarios. As long as it is necessary to deploy a pre-trained model in an edge device, the inference acceleration method proposed in this application for edge devices can be adopted.

[0091] Please refer to Figure 4 , Figure 4 This is a schematic diagram of the reasoning of the abnormal behavior monitoring system under the conventional method provided in the embodiment of the present application.

[0092] like Figure 4 As shown, when the abnormal behavior monitoring system in the related art receives the sample to be detected, it directly performs inference calculation based on the weight of the large model and outputs the inference result. This method has a large overall computational load and requires a lot of resources.

[0093] Please refer to Figure 5 , Figure 5 This is a schematic diagram of the reasoning of the abnormal behavior monitoring system after quantization parameters provided in an embodiment of the present application.

[0094] like Figure 5 As shown, the abnormal behavior monitoring system after quantization of parameters in this application pre-stores the quantized model parameters and the corresponding quantization hyperparameters. When receiving the sample to be detected, the sample to be detected can be directly quantized according to the stored quantization hyperparameters, and the detection results can be further output. The detection results are then dequantized to obtain the final inference results, which greatly reduces resource usage.

[0095] Please refer to Figure 6 , Figure 6 This is a structural diagram of an inference acceleration device applied to edge devices proposed in an embodiment of the present application. Figure 6 As shown, the apparatus may include a first division unit 601, a second division unit 602, and a quantization processing unit 603. Specifically, the apparatus includes: A first dividing unit 601 is configured to divide the weight matrix of the pre-trained model into N weight sub-blocks; wherein N is greater than 1, and each weight sub-block has a corresponding quantization bit width; The second division unit 602 is configured to divide the activation matrix of the pre-trained model into M activation sub-blocks; wherein M is less than or equal to N; each activation sub-block is configured with a corresponding quantization bit width; and any activation sub-block corresponds to at least one weight sub-block; The quantization processing unit 603 is used to quantize the weight sub-block and the activation sub-block corresponding to the weight sub-block in the pre-trained model to obtain a target model; the target model is deployed on the edge device for accelerated inference of the task; wherein, when the weight sub-block and the activation sub-block corresponding to the weight sub-block are quantized, if the quantization bit widths of at least two weight sub-blocks are the same, then based on the weight value quantization hyperparameters corresponding to the at least two weight sub-blocks and the activation value quantization hyperparameters corresponding to the activation sub-blocks corresponding to each weight sub-block, the at least two weight sub-blocks are quantized as a whole.

[0096] Optionally, the quantization bit width corresponding to each weight sub-block is determined according to the following method: Determine a reference weight range corresponding to the weight sub-block according to the weight value in the weight sub-block; Determine the quantization bit width corresponding to the reference weight range as the quantization bit width corresponding to the weight sub-block; And / or, the N weight sub-blocks have the same size; the activation matrix of the pre-trained model is divided into M activation sub-blocks, including: According to the size of the weight sub-block, determine the number M by which the weight matrix is ​​divided in the input channel dimension; According to M, the activation matrix corresponding to the weight matrix is ​​divided into M activation sub-blocks in the input channel dimension; And / or, the weight value quantization hyperparameter corresponding to any weight sub-block is determined according to the following method: Determining a quantization interval of the weight value corresponding to the weight sub-block according to the weight value in the weight sub-block and the quantization bit width corresponding to the weight sub-block; The weight value quantization interval is used to determine the weight value quantization hyper parameter corresponding to the weight sub-block; And / or, the activation value quantization hyperparameter corresponding to any activation sub-block is determined according to the following method: Determining an activation value quantization interval corresponding to the activation sub-block according to the activation value in the activation sub-block and the quantization bit width corresponding to the activation sub-block; The activation value quantization interval is used to determine the activation value quantization hyperparameter corresponding to the activation sub-block; And / or, the pre-trained model includes at least one MHSA layer structure and at least one FFN layer structure, the MHSA layer structure and the FFN layer structure include multiple sorting operators, and the sorting operators are used to sort the weight sub-blocks based on the quantization bit width; wherein, the weight sub-blocks with the same quantization bit width are sorted together; And / or, the quantization processing unit 603 is further configured to: If it is found that the first sorting operator in any layer structure is used to sort the first weight matrix, and the second sorting operator is used to sort the calculation results of the first weight matrix, then the first sorting operator and the second sorting operator are merged into a whole to sort the first weight matrix and the calculation results of the first weight matrix.

[0097] So far, completed Figure 6 Description of the inference acceleration device applied to edge devices.

[0098] The present application also provides Figure 6 The hardware structure of the device is described in the following figure. Figure 7 The structure of the electronic device shown. Figure 7 , Figure 7 This is a structural diagram of an electronic device provided in an embodiment of the present application. Figure 7 As shown, the hardware structure may include: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the method disclosed in the above example of this application.

[0099] Based on the same application concept as the above method, an embodiment of the present application also provides a machine-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed by a processor, the method disclosed in the above example of the present application can be implemented.

[0100] Exemplarily, the machine-readable storage medium may be any electronic, magnetic, optical, or other physical storage device that may contain or store information, such as executable instructions, data, and the like. For example, the machine-readable storage medium may be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, a storage drive (such as a hard disk drive), a solid-state drive, any type of storage disk (such as a CD, DVD, etc.), or similar storage media, or a combination thereof.

[0101] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A method for accelerating inference on edge devices, characterized in that: The method includes: Divide the weight matrix of the pre-trained model into N weight sub-blocks; where N is greater than 1, and each weight sub-block has a corresponding quantization bit width; The activation matrix of the pre-trained model is divided into M activation sub-blocks; where M is less than or equal to N; each activation sub-block is configured with a corresponding quantization bit width; any activation sub-block corresponds to at least one weight sub-block; The weight sub-block and the activation sub-block corresponding to the weight sub-block in the pre-trained model are quantized to obtain a target model; the target model is deployed on an edge device for accelerated inference of the task; wherein, when the weight sub-block and the activation sub-block corresponding to the weight sub-block are quantized, if the quantization bit widths of at least two weight sub-blocks are the same, then based on the weight value quantization hyperparameters corresponding to the at least two weight sub-blocks and the activation value quantization hyperparameters corresponding to the activation sub-blocks corresponding to each weight sub-block, the at least two weight sub-blocks are quantized as a whole.

2. The method according to claim 1, characterized in that The quantization bit width corresponding to each weight sub-block is determined according to the following method: Determining a reference weight range corresponding to the weight sub-block according to the weight value in the weight sub-block; The quantization bit width corresponding to the reference weight range is determined as the quantization bit width corresponding to the weight sub-block.

3. The method according to claim 1, characterized in that The N weight sub-blocks have the same size; the activation matrix of the pre-trained model is divided into M activation sub-blocks, including: Determining the number M by which the weight matrix is ​​divided in the input channel dimension according to the size of the weight sub-block; According to M, the activation matrix corresponding to the weight matrix is ​​divided into M activation sub-blocks in the input channel dimension.

4. The method according to claim 1, wherein The weight quantization hyperparameter corresponding to any weight sub-block is determined according to the following method: Determining a quantization interval of the weight value corresponding to the weight sub-block according to the weight value in the weight sub-block and the quantization bit width corresponding to the weight sub-block; The weight value quantization interval is used to determine the weight value quantization hyper-parameter corresponding to the weight sub-block.

5. The method according to claim 1, wherein The activation value quantization hyperparameter corresponding to any activation sub-block is determined according to the following method: Determining an activation value quantization interval corresponding to the activation sub-block according to the activation value in the activation sub-block and the quantization bit width corresponding to the activation sub-block; The activation value quantization interval is used to determine the activation value quantization hyper-parameter corresponding to the activation sub-block.

6. The method according to claim 1, characterized in that The pre-trained model includes at least one MHSA layer structure and at least one FFN layer structure, the MHSA layer structure and the FFN layer structure include multiple sorting operators, and the sorting operators are used to sort the weight sub-blocks based on the quantization bit width; wherein, the weight sub-blocks with the same quantization bit width are sorted together.

7. The method according to claim 6, characterized in that The method further comprises: If it is found that the first sorting operator in any layer structure is used to sort the first weight matrix, and the second sorting operator is used to sort the calculation results of the first weight matrix, then the first sorting operator and the second sorting operator are merged into a whole to sort the first weight matrix and the calculation results of the first weight matrix.

8. An inference acceleration device applied to an edge device, characterized in that: The device includes: A first division unit is used to divide the weight matrix of the pre-trained model into N weight sub-blocks; wherein N is greater than 1, and each weight sub-block has a corresponding quantization bit width; The second division unit is configured to divide the activation matrix of the pre-trained model into M activation sub-blocks; wherein M is less than or equal to N; each activation sub-block is configured with a corresponding quantization bit width; and any activation sub-block corresponds to at least one weight sub-block; A quantization processing unit is used to quantize the weight sub-block and the activation sub-block corresponding to the weight sub-block in the pre-trained model to obtain a target model; the target model is deployed on an edge device for accelerated inference of the task; wherein, when the weight sub-block and the activation sub-block corresponding to the weight sub-block are quantized, if the quantization bit widths of at least two weight sub-blocks are the same, then based on the weight value quantization hyperparameters corresponding to the at least two weight sub-blocks and the activation value quantization hyperparameters corresponding to the activation sub-blocks corresponding to each weight sub-block, the at least two weight sub-blocks are quantized as a whole.

9. The device according to claim 8, characterized in that The quantization bit width corresponding to each weight sub-block is determined according to the following method: Determining a reference weight range corresponding to the weight sub-block according to the weight value in the weight sub-block; Determining the quantization bit width corresponding to the reference weight range as the quantization bit width corresponding to the weight sub-block; And / or, the N weight sub-blocks have the same size; Divide the activation matrix of the pre-trained model into M activation sub-blocks, including: Determining the number M by which the weight matrix is ​​divided in the input channel dimension according to the size of the weight sub-block; Dividing the activation matrix corresponding to the weight matrix into M activation sub-blocks in the input channel dimension according to M; And / or, the weight value quantization hyperparameter corresponding to any weight sub-block is determined according to the following method: Determining a quantization interval of the weight value corresponding to the weight sub-block according to the weight value in the weight sub-block and the quantization bit width corresponding to the weight sub-block; Determine the weight value quantization hyper-parameter corresponding to the weight sub-block by quantizing the weight value interval; And / or, the activation value quantization hyperparameter corresponding to any activation sub-block is determined according to the following method: Determining an activation value quantization interval corresponding to the activation sub-block according to the activation value in the activation sub-block and the quantization bit width corresponding to the activation sub-block; Determine the activation value quantization hyper-parameter corresponding to the activation sub-block by quantizing the activation value interval; And / or, the pre-trained model includes at least one MHSA layer structure and at least one FFN layer structure, the MHSA layer structure and the FFN layer structure include multiple sorting operators, and the sorting operators are used to sort the weight sub-blocks based on the quantization bit width; wherein, the weight sub-blocks with the same quantization bit width are sorted together; And / or, the quantization processing unit is further configured to: If it is found that the first sorting operator in any layer structure is used to sort the first weight matrix, and the second sorting operator is used to sort the calculation results of the first weight matrix, then the first sorting operator and the second sorting operator are merged into a whole to sort the first weight matrix and the calculation results of the first weight matrix.

10. An electronic device, characterized in that: include: a processor and a machine-readable storage medium storing machine-executable instructions capable of being executed by the processor; The processor is configured to execute machine-executable instructions to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data compression transmitting and decompression method and apparatus

    CN103634273A

  • Convolutional neural network operation method and related equipment

    CN114418057A

  • Matrix vector multiplier implementation method and device supporting mixed bit quantization

    CN118035628A

  • Quantization method and reasoning method and device of large language model, equipment and medium

    CN118036755A

  • Hybrid precision weight processing method, apparatus and device, and computer program product

    CN118378005A