Deep learning model quantification method, electronic equipment and storage medium

By grouping normalizing and segmenting the weight data of deep learning models, the problem of difficult to balance model accuracy and efficiency in embedded AI is solved, which improves the quantization accuracy and avoids fixed-point model inference losses.

CN120218152APending Publication Date: 2025-06-27ZHEJIANG SUNNY INTELLIGENT OPTICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311754683.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to take into account both model accuracy and efficiency in the embedded AI deployment of deep learning models. The accuracy is high but the efficiency is low when quantizing high bit width, and the accuracy is low but the efficiency is high when quantizing low bit width.

Method used

By grouping normalizing and segmenting the weight data of the deep learning model, the quantization parameters of each segment are determined and the model is quantized based on these parameters.

Benefits of technology

The quantization accuracy is improved, the local optimal problem of quantization parameters caused by the missing segmented quantization boundary is avoided, and the fixed-point model inference loss caused by excessive differences in the range of weight data between adjacent layers is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218152A_ABST
    Figure CN120218152A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a deep learning model quantification method, electronic equipment and a storage medium, a deep learning model comprises a plurality of layers, and a plurality of channels in each layer are divided into a plurality of groups; the method comprises the following steps: for each layer, carrying out grouping normalization processing on weight data of the layer, segmenting the weight data of each group, and quantifying the weight data of each segment to determine a quantization parameter of each segment; and performing quantization processing on the deep learning model according to the quantization parameter of each segment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of deep learning model quantization, and in particular, to a deep learning model quantization method, an electronic device, and a storage medium. Background Art

[0002] Currently, with the rapid development of artificial intelligence, more and more embedded-based applications have received extensive attention in the industrial and academic fields.

[0003] However, the computational complexity of deep learning models is the main bottleneck for the deployment of embedded AI (Artificial Intelligence). Although various model compression or model quantization algorithms have been continuously proposed, it is still difficult to solve the dual problems of accuracy loss and low efficiency. When choosing high-bitwidth (such as 32bit / 16bit) quantization, the accuracy of the quantized model is relatively high, but the model scale is relatively large and the execution efficiency is relatively low; while when choosing low-bitwidth (such as 8bit / 4bit) quantization, the accuracy of the quantized model is relatively low, but the model scale is relatively small and the execution efficiency is relatively high. Therefore, it is often difficult to balance the accuracy and efficiency of the quantized model in related technologies. Summary of the Invention

[0004] The technical solution provided by this application at least partially solves the above technical problems.

[0005] According to a first aspect of this application, a deep learning model quantization method is provided. The deep learning model includes multiple layers, and multiple channels in each layer are divided into multiple groups. The method includes: for each layer, performing grouped normalization processing on the weight data of the layer, segmenting the weight data of each group, and determining the quantization parameters of each segment by quantizing the weight data of each segment; and performing quantization processing on the deep learning model according to the quantization parameters of each segment.

[0006] In some embodiments, the deep learning model includes multiple blocks, and each block includes one layer or multiple adjacent layers.

[0007] In some embodiments, the above method further includes: before performing grouped normalization processing on the weight data of the layer, performing block normalization processing on the weight data of each block.

[0008] In some embodiments, the above method further includes: partitioning the deep learning model according to the weight data distribution of adjacent layers in the deep learning model, taking multiple adjacent layers whose weight data distribution satisfies a first preset condition as one block, where the first preset condition includes at least one of the following: the difference in the data ranges of the weight data of adjacent layers is less than a predetermined first threshold, the difference in the data distribution centers of adjacent layers is less than a predetermined second threshold; wherein the predetermined second threshold is not less than the predetermined first threshold.

[0009] In some embodiments, the above method further includes: grouping each layer according to the weight data distribution of each channel within the same layer, taking multiple channels within the same layer whose weight data distribution satisfies a second preset condition as one group, where the second preset condition includes at least one of the following: the difference between the data ranges of the weight data of different channels is less than a predetermined third threshold, the difference between the data distribution centers of different channels is less than a predetermined fourth threshold; wherein the predetermined fourth threshold is not less than the predetermined third threshold.

[0010] In some embodiments, performing block normalization processing on the weight data of each block includes: removing outliers from the weight data of each block to obtain the maximum value and the minimum value of the weight data of the block; and performing block normalization processing by dividing the weight data of the block by the difference between the maximum value and the minimum value.

[0011] In some embodiments, performing group normalization processing on the weight data of a layer includes: obtaining the data distribution of the average group according to the weight data distribution of each group within the layer; and for each group in the layer, determining a scale factor and a translation factor according to the data distribution of the average group, and mapping the weight data of the group based on the scale factor and the translation factor to perform group normalization processing; wherein, the data distribution of the average group includes the data range and the data distribution center of the average group, the scale factor is used to represent the ratio of the data range of the weight data of each group to the data range of the average group, and the translation factor is used to represent the difference between the data distribution center of the weight data of each group and the data distribution center of the average group.

[0012] In some embodiments, the above method further includes: performing an inverse transformation process on the group normalization process of the weight data of a layer before inputting the output value of each layer to the next layer.

[0013] In some embodiments, the weight data of each group is segmented, and by quantifying the weight data of each segment, quantization parameters for each segment are determined, including: for each group, clustering the weight data of the group by a clustering algorithm, obtaining multiple segments according to the clustering result, and setting an overlapping extended region between two adjacent segments; determining the quantization parameters of the segment according to the data range of the weight data of each segment and the extended region between two segments adjacent to the segment, wherein the quantization parameters of the segment include a fixed-point scaling factor, and the fixed-point scaling factor is used to convert floating-point weight data into fixed-point weight data.

[0014] In some embodiments, the above method further includes: obtaining a first quantization model according to the quantization process performed on the deep learning model; performing a second quantization process on the deep learning model to obtain a second quantization model; comparing the accuracy of each first subgraph corresponding to each block of the first quantization model with the second subgraph corresponding to the block in the second quantization model respectively, and taking the subgraph with higher accuracy between the first subgraph and the second subgraph as the target subgraph corresponding to the block; and obtaining a target quantization model according to the target subgraph corresponding to each block.

[0015] In some embodiments, the above method further includes: in response to detecting that the sample data meets a preset calibration condition, online quantization calibration is performed on the activation values of the deep learning model by updating the activation quantization parameters, wherein the activation quantization parameters include a scaling factor of the data range of the activation values, and the preset calibration condition includes: the difference between the activation value of the sample data and the historical average activation value is less than a preset threshold range.

[0016] In some embodiments, updating the activation quantization parameters includes: recalculating the scaling factor according to the maximum activation value and the minimum activation value of the sample data, or according to the historical average maximum activation value and the historical average minimum activation value.

[0017] In some embodiments, the above method further includes: for the deep learning model on the embedded side, direct memory access and cache memory are used for caching and online quantization calibration of the sample data.

[0018] In some embodiments, the deep learning network model is used to process one of the following sample data: image data, audio-visual data, text data.

[0019] According to a second aspect of the present application, an electronic device is further provided, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method of the first aspect of the present application.

[0020] According to a third aspect of the present application, there is also provided a computer-readable storage medium, on which machine-executable instructions are stored, and when the machine-executable instructions are executed, the machine is caused to execute the method according to the first aspect of the present application.

[0021] According to the deep learning model quantization method provided by the embodiments of the present application, through the group mapping based on the block data and the segmented quantization method based on the extended data, the quantization accuracy is improved; the problem of local optimality of quantization parameters caused by the lack of segmented quantization boundaries is avoided. In addition, according to some embodiments of the present application, the problem of inference loss of the fixed-point model caused by the too large difference in the weight data range between adjacent layers is also solved by proposing a block scheme for network layers.

[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more obvious. The drawings are used to better understand the solution and do not constitute a limitation to the present application. In the drawings:

[0024] Figure 1 is an exemplary system architecture diagram to which an embodiment of the present application can be applied;

[0025] Figure 2 is a schematic diagram of a deep learning model to be quantized according to an embodiment of the present application;

[0026] Figure 3 is a flowchart of a deep learning model quantization method according to an embodiment of the present application;

[0027] Figure 4 is a schematic diagram of performing grouped normalization processing in a deep learning model quantization method according to an embodiment of the present application;

[0028] Figure 5 is a schematic diagram of quantizing segmented weight data in a deep learning model quantization method according to an embodiment of the present application;

[0029] Figure 6A is a schematic diagram of performing optimal subgraph screening in a deep learning model quantization method according to an embodiment of the present application;

[0030] Figure 6B is a schematic diagram of a target quantization model according to an embodiment of the present application;

[0031] Figure 7It is a schematic flow diagram for updating activation quantization parameters in the deep learning model quantization method according to an embodiment of the present application;

[0032] Figure 8A It is a schematic diagram for online quantization calibration according to the asymmetric quantization method of an embodiment of the present application;

[0033] Figure 8B It is a schematic diagram for processing different segments using a multi-bit width coprocessor according to an embodiment of the present application; and

[0034] Figure 9 It is a schematic structural diagram of a computer system of an electronic device suitable for implementing an embodiment of the present application. Detailed implementation manners

[0035] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present application, rather than limiting the present application. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.

[0036] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.

[0037] Figure 1 An exemplary system architecture 100 for an embodiment of the quantization method and apparatus for a deep learning neural network to which the present application can be applied is shown.

[0038] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0039] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as a web browser application, a shopping application, a search application, an instant messaging tool, an email client, a social platform software, etc.

[0040] The terminal devices 101, 102, and 103 can be hardware or software. When the terminal devices 101, 102, and 103 are hardware, they can be various electronic devices, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, desktop computers, and so on. When the terminal devices 101, 102, and 103 are software, they can be installed in the above-listed electronic devices. They can be implemented as multiple software or software modules (such as multiple software or software modules for providing distributed services), or can be implemented as a single software or software module. No specific limitation is made here.

[0041] The server 105 can be a server that provides various services, such as a model processing server for processing the deep learning models sent by the terminal devices 101, 102, and 103. The model processing server can analyze and process data such as the received deep learning models, and feedback the processing results (such as the quantized deep learning models) to the terminal devices.

[0042] It should be noted that the deep learning model quantization method provided by the embodiments of the present application is generally executed by the server 105. Correspondingly, the deep learning model quantization device is generally set in the server 105. In addition, for the deep learning models at the embedded end, the deep learning model quantization method provided by the embodiments of the present application can also be executed by the terminal devices 101, 102, and 103. Correspondingly, the deep learning model quantization device can be set in 101, 102, and 103.

[0043] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules (such as multiple software or software modules for providing distributed services), or can be implemented as a single software or software module. No specific limitation is made here.

[0044] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in

[0045] Figure 2 shows a deep learning model according to an embodiment of the present application. In combination withFigure 2 As shown, a deep learning model may include multiple blocks, and each block may include one layer or multiple adjacent layers. For example, the deep learning model may be divided into two blocks, Block1 and Block2. Block1 includes three adjacent layers, Layer1, Layer2, and Layer3, and Block2 includes two adjacent layers, Layer4 and Layer5. At the same time, multiple channels within the same layer may be divided into multiple groups, and each group includes one or more channels. For example, the five channels in Layer2 can be divided into two groups, Group1 and Group2.

[0046] It should be noted that the number and types of adjacent layers included in each block may be the same or different. For example, each layer may be, but is not limited to, a convolutional layer, a pooling layer, a fully connected layer, an activation layer, etc.; the number of channels included in each group may be the same or different, and the present application does not make specific restrictions on this.

[0047] In this embodiment, the execution subject of the deep learning model quantization method (such as Figure 1 the server shown) can obtain the above-mentioned deep learning model to be quantized from remote or local through a wired connection method or a wireless connection method. The deep learning model to be quantized can be an untrained network model or a trained network model. The function or input and output of the deep learning model to be quantized can be predetermined.

[0048] Among them, the deep learning model can be used to process at least one of the following sample data: image data, audio-visual data, and text data. For example, the deep learning model can be a neural network model for image classification, and the sample data can be preprocessed image data. The sample data is input into the trained deep learning model, and the deep learning model processes the sample data and outputs an image classification result.

[0049] In this embodiment, the deep learning model to be quantized includes initial model parameters, and the initial model parameters can be floating-point model parameters, such as floating-point weight values and floating-point activation values. The quantization process of the deep learning model refers to the operation process of converting the weight values, activation values, etc. of the deep neural network to be quantized from high precision to low precision. For example, converting 32-bit floating-point numbers to 8-bit integers int8, while expecting the accuracy or accuracy rate of the converted model to be similar to that before conversion.

[0050] Figure 3Flow 300 of an embodiment of the deep learning model quantization method according to the present application is shown. The deep learning model quantization method 300 includes the following steps:

[0051] Step 301, perform group normalization processing on the weight data of each layer;

[0052] Step 302, segment the weight data of each group;

[0053] Step 303, quantize the weight data of each segment, and determine the quantization parameters of each segment;

[0054] Step 304, perform quantization processing on the deep learning model according to the quantization parameters of each segment.

[0055] The above steps 301-304 will be further described below.

[0056] As Figure 3 shown, the deep learning model quantization method 300 starts from step 301, and performs group normalization processing on the weight data of each layer.

[0057] Referring to Figure 4 shown, in an exemplary embodiment, the group normalization processing performed on the weight data within the same layer may include the following sub-steps:

[0058] Step 3011, obtain the data distribution of the average group according to the weight data distribution of each group within the layer.

[0059] In this embodiment, the data distribution of the average group may include the data range of the average group (the data range is defined by the minimum value VgMin and VgMax), and the data distribution center VgAvgCenter of the average group. Specifically, the data range of the average group may be obtained based on the data statistical processing of the data of each group within the current layer (which can be regarded as a data set), for example, it can be obtained by averaging according to the numerical value and distribution of each weight data. The data distribution center may be a numerical value where the weight data distribution is the most concentrated, and the present application does not limit this.

[0060] Step 3012, for each group in the layer, determine the scale scaling factor and the translation factor according to the data distribution of the average group, and map the weight data of the group based on the scale scaling factor and the translation factor to perform group normalization processing.

[0061] In this embodiment, mapping each group in the current layer to the average group data includes scaling and translation. The scaling factor is used to represent the ratio of the data range of the weight data of each group to the data range of the average group, and the translation factor is used to represent the difference between the data distribution center of the weight data of each group and the data distribution center of the average group. Thus, the loss of quantization accuracy caused by too large a difference in data range between different groups can be avoided.

[0062] Exemplarily, the mapped weight data can be obtained through the following method:

[0063]

[0064] Among them, V g1Cur is the mapped weight data, V g1 is the initial weight data, α is the scaling factor, and β is the translation factor.

[0065] Combined with Figure 4 shown, taking Group 1 as an example, the scaling factor α and the translation factor β can be determined through the following method:

[0066]

[0067] β = V gAvgCenter - V g1Center3

[0068] Among them, V g1Min1 and V g1Max1 are the initial data distribution ranges of the weight data of Group 1, V g1Center1 is the initial data distribution center; V g1Min2 and V g1Max2 are the data distribution ranges of the weight data of Group 1 after translation, V g1Center2 is the data distribution center after translation; V g1Min3 and V g1Max3 are the data distribution ranges of the weight data of Group 1 after scaling and translation (i.e., after mapping), V g1Center3 is the data distribution center of Group 1 after mapping.

[0069] Next, step 302 is executed, and the weight data of each group of each layer (such as layer1) is segmented.

[0070] Referring to Figure 5 shown, the weight data within the same group can be clustered through a clustering algorithm. According to the clustering results, each class of data is taken as a separate segment, and thus multiple segments (Segments) are obtained, and an overlapping extended area is set between two adjacent segments. For example, for Figure 5Among three consecutive segments (segment A, segment B, and segment C), an extended area (Span) is set at the boundary between two adjacent segments.

[0071] Taking the quantization of the weight data of segment B as an example, according to the data range of the weight data of segment B (bounded by the minimum value VMinB and VMaxB), and the extended areas between the two segments adjacent to segment B (including the extended area SpanBA between segment B and segment A, and the extended area SpanBC between segment B and segment C), the quantization parameters of segment B are determined, such as the fixed-point scaling factor SB, which is used to convert the floating-point weight data into fixed-point weight data.

[0072] By performing the above segmentation and quantization processing based on the segment data, it is possible to avoid the situation where the local data range has an unreasonable distribution (such as a multi-distribution situation), resulting in non-optimal local quantization parameters. Adopting this segmentation method will cause a certain degree of reduction in quantization efficiency, but it does not affect the overall execution efficiency of model quantization. At the same time, it can optimize the data range to a certain extent, and thus can improve the local quantization accuracy.

[0073] Then continue with step S303, and determine the quantization parameters of each segment by quantizing the weight data of each segment.

[0074] Exemplarily, taking the quantization of the weight data of segment B as an example, in step S303, the weight data of segment B can be quantized and the quantization parameters of segment B, that is, the fixed-point scaling factor SB, can be determined in the following way:

[0075]

[0076]

[0077]

[0078] Among them, V fB is the floating-point weight data of segment B, V iB is the fixed-point weight data of segment B, ξ f is the floating-point random noise, SB is the fixed-point scaling factor, Clamp B is the truncated data of segment B, Zoffset is the zero-point offset factor, SpanBA is the extended area between segment B and segment A, SpanBC is the extended area between segment B and segment C, EBA is the minimum value of segment B in the extended area SpanBA, and EBC is the maximum value of segment B in the extended area SpanBC.

[0079] Specifically, the quantization of weight data refers to converting floating-point data into fixed-point data within a certain value range. Here, the value range is limited by the number of bits of the fixed-point data. For example, if the integer data to be converted is 8 bits (i.e., 8bit), the value range is (0, 255). It can be understood that for floating-point data and fixed-point data with the same number of bits, since floating-point data can record data information after the decimal point, it has higher precision. While fixed-point data can occupy less storage space and has a faster calculation speed.

[0080] By performing this step 303 on each layer, that is, through layer-by-layer quantization, quantization parameters for each segment of each layer can be obtained. It should be noted that before inputting the output value of each layer into the next layer, the inverse transformation of the above-mentioned grouped normalization process needs to be performed on the weight data of the current layer, or before performing the operation of step 303 on the next layer, the corresponding inverse transformation or inverse mapping process needs to be performed on the weight data of the previous layer. For example, for layer2 in Block1, the weight data of its corresponding layer1 channels can be multiplied by α*λ and subtracted by β to match the scale scaling and translation mapping of the previous layer1, so as to match the scale scaling and translation of the previous layer1. Where λ is the scale adjustment term, and the value of this λ can be 1 or other values.

[0081] Then in step 304, the deep learning model is quantized according to the quantization parameters of each segment, and the first quantized model can be obtained.

[0082] After obtaining the quantization parameters of each segment of the deep learning model by performing the above steps 301 - 303 on each layer, the deep learning model can be quantized according to the quantization parameters of each segment to obtain the first quantized model after quantization. In some embodiments, the obtained quantization parameters can also be used for internal conversion of the inference device when executing the model. In some embodiments, the obtained quantization parameters and the first quantized model can be used together as input for the inference program.

[0083] In addition, in some embodiments, the above deep learning model quantization method 300 may further include: dividing the deep learning model into blocks, and grouping multiple channels of each layer.

[0084] Specifically, partitioning the deep learning model may include: partitioning the deep learning model according to the weight data distribution of adjacent layers in the deep learning model, and taking multiple adjacent layers whose weight data distribution meets the first preset condition as one block. Among them, the first preset condition may include at least one of the following: the difference in the data range of the weight data of adjacent layers is less than a predetermined first threshold, and the difference in the data distribution centers of adjacent layers is less than a predetermined second threshold. Wherein the predetermined second threshold is not less than the predetermined first threshold. If there is a layer that does not meet the first preset condition with adjacent layers, then this layer can be taken as a separate block.

[0085] Exemplarily, two or more adjacent layers whose maximum value, minimum value, and data distribution center in the weight data meet the first preset condition can be divided into one block (as a sub-network). For example, by calculating the mean of the data ranges of different layers as the proposed distribution center, layers that deviate from this proposed distribution center or the difference in the data ranges between different layers exceeds a predetermined threshold do not belong to this block. For the partitioned data, its data range should be greater than or equal to the data range of a single layer and less than or equal to the data range of the entire model, with the best case being that the accuracy loss is not greater than the accuracy of a single-layer network.

[0086] Specifically, grouping the multiple channels of each layer may include: grouping each layer according to the weight data distribution of the channels within the same layer, and taking multiple channels within the same layer whose weight data distribution meets the second preset condition as one group. Among them, the second preset condition may include at least one of the following: the difference between the data ranges of the weight data of different channels is less than a predetermined third threshold, and the difference between the data distribution centers of different channels is less than a predetermined fourth threshold. Wherein the predetermined fourth threshold is not less than the predetermined third threshold.

[0087] In some embodiments, the above deep learning model quantization method 300 may further include: before performing group normalization processing on the weight data of each layer (such as the above step 301), performing block normalization processing on the weight data of each block.

[0088] Exemplarily, for each block, the group normalization processing may include: removing the outliers in the weight data of the current block to obtain the maximum value MaxV1 and the minimum value MinV2 of the weight data of this block; performing block normalization processing by dividing the weight data of the current block by the difference (MaxV1 - MinV2) between the maximum value MaxV1 and the minimum value MinV2 to obtain the weight data after block normalization processing.

[0089] According to the above embodiments, by "blocking" the deep learning model, it is possible to avoid the situation where the quantization parameters are locally optimal due to the overly concentrated data distribution in a single layer. For example, the quantization parameters of a single layer are optimal, but they are not ideal for multiple layers or even the entire network. By taking adjacent layers with similar distributions as a block and performing normalization, the data range can be expanded to a certain extent, and the global quantization accuracy can be improved. Further, by grouping multiple channels of each layer, channels with similar distributions are grouped together and grouped normalization processing is performed, such as scale scaling and translation operations, which can avoid the loss of quantization accuracy caused by too large a difference in data range between different groups.

[0090] Reference Figure 6A and Figure 6B As shown, after obtaining the first quantized model through the deep learning model quantization method 300 according to the above embodiments of the present application, the method may further include:

[0091] Step 601, performing a second quantization process on the deep learning model to obtain a second quantized model;

[0092] Step 602, respectively comparing the accuracy of the first subgraph corresponding to each block of the first quantized model with the second subgraph corresponding to the same block in the second quantized model, and taking the subgraph with higher accuracy between the first subgraph and the second subgraph as the target subgraph corresponding to the block; and

[0093] Step 603, obtaining the final target quantized model according to the target subgraph corresponding to each block.

[0094] As Figure 6A shown, the quantized model A is a quantized model obtained by using the quantization method of related technologies, and the quantized model B is a quantized model obtained by using the deep learning model quantization method of the present application. By comparing the subgraphs corresponding to each block (Block1 - Block3), subgraphs with higher accuracy are selected for the best subgraph screening. Thus, as Figure 6B shown, the target quantized model composed of the best subgraphs (best sub Figure 1 - best sub Figure 3 ) of each block is obtained.

[0095] In addition, in some embodiments, the above deep learning model quantization method 300 may further include: in response to detecting that the sample data meets a preset calibration condition, dynamically online quantizing and calibrating the activation values of the deep learning model by updating the activation quantization parameters. Wherein, the activation quantization parameters may include a scale factor of the data range of the activation values, and the preset calibration condition may include: the difference between the activation value of the sample data and the historical average activation value is less than a preset threshold range.

[0096] Figure 7shows an online quantization calibration process 700 according to an exemplary embodiment. As Figure 7 shown, for the deep learning model on the embedded side, the online quantization calibration process 700 may include:

[0097] Step 701, receiving real-time data, which may be data screened based on actual inference services.

[0098] Step 702, determining whether there is an inference task currently. If yes, execute Step 703; otherwise, execute Step 708.

[0099] Step 703, performing an inference process. For embedded AI, its AI core may include a single CPU or multiple CPUs, and the model inference is performed through the AI core and coprocessor.

[0100] Step 704, the CPU performs FP (Forward Propagation) and BP (Back Propagation) propagation processing.

[0101] Step 705, caching the sample data.

[0102] In this embodiment, DMA (Direct Memory Access) can be used for the transmission and caching of sample data. The DMA data transmission can bypass the CPU, reducing the consumption of CPU resources and saving CPU resources for other operations.

[0103] Step 706, obtaining the screened calibration samples.

[0104] Step 707, determining whether there is an idle CPU currently. If yes, execute Step 708; otherwise, return to Step 705.

[0105] Step 708, using the idle CPU for online quantization calibration and executing the next Step 709.

[0106] Step 709, updating the activation quantization parameters.

[0107] Specifically, the CPU can determine that when the difference between the current sample historical average activation value and the current sample activation value is within the threshold range according to the preset calibration conditions, use the latest distribution of the activation values (such as the maximum activation value and minimum activation value of the current sample, or the average maximum activation value and minimum activation value of multiple cached samples) to recalculate the scale factor S of the data range of the activation values for online quantization calibration. As Figure 8A shown, an asymmetric quantization method can be specifically used for online quantization calibration.

[0108] Exemplarily, the scale factor S can be calculated in the following manner:

[0109]

[0110] Among them, V afMax is the maximum value of the floating-point activation value, and V afMin is the minimum value of the floating-point activation value, and V aiMax is the maximum value of the fixed-point activation value, and V aiMin is the minimum value of the fixed-point activation value, and ξ af is the random compensation value of the activation value.

[0111] According to the deep learning model quantization method of the above embodiments of the present application, during the sample data caching and online quantization calibration process, by using DMA and Cache (cache memory), the online quantization efficiency of the embedded AI can be improved.

[0112] In addition, as Figure 8B shown, the embedded AI can also adopt a multi-bit-width coprocessor design. For the different bit widths corresponding to different segments in each group (Group), for example, the bit width corresponding to segment A is 2 bits, the bit width corresponding to segment B is 4 bits, etc., a multi-bit-width coprocessor can be used for parallel execution, thereby accelerating the inference speed of the above quantization model.

[0113] The deep learning model quantization method provided according to the embodiments of the present application has at least one of the following beneficial effects:

[0114] 1. By proposing a block scheme and a block normalization scheme for the network layer, the problem of inference loss of the fixed-point model caused by the too large difference in the weight data range between adjacent layers is solved;

[0115] 2. A group mapping scheme based on block data is proposed, which improves the quantization accuracy;

[0116] 3. By using a segmented quantization method based on extended data, the problem of local optimality of quantization parameters caused by the lack of segmented quantization boundaries is avoided;

[0117] 4. By means of the quantization model subgraph screening and model reconstruction mechanism, the accuracy loss caused by excessive segmented quantization is avoided;

[0118] 5. Through dynamic quantization calibration (online quantization calibration) on the embedded side, while adaptively improving the online inference accuracy, the online upgrade frequency of the quantization model is reduced;

[0119] 6. Through the collaborative processing of the embedded multi-bit-width coprocessor, the model inference efficiency of the embedded side is improved.

[0120] Figure 9 Illustrates a device suitable for implementing the embodiments of the present application (for example Figure 1Schematic diagram of the computer system 900 of the server or terminal device) in []. The terminal device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The devices shown are merely examples and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0121] Such as Figure 9 As shown, the computer system 900 may include a processor (such as a CPU, Central Processing Unit) 901, which can perform various appropriate actions and processes according to the program stored in the Read-Only Memory (ROM) 902 or the program loaded from the storage section 908 into the Random Access Memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the system 900 are also stored. The processor 901, ROM 902, and RAM 903 are connected to each other via a bus 904. The Input / Output (I / O) interface 905 is also connected to the bus 904.

[0122] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed so that the computer program read from it can be installed into the storage section 908 as needed.

[0123] In particular, according to the embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 909, and / or installed from the removable medium 911. When the computer program is executed by the processor 901, the above-mentioned functions defined in the methods of the embodiments of the present application are executed.

[0124] It should be noted that the computer-readable medium described in the embodiments of the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments of the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the embodiments of the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0125] The above computer-readable medium may be included in the above electronic device; or it may exist separately and not be assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to perform the above functions defined in the method of the embodiments of the present application.

[0126] Computer program code for performing the operations of the embodiments of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0127] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0128] The modules described in the embodiments of the present application may be implemented in software or in hardware. The described modules may also be provided in a processor, and the names of these modules do not, in some cases, constitute a limitation on the module itself.

[0129] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the embodiments of the present application.

Claims

1. A deep learning model quantization method, characterized in that, The deep learning model includes multiple layers, and multiple channels in each layer are divided into multiple groups; The method includes: For each layer, perform group normalization processing on the weight data of the layer, segment the weight data of each group, and determine the quantization parameter of each segment by quantizing the weight data of each segment; Quantize the deep learning model according to the quantization parameter of each segment.

2. The method according to claim 1, wherein The deep learning model includes multiple blocks, and each block includes one of the layers or multiple adjacent layers.

3. The method according to claim 2, wherein The method further includes: Before performing group normalization processing on the weight data of the layer, perform block normalization processing on the weight data of each block.

4. The method according to claim 2, wherein, The method further includes: Divide the deep learning model according to the weight data distribution of adjacent layers in the deep learning model, and use multiple adjacent layers whose weight data distribution satisfies the first preset condition as one block, The first preset condition includes at least one of the following: the difference between the data ranges of the weight data of adjacent layers is less than a predetermined first threshold, and the difference between the data distribution centers of adjacent layers is less than a predetermined second threshold; where the predetermined second threshold is not less than the predetermined first threshold.

5. The method according to claim 1, wherein The method further includes: Group each layer according to the weight data distribution of each channel within the layer, and use multiple channels whose weight data distribution within the same layer satisfies the second preset condition as one group, The second preset condition includes at least one of the following: the difference between the data ranges of the weight data of different channels is less than a predetermined third threshold, and the difference between the data distribution centers of different channels is less than a predetermined fourth threshold; where the predetermined fourth threshold is not less than the predetermined third threshold.

6. The method according to claim 3, wherein The performing block normalization processing on the weight data of each block includes: Remove the outliers from the weight data of each block to obtain the maximum value and minimum value of the weight data of the block; Perform the block normalization processing by dividing the weight data of the block by the difference between the maximum value and the minimum value.

7. The method according to claim 1, wherein The performing group normalization processing on the weight data of the layer includes: Obtain the data distribution of the average group according to the weight data distribution of each group within the layer; and For each group in the layer, determine the scale scaling factor and the translation factor according to the data distribution of the average group, and map the weight data of the group based on the scale scaling factor and the translation factor to perform the group normalization processing; Wherein, the data distribution of the average group includes the data range and the data distribution center of the average group, the scale scaling factor is used to represent the ratio of the data range of the weight data of each group to the data range of the average group, and the translation factor is used to represent the difference between the data distribution center of the weight data of each group and the data distribution center of the average group.

8. The method according to claim 7, wherein, The method further includes: Before inputting the output value of each layer into the next layer, perform the inverse transformation processing of the group normalization processing on the weight data of the layer.

9. The method according to claim 1, wherein Segmenting the weight data of each of the groups, and determining the quantization parameter of each segment by quantizing the weight data of each segment, including: For each of the groups, clustering the weight data of the group by a clustering algorithm, obtaining a plurality of segments according to the clustering result, and setting an overlapping extended region between two adjacent segments; Determining the quantization parameter of each segment according to the data range of the weight data of each segment and the extended region between two segments adjacent to the segment, wherein the quantization parameter of the segment includes a fixed-point scaling factor for converting floating-point weight data into fixed-point weight data.

10. The method according to claim 2, wherein, The method further includes: Obtaining a first quantized model according to the quantization processing performed on the deep learning model; Performing a second quantization process on the deep learning model to obtain a second quantized model; Comparing the accuracy of the first subgraph corresponding to each block of the first quantized model with the second subgraph corresponding to the block in the second quantized model respectively, and using the subgraph with higher accuracy in the first subgraph and the second subgraph as the target subgraph corresponding to the block; and Obtaining a target quantized model according to the target subgraph corresponding to each block.