Lightweight network quantification method, computing device and computer storage medium

By quantizing and optimizing the network blocks of lightweight networks, the problem of insufficient quantization accuracy is solved, and more efficient storage and computing performance is achieved, making the network more suitable in resource-constrained environments.

CN119990192APending Publication Date: 2025-05-13NANJING UNIV OF INFORMATION SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510061427.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

How to improve the quantization accuracy of lightweight networks and solve the problem of insufficient quantization accuracy in the prior art.

Method used

By obtaining the quantization request, determining the network block to be quantized, constructing the initial equalization vector, deviation absorption vector and quantization coefficient, quantizing the network blocks, and optimizing these vectors and coefficients to minimize network errors and block reconstruction errors.

Benefits of technology

It improves the quantization accuracy of lightweight networks, reduces storage requirements and computing complexity, and makes the network more suitable in resource-constrained environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990192A_ABST
    Figure CN119990192A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a lightweight network quantification method, computing equipment and a computer storage medium. The lightweight network quantification method comprises the following steps: constructing an initialization equalization vector, an initialization deviation absorption vector and an initialization quantification coefficient of a first network block; quantizing the first network block to obtain a first quantized network block; constructing a network error based on the first output data and the second output data, and constructing a block reconstruction error based on the third output data and the fourth output data; and optimizing the initialized equalization vector, the initialized deviation absorption vector and the initialized quantization coefficient of the first network block by taking minimization of the block reconstruction error and the network error as an optimization target, generating a target equalization vector, a target deviation absorption vector and a target quantization coefficient, and obtaining a quantized second quantization network block. According to the technical scheme provided by the embodiment of the invention, the quantization precision of the lightweight network can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a lightweight network quantization method, computing device, and computer storage medium. Background Art

[0002] Lightweight networks can include convolutional neural networks that focus on achieving efficient operation in resource-constrained environments. Their core design concept revolves around reducing model size, reducing computational burden, and improving inference speed to meet the needs of real-time inference in scenarios such as mobile devices, embedded systems, or edge computing devices.

[0003] By quantizing the lightweight network, the memory space required to store these parameters (such as weights and activation values) can be effectively reduced by converting the parameters in the network (such as weights and activation values) from high-precision data types (such as 32-bit floating-point numbers) to low-precision data types (such as 8-bit integers, 4-bit integers, etc.).

[0004] In the process of realizing the concept of this application, the inventors found that how to improve the quantization accuracy of lightweight networks has become a technical problem that needs to be solved urgently. Summary of the invention

[0005] Embodiments of the present application provide a lightweight network quantization method, apparatus, computing device, and computer storage medium.

[0006] In a first aspect, an embodiment of the present application provides a quantization method for a lightweight network, wherein the lightweight network includes at least one network block, and the method includes:

[0007] Obtaining a quantization request for the lightweight network;

[0008] In response to the quantization request, determining a first network block to be quantized from the at least one network block;

[0009] Constructing an initialization equalization vector, an initialization deviation absorption vector, and an initialization quantization coefficient of the first network block;

[0010] quantizing the first network block using the initialized equalization vector, the initialized deviation absorption vector, and the initialized quantization coefficient to obtain a first quantized network block;

[0011] Acquire first output data of the lightweight network, second output data of the quantized lightweight network after quantization, third output data of the first network block, and fourth output data of the first quantized network block;

[0012] Constructing a network error based on the first output data and the second output data, and constructing a block reconstruction error based on the third output data and the fourth output data;

[0013] Taking minimizing the block reconstruction error and the network error as the optimization goal, the initialization equalization vector, the initialization deviation absorption vector and the initialization quantization coefficient of the first network block are optimized to generate a target equalization vector, a target deviation absorption vector and a target quantization coefficient, and obtain a second quantized network block that has been quantized.

[0014] In a second aspect, an embodiment of the present application provides a quantization device for a lightweight network, wherein the lightweight network includes at least one network block, and the device includes:

[0015] A request acquisition module, used to acquire a quantization request for the lightweight network;

[0016] A network block determination module, configured to determine, in response to the quantization request, a first network block to be quantized from the at least one network block;

[0017] A vector construction module, used to construct an initialization equalization vector, an initialization deviation absorption vector and an initialization quantization coefficient of the first network block;

[0018] A first quantization module, configured to quantize the first network block using the initialization equalization vector, the initialization deviation absorption vector, and the initialization quantization coefficient to obtain a first quantized network block;

[0019] An output acquisition module, used to acquire first output data of the lightweight network, second output data of the quantized lightweight network after quantization, third output data of the first network block, and fourth output data of the first quantized network block;

[0020] an error construction module, configured to construct a network error based on the first output data and the second output data, and to construct a block reconstruction error based on the third output data and the fourth output data;

[0021] An optimization module is used to optimize the initialization equalization vector, the initialization deviation absorption vector and the initialization quantization coefficient of the first network block with the optimization goal of minimizing the block reconstruction error and the network error, generate a target equalization vector, a target deviation absorption vector and a target quantization coefficient, and obtain a second quantized network block that has been quantized.

[0022] In an embodiment of the present application, by adopting: obtaining a quantization request for the lightweight network; determining a first network block to be quantized from the at least one network block in response to the quantization request; constructing an initialization equalization vector, an initialization deviation absorption vector and an initialization quantization coefficient of the first network block; quantizing the first network block using the initialization equalization vector, the initialization deviation absorption vector and the initialization quantization coefficient to obtain a first quantized network block; obtaining the first output data of the lightweight network, the second output data of the quantized lightweight network after quantization, the third output data of the first network block and the fourth output data of the first quantized network block; constructing a network error based on the first output data and the second output data, and constructing a block reconstruction error based on the third output data and the fourth output data; optimizing the initialization equalization vector, the initialization deviation absorption vector and the initialization quantization coefficient of the first network block with minimizing the block reconstruction error and the network error as the optimization goal, generating a target equalization vector, a target deviation absorption vector and a target quantization coefficient, and obtaining a technical solution for a quantized second quantized network block, the technical effect of improving the quantization accuracy of the lightweight network is achieved.

[0023] These and other aspects of the present application will become more clearly understood in the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0025] Figure 1 A flowchart of a lightweight network quantization method provided by an embodiment of the present application is shown;

[0026] Figure 2 A schematic diagram of a lightweight network quantization method provided in an embodiment of the present application is shown;

[0027] Figure 3 A block diagram of a lightweight network quantization device provided in an embodiment of the present application is shown;

[0028] Figure 4 A block diagram of a computing device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0029] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0030] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit the "first" and "second" to be different types.

[0031] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0032] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0033] The implementation details of the technical solution of the embodiment of the present application are elaborated in detail below.

[0034] Figure 1 A flowchart of a quantization method for a lightweight network provided by an embodiment of the present application is shown. The lightweight network includes at least one network block, such as Figure 1 As shown, the method may specifically include the following steps:

[0035] 101, obtain a quantization request for a lightweight network.

[0036] Among them, lightweight networks can include convolutional neural networks that focus on achieving efficient operation in resource-constrained environments. Their core design concept revolves around reducing model size, reducing computational burden, and improving reasoning speed to meet the needs of real-time reasoning in scenarios such as mobile devices, embedded systems, or edge computing devices.

[0037] Compared with traditional deep neural networks, lightweight networks have significantly fewer parameters, which means that less memory space can be occupied when storing the model. For example, on mobile devices, limited storage resources can accommodate lightweight network models without worrying about insufficient storage due to overly large models.

[0038] Through specific network structure design, lightweight networks require fewer computing resources when performing inference calculations, which enables them to quickly complete computing tasks on devices with limited computing power (such as embedded systems), improve response speed, and implement applications with high real-time requirements, such as real-time image recognition and video processing.

[0039] Due to the small number of parameters and low computational complexity, lightweight networks can quickly produce results when performing inference (operations such as prediction or classification of input data). In actual application scenarios, such as real-time target detection in smart security, fast inference speed can detect abnormal situations in a timely manner and respond to them.

[0040] The lightweight network itself has reduced the number of parameters and computation through structural design, but further optimization is still needed in some devices with extremely limited resources (such as some low-end embedded devices or mobile devices with small memory).

[0041] By quantizing the lightweight network, the parameters in the network (such as weights and activation values) can be converted from high-precision data types (such as 32-bit floating-point numbers) to low-precision data types (such as 8-bit integers, 4-bit integers, etc.), which can effectively reduce the memory space required to store these parameters.

[0042] In addition to reducing storage, the quantized network can also be more efficient in the calculation process. In particular, when the quantization coefficient is constrained to a power of 2, shift operations can be used on the hardware to replace some multiplication operations, thereby speeding up the inference speed. This is very important for application scenarios that require real-time response (such as real-time image recognition, target detection in video surveillance, etc.), which enables lightweight networks to process input data faster on devices with limited resources.

[0043] Among them, the quantization request can come from the user's desire to deploy a lightweight network on a specific device (such as a mobile terminal, embedded system, or other resource-constrained device), and it needs to be quantized to improve computing efficiency and reduce storage requirements. For example, when developing a mobile-based image recognition application, in order to make the model run quickly and not occupy too much memory, a quantization request for a lightweight network can be issued.

[0044] 102. In response to a quantization request, determine a first network block to be quantized from at least one network block.

[0045] In the embodiment of the present application, when quantizing a lightweight network, a method of sequentially quantizing multiple network blocks can be adopted. This sequential processing method helps to systematically and gradually complete the quantization process of the entire network, making the quantization operation more organized and easy to manage and monitor.

[0046] The first network block may include a network block that has not been subjected to a quantization operation.

[0047] In other implementations of the present application, the first network block may be determined based on factors such as the importance of the network block, computational complexity, or data flow order. For example, the initial network block responsible for feature extraction in the network is first selected for quantization, because it has an important impact on the quality of input data of subsequent network blocks, and its quantized effect may play a key role in the performance of the entire network.

[0048] 103 , construct an initialization equalization vector, an initialization deviation absorption vector, and an initialization quantization coefficient of the first network block.

[0049] In the embodiment of the present application, the initialization equalization vector can be constructed by analyzing the amplitude range of the weights of two adjacent convolutional layers. For example, if there are two convolutional layers, the initialization equalization vector can be constructed based on the amplitude range of their weights. (The weight of the first convolutional layer w (1) The amplitude range of the i-th channel) and (The weight of the second convolutional layer w (2) The amplitude range of the i-th input channel is calculated using the following formula (1) to calculate the initialization equalization vector S corresponding to each channel. The initialization equalization vector S can be used to adjust the weight distribution between two adjacent convolutional layers so that the quantization error can be more evenly distributed between the two layers.

[0050]

[0051] In an embodiment of the present application, the mean β of each output channel stored in the BN layer (batch normalization layer) can be used. i With variance γ i, the initialization deviation absorption vector C is calculated by the following formula (2). The initialization deviation absorption vector C can be used to suppress the problem of increased activation range caused by the adjustment of the equalization vector, thereby reducing the impact of activation quantization error on network performance.

[0052] C i =max(0,β i -3γ i ); (2)

[0053] In an embodiment of the present application, an initialization quantization coefficient can be constructed according to the quantization bit width and data distribution, and the initialization quantization coefficient can be used to map the original parameters of the lightweight network to a low-precision data representation. In a possible implementation, the quantization coefficient can be constrained to be a power of 2.

[0054] 104. quantize the first network block by using the initialized equalization vector, the initialized deviation absorption vector, and the initialized quantization coefficient to obtain a first quantized network block.

[0055] In an embodiment of the present application, the initial equalization vector, the initial deviation absorption vector and the initial quantization coefficient can be used to perform a quantization operation on the initial weights and activation values ​​of the first network block to obtain the quantized weights and activation values, and obtain the first quantized network block with the quantized weights and activation values ​​as parameters.

[0056] 105 , obtaining first output data of the lightweight network, second output data of the quantized lightweight network after quantization, third output data of the first network block, and fourth output data of the first quantized network block.

[0057] In some embodiments, obtaining the first output data of the lightweight network and the quantized second output data of the lightweight network can be specifically implemented as follows:

[0058] Get input data;

[0059] Inputting the input data into the lightweight network to obtain first output data;

[0060] Input data is input into a quantized lightweight network to obtain second output data, wherein the quantized lightweight network is obtained by quantizing at least one network block in the lightweight network.

[0061] Among them, the first output data may refer to the output result obtained by inputting the original input data into the lightweight network before any quantization operation is performed on the lightweight network and going through the normal calculation process of the network (i.e., using the original high-precision parameters, such as weights, biases, and activation values ​​represented by 32-bit floating point numbers, etc.).

[0062] The second output data may refer to output data obtained by inputting the same original input data as that used to obtain the first output data into the quantized lightweight network after quantizing at least one network block in the lightweight network. If there are still unquantized network blocks in the quantized lightweight network, the input data may be processed by the quantized network block and the original network block respectively and then output.

[0063] The third output data may refer to inputting the original input data into the lightweight network alone before quantizing the first network block, and only focusing on the processing result of the first network block on the input data.

[0064] The fourth output data may refer to the output result obtained by inputting the original input data into the first quantization network block after the quantization operation is completed on the first network block to obtain the first quantization network block (at this time, the parameters inside the first quantization network block have been converted from high precision to low precision).

[0065] 106. Construct a network error based on the first output data and the second output data, and construct a block reconstruction error based on the third output data and the fourth output data.

[0066] 107, with minimizing the block reconstruction error and the network error as the optimization goal, the initialization equalization vector, the initialization deviation absorption vector and the initialization quantization coefficient of the first network block are optimized to generate the target equalization vector, the target deviation absorption vector and the target quantization coefficient, and obtain the quantized second quantized network block.

[0067] In the embodiment of the present application, by constructing a network error based on the first output data and the second output data, the degree of difference in performance between the entire lightweight network after quantization and the original lightweight network without quantization can be measured.

[0068] In an embodiment of the present application, by adopting: obtaining a quantization request for the lightweight network; determining a first network block to be quantized from the at least one network block in response to the quantization request; constructing an initialization equalization vector, an initialization deviation absorption vector and an initialization quantization coefficient of the first network block; quantizing the first network block using the initialization equalization vector, the initialization deviation absorption vector and the initialization quantization coefficient to obtain a first quantized network block; obtaining the first output data of the lightweight network, the second output data of the quantized lightweight network after quantization, the third output data of the first network block and the fourth output data of the first quantized network block; constructing a network error based on the first output data and the second output data, and constructing a block reconstruction error based on the third output data and the fourth output data; optimizing the initialization equalization vector, the initialization deviation absorption vector and the initialization quantization coefficient of the first network block with minimizing the block reconstruction error and the network error as the optimization goal, generating a target equalization vector, a target deviation absorption vector and a target quantization coefficient, and obtaining a technical solution for a quantized second quantized network block, the technical effect of improving the quantization accuracy of the lightweight network is achieved.

[0069] In some embodiments, constructing the network error based on the first output data and the second output data may be specifically implemented as follows:

[0070] Determine the lightweight network network type;

[0071] Determine the distance metric function that matches the network type;

[0072] A network error is constructed based on the first output data and the second output data using a distance metric function.

[0073] The specific method of constructing the network error can vary depending on the type of application task of the network. For example, for classification tasks, the KL divergence (Kullback-Leibler Divergence) can be used to construct the network error. KL divergence can measure the difference between two probability distributions, which is the difference between the quantized network output probability distribution (represented by the second output data) and the original network output probability distribution (represented by the first output data). By calculating the value of KL divergence, the degree of influence of the quantization operation on the network classification performance can be quantitatively expressed.

[0074] For other application scenarios such as detection tasks, the L2 norm can be used to construct the network error. For example, the square root of the sum of the squares of the differences between the first output data and the second output data in the corresponding dimension is calculated to measure the distance between the two, thereby reflecting the impact of quantization on network detection performance and other aspects.

[0075] In the embodiment of the present application, by constructing a block reconstruction error based on the third output data and the fourth output data, the degree of change in the function of the first network block itself caused by the quantization operation can be evaluated.

[0076] In an embodiment of the present application, the block reconstruction error can be constructed by calculating the L2 distance of the outputs of the two. That is, the sum of the squares of the differences between the third output data and the fourth output data in each dimension is first calculated, and then the square root of this sum is taken, and the result is the value of the block reconstruction error. The smaller this value is, the closer the quantized first network block is to the unquantized first network block in function, which means that the quantization operation has less impact on the first network block itself, which reflects to a certain extent that the quantization process has a better processing effect on the first network block, and can better maintain its original functional characteristics while quantizing.

[0077] As mentioned above, the block reconstruction error reflects the degree of difference in function between the first network block after quantization (initial quantized network block) and the first network block without quantization, and the network error measures the degree of difference in performance between the entire lightweight network after quantization and the original lightweight network without quantization. Therefore, the optimization goal can be set to minimize the block reconstruction error and the network error, so that the quantized network can be close to the original network without quantization in overall performance, and each network block can maintain its original function as much as possible after quantization.

[0078] In an embodiment of the present application, during the optimization process, an optimization algorithm such as gradient descent can be used to calculate the gradient of the block reconstruction error and the network error with respect to the initialization equalization vector, the initialization deviation absorption vector, and the initialization quantization coefficient, and then update the parameter value according to the gradient direction and size. After multiple iterations of optimization, when the error reaches a minimum value or meets certain convergence conditions, the target equalization vector, the target deviation absorption vector, and the target quantization coefficient are generated. The second quantized network block obtained at this time is the final optimized result. This quantized network block has lower storage requirements and higher computing efficiency while maintaining certain performance, and is more suitable for deployment and operation in resource-constrained environments.

[0079] In some embodiments, the method may further include:

[0080] Determining bias parameters of a second quantization network block;

[0081] Taking minimizing the block reconstruction error and network error as the optimization goal, the bias parameters are optimized to generate the target bias parameters and obtain the third quantization network block.

[0082] In an embodiment of the present application, after the first network block is quantized to generate a second quantized network block, its existing bias parameters can be determined. These bias parameters may be determined in the process of quantizing the network and optimizing related parameters, and they play an important role in the calculation and function realization of the network. For example, in the calculation of a neural network, the output calculation of a neuron is usually the weighted sum of the input data and the weight plus the bias parameter. The bias parameter can adjust the activation threshold of the neuron and affect the output result of the network. Therefore, it is necessary to quantize the bias parameter as well.

[0083] In an embodiment of the present application, for the bias parameter, the block reconstruction error and the network error are again used as the loss function, and the target bias parameter is generated by optimizing with minimizing the loss function as the optimization goal.

[0084] By further optimizing the bias parameters, the second quantization network block can more accurately compensate for problems such as activation mean shift caused by quantization, further improving the quantization accuracy.

[0085] In some embodiments, the method may further include:

[0086] Constructing an objective function based on a block reconstruction error, a network error, and an adaptive rounding function, wherein the adaptive rounding function is used to determine a rounding method for weights of a second quantized network block;

[0087] The optimization goal is to minimize the block reconstruction error and network error, optimize the bias parameters, and generate the target bias parameters including:

[0088] Taking minimizing the objective function as the optimization goal, the bias parameters are optimized to generate target bias parameters and target adaptive rounding function.

[0089] When quantizing network weights, high-precision weights need to be converted into low-precision data types (such as 8-bit integers, 4-bit integers, etc.), and this process involves rounding operations. The adaptive rounding function can be used to determine the rounding method of the weights of the second quantized network block. It is different from the traditional fixed rounding method (such as rounding up, rounding down, or rounding up). Instead, it introduces mechanisms such as trainable tensors, which can dynamically adjust the rounding method according to the characteristics of the data and the needs of the network to improve the quantization accuracy. For example, the adaptive rounding function may determine the most appropriate rounding strategy based on factors such as the input and output data distribution of the current network block and the numerical range of the weights, so that the weights can be converted more accurately during the quantization process and reduce the quantization error.

[0090] Therefore, in another embodiment of the present application, by combining the three parts of block reconstruction error, network error and adaptive rounding function into an objective function, multiple key factors in the quantization process can be comprehensively considered under a unified framework. By minimizing this objective function, it is possible to simultaneously maintain the network block function (by controlling the block reconstruction error), improve the performance of the entire network (by controlling the network error) and optimize the weight rounding method (through an adaptive rounding function), thereby achieving a better quantization effect, so that the quantized network can meet the requirements of storage and computing efficiency in a resource-constrained environment, while maintaining good performance.

[0091] After multiple iterations of optimization, when the objective function reaches the minimum value or meets certain convergence conditions, the bias parameters at this time become the target bias parameters. These target bias parameters are bias settings that can achieve better quantization effects while minimizing the objective function after optimization. They can better adapt to the quantized network calculations and further reduce the impact of quantization errors on network performance. At the same time, they also play a more precise role in adjusting the output of the network block based on comprehensive consideration of the block reconstruction error, network error, and adaptive rounding function.

[0092] Moreover, while optimizing the bias parameters, since the objective function contains an adaptive rounding function and it is related to network elements such as the bias parameters, the adaptive rounding function will also change with the optimization of the bias parameters during the optimization process. When the objective function reaches the minimum value or meets the convergence condition, the adaptive rounding function at this time becomes the target adaptive rounding function. It is a weight rounding method that can achieve better quantization effect based on comprehensive consideration of factors such as block reconstruction error, network error, and bias parameter optimization. It can further improve the quantization accuracy, so that the weights can be converted more accurately during the quantization process, reduce the quantization error, and thus provide a better quantization solution for the entire quantization network.

[0093] In some embodiments, initializing the quantization coefficients includes initializing the weight quantization coefficients and initializing the activation quantization coefficients.

[0094] In the process of quantization-aware equalization, there is a mutually coupled relationship between the weight quantization coefficient and the equalization vector. The sudden change of the weight quantization coefficient may cause a drastic change in the loss function composed of block reconstruction error and network error, and will seriously interfere with the optimization process of the equalization vector.

[0095] In some embodiments, the initialization equalization vector, the initialization deviation absorption vector, and the initialization quantization coefficient of the first network block are optimized with minimization of the block reconstruction error and the network error as the optimization goal, and the generation of the target equalization vector, the target deviation absorption vector, and the target quantization coefficient can be specifically implemented as follows:

[0096] The initialized quantization weights are quantized and fixed to initial values, and the initialization equalization vector, the initialization deviation absorption vector, and the initialization activation quantization coefficient are optimized with minimization of the difference between the first output data and the second output data as the optimization goal.

[0097] In an embodiment of the present application, the initialization weight quantization coefficient can be used for the relevant settings of the quantization operation of the weights in the network block. In the process of converting the weights in the network from high-precision data types (such as 32-bit floating point numbers) to low-precision data types (such as 8-bit integers, etc.), the initialization weight quantization coefficient plays a key role. Through a specific calculation method (for example, determined according to factors such as the numerical range of the weights, the quantization bit width, etc.), it can determine how to map the original weights to a low-precision quantization representation so as to reduce storage requirements and computational complexity in subsequent calculations.

[0098] Initializing the activation quantization coefficient can set the relevant parameters for quantizing the activation values ​​in the network block. In the calculation process of the neural network, the activation value is the output result of the neuron after being processed by the activation function. Similar to the weight, in order to achieve quantitative optimization of the network, the activation value also needs to be converted from a high-precision data type to a low-precision data type. Initializing the activation quantization coefficient is to determine the specific quantization method based on the characteristics of the activation value (such as distribution, value range, etc.) and the target requirements of quantization (such as the determined quantization bit width), so that the activation value can participate in the quantized network calculation in a suitable low-precision form, which also helps to reduce storage requirements and improve computing efficiency.

[0099] In the embodiment of the present application, during the entire optimization process, the weight quantization coefficients may be initialized to be fixedly set to their initial values ​​and not participate in the optimization adjustment, thereby improving the stability of the optimization process.

[0100] In some embodiments, during the optimization process of the initialization activation quantization coefficient, the gradient of the initialization activation quantization coefficient is calculated by the following formula:

[0101]

[0102] Where s represents the equalization vector, x represents the input data, n represents the quantization bit width, and q(·) represents the quantization function.

[0103] Figure 2A schematic diagram of a lightweight network quantization method provided in an embodiment of the present application is shown.

[0104] like Figure 2 As shown, 201 may represent a lightweight network, and 202 may represent a quantized lightweight network generated by quantizing the lightweight network.

[0105] The quantized lightweight network 202 includes a first quantized network block 2021, a second quantized network block 2022, a third network block 2013, and a fourth network block 2014. The first quantized network block 2021 and the second quantized network block 2022 can be obtained by quantizing the first network block 2011 and the second network block 2012 in the lightweight network 201. The third network block 2013 and the fourth network block 2014 have not been quantized.

[0106] In the embodiment of the present application, the third network block 2013 may be determined as the first network block to be quantized.

[0107] The third network block 2013 may include a first network layer and a second network layer.

[0108] After inputting data x' to the first network layer, h = f(w (1) ·x'+b 1 ), where w (1) Can refer to the weight parameter of the first network layer, b 1 It can represent the bias parameters of the first network layer.

[0109] The output h' of the first network layer can be further input into the second network layer, and the second network layer can output y=f(w (2) ·h+b 2 ), where w (2) It can represent the weight parameter of the second network layer, b 2 It can represent the bias parameters of the second network layer.

[0110] After constructing the initial equalization vector, initializing the deviation absorption vector, initializing the weight quantization coefficient, and initializing the activation quantization coefficient, the first network layer can be quantized first. After quantization, the output of the first network layer can be The output of the second network layer can be y'=(q(w (2) ·S,S w(2) )q(h',S h‘ )+b 2 +w (2) *C).

[0112] After the quantization of the third network block 2013 is completed, a third quantization network block (not shown in the figure) can be generated.

[0113] Furthermore, the fourth network block 2014 may be further quantized to obtain a fourth quantized network block (not shown in the figure) to complete the quantization of the lightweight network 201 .

[0114] Figure 3 A block diagram of a quantization device for a lightweight network provided in an embodiment of the present application is shown. The lightweight network includes at least one network block, and the device includes:

[0115] A request acquisition module 301, used to acquire a quantization request for a lightweight network;

[0116] A network block determining module 302, configured to determine, in response to a quantization request, a first network block to be quantized from at least one network block;

[0117] A vector construction module 303, configured to construct an initialization equalization vector, an initialization deviation absorption vector, and an initialization quantization coefficient of the first network block;

[0118] A first quantization module 304 is used to quantize the first network block by using the initialization equalization vector, the initialization deviation absorption vector and the initialization quantization coefficient to obtain a first quantized network block;

[0119] An output acquisition module 305, used to acquire the first output data of the lightweight network, the second output data of the quantized lightweight network after quantization, the third output data of the first network block, and the fourth output data of the first quantized network block;

[0120] An error construction module 306 is used to construct a network error based on the first output data and the second output data, and to construct a block reconstruction error based on the third output data and the fourth output data;

[0121] The optimization module 307 is used to optimize the initialization equalization vector, the initialization deviation absorption vector and the initialization quantization coefficient of the first network block with the optimization goal of minimizing the block reconstruction error and the network error, generate the target equalization vector, the target deviation absorption vector and the target quantization coefficient, and obtain the second quantized network block after quantization.

[0122] In some embodiments, the apparatus further comprises:

[0123] A first bias determination module, used to determine a bias parameter of a second quantization network block;

[0124] The bias optimization module is used to optimize the bias parameters with the minimization of the block reconstruction error and the network error as the optimization goal, generate the target bias parameters, and obtain the third quantization network block.

[0125] In some embodiments, initializing the quantization coefficients includes initializing the weight quantization coefficients and initializing the activation quantization coefficients.

[0126] In some embodiments, the optimization module 307 is specifically used to:

[0127] The initialized quantization weights are quantized and fixed to initial values, and the initialization equalization vector, the initialization deviation absorption vector, and the initialization activation quantization coefficient are optimized with minimization of the difference between the first output data and the second output data as the optimization goal.

[0128] In some embodiments, the output acquisition module 305 is specifically used to:

[0129] Get input data;

[0130] Inputting the input data into the lightweight network to obtain first output data;

[0131] Input data is input into a quantized lightweight network to obtain second output data, wherein the quantized lightweight network is obtained by quantizing at least one network block in the lightweight network.

[0132] In some embodiments, during the optimization process of the initialization activation quantization coefficient, the gradient of the initialization activation quantization coefficient is calculated by the following formula:

[0133]

[0134] Where s represents the equalization vector, x represents the input data, n represents the quantization bit width, and q(·) represents the quantization function.

[0135] In some embodiments, the device further comprises:

[0136] An objective function construction module, used to construct an objective function based on a block reconstruction error, a network error and an adaptive rounding function, wherein the adaptive rounding function is used to determine a rounding method for a weight of a second quantized network block;

[0137] In some embodiments, the bias optimization module is specifically used to:

[0138] Taking minimizing the objective function as the optimization goal, the bias parameters are optimized to generate target bias parameters and target adaptive rounding function.

[0139] In some embodiments, the error construction module 306 is specifically configured to:

[0140] Determine the lightweight network network type;

[0141] Determine the distance metric function that matches the network type;

[0142] A network error is constructed based on the first output data and the second output data using a distance metric function.

[0143] Figure 3 The quantization device of the lightweight network can perform Figure 1 The implementation principle and technical effect of the lightweight network quantization method described in the illustrated embodiment will not be described in detail. The specific manner in which each module and unit performs operations in the lightweight network quantization device in the above embodiment has been described in detail in the embodiment of the method, and will not be described in detail here.

[0144] In a possible design, the quantization device of the lightweight network provided in the embodiment of the present application can be implemented as a computing device, such as Figure 4 As shown, the computing device may include a storage component 401 and a processing component 402;

[0145] The storage component 401 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 402 to implement the lightweight network quantization method provided in the embodiment of the present application.

[0146] Of course, the computing device may also include other components, such as input / output interfaces, communication components, etc. The input / output interface provides an interface between the processing component and the peripheral interface module, which may be an output device, an input device, etc. The communication component is configured to facilitate wired or wireless communication between the computing device and other devices.

[0147] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.

[0148] When the computing device is a physical device, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device.

[0149] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, which, when executed by a computer, can implement the lightweight network quantization method provided in the embodiment of the present application.

[0150] The embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a computer, can implement the lightweight network quantization method provided in the embodiment of the present application.

[0151] The processing components in the above corresponding embodiments may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing components may also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above method.

[0152] The storage component is configured to store various types of data to support operations in the device. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0153] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0154] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0155] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A lightweight network quantization method, characterized in that: The lightweight network includes at least one network block, and the method includes: Obtaining a quantization request for the lightweight network; In response to the quantization request, determining a first network block to be quantized from the at least one network block; Constructing an initialization equalization vector, an initialization deviation absorption vector, and an initialization quantization coefficient of the first network block; quantizing the first network block using the initialized equalization vector, the initialized deviation absorption vector, and the initialized quantization coefficient to obtain a first quantized network block; Acquire first output data of the lightweight network, second output data of the quantized lightweight network after quantization, third output data of the first network block, and fourth output data of the first quantized network block; Constructing a network error based on the first output data and the second output data, and constructing a block reconstruction error based on the third output data and the fourth output data; Taking minimizing the block reconstruction error and the network error as the optimization goal, the initialization equalization vector, the initialization deviation absorption vector and the initialization quantization coefficient of the first network block are optimized to generate a target equalization vector, a target deviation absorption vector and a target quantization coefficient, and obtain a second quantized network block that has been quantized.

2. The method according to claim 1, characterized in that The method further comprises: Determining a bias parameter of the second quantization network block; Taking minimizing the block reconstruction error and the network error as the optimization goal, the bias parameter is optimized to generate a target bias parameter, and a third quantization network block is obtained.

3. The method according to claim 1, characterized in that: The initialization quantization coefficients include initialization weight quantization coefficients and initialization activation quantization coefficients; The optimizing step of optimizing the initialization equalization vector, the initialization deviation absorption vector, and the initialization quantization coefficient of the first network block with minimizing the block reconstruction error and the network error to generate a target equalization vector, a target deviation absorption vector, and a target quantization coefficient includes: The initialized quantization weight is quantized and fixed to an initial value, and the initialized equalization vector, the initialized deviation absorption vector, and the initialized activation quantization coefficient are optimized with minimizing the difference between the first output data and the second output data as an optimization goal.

4. The method according to claim 1, characterized in that: The step of obtaining the first output data of the lightweight network and the quantized second output data of the lightweight network includes: Get input data; Inputting the input data into the lightweight network to obtain the first output data; The input data is input into the quantized lightweight network to obtain the second output data, wherein the quantized lightweight network is obtained by quantizing at least one network block in the lightweight network.

5. The method according to claim 3, characterized in that: In the process of optimizing the initial activation quantization coefficient, the gradient of the initial activation quantization coefficient is calculated by the following formula: ; in, represents the equalization vector, Represents input data, Indicates the quantization bit width, Represents a quantization function.

6. The method according to claim 2, characterized in that: The method further comprises: Constructing an objective function based on the block reconstruction error, the network error and an adaptive rounding function, wherein the adaptive rounding function is used to determine a rounding method for the weight of the second quantization network block; The method of optimizing the bias parameter with minimizing the block reconstruction error and the network error as the optimization goal, and generating the target bias parameter includes: Taking minimizing the objective function as the optimization goal, the bias parameter is optimized to generate a target bias parameter and a target adaptive rounding function.

7. The method according to claim 1, characterized in that The constructing a network error based on the first output data and the second output data comprises: Determining the network type of the lightweight network; determining a distance metric function that matches the network type; The network error is constructed based on the first output data and the second output data using the distance metric function.

8. A computing device, characterized in that including a processing component and a storage component; The storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the lightweight network quantization method as described in any one of claims 1 to 7.

9. A computer storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a computer, the lightweight network quantization method according to any one of claims 1 to 7 is implemented.

10. A computer program product, characterized in that The computer program product comprises computer program code, and when the computer program code is executed by a computer, the lightweight network quantization method according to any one of claims 1 to 7 is implemented.