Model quantization method, apparatus, device, storage medium, and product
Patent Information
- Application Number
- CN202310745288.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-21
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-06-21
AI Technical Summary
[0002]目前,深度学习模型已经在各行各业中占据了主导应用,然而为了提高深度学习模型的精度,深度学习模型往往比较大,从而导致深度学习模型很难部署到资源受限的边缘设备上(例如,手机或者智能手表等);因此,需要对深度学习模型进行量化
[0017]在本申请实施例中,由于第一样本数据中包括了伪造数据和代理数据;而仿造数据是基于深度学习模型对应的标签生成的,因此仿造数据是符合标签要求的数据;而代理数据是真实获取到的开源数据,而开源数据具有丰富的特征;因此代理数据在数据特征方面具有优势;因此,基于第一样本数据,对深度学习模型进行量化是结合了仿造数据和代理数据的优势,能够提高量化深度学习模型的精度。
Smart Images

Figure CN116776924B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning technology, and in particular to a model quantization method, apparatus, device, storage medium, and product. Background Technology
[0002] Currently, deep learning models have become the dominant application in various industries. However, in order to improve the accuracy of deep learning models, they are often quite large, making it difficult to deploy them on resource-constrained edge devices (such as mobile phones or smartwatches). Therefore, it is necessary to quantize deep learning models. Summary of the Invention
[0003] This application provides a model quantization method, apparatus, device, storage medium, and product, which can improve the accuracy of quantizing deep learning models. The technical solution is as follows:
[0004] On the one hand, a model quantization method is provided, the method comprising:
[0005] Based on the target label and Gaussian noise corresponding to the deep learning model, a generator generates fake data, wherein the label of the fake data is the target label.
[0006] Based on multiple first-batch normalized BN layers in the deep learning model, proxy data is determined from the proxy data set, which includes multiple proxy data, and the proxy data is open-source real data;
[0007] The forged data and the proxy data are mixed to obtain the first sample data;
[0008] The deep learning model is quantized based on the first sample data.
[0009] On the other hand, a model quantization apparatus is provided, the apparatus comprising:
[0010] The generation module is used to generate fake data based on the target label and Gaussian noise corresponding to the deep learning model, wherein the label of the fake data is the target label.
[0011] The first determining module is used to determine proxy data from a proxy data set based on multiple first batch normalized BN layers in the deep learning model. The proxy data set includes multiple proxy data, and the proxy data is open-source real data.
[0012] A mixing module is used to mix the forged data and the proxy data to obtain first sample data;
[0013] The quantization module is used to quantize the deep learning model based on the first sample data.
[0014] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one piece of program code, the at least one piece of program code being loaded and executed by the processor to implement the model quantization method described above.
[0015] On the other hand, a computer-readable storage medium is provided, wherein at least one piece of program code is stored in the storage medium, the at least one piece of program code being loaded and executed by a processor to implement the model quantization method described above.
[0016] On the other hand, a computer program product is provided, which stores at least one piece of program code for execution by a processor to implement the model quantization method described above.
[0017] In this embodiment, the first sample data includes both forged data and proxy data. The forged data is generated based on the labels corresponding to the deep learning model, and therefore conforms to the label requirements. The proxy data is real open-source data, which has rich features. Therefore, the proxy data has advantages in terms of data features. Thus, quantizing the deep learning model based on the first sample data combines the advantages of both forged data and proxy data, and can improve the accuracy of quantizing the deep learning model. Attached Figure Description
[0018] Figure 1 A schematic diagram illustrating the implementation environment of the model quantization method shown in an exemplary embodiment of this application is provided.
[0019] Figure 2 A flowchart illustrating a model quantization method in an exemplary embodiment of this application is shown;
[0020] Figure 3 A schematic diagram illustrating a model quantization method according to an exemplary embodiment of this application is shown;
[0021] Figure 4 A flowchart illustrating a model quantization method in an exemplary embodiment of this application is shown;
[0022] Figure 5 A flowchart illustrating a model quantization method in an exemplary embodiment of this application is shown;
[0023] Figure 6 A flowchart illustrating a model quantization method in an exemplary embodiment of this application is shown;
[0024] Figure 7 A flowchart illustrating a model quantization method in an exemplary embodiment of this application is shown;
[0025] Figure 8 A block diagram illustrating a model quantization apparatus as shown in an exemplary embodiment of this application is presented;
[0026] Figure 9 A block diagram of a terminal illustrated in an exemplary embodiment of this application is shown;
[0027] Figure 10 A block diagram of a server illustrated in an exemplary embodiment of this application is shown. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0029] In this article, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0030] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the forged data and proxy data involved in this application were obtained with full authorization.
[0031] Please refer to Figure 1 This diagram illustrates an implementation environment for a model quantization method according to an exemplary embodiment of this application. See also... Figure 1 The implementation environment includes a computer device 101 and a terminal 102. The computer device 101 is used to quantize the deep learning model and deploy the quantized deep learning model to the terminal 102. This embodiment of the application is a data-free quantization method for the model, that is, a method that does not rely on the original sample data used to train the deep learning model for quantization. The computer device 101 can be a terminal or a server; Figure 1 The example described uses computer device 101 as a server. Terminal 102 can be a resource-constrained edge terminal, such as a mobile phone or wearable device.
[0032] In some embodiments, the deep learning model can be an image recognition model, that is, the computer device 101 quantizes the image recognition model and deploys the quantized image recognition model to the terminal 102, enabling the terminal 102 to perform image recognition using the quantized image recognition model. In some embodiments, the deep learning model can be a speech recognition model, that is, the computer device 101 quantizes the speech recognition model and deploys the quantized speech recognition model to the terminal 102, enabling the terminal 102 to perform speech recognition using the quantized speech recognition model; for example, the terminal 102 recognizes voice control commands in the speech signal, thereby executing the operation corresponding to the voice control command, and thus achieving the purpose of controlling the terminal 102 through the voice signal. In some embodiments, the deep learning model can be an image classification model, that is, the computer device 101 quantizes the image classification model and deploys the quantized image classification model to the terminal 102, enabling the terminal 102 to perform image classification using the quantized image classification model. Furthermore, the model quantization method provided in this application embodiment can quantize any deep learning model; the above are merely examples of deep learning models and do not limit the scope of deep learning models.
[0033] Please refer to Figure 2 The diagram illustrates a flowchart of a model quantization method according to an exemplary embodiment of this application. (Reference) Figure 2 The method includes:
[0034] Step 201: Based on the target label and Gaussian noise corresponding to the deep learning model, generate fake data through a generator. The label of the fake data is the target label.
[0035] The forged data is the training data corresponding to the deep learning model; in some embodiments, when the deep learning model is an image recognition model, the forged data can be a forged image; when the deep learning model is a speech recognition model, the forged data can be a forged speech signal; when the deep learning model is an image classification model, the forged data can be a forged image.
[0036] In some embodiments, this step can be implemented by the following steps (1) to (3):
[0037] (1) Determine the target label from the label set. The label set stores the label of the second sample data. The second sample data is the sample data for training the deep learning model.
[0038] For example, a label y is sampled from the label set {0,1,…,N-1}, and this label y is the target label, where N is the number of labels. Here, {0,1,…,N-1} in the label set represents the numerical identifiers corresponding to the labels of the second sample data. For instance, if the deep learning model is an image recognition model, and the labels of the image recognition model are "dog," "cat," and "rabbit," then the label set is {0,1,2,3}, where 0 corresponds to the label "dog," 1 corresponds to the label "cat," 2 corresponds to the label "rabbit," and 3 corresponds to the label "mouse."
[0039] (2) Determine the Gaussian noise.
[0040] The computer device randomly selects a Gaussian noise from the normal distribution parameters of the labels; for example, the computer device samples a random vector z, z∈R from N(0,1). 1×N R represents a natural number.
[0041] (3) Based on the target label and Gaussian noise, a generator is used to generate fake data, and the label of the fake data is the target label.
[0042] The computer device inputs the target label and Gaussian noise into the generator and outputs forged data. High-speed noise is used to increase the difference between the generated forged data; that is, there are multiple forged data sets, and all forged data sets are labeled with the target label, but each forged data set is different. The generator can generate forged data based on the target label and Gaussian noise using the following formula: Formula 1: I FD =G(z|y),z~N(0,1); among them, I FD This indicates fabricated data; G(·) represents the generator, y represents the target label, and z represents Gaussian noise. For example, see [link to example]. Figure 3 The computer device selects the target label and Gaussian noise as y and z, respectively. y+z is input into the generator G, and the output is the forged data I. FD I FD ∈R c×h×w Where c represents I FD The number of channels, where h represents the height and w represents the width.
[0043] Step 202: Based on multiple first-batch normalized BN layers in the deep learning model, determine the proxy data from the proxy data set. The proxy data set includes multiple proxy data, which are open-source real data.
[0044] The proxy data is the training data corresponding to the deep learning model; in some embodiments, when the deep learning model is an image recognition model, the proxy data can be an open-source real image; when the deep learning model is a speech recognition model, the proxy data can be an open-source real speech signal; when the deep learning model is an image classification model, the proxy data can be an open-source real image.
[0045] Since not all open-source data, when used as proxy data, can ultimately improve the accuracy of the dataless quantization method, selecting suitable proxy data is crucial. In this embodiment, a Batch Normalization (BN) layer is used to select proxy data, thereby improving the final accuracy of the dataless quantization method.
[0046] Step 203: Mix the forged data and the proxy data to obtain the first sample data.
[0047] There are multiple instances of both forged data and proxy data. In some embodiments, the computer device can directly mix the forged data and proxy data determined in step 201 to obtain the first sample data. In other embodiments, the computer device can control the mixing ratio of forged data and proxy data; correspondingly, step 202 can be: determining a first hyperparameter, which is used to constrain the mixing ratio of forged data and proxy data; and mixing the forged data and proxy data based on the first hyperparameter to obtain the first sample data.
[0048] In some embodiments, the first hyperparameter is the weight of the forged data; then, the step of mixing the forged data and the proxy data based on the first hyperparameter to obtain the first sample data can be achieved by the following formula two: Formula two: in, Let I represent the first sample data, γ represent the first hyperparameter, and I represent the first sample data. FD Indicates falsified data, I PD The first hyperparameter represents proxy data, and Concat represents the blending function. The first hyperparameter can be pre-configured in the deep learning model; and the first hyperparameter can be set and changed as needed. In this embodiment, the first hyperparameter is not specifically limited.
[0049] In this embodiment, proxy data can be inserted into the dataless quantization method at minimal cost, thereby reducing the cost of model quantization.
[0050] Step 204: Quantize the deep learning model based on the first sample data.
[0051] The computer equipment is based on a deep learning model to determine a quantization model. The quantization model is a model obtained by quantizing the deep learning model. Based on the first sample data, the quantization model is optimized in multiple rounds until the iteration stopping condition is met, and finally the quantization process of the deep learning model is completed.
[0052] Since the first sample data includes both fake data and proxy data, and the fake data is generated based on the labels corresponding to the deep learning model, it meets the label requirements. The proxy data is real open-source data, which has rich features. Therefore, the proxy data has an advantage in terms of data features. Thus, quantizing the deep learning model based on the first sample data combines the advantages of both fake and proxy data, which can improve the accuracy of quantizing the deep learning model.
[0053] Please refer to Figure 4 The diagram illustrates a flowchart of a model quantization method according to an exemplary embodiment of this application. (Reference) Figure 4 The method includes:
[0054] Step 401: The computer device generates fake data based on the target label and Gaussian noise corresponding to the deep learning model through a generator. The label of the fake data is the target label.
[0055] In some embodiments, this step is the same as step 201, and will not be described again here.
[0056] Step 402: The computer device determines a proxy data set, which includes multiple candidate proxy data sets.
[0057] The candidate proxy data is open-source data; correspondingly, the computer device determines the proxy data set from the open-source data set. In some embodiments, the computer device determines open-source data of the data type corresponding to the deep learning model from the open-source data set, and the determined open-source data forms the proxy data set. For example, if the deep learning model is an image recognition model, the computer device determines multiple images from the open-source data set to form the proxy data set; as another example, if the deep learning model is a speech recognition model, the computer device determines multiple speech signals from the open-source data set to form the proxy data set; as yet another example, if the deep learning model is an image classification model, the computer device determines multiple images from the open-source data set to form the proxy data set.
[0058] In some embodiments, the computer device randomly selects multiple open-source data from an open-source dataset and combines these multiple open-source data into a proxy dataset, thereby ensuring the randomness of the data and improving the accuracy of model quantization.
[0059] Step 403: The computer device determines the normalized distances corresponding to multiple candidate proxy data in the proxy data set based on multiple first-batch normalized BN layers in the deep learning model. The normalized distances corresponding to the candidate proxy data are used to characterize the correlation between the candidate proxy data and the second sample data, which are the sample data used to train the deep learning model.
[0060] A larger normalized distance between the candidate surrogate data and the second sample data indicates a weaker correlation, while a smaller normalized distance indicates a stronger correlation. It's important to note that when quantifying the deep learning model, the second sample data is not actually obtained; the correlation is represented solely by the normalized distance.
[0061] In some embodiments, this step can be implemented by the following steps (1) to (3):
[0062] (1) For any candidate agent data, the computer device inputs the candidate agent data into the deep learning model.
[0063] The computer device can input multiple candidate agent data from the agent data set into the deep learning model at once, and the deep learning model will sequentially perform step (2) on the multiple candidate agent data.
[0064] (2) The computer device determines the first mean and first variance of the candidate agent data in multiple BN layers through multiple first BN layers in the deep learning model.
[0065] Deep learning models include multiple first batch normalization (BN) layers, which are used to determine the normalized distance.
[0066] (3) The computer device determines the normalized distance corresponding to the candidate agent data based on the first mean and first variance of the candidate agent data in multiple first BN layers and the second mean and second variance pre-stored in multiple first BN layers.
[0067] This step can be achieved through the following steps (3-1) to (3-2), including:
[0068] (3-1) For any first BN layer, the computer device determines a first normalized distance based on the first mean of the candidate agent data in the first BN layer and the second mean pre-stored in the first BN layer, and determines a second normalized distance based on the first variance of the candidate agent data in the first BN layer and the second variance pre-stored in the first BN layer, and determines the sum of the first normalized distance and the second normalized distance to obtain a third normalized distance.
[0069] (3-2) The computer device determines the average value of the third normalized distances corresponding to multiple first BN layers to obtain the normalized distances corresponding to the candidate agent data.
[0070] In some embodiments, steps (3-1) and (3-2) can be implemented using the following formula:
[0071] Formula 3:
[0072] Where, d BN This represents the normalized distance corresponding to the candidate proxy data. and Let μ1 and σ1 be the mean and variance calculated for the candidate proxy data in the i-th first BN layer, respectively, and let μ1 and σ1 be the corresponding mean and variance stored in the i-th first BN layer of the deep learning model. L1 represents the total number of first BN layers in the deep learning model.
[0073] Step 404: The computer device determines the proxy data whose normalized distance meets the condition from the proxy data set based on the normalized distance corresponding to multiple candidate proxy data in the proxy data set.
[0074] A larger normalized distance between candidate proxy data indicates a weaker correlation between the candidate proxy data and the second sample data, while a smaller normalized distance indicates a stronger correlation between the candidate proxy data and the second sample data. Therefore, the computer device determines multiple proxy data sets with the smallest corresponding normalized distances from the proxy data set based on the normalized distances between multiple candidate proxy data sets; or, it determines multiple proxy data sets with corresponding normalized distances less than a preset distance from the proxy data set.
[0075] Step 405: The computer device mixes the forged data and the proxy data to obtain the first sample data.
[0076] In some embodiments, this step is the same as step 203, and will not be described again here.
[0077] Step 406: The computer device quantizes the deep learning model based on the first sample data.
[0078] In some embodiments, this step is the same as step 204, and will not be described again here.
[0079] In this embodiment, since not all open-source data used as proxy data can ultimately improve the accuracy of the dataless quantization method, selecting suitable proxy data is crucial. This embodiment utilizes a Batch Normalization (BN) layer to select proxy data, thereby improving the final accuracy of the dataless quantization method.
[0080] Please refer to Figure 5 The diagram illustrates a flowchart of a model quantization method according to an exemplary embodiment of this application. (Reference) Figure 5 The method includes:
[0081] Step 501: The computer device generates fake data based on the target label and Gaussian noise corresponding to the deep learning model through a generator. The label of the fake data is the target label.
[0082] In some embodiments, this step is the same as step 201, and will not be described again here.
[0083] Step 502: The computer device determines proxy data from the proxy data set based on multiple first-batch normalized BN layers in the deep learning model. The proxy data set includes multiple proxy data, which are open-source real data.
[0084] In some embodiments, this step can be implemented through steps 402-404, which will not be described in detail here.
[0085] Step 503: The computer device labels the agent data.
[0086] Since the forged data is generated based on the target labels of the deep learning model, the forged data and the original data (the second sample data for training the deep learning model) have the same labels, and the forged data is labeled with its corresponding labels; while the proxy data is open source data obtained from nature, and the open source data may not have the same labels as the original data (the second sample data for training the deep learning model); therefore, it is necessary to construct labels for the proxy data, and then label the proxy data with its corresponding labels.
[0087] In some embodiments, the computer device inputs surrogate data into the deep learning model and outputs labels for the surrogate data, thus annotating the surrogate data with these labels. Since the deep learning model is quantized using surrogate data, it can determine a label from its multiple corresponding labels as the label for the surrogate data. This determined label has the same label as the original data (the second sample data used to train the deep learning model), improving the accuracy of subsequent quantization.
[0088] In some embodiments, the labels generated by the deep learning model for the proxy data can be achieved using the following formula: Formula 4: in, The label represents the proxy data, F(·) represents the original deep learning model, and I PD This represents proxy data. For example, see [reference to...] Figure 3 Computer equipment will I PDInput the original deep learning model and output the labels of the proxy data.
[0089] Step 504: The computer device mixes the forged data and the proxy data to obtain the first sample data.
[0090] In some embodiments, this step is the same as step 201, and will not be repeated here. The process of quantizing the deep learning model by the computer device is as follows: the computer device determines the quantization model, which is the model obtained by quantizing the deep learning model. Then, through multiple rounds of iterative optimization, the quantization model is optimized to finally complete the quantization of the deep learning model; each round of iterative optimization is implemented through steps 505-508. In some embodiments, the process of the computer device determining the quantization model can be: the computer device quantizes the floating-point data in the deep learning model into fixed-point data to obtain the quantization model. The computer device can also determine the quantization model through other quantization methods. In the embodiments of this application, the process of the computer device determining the quantization model is not specifically limited.
[0091] Step 505: For any round of iterative optimization, the computer device determines the first loss function based on the first sample data and the label of the first sample data through a quantization model.
[0092] A quantization model is a model obtained by quantizing a deep learning model. Based on the first sample data and its labels, the computer device uses the quantization model and the cross-entropy loss function to determine the first loss function. For example, the computer device inputs the labeled first sample data into the quantization model and determines the first loss function using the cross-entropy loss function. This process can be achieved using Formula 5: Formula 5: Among them, L CE (Q) represents the first loss function, and Q(·) represents the quantization model. This represents the first sample data. Let represent the label of the first sample data, and CE(·,·) represent the cross-entropy loss function.
[0093] Step 506: The computer device determines the second loss function based on the first sample data, using a quantization model and a deep learning model.
[0094] The computer device, based on the first sample data, determines the second loss function using a quantization model and a deep learning model, based on the relative entropy function. For example, the computer device inputs the first sample data into the quantization model and the deep learning model respectively, and determines the second loss function using the relative entropy function. This process can be achieved using Formula Six: Formula Six: Among them, L KD(Q) represents the second loss function, and Q(·) represents the quantization model. Let F(·) represent the first sample data, and F(·) represent the deep learning model.
[0095] Step 507: The computer device performs a weighted summation of the first loss function and the second loss function based on the second hyperparameter to obtain the first total loss function. The second hyperparameter is used to constrain the weights of the second loss function.
[0096] In some embodiments, the computer device, based on a second hyperparameter, performs a weighted summation of the first loss function and the second loss function using the following formula (Formula 7) to obtain the first total loss function. Formula 7: L Q =L CE (Q)+β·L KD (Q); where L Q Let L represent the first total loss function, β represent the second hyperparameter, and L represent the second total loss function. KD (Q) represents the second loss function, L CE (Q) represents the first loss function.
[0097] In some embodiments, the computer device, based on a second hyperparameter and a fourth hyperparameter, performs a weighted summation of the first loss function and the second loss function using the following formula (Equation 8) to obtain a first total loss function. The second hyperparameter is used to constrain the weights of the second loss function, and the fourth hyperparameter is used to constrain the weights of the first loss function. Formula 8: L Q =εL CE (Q)+β·L KD (Q); where L Q Let L represent the first total loss function, β represent the second hyperparameter, ε represent the fourth hyperparameter, and L represent the fourth hyperparameter. KD (Q) represents the second loss function, L CE (Q) represents the first loss function.
[0098] In this embodiment, the total loss function is obtained by weighted summation of the loss functions obtained by the two methods. This approach balances the impact of the two loss functions on the overall loss function, thereby improving the accuracy of model quantization based on the total loss function.
[0099] Step 508: The computer device performs one round of optimization on the quantization model based on the first total loss function.
[0100] If the value corresponding to the first total loss function is less than a first preset threshold, the optimization of the quantization model ends; if the value corresponding to the first total loss function is not less than the first preset threshold, the parameter values of the quantization model are adjusted, and then the next round of optimization is performed. Alternatively, if the difference between the first total loss function and the value corresponding to the first total loss function in the previous round is less than a second preset threshold, the optimization of the quantization model ends; if the difference between the first total loss function and the value corresponding to the first total loss function in the previous round is not less than the second preset threshold, then the next round of optimization is performed until the quantization model optimization is complete. For example, continue to refer to... Figure 3 Computer devices will act as agents for data I PD and falsified data I FD Input into the quantization model to optimize the quantization model.
[0101] In this embodiment, the first and second loss functions are determined directly using first sample data combining proxy data and forged data. This allows the method of this application to be incorporated into other data-free quantization methods at zero cost, thereby helping to improve the quantization accuracy of other data-free quantization methods. Furthermore, the data-free quantization method based on proxy data proposed in this application can alleviate the problem of a lack of original datasets in data-free quantization methods, which further leads to degradation in model quantization accuracy. Simultaneously, this embodiment verifies that relying entirely on forged datasets is unnecessary and also verifies that proxy data can be cost-effectively incorporated into other data-free quantization methods.
[0102] Please refer to Figure 6 The diagram illustrates a flowchart of a model quantization method according to an exemplary embodiment of this application. (Reference) Figure 6 The method includes:
[0103] Step 601: The computer device generates fake data based on the target label and Gaussian noise corresponding to the deep learning model through a generator. The label of the fake data is the target label.
[0104] In some embodiments, this step is the same as step 201, and will not be described again here.
[0105] Step 602: The computer device determines proxy data from the proxy data set based on multiple first-batch normalized BN layers in the deep learning model. The proxy data set includes multiple proxy data, which are open-source real data.
[0106] In some embodiments, this step can be implemented through steps 402-404, which will not be described in detail here.
[0107] Step 603: The computer device mixes the forged data and the proxy data to obtain the first sample data.
[0108] In some embodiments, this step is the same as step 203, and will not be described again here.
[0109] Step 604: The computer device quantizes the deep learning model based on the first sample data.
[0110] In some embodiments, this step can be implemented through steps 505-508, which will not be described in detail here.
[0111] Step 605: The computer device determines the generator's third loss function based on the forged data and target labels using a deep learning model.
[0112] The computer device, based on forged data and target labels, uses a deep learning model and a cross-entropy loss function to determine a third loss function. For example, the computer device inputs forged data labeled with target labels into a deep learning model and determines the third loss function using the cross-entropy loss function. This process can be implemented using Equation Nine: Equation Nine: L CE (G)=CE(F(G(z|y)),y); where, L Ce (G) represents the third loss function, F(·) represents the deep learning model, G(z|y) represents the fake data, y represents the label of the fake data, and CE(·,·) represents the cross-entropy loss function.
[0113] Step 606: The computer device determines the fourth loss function of the generator based on the simulated data through multiple second BN layers in the deep learning model.
[0114] The computer device introduces a Batch Normalization (BN) loss function to constrain the generator, thereby making the fabricated data closer to the real distribution and improving the accuracy of data quantization without data. Accordingly, the computer device inputs the fabricated data into the generator, and determines the generator's fourth loss function through multiple second BN layers in the deep learning model. This process can be achieved using Equation 10: Equation 10: in, and σi and σi are the mean and variance calculated from the simulated data in the i-th second BN layer, respectively, and μ2 and σ2 are the corresponding mean and variance of the original model stored in the i-th second BN layer, respectively. L2 represents the total number of second BNs in the network. In some embodiments, the second BN layer and the first BN layer can be the same BN layer, or they can be different BN layers in the deep learning model.
[0115] Step 607: The computer device performs a weighted summation of the third loss function and the fourth loss function based on the third hyperparameter to obtain the second total loss function. The third hyperparameter is used to constrain the weight of the fourth loss function.
[0116] In some embodiments, the computer device, based on a third hyperparameter, performs a weighted summation of the third loss function and the fourth loss function using the following formula eleven to obtain the second total loss function. Formula eleven: L(G) = L CE (G)+α·L BN (G); where L(G) represents the second total loss function, α represents the third hyperparameter, and L CE (G) represents the third loss function, L NN (G) represents the fourth loss function.
[0117] In some embodiments, the computer device, based on a third hyperparameter, performs a weighted summation of the third loss function and the fourth loss function using the following formula (Equation Twelve) to obtain the second total loss function. Formula Twelve: L(G)=(1-α)L CE (G)+α·L BN (G); where L(G) represents the second total loss function, α represents the third hyperparameter, and L CE (G) represents the third loss function, L BN (G) represents the fourth loss function.
[0118] In this embodiment, the total loss function is obtained by weighted summation of the loss functions obtained by the two methods. This approach balances the impact of the two loss functions on the overall loss function, thereby improving the accuracy of model quantization based on the total loss function.
[0119] Step 608: The computer device updates the generator based on the second total loss function.
[0120] If the value corresponding to the second total loss function is less than the third preset threshold, the generator update ends; if the value corresponding to the second total loss function is not less than the third preset threshold, the generator parameter values are adjusted, and then the next round of update process begins. Alternatively, if the difference between the value corresponding to the second total loss function in the second round and the value corresponding to the second total loss function in the previous round is less than the fourth preset threshold, the generator update ends; if the difference between the value corresponding to the second total loss function in the second round and the value corresponding to the second total loss function in the previous round is not less than the fourth preset threshold, then the next round of update process begins, until the generator update ends.
[0121] In some embodiments, after updating the generator, the computer device selects a new target label from the label set corresponding to the deep learning model. Based on the new target label and the redefined Gaussian noise, new fake data is generated by the updated generator, and then the subsequent model quantization process is performed. In this embodiment, the generator is updated after generating fake data once. This constrains the generator, making the generated fake data closer to the real distribution, thereby improving the accuracy of data-free quantization.
[0122] Please refer to Figure 7 The diagram illustrates a flowchart of a model quantization method according to an exemplary embodiment of this application. (Reference) Figure 7 The method includes:
[0123] Step 701: The computer device generates fake data based on the target label and Gaussian noise corresponding to the deep learning model through a generator. The label of the fake data is the target label.
[0124] In some embodiments, this step is the same as step 201, and will not be described again here.
[0125] Step 702: The computer device determines proxy data from the proxy data set based on multiple first-batch normalized BN layers in the deep learning model. The proxy data set includes multiple proxy data, which are open-source real data.
[0126] In some embodiments, this step can be implemented through steps 402-404, which will not be described in detail here.
[0127] Step 703: The computer device mixes the forged data and the proxy data to obtain the first sample data.
[0128] In some embodiments, this step is the same as step 203, and will not be described again here.
[0129] Step 704: The computer device performs one round of optimization on the quantization model based on the first sample data.
[0130] In some embodiments, this step can be implemented through steps 505-508, which will not be described in detail here.
[0131] Step 705: The computer device updates the generator once, generating new forged data using the updated generator.
[0132] In some embodiments, the process of updating the generator by the computer device can be implemented through steps 605-608. The process of the computer device generating new forged data through the updated generator is similar to step 201, except that the target label is different. Of course, the Gaussian noise can also be different. The specific process will not be described in detail here.
[0133] Step 706: The computer device mixes the new forged data and proxy data to obtain new first sample data.
[0134] In some embodiments, this step is the same as step 203, and will not be described again here.
[0135] Step 707: The computer equipment performs another round of optimization on the quantization model based on the new first sample data.
[0136] In some embodiments, this step can be implemented through steps 504-507, which will not be described in detail here.
[0137] Step 708: The computer equipment updates the generator again.
[0138] In some embodiments, this step can be implemented through steps 504-507, which will not be described in detail here. After the computer updates the generator again, it generates new forged data based on the updated generator, and then mixes the new forged data and proxy data again, that is, executes steps 706-708, until the quantization model optimization is completed.
[0139] In the embodiments of this application, the generator and quantization model are updated alternately in each training cycle, so that when optimizing the quantization model, fake data can be generated based on the latest generator, which can improve the accuracy of the generated fake data and thus improve the accuracy of the quantization model.
[0140] Please refer to Figure 8 This application illustrates a model quantization apparatus according to an exemplary embodiment, the apparatus comprising:
[0141] The generation module 801 is used to generate fake data based on the target label and Gaussian noise corresponding to the deep learning model, and the label of the fake data is the target label.
[0142] The first determining module 802 is used to determine proxy data from a proxy data set based on multiple first batch normalized BN layers in a deep learning model. The proxy data set includes multiple proxy data, and the proxy data is open-source real data.
[0143] The mixing module 803 is used to mix the forged data and the proxy data to obtain the first sample data;
[0144] The quantization module 804 is used to quantize the deep learning model based on the first sample data.
[0145] In some embodiments, the first determining module 802 is configured to determine a proxy data set, which includes multiple candidate proxy data; determine the normalized distances corresponding to the multiple candidate proxy data in the proxy data set based on multiple first-batch normalized BN layers in the deep learning model, wherein the normalized distances corresponding to the candidate proxy data are used to characterize the correlation between the candidate proxy data and the second sample data, wherein the second sample data are the sample data used to train the deep learning model; and determine the proxy data whose corresponding normalized distances satisfy the conditions from the proxy data set based on the normalized distances corresponding to the multiple candidate proxy data in the proxy data set.
[0146] In some embodiments, the first determining module 802 is configured to input any candidate proxy data into a deep learning model; determine the first mean and first variance of the candidate proxy data in multiple first batch normalization (BN) layers through multiple first batch normalization (BN) layers in the deep learning model; and determine the normalized distance corresponding to the candidate proxy data based on the first mean and first variance of the candidate proxy data in multiple first BN layers and the second mean and second variance pre-stored in multiple first BN layers.
[0147] In some embodiments, the first determining module 802 is configured to, for any first BN layer, determine a first normalized distance based on a first mean of the candidate proxy data in the first BN layer and a second mean pre-stored in the first BN layer, and determine a second normalized distance based on a first variance of the candidate proxy data in the first BN layer and a second variance pre-stored in the first BN layer; determine the sum of the first normalized distance and the second normalized distance to obtain a third normalized distance; and determine the average of the third normalized distances corresponding to multiple first BN layers to obtain the normalized distance corresponding to the candidate proxy data.
[0148] In some embodiments, the mixing module 803 is used to determine a first hyperparameter, which is used to constrain the mixing ratio of fake data and proxy data; based on the first hyperparameter, the fake data and proxy data are mixed to obtain first sample data.
[0149] In some embodiments, the quantization module 804 is used for any round of iterative optimization to determine a first loss function based on the first sample data and the labels of the first sample data, using a quantization model, where the quantization model is a model obtained by quantizing a deep learning model; to determine a second loss function based on the first sample data, using the quantization model and the deep learning model; to obtain a first total loss function by weighted summation of the first loss function and the second loss function based on a second hyperparameter, where the second hyperparameter is used to constrain the weights of the second loss function; and to optimize the quantization model based on the first total loss function.
[0150] In some embodiments, the apparatus further includes:
[0151] The annotation module is used to input proxy data into the deep learning model and output the labels of the proxy data; it also annotates the proxy data with the corresponding labels.
[0152] In some embodiments, the apparatus further includes:
[0153] The second determination module is used to determine the generator's third loss function based on the forged data and target labels using a deep learning model;
[0154] The third determination module is used to determine the fourth loss function of the generator based on the simulated data and through multiple second BN layers in the deep learning model.
[0155] The weighting module is used to perform a weighted summation of the third loss function and the fourth loss function based on the third hyperparameter to obtain the second total loss function. The third hyperparameter is used to constrain the weights of the fourth loss function.
[0156] The update module is used to update the generator based on the second total loss function.
[0157] In this embodiment, the first sample data includes both forged data and proxy data. The forged data is generated based on the labels corresponding to the deep learning model, and therefore conforms to the label requirements. The proxy data is real open-source data, which has rich features. Therefore, the proxy data has advantages in terms of data features. Thus, quantizing the deep learning model based on the first sample data combines the advantages of both forged data and proxy data, and can improve the accuracy of quantizing the deep learning model.
[0158] It should be noted that the model quantization device provided in the above embodiments is only illustrated by the division of the above functional modules when performing model quantization. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the terminal can be divided into different functional modules to complete all or part of the functions described above. In addition, the model quantization device and the model quantization method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0159] In some embodiments, the computer device may be a terminal; please refer to Figure 9 The diagram illustrates a block diagram of a terminal 900 according to an exemplary embodiment of this application. The terminal 900 in this application may include one or more components such as a processor 910, a memory 920, a display screen 930, and a computer device communication module 940.
[0160] The processor 910 is electrically connected to the computer device communication module 940 via a bus, and the processor 910 communicates with the computer device through the computer device communication module 940 to obtain instruction information.
[0161] The processor 910 may include one or more processing cores. The processor 910 connects to various parts within the terminal 900 using various interfaces and lines, and performs various functions and processes data of the terminal 900 by running or executing instructions, programs, code sets, or instruction sets stored in the memory 920, and by calling data stored in the memory 920. Optionally, the processor 910 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 910 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), Neural-network Processing Unit (NPU), and modem. Specifically, the CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required to be displayed on the display screen 930; the NPU is used to implement Artificial Intelligence (AI) functions; and the modem is used to handle wireless communication. Understandably, the aforementioned modem may also be implemented separately as a single chip, rather than being integrated into the processor 910.
[0162] The memory 920 may include random access memory (RAM) or read-only memory. Optionally, the memory 920 may include a non-transitory computer-readable storage medium. The memory 920 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 920 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the various method embodiments described below, etc.; the data storage area may store data created according to the use of the terminal 900 (such as audio data, phone book, etc.).
[0163] Display screen 930 is a display component used to display a user interface. Optionally, display screen 930 is a touch-enabled display screen, through which users can use their fingers, styluses, or any suitable object to perform touch operations on display screen 930.
[0164] The display screen 930 is typically located on the front panel of the terminal 900. The display screen 930 can be designed as a full-screen, curved screen, irregularly shaped screen, dual-sided screen, or foldable screen. The display screen 930 can also be designed as a combination of a full-screen and a curved screen, or a combination of an irregularly shaped screen and a curved screen, etc., but this embodiment does not limit it in this way.
[0165] In addition, those skilled in the art will understand that the structure of the terminal 900 shown in the above figures does not constitute a limitation on the terminal 900. The terminal 900 may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the terminal 900 may also include audio acquisition devices, speakers, radio frequency circuits, input units, sensors, audio circuits, a Wireless Fidelity (Wi-Fi) module, a power supply, a Bluetooth module, and other components, which will not be described in detail here.
[0166] In some embodiments, the computer device may be a server; please refer to Figure 10This diagram illustrates a block diagram of a server 1010 according to an exemplary embodiment of this application. The server 1010 can vary significantly due to different configurations or performance characteristics, and may include a processor (central processing unit, CPU) 1001 and a memory 1002. The memory 1002 stores at least one line of program code, which is loaded and executed by the processor 1001 to implement the methods provided in the various method embodiments described above. Of course, the server 1010 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input / output. The server 1010 may also include other components for implementing device functions, which will not be elaborated upon here.
[0167] This application also provides a computer-readable medium storing at least one piece of program code, which is loaded and executed by a processor to implement the model quantization method shown in the above embodiments.
[0168] This application also provides a computer program product that stores at least one piece of program code, which is loaded and executed by a processor to implement the model quantization method shown in the above embodiments.
[0169] In some embodiments, the computer program product involved in the present application can be deployed and executed on a computer device, or on multiple computer devices located in one location, or on multiple computer devices distributed in multiple locations and interconnected through a communication network. Multiple computer devices distributed in multiple locations and interconnected through a communication network can form a blockchain system.
[0170] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0171] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A model quantization method, characterized in that, The method includes: Based on the target label and Gaussian noise corresponding to the image recognition model, a generator generates a fake image, and the label of the fake image is the target label; Determine a set of real images, which includes multiple candidate real images, and the candidate real images are open-source real images; Based on multiple first-batch normalized BN layers in the image recognition model, the normalized distances corresponding to multiple candidate real images in the real image set are determined. The normalized distances corresponding to the candidate real images are used to characterize the correlation between the candidate real images and the second sample images, where the second sample images are sample images used to train the image recognition model. Based on the normalized distances corresponding to multiple candidate real images in the real image set, the real images whose normalized distances satisfy the conditions are determined from the real image set. The forged image and the real image are mixed to obtain a first sample image; Based on the first sample image, the image recognition model is quantized, and the quantized image recognition model is deployed to the terminal, so that the terminal can perform image recognition through the quantized image recognition model.
2. The method according to claim 1, characterized in that, The step of determining the normalized distances corresponding to multiple candidate real images in the set of real images based on multiple first-batch normalized BN layers in the image recognition model includes: For any candidate real image, the candidate real image is input into the image recognition model; The first mean and first variance of the candidate real image in the multiple first BN layers in the image recognition model are determined. Based on the first mean and first variance of the candidate real image in the plurality of first BN layers, and the second mean and second variance pre-stored in the plurality of first BN layers, the normalized distance corresponding to the candidate real image is determined.
3. The method according to claim 2, characterized in that, The step of determining the normalized distance corresponding to the candidate real image based on the first mean and first variance of the candidate real image in the plurality of first BN layers and the second mean and second variance pre-stored in the plurality of first BN layers includes: For any first BN layer, a first normalized distance is determined based on the first mean of the candidate real image in the first BN layer and the second mean pre-stored in the first BN layer, and a second normalized distance is determined based on the first variance of the candidate real image in the first BN layer and the second variance pre-stored in the first BN layer. The sum of the first normalized distance and the second normalized distance is determined to obtain a third normalized distance. The average value of the third normalized distances corresponding to the plurality of first BN layers is determined to obtain the normalized distance corresponding to the candidate real image.
4. The method according to claim 1, characterized in that, The step of mixing the forged image and the real image to obtain the first sample image includes: A first hyperparameter is determined, which is used to constrain the mixing ratio of the fake image and the real image; Based on the first hyperparameter, the fake image and the real image are mixed to obtain the first sample image.
5. The method according to claim 1, characterized in that, The step of quantizing the image recognition model based on the first sample image includes: For any round of iterative optimization, based on the first sample image and the label of the first sample image, a first loss function is determined by quantizing the image recognition model, wherein the quantized image recognition model is a model obtained by quantizing the image recognition model; Based on the first sample image, a second loss function is determined by quantizing the image recognition model and the image recognition model; Based on the second hyperparameter, the first loss function and the second loss function are weighted and summed to obtain the first total loss function, and the second hyperparameter is used to constrain the weight of the second loss function; The quantized image recognition model is optimized based on the first total loss function.
6. The method according to claim 1, characterized in that, Before mixing the forged image and the real image to obtain the first sample image, the method further includes: The real image is input into the image recognition model, and the label of the real image is output. Label the real image with tags.
7. The method according to any one of claims 1-6, characterized in that, The method further includes: Based on the forged image and the target label, the third loss function of the generator is determined using the image recognition model; Based on the forged image, the fourth loss function of the generator is determined through multiple second BN layers in the image recognition model; Based on the third hyperparameter, the third loss function and the fourth loss function are weighted and summed to obtain the second total loss function. The third hyperparameter is used to constrain the weight of the fourth loss function. The generator is updated based on the second total loss function.
8. A model quantization device, characterized in that, The device includes: The generation module is used to generate a fake image based on the target label corresponding to the image recognition model and Gaussian noise, wherein the label of the fake image is the target label; A first determining module is used to determine a set of real images, the set of real images including multiple candidate real images, the candidate real images being open-source real images; based on multiple first-batch normalized BN layers in the image recognition model, determine the normalized distances corresponding to the multiple candidate real images in the set of real images, the normalized distances corresponding to the candidate real images being used to characterize the correlation between the candidate real images and second sample images, the second sample images being sample images used to train the image recognition model; based on the normalized distances corresponding to the multiple candidate real images in the set of real images, determine the real images whose corresponding normalized distances satisfy the conditions from the set of real images; A mixing module is used to mix the forged image and the real image to obtain a first sample image; The quantization module is used to quantize the image recognition model based on the first sample image, and deploy the quantized image recognition model to the terminal, so that the terminal can perform image recognition through the quantized image recognition model.
9. The apparatus according to claim 8, characterized in that, The first determining module is configured to input any candidate real image into the image recognition model; and determine the first mean and first variance of the candidate real image in the multiple first BN layers of the image recognition model. Based on the first mean and first variance of the candidate real image in the plurality of first BN layers, and the second mean and second variance pre-stored in the plurality of first BN layers, the normalized distance corresponding to the candidate real image is determined.
10. The apparatus according to claim 9, characterized in that, The first determining module is configured to, for any first BN layer, determine a first normalized distance based on a first mean of the candidate real image in the first BN layer and a second mean pre-stored in the first BN layer; determine a second normalized distance based on a first variance of the candidate real image in the first BN layer and a second variance pre-stored in the first BN layer; determine the sum of the first normalized distance and the second normalized distance to obtain a third normalized distance; and determine the average of the third normalized distances corresponding to the plurality of first BN layers to obtain the normalized distance corresponding to the candidate real image.
11. The apparatus according to claim 8, characterized in that, The mixing module is used to determine a first hyperparameter, which constrains the mixing ratio of the forged image and the real image; based on the first hyperparameter, the forged image and the real image are mixed to obtain the first sample image.
12. The apparatus according to claim 8, characterized in that, The quantization module is used for any round of iterative optimization to determine a first loss function based on the first sample image and its label, using a quantized image recognition model, where the quantized image recognition model is a model obtained by quantizing the image recognition model; to determine a second loss function based on the first sample image, using the quantized model and the image recognition model; and to obtain a first total loss function by weighted summation of the first and second loss functions based on a second hyperparameter, where the second hyperparameter is used to constrain the weights of the second loss function. The quantized image recognition model is optimized based on the first total loss function.
13. The apparatus according to claim 8, characterized in that, The device further includes: The annotation module is used to input the real image into the image recognition model and output the label of the real image; and to annotate the real image with the label.
14. The apparatus according to any one of claims 8-13, characterized in that, The device further includes: The second determining module is used to determine the third loss function of the generator based on the forged image and the target label, through the image recognition model; The third determining module is used to determine the fourth loss function of the generator based on the forged image and through multiple second BN layers in the image recognition model. The weighting module is used to perform a weighted summation of the third loss function and the fourth loss function based on the third hyperparameter to obtain the second total loss function. The third hyperparameter is used to constrain the weight of the fourth loss function. An update module is used to update the generator based on the second total loss function.
15. A terminal, characterized in that, The terminal includes a processor and a memory, the memory storing at least one piece of program code, which is loaded and executed by the processor to implement the model quantization method as described in any one of claims 1 to 7.
16. A computer-readable storage medium, characterized in that, The storage medium stores at least one piece of program code, which is loaded and executed by a processor to implement the model quantization method as described in any one of claims 1 to 7.
17. A computer program product, characterized in that, The computer program product stores at least one line of program code, which is executed by a processor to implement the model quantization method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Electric cooker liner image data enhancement method based on mask generative adversarial network
CN115908379A