Quantization Method and System for Binary Input Model, and Computer-Readable Storage Medium

By performing balanced processing of binary input images and training and quantization of deep learning models, combined with equivalent scaling technology, the problem of serious accuracy loss in the quantization process of binary input models is solved, and more efficient and accurate model quantization is achieved.

CN114444679BActive Publication Date: 2025-06-03SHANDONG IND RES KUNYUN ARTIFICIAL INTELLIGENCE RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011235611.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-06
Publication Date
2025-06-03
Estimated Expiration
2040-11-06

AI Technical Summary

Technical Problem

In the prior art, the quantization process of binary input models has serious accuracy losses, and training fine-tuning and data calibration methods is difficult to effectively reduce quantization losses.

Method used

The image data distribution is equalized by equalizing the preprocessed binarized input image, including Gaussian filtering, inverting color processing, histogram equalization and random linear perturbation. Then, the image after equalization is input to the deep learning model for training, the weight copy is saved, the model is quantized, and the model parameters are adjusted by equivalent scaling to achieve a more balanced output distribution.

Benefits of technology

It effectively reduces the accuracy loss in the quantization process of binary input model, improves the running speed and accuracy of the model, and solves the problems of many parameters and slow running speed of binary network model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114444679B_ABST
    Figure CN114444679B_ABST
Patent Text Reader

Abstract

The present invention discloses a quantization method for a binarized input model, comprising the following steps: performing equalization processing on the preprocessed binarized input image; inputting the input image after equalization processing into a deep learning model for training to obtain model parameters; performing equivalent scaling on adjacent layers of the model parameters to obtain scaled model parameters; performing model quantization on the deep learning model with the scaled model parameters; The present invention also discloses a binarized input model system and a computer-readable storage medium, which solve the problem of precision loss in the quantization process of the binarized input model in the prior art, reduce model precision loss and improve the running speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of binarized input networks, and in particular, to a quantization method and system for a binarized input model. Background Art

[0002] With the development of deep learning, neural models trained using a large amount of real-scene data have been deployed in real scenarios. However, the huge number of model parameters and computational complexity that follow have posed significant challenges to model deployment, especially on embedded devices and mobile devices. These devices usually have low computing power and limited storage resources, resulting in excessive resource consumption by neural models and inability to run in real time.

[0003] Currently, most trained neural models are based on a float32 floating-point data type. To better adapt neural models to embedded devices and mobile devices, researchers have invented a technique that converts floating-point calculations into low-bit fixed-point calculations, which can effectively reduce the model parameter size and improve the running speed. However, this technique often incurs a certain degree of precision loss. The quantization of existing binarized input models has the following disadvantages: 1. The training fine-tuning method requires modifying the model training source code and increasing the model deployment cycle; 2. The data calibration method suffers from abnormal data distribution, resulting in unsaturated quantization mapping and serious precision loss. These methods cannot effectively reduce the quantization loss of models with binarized data as input.

[0004] Therefore, it is crucial to provide a method for reducing the quantization loss of binarized input models for which general quantization methods are ineffective. Summary of the Invention

[0005] The main objective of the present invention is to provide a quantization method and system for a binarized input model, as well as a computer-readable storage medium, aiming to solve the problem of precision loss during the quantization process of binarized input models in the prior art.

[0006] To achieve the above objective, the present invention provides a quantization method for a binarized input model, and the quantization method for the binarized input model includes the following steps:

[0007] In one embodiment, perform equalization processing on the preprocessed binarized input image;

[0008] Input the input image after equalization processing into a deep learning model for training to obtain model parameters;

[0009] Perform equivalent scaling on adjacent layers of the model parameters to obtain scaled model parameters;

[0010] Perform model quantization on the deep learning model with the scaled model parameters.

[0011] In one embodiment, the equalization processing includes: obtaining a three-channel image of the input image and performing augmentation processing on the images of each channel.

[0012] In one embodiment, performing augmentation processing on the images of each channel includes:

[0013] Performing Gaussian filtering on the channel images to reduce the pixel noise of each image;

[0014] Performing color inversion and histogram equalization on the channel images to enhance the overall contrast effect of the images;

[0015] Performing random linear perturbation on the channel images to balance the distribution of each image pixel.

[0016] In one embodiment, inputting the input image after equalization processing into a deep learning model for training includes:

[0017] During the training process, saving a weight copy of each channel image;

[0018] During the forward propagation of the binary network model, using the sign function to obtain the binary activation value and the binary weight;

[0019] Performing model inference according to the binary weight and the binary activation value;

[0020] During the backpropagation of the binary network model, using the gradient to update the weight copy, wherein when updating the weight copy, the weights that do not meet the preset value are clipped.

[0021] In one embodiment, before performing model inference according to the binary weight and the binary activation value, it further includes: normalizing the channel image data.

[0022] In one embodiment, normalizing the channel image data includes:

[0023] Obtaining the image data of each channel and calculating the mean value of the image data in units of convolution kernels;

[0024] Performing mean subtraction and normalization operations on the image data in each convolution kernel;

[0025] Obtaining the image data of each channel, calculating the mean value of each image channel, and performing mean subtraction operation on all pixels in the corresponding channel.

[0026] In one embodiment, the image pixels in the binary weight follow a Bernoulli distribution with a variance of 1, and the image pixels in the binary activation value follow a Bernoulli distribution.

[0027] In one embodiment, the equivalent scaling of adjacent layers of the model parameters to obtain the scaled model parameters includes:

[0028] Obtain the scale factors for equivalent scaling of two adjacent layers, and adjust the weights of different channels according to the maximum and minimum values in the two adjacent layers;

[0029] Traverse to the output layer to adjust the weights of the outputs that can be equivalently scaled in the entire deep learning model.

[0030] To achieve the above object, the present invention also provides a binary input model system, which at least includes one or more processors, a memory, and a quantization program of the binary input model stored on the memory and executable on the processor. When the processor executes the quantization program of the binary input model, each step of the quantization method of the binary input model as described above is implemented.

[0031] To achieve the above object, the present invention also provides a computer-readable storage medium, which stores a quantization program of the binary input model. When the quantization program of the binary input model is executed by a processor, each step of the quantization method of the binary input model as described above is implemented.

[0032] The technical solutions of the quantization method, system, and computer-readable storage medium of the binary input model provided in the embodiments of the present application at least have the following technical effects:

[0033] 1. Since the technical solution of performing equalization processing on the preprocessed binary input image, obtaining the three-channel image of the input image, performing Gaussian filtering on the channel image to reduce the pixel noise of each image, performing inverse color processing and histogram equalization processing on the channel image to enhance the overall contrast effect of the image, performing random linear perturbation processing on the channel image to make the pixel distribution of each image balanced, and performing augmentation processing on the images of each channel is adopted, the problem of serious progress loss caused by the abnormal distribution of image data in the prior art is solved, the balanced distribution of image data is realized, and the accuracy loss is reduced.

[0034] 2. By inputting the input image after equalization processing into the deep learning model for training to obtain model parameters, during the training process, saving the weight copies of each channel image, normalizing the channel image data, in the forward propagation process of the binary network model, using the sign function to obtain the binary activation value and binary weight, and performing model inference according to the binary weight and binary activation value; in the backward propagation process of the binary network model, using the gradient to update the weight copy, wherein when updating the weight copy, trimming the weights that do not meet the preset value, the technical solution solves the problems of many parameters and slow running speed of the binary network model in the prior art, reduces the accuracy loss, and improves the running speed.

[0035] 3. By equivalently scaling adjacent layers of the model parameters to obtain the scaled model parameters, by obtaining the scale factors available for equivalent scaling of adjacent two layers, adjusting the weights of different channels according to the maximum and minimum values in the adjacent two layers, traversing to the output layer to adjust the weights of the output that can be equivalently scaled by the entire deep learning model, and then quantifying the deep learning model with the scaled model parameters, the technical solution solves the problem of serious quantization accuracy loss caused by uneven output distribution of the binary network model in the prior art, and realizes the balanced distribution of the binary network model. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a schematic structural diagram of a binary input model system according to an embodiment of the present invention;

[0037] Figure 2 It is a schematic flowchart of the first embodiment of the quantization method of the binary input model of the present invention;

[0038] Figure 3 It is a detailed flowchart of the first embodiment of the quantization method of the binary input model of the present invention;

[0039] Figure 4 It is a detailed flowchart of step S111 of the first embodiment of the quantization method of the binary input model of the present invention;

[0040] Figure 5 It is a detailed flowchart of step S120 of the first embodiment of the quantization method of the binary input model of the present invention;

[0041] Figure 6 It is a detailed flowchart of step S130 of the first embodiment of the quantization method of the binary input model of the present invention;

[0042] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] It should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0044] To solve the problem of precision loss in the quantization process of the binary input model in the prior art, the present application adopts the following steps: performing equalization processing on the preprocessed binary input image; inputting the equalized input image into a deep learning model for training to obtain model parameters; performing equivalent scaling on adjacent layers of the model parameters to obtain scaled model parameters; performing model quantization on the deep learning model with the scaled model parameters. The present invention also adopts a binary input model system and a computer-readable storage medium to reduce model precision loss and improve the running speed.

[0045] To better understand the above technical solutions, the exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0046] Those skilled in the art can understand that Figure 1 the hardware structure of the binary input model system shown does not constitute a limitation on the binary input model system. The binary input model system may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0047] As an implementation manner, it can be as Figure 1 shown Figure 1 is a schematic structural diagram of the binary input model system according to an embodiment of the present invention.

[0048] The processor 1100 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 1100 or the instructions in the form of software. The above-mentioned processor 1100 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The software module may be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory 1200, and the processor 1100 reads the information in the memory 1200 and combines its hardware to complete the steps of the above method.

[0049] It can be understood that the memory 1200 in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM). The memory 1200 of the systems and methods described in the embodiments of the present invention is intended to include but not be limited to these and any other suitable types of memory.

[0050] For software implementation, the techniques described in the embodiments of the present invention can be implemented by modules (such as procedures, functions, etc.) that execute the functions described in the embodiments of the present invention. The software code can be stored in a memory and executed by a processor. The memory can be implemented inside or outside the processor.

[0051] Based on the above structure, embodiments of the present invention are proposed.

[0052] Referring to Figure 2 , Figure 2 is a flowchart of the first embodiment of the quantization method for the binarized input model of the present invention. The method includes the following steps:

[0053] Step S110: Perform equalization processing on the preprocessed binarized input image.

[0054] In this embodiment, before performing equalization processing on the binarized image, it is necessary to preprocess the input image. The input image is generally a color image. The preprocessing process includes: First, convert the color image to a grayscale image through grayscale conversion. The grayscale image is also called a gray-level image. Each pixel in the grayscale image can be represented by a brightness value from 0 (black) to 255 (white), and different gray levels are represented between 0 - 255; Finally, after performing grayscale processing on the color image, perform binarization processing to convert it into a binarized input image. The binarization processing can adopt the threshold method. Utilizing the difference between the target and the background in the image, set the image to two different levels respectively, select a suitable threshold to determine whether the pixel is a target or a background, thereby obtaining a binarized image, and then perform equalization processing on the preprocessed binarized input image to make the pixel distribution of the input image balanced.

[0055] Referring to Figure 3 , Figure 3 Step S111 is a refinement of step S110 of the quantization method for the binarized input model of the present invention. The other steps are the same as Figure 1 and will not be elaborated here.

[0056] Step S111: Obtain the three-channel image of the input image and perform augmentation processing on the image of each channel.

[0057] In this embodiment, color images are mainly divided into two types, RGB and CMYK. Among them, the RGB color image contains three channels, which are respectively composed of three different color components, one is red, one is green, and the other is blue; while the CMYK type of image is composed of four color components: cyan C, magenta M, yellow Y, and black K. This application takes the operation on the RGB three-channel image as an example. First, the RGB three-channel image is converted into a grayscale image to reduce the amount of calculation, and then augmentation processing operations are independently performed for each channel, so as to enrich the diversity of samples and make the image pixels evenly spread throughout the color space.

[0058] Referring to Figure 4 , Figure 4 FIG. is a detailed flowchart of step S111 of the quantization method of the binarization input model of the present invention in the first embodiment, including the following steps:

[0059] Step S1111, perform Gaussian filtering on the channel image to reduce the noise of each image pixel.

[0060] In this embodiment, the Gaussian filtering process is a process of weighted averaging of the pixels of the entire image. The value of each image pixel point is obtained by weighted averaging of itself and other image pixel values in the neighborhood. The specific operation of Gaussian filtering is: scan each image pixel in the image with a template, and replace the value of the pixel point at the center of the template with the weighted average gray value of the image pixels in the neighborhood determined by the template, so as to effectively reduce the difference between the noise of the binarized image pixel set and the surrounding image pixels, make the image smoother to reduce false detection.

[0061] Step S1112, perform color inversion processing and histogram equalization processing on the channel image to enhance the overall contrast effect of the image.

[0062] In this embodiment, the color inversion processing refers to inverting the three channels, that is, the R, G, and B values. For example, if the quantization level of the original input image color is 256, the R, G, and B values of the new image are 255 minus the R, G, and B values of the original input image, and the inverted R, G, and B values are written into the new image. For example, the color of an image pixel point is (0, 0, 0), and after color inversion, it is (255, 255, 255), so as to change the contrast form between the image background and the target; the histogram equalization processing refers to broadening the gray values with a large number of image pixels and merging the gray values with a small number of image pixels, so as to increase the contrast, change the histogram of the image to change the gray level of each image pixel, and achieve the effect of enhancing the overall contrast of the image.

[0063] Step S1113, perform random linear perturbation processing on the channel image to make the distribution of each image pixel balanced.

[0064] In this embodiment, random linear perturbation processing is used to make the model adapt to more brightness changes, aiming to make the image pixel distribution more balanced, the morphology more diverse, and reduce false detections caused by the difference between noise and background.

[0065] Since the technical solution of performing equalization processing on the preprocessed binarized input image, obtaining the three-channel image of the input image, performing Gaussian filtering on the channel image to reduce the noise of each image pixel, performing color inversion processing and histogram equalization processing on the channel image to enhance the overall contrast effect of the image, performing random linear perturbation processing on the channel image to make the distribution of each image pixel balanced, and performing augmentation processing on the image of each channel, the problem of serious progress loss caused by the abnormal distribution of image data in the prior art is solved, the balanced distribution of image data is realized, and the precision loss is reduced.

[0066] Step S120, input the input image after equalization processing into a deep learning model for training to obtain model parameters.

[0067] In this embodiment, the binarized data of the input image after equalization processing is input into a deep learning model for training. During the training process, during the forward propagation of the binarized network model, a sign function is used to obtain the binarized activation value and the binarized weight. Among them, the binarized weight and the binarized activation value are limited to +1 and -1. At the same time, the binarized weight that does not meet the preset value is clipped. During the backward propagation of the binarized network model, the binarized weight and the binarized activation value are used to calculate the parameter gradient, and the gradient is used to update the weight copy.

[0068] Refer to Figure 5 , Figure 5 FIG. is a detailed flowchart of step S120 of the first embodiment of the quantization method of the binarized input model of the present invention, including the following steps:

[0069] Step S121, during the training process, save the weight copy of each channel image.

[0070] In this embodiment, during the training process of the binarized model, each channel image corresponds to a weight. During the forward propagation of the binarized network model, the generated binarized weight and the binarized activation value are in a binarized data format; while during the backward propagation of the binarized network model, the calculated gradient value is a real number rather than binarized. The weight copy is saved before the forward propagation of the binarized network model, and the weight copy can be used for updating when the gradient is updated during the backward propagation of the binarized network model.

[0071] Step S122, during the forward propagation of the binarized network model, use a sign function to obtain the binarized activation value and the binarized weight.

[0072] In this embodiment, a neuron has n binarized activation values as inputs, and each binarized activation value corresponds to a weight w. Inside the neuron, the binarized activation values and the binarized weights are multiplied and then summed. The result of the summation is subtracted from the bias, and finally the result is put into the sign function, which gives the final output. Among them, the binarized activation values and the binarized weights are represented in binary form, and the output is also in binary form. The output state of 0 represents inhibition, and the state of 1 represents activation. During the forward propagation process of the binarized network model, given the input data, it is calculated layer by layer. The result of the sign function of the previous layer is used as the input of the next layer. The sign function is used to obtain the binarized weights and the binarized activation values. The binarized activation values of the previous layer are multiplied by the binarized weights and summed. The result of the summation is subtracted from the bias to obtain a bias parameter, which is also a binarized parameter.

[0073] Step S123: Perform normalization processing on the channel image data.

[0074] In this embodiment, the binarized weights are sent to the normalization processing layer for normalization operations. The normalization, also known as standard deviation normalization, uses z-score to perform normalization processing on the channel image data. First, the image data of each channel is obtained, and the mean of the image data is calculated in units of convolution kernels. Then, the mean is subtracted from the image data in each convolution kernel and a normalization operation is performed. Finally, the image data of each channel is obtained, the mean of each image channel is calculated, and the mean is subtracted from all pixels in the corresponding channel. Among them, the image pixels in the binarized weights follow a Bernoulli distribution with a variance of 1, and the image pixels in the binarized activation values follow a Bernoulli distribution. The normalization processing is usually used before entering the activation layer during the forward propagation process of the binarized network model, which can accelerate model training and reduce the influence of the channel image weight scale to improve the convergence speed and accuracy of the model.

[0075] Step S124: Perform model inference according to the binarized weights and the binarized activation values.

[0076] In this embodiment, the binarized weights and the binarized activation values corresponding to the normalized image data are input into the model for binarized input model inference. The model inference process includes the backward propagation process of the binarized network model and the forward propagation process of the binarized network model.

[0077] Step S125: During the backward propagation process of the binarized network model, use the gradient to update the weight copy. Among them, when updating the weight copy, the weights that do not meet the preset value are clipped.

[0078] In this embodiment, during the backpropagation process of the binarized network model, the gradient of each layer is calculated. Starting from the output layer, the previous layer is calculated backward until the gradient value of the first layer is calculated. Among them, the model gradient is calculated according to the binarized weights and binarized activation values. The gradient is a real number rather than binarized. The weight copy is used to update the gradient. When updating the weight copy, the weights that do not meet the preset values of +1 and -1 are clipped to reduce the channel image weight scale.

[0079] Since the input image after equalization processing is input into the deep learning model for training to obtain model parameters, by saving the weight copies of each channel image during the training process, during the forward propagation process of the binarized network model, the sigmoid function is used to obtain the binarized activation values and binarized weights, the channel image data is normalized, and model inference is performed according to the binarized weights and binarized activation values, accelerating model training and reducing the channel image weight scale; during the backpropagation process of the binarized network model, the weight copy is updated using the gradient. Among them, when updating the weight copy, the weights that do not meet the preset values are clipped, which solves the problems of many parameters and slow running speed of the binarized network model in the prior art, reduces the accuracy loss, and improves the running speed.

[0080] Step S130, equivalently scale adjacent layers of the model parameters to obtain the scaled model parameters.

[0081] In this embodiment, during the forward and backward propagation processes of the model, the maximum and minimum values of the output of each layer are saved and updated, the scale factors for equivalent scaling available for adjacent layers are obtained, the weights of different channels are adjusted according to the obtained maximum and minimum values, and the entire process will traverse to the output layer until the outputs that can be equivalently scaled for the entire model are adjusted.

[0082] Refer to Figure 6 , Figure 6 is a detailed flowchart of step S130 of the first embodiment of the quantization method of the binarized input model of the present invention, including the following steps:

[0083] Step S131, obtain the scale factors for equivalent scaling available for adjacent layers, and adjust the weights of different channels according to the maximum and minimum values in the adjacent layers.

[0084] In this embodiment, in the binarization network, the output of the previous layer can be used as the input of the next layer until the traversal reaches the output layer and terminates; for example: the binarization network model includes three layers, the last layer is the output layer, and each input layer includes three nodes, and each node corresponds to an activation value, which may include a maximum value, a minimum value, and an intermediate value. The output of the first layer nodes can be used as the input of the second layer nodes, and so on until the traversal reaches the output layer and terminates. The nodes between each layer can be arbitrarily combined, and the connection of each layer of nodes corresponds to a channel weight. During the training process, the maximum value and the minimum value of the output of the first layer are saved, and the weights of the corresponding channels are adjusted by obtaining the maximum value and the minimum value of each layer.

[0085] Step S132, traverse to the output layer to adjust the weights of the equivalently scalable outputs of the entire deep learning model.

[0086] In this embodiment, traverse to the output layer, save the maximum value and the minimum value of each layer to adjust the weights of the corresponding channels, find the scale factors that can be used for equivalent scaling between adjacent layers, obtain the maximum value and the minimum value of each layer, and multiply the weights of the channels of the equivalently scalable outputs of each layer by a scale factor. Among them, the weights of the corresponding channels are adjusted by the corresponding maximum value and minimum value of each layer, so that the weights of each equivalently scalable output in the entire deep learning model are adjusted. For example: assume that there are adjacent y = f(w 1 x + b 1 ) and h = f(w 2 x + b 2 ) layers in the model, where y and h represent different layers, w 1 and w 2 correspond to the weights of each layer, x represents the output, b 1 and b 2 represent biases. Multiply w 1 and b 1 by a scale factor scale matrix to balance the output distribution. Then, in order to ensure that the output of h is equivalent to that before multiplying by the scale factor, it is necessary to multiply w 2 by a scale diagonal matrix to restore the result.

[0087] Since the adjacent layers of the model parameters are equivalently scaled to obtain the scaled model parameters, by obtaining the scale factors that can be used for equivalent scaling between adjacent layers, adjusting the weights of different channels according to the maximum value and the minimum value in the adjacent two layers, traversing to the output layer to adjust the weights of the equivalently scalable outputs of the entire deep learning model, and then quantifying the deep learning model with the scaled model parameters, the technical solution solves the problem of serious loss of quantization accuracy caused by the unbalanced output distribution of the binarization network model in the prior art, and realizes the balanced distribution of the binarization network model.

[0088] Step S140: Quantize the deep learning model with the scaled model parameters.

[0089] In this embodiment, the scaled model parameters can represent the binarized weights and binarized activation values corresponding to the scaled output. The principle of model quantization is a process of approximately representing the continuously valued floating-point model weights or tensor data flowing through the model as a finite number of discrete values with a relatively low inference accuracy loss. It is a process of using a data type with fewer bits to approximately represent 32-bit finite-range floating-point data, while the input and output of the model remain floating-point, so as to achieve goals such as reducing the model size, reducing the model memory consumption, and accelerating the model inference speed. The model quantization in this application compresses the binarized network by reducing the number of bits required for each binarized weight, quantizes the deep learning model with the binarized weights and binarized activation values corresponding to the scaled output, and realizes binarized weight sharing through model quantization to accelerate the running speed of the binarized input network.

[0090] Due to the technical solution of performing equalization processing on the preprocessed binarized input image, inputting the equalized input image into the deep learning model for training to obtain model parameters, equivalently scaling adjacent layers of the model parameters to obtain scaled model parameters, and quantizing the deep learning model with the scaled model parameters, the problem of serious progress loss caused by the abnormal distribution of image data in the prior art is solved, the balanced distribution of the binarized network model is realized, the precision loss of binarization is reduced, and the running speed is improved.

[0091] Based on the same inventive concept, the embodiment of the present application also provides a binarized input model system. The binarized input model system includes one or more processors, a memory, and a quantization program of the binarized input model stored in the memory and executable on the processors. When the processors execute the quantization program of the binarized input model, each step of the quantization method of the binarized input model as described above is implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.

[0092] Since the binarized input model system provided by the embodiment of the present application is the binarized input model system adopted for implementing the method of the embodiment of the present application, based on the method introduced in the embodiment of the present application, those skilled in the art can understand the specific structure and variations of the binarized input model system, so it will not be elaborated here. Any binarized input model system adopted by the method of the embodiment of the present application belongs to the scope to be protected by the present application. The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.

[0093] Based on the same inventive concept, an embodiment of the present application further provides a computer-readable storage medium, on which a quantization program for a binarized input model is stored. When the quantization program of the binarized input model is executed by a processor, it implements each step of the quantization method of the binarized input model as described above and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0094] Since the computer-readable storage medium provided by the embodiment of the present application is the computer-readable storage medium used to implement the method of the embodiment of the present application, based on the method introduced in the embodiment of the present application, those skilled in the art can understand the specific structure and variations of the computer-readable storage medium, so it will not be elaborated here. Any computer-readable storage medium used in the method of the embodiment of the present application belongs to the scope to be protected by the present application.

[0095] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0096] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0097] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one or more flows or multiple flows and / or blocks Figure 1 one or more blocks or multiple blocks.

[0098] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the specified functions in Figure 1 one or more flows or multiple flows and / or blocks Figure 1 one or more blocks or multiple blocks.

[0099] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or one block or a plurality of blocks. Figure 1 one process or a plurality of processes and / or Figure 1 steps for realizing the functions specified in one block or a plurality of blocks.

[0100] It should be noted that, in the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of other elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims listing several means, several of these means can be embodied by one and the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.

[0101] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0102] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A quantization method for a binarized input model, characterized in that, the method includes: Performing equalization processing on the preprocessed binarized input image, where the equalization processing includes: obtaining the three-channel image of the input image; performing Gaussian filtering on the channel images to reduce the noise of each image pixel; performing color inversion processing and histogram equalization processing on the channel images to enhance the overall contrast effect of the image; performing random linear perturbation processing on the channel images to make the distribution of each image pixel balanced; Inputting the input image after equalization processing into a deep learning model for training to obtain model parameters; During the forward propagation and backward propagation of the model, saving and updating the maximum and minimum values of the output activation values of each layer, obtaining the scaling factors that can be used for equivalent scaling between adjacent layers, adjusting the weights of different channels according to the obtained maximum and minimum values, and traversing to the output layer to adjust the weights of the equivalently scalable output of the entire deep learning model to obtain the scaled model parameters; Quantizing the deep learning model with the scaled model parameters.

2. The quantization method for a binarized input model according to claim 1, characterized in that, The inputting the input image after equalization processing into a deep learning model for training includes: During the training process, saving the weight copies of each channel image; During the forward propagation of the binarized network model, using the sign function to obtain the binarized activation values and binarized weights; Performing model inference according to the binarized weights and binarized activation values; During the backward propagation of the binarized network model, using the gradient to update the weight copies, where when updating the weight copies, the weights that do not meet the preset value are clipped.

3. The quantization method for a binarized input model according to claim 2, characterized in that, Before performing model inference according to the binarized weights and binarized activation values, it further includes: performing normalization processing on the channel image data.

4. The quantization method for a binarized input model according to claim 3, characterized in that, The performing normalization processing on the channel image data includes: Obtaining the channel image data of each channel and calculating the mean value of the image data in units of convolution kernels; Performing mean subtraction and standardization operations on the image data in each convolution kernel; Obtaining the channel image data of each channel, calculating the mean value of each image channel and performing mean subtraction operations on all pixels in the corresponding channel.

5. The quantization method for a binarized input model according to claim 4, characterized in that, The image pixels in the binarized weights follow a Bernoulli distribution with a variance of 1, and the image pixels in the binarized activation values follow a Bernoulli distribution.

6. A binarized input model system, characterized in that, the system at least includes one or more processors, a memory, and a quantization program for the binarized input model stored on the memory and executable on the processors, and when the processors execute the quantization program for the binarized input model, the method according to any one of claims 1-5 is implemented.

7. A computer-readable storage medium, characterized in that, A quantization program of a binarized input model is stored thereon. It is characterized in that when the quantization program of the binarized input model is executed by a processor, the method according to any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Model training method, model training device, electronic equipment and computer readable storage medium

    CN110414679A

  • Balanced binarization neural network quantification method and system

    CN110472725A