A computer vision processing method, device, equipment and storage medium
By adjusting and quantifying the parameter distribution interval of the computer vision model, a smaller target model is obtained, which solves the problem of high requirements for hardware devices by computer vision processing and achieves more efficient computer vision processing.
Patent Information
- Application Number
- CN202411839808.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-13
AI Technical Summary
How to reduce the requirements of computer vision processing on hardware devices, especially to reduce the size and computing volume of models, and reduce storage space usage and operation consumption.
By obtaining the quantization accuracy and several substructures of the original model, determining the equalization factor based on the reference processing results, adjusting the model parameter distribution interval, and performing target quantization processing based on the quantization accuracy, a smaller target model is obtained.
The target model used is smaller in size, less computational volume, less storage space, and less consumption of model operation, reducing the requirements for computer vision processing on hardware devices.
Smart Images

Figure CN119295896B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and in particular to a computer vision processing method, device, equipment and storage medium. Background Art
[0002] Computer vision is an important branch of artificial intelligence and has been widely used in many fields, such as image recognition and object detection.
[0003] The application of neural network models in the field of computer vision has also gradually developed, and the application of neural network models has put forward higher requirements on hardware devices. How to reduce the requirements of computer vision processing on hardware devices has become an urgent problem to be solved. Summary of the invention
[0004] The present application at least provides a computer vision processing method, apparatus, device and storage medium.
[0005] The present application provides a computer vision processing method, including: a processing device acquires image data; a target model stored in the processing device is used to perform computer vision processing on the image data to obtain a computer vision processing result; wherein the target model is obtained by the following processing: acquiring a quantization accuracy corresponding to an original model; acquiring a plurality of first substructures from the original model, and determining an equalization factor corresponding to each first substructure based on a first reference processing result of each first substructure on a first sample image; adjusting model parameters of the corresponding first substructure in the original model based on each equalization factor to obtain a model to be quantized; the equalization factor is a parameter used to adjust the distribution interval of the model parameter; and performing target quantization processing on the model parameters of the model to be quantized based on the quantization accuracy to obtain the target model.
[0006] The present application provides a computer vision processing device, including an acquisition module and a computer vision processing module. The acquisition module is used to acquire image data; the computer vision processing module is used to perform computer vision processing on the image data using a stored target model to obtain a computer vision processing result; wherein the target model is obtained by the following processing: obtaining the quantization accuracy corresponding to the original model; obtaining a number of first substructures from the original model, and determining the equalization factor corresponding to each first substructure based on the first reference processing result of each first substructure on the first sample image; adjusting the model parameters of the corresponding first substructure in the original model based on each equalization factor to obtain a model to be quantized; the equalization factor is a parameter used to adjust the distribution interval of the model parameters; performing target quantization processing on the model parameters of the model to be quantized based on the quantization accuracy to obtain the target model.
[0007] The present application provides an electronic device, including a memory and a processor, wherein the processor is used to execute program instructions stored in the memory to implement any of the above methods.
[0008] The present application provides a computer-readable storage medium on which program instructions are stored. When the program instructions are executed by a processor, any of the above methods is implemented.
[0009] In the above scheme, the processing device uses the target model stored in it to perform computer vision processing on the image data to obtain computer vision processing results. Using a target model with smaller size and computational complexity occupies less storage space and consumes less energy to run the model, thereby reducing the requirements of computer vision processing on hardware devices.
[0010] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to illustrate the technical solution of the present application.
[0012] Figure 1 It is a flowchart of an embodiment of the computer vision processing method of the present application;
[0013] Figure 2 This is a flow chart of an embodiment of the quantitative accuracy acquisition step 1 of the present application;
[0014] Figure 3 It is a schematic diagram of the framework of an embodiment of the computer vision processing device of the present application;
[0015] Figure 4 It is a schematic diagram of the framework of an embodiment of the electronic device of the present application;
[0016] Figure 5 It is a schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION
[0017] The scheme of the embodiment of the present application is described in detail below in conjunction with the drawings of the specification.
[0018] In the following description, for the purpose of explanation rather than limitation, specific details such as specific subsystem structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.
[0019] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there may be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the objects associated before and after are in an "or" relationship. In addition, "many" in this article means two or more than two. In addition, the term "at least one" in this article means any combination of at least two of any one or more of a plurality of, for example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C.
[0020] See also Figure 1 , Figure 1 1 is a flowchart of an embodiment of a computer vision processing method of the present application. The method may be executed by a processing device. Specifically, the method may include:
[0021] Step S110: Acquire image data.
[0022] Step S120: Perform computer vision processing on the image data using the target model stored in the processing device to obtain a computer vision processing result.
[0023] Among them, computer vision processing can be target detection, target recognition, target classification, etc.
[0024] In a specific application scenario, the target model is used to perform object detection on image data to detect animals present in the image data.
[0025] Specifically, the target model can be obtained through the following processing: obtaining the quantization accuracy corresponding to the original model; obtaining several first substructures from the original model, determining the equalization factors corresponding to each first substructure, and adjusting the model parameters of the corresponding first substructures in the original model based on each equalization factor to obtain the model to be quantized; performing target quantization processing on the model parameters of the model to be quantized based on the quantization accuracy to obtain the target model.
[0026] The first substructure may be selected from the original model.
[0027] The equalization factors corresponding to the respective first substructures may be determined separately. Specifically, the equalization factors corresponding to the first substructures are determined based on the first reference processing result of the first substructure on the first sample image.
[0028] The balancing factor may be a parameter used to adjust the distribution interval of the model parameters.
[0029] In some embodiments, the equalization factor can narrow the distribution range of the model parameters.
[0030] The first substructure may include one or more layers. If the first substructure includes one layer, the corresponding equalization factor is the equalization factor of the layer. If the first substructure includes multiple layers, the corresponding equalization factor includes the equalization factors of each layer.
[0031] In some implementation scenarios, a first substructure may include multiple layers, wherein the equalization factors of different layers may be the same or different.
[0032] It can be understood that the equalization factor can be configured as a learnable parameter and learned based on the first sample image.
[0033] In some embodiments, the processing device may pre-acquire the target model and store the target model. Compared with the original model, the target model is smaller in size, occupies less storage space, and has a lower cost for transmitting the model. The model also requires less computation, which enables more efficient computing, improves the efficiency of computer vision processing, reduces the consumption of model operation, and reduces the requirements of computer vision processing on hardware devices.
[0034] Furthermore, in the process of obtaining the target model, a corresponding equalization factor is set for each first substructure, and the model parameters of the first substructure are adjusted, thereby adjusting the distribution range of the model parameters and reducing the quantization error.
[0035] In some embodiments, for each first substructure, an initial value of an equalization factor corresponding to the first substructure is determined; the first calibration input is processed using the first substructure to obtain a first reference processing result, and the first calibration input is processed using a first quantization structure to obtain a first quantization processing result; based on the difference between the first reference processing result and the first quantization processing result, the equalization factor corresponding to the first substructure is adjusted.
[0036] The first calibration input is the first sample image, or the preprocessing result of the first sample image. The first quantization structure is obtained by adjusting the model parameters in the first substructure using the current equalization factor.
[0037] In some embodiments, if the first substructure includes the first layer of the original model, the first calibration input of the first substructure and the corresponding first quantization structure is the first sample image. If the first substructure is an intermediate structure of the original model, after the first sample image is input into the original model, the first calibration input is obtained after being processed by the previous module, as the input of the first substructure and the corresponding first quantization structure. The processing by the previous module can be used as preprocessing, and the first calibration input can be the preprocessing result of the first sample image. The first substructure corresponds to different previous modules, and the specific steps of the corresponding preprocessing are different.
[0038] It can be understood that the equalization factor is a parameter that affects the model parameters in the first quantization structure. After the equalization factor is adjusted once, the model parameters in the first quantization structure can be updated accordingly following the update of the equalization factor.
[0039] In some implementation scenarios, the initial value of the equalization factor is first determined, and a first quantization structure is constructed based on the initial value. The first reference processing result is obtained using the first substructure, and the current first quantization structure is used to obtain the current first quantization processing result. The equalization factor is adjusted based on the difference between the two, and the first quantization structure is updated accordingly. After the first quantization structure is updated, the steps of obtaining the first reference processing result and the first quantization processing result, and adjusting the equalization factor based on the difference between the two are returned to execute. When the adjustment stop condition is met, the adjustment of the equalization factor is stopped. The adjustment stop condition can be set according to the actual application needs. For example, the number of adjustments reaching a number threshold can be set as the adjustment stop condition.
[0040] In some embodiments, the first quantization structure can be obtained by the following processing: first adjusting the model parameters in the first substructure based on the current equalization factor to obtain the equalization substructure corresponding to the first substructure; quantizing the model parameters in the equalization substructure based on the quantization precision to obtain the first quantization structure. The quantization precision used is the quantization precision corresponding to the first substructure.
[0041] The step of quantizing the model parameters in the equalization substructure can be set according to actual application requirements. For example, it can be symmetric quantization or asymmetric quantization.
[0042] It is understandable that the original model includes multiple layers. The quantization accuracy corresponding to the original model may include the quantization accuracy of each layer. The quantization accuracy of different layers may be different or the same. By obtaining the quantization accuracy corresponding to the original model, the quantization accuracy corresponding to the first substructure and the quantization accuracy corresponding to each layer in the first substructure are obtained.
[0043] In some implementation scenarios, after determining the initial value of the equalization factor, a first adjustment is made to the model parameters in the first substructure based on the current equalization factor to obtain an equalization substructure corresponding to the first substructure, and the equalization substructure is quantized to obtain an initial first quantized structure. After adjusting the equalization factor, the model parameters of the equalization substructure and the model parameters of the first quantized structure are updated.
[0044] The first adjustment object may be a weight in the first substructure.
[0045] In some implementation scenarios, the first substructure includes a layer, and the equalization factor of the first substructure is the equalization factor of the layer. The equalization factor of the layer is used to perform a first adjustment on the weight of the layer to obtain a corresponding equalization substructure.
[0046] In some implementation scenarios, the first substructure includes multiple layers, and the equalization factor of the first substructure includes the equalization factors of the multiple layers. Each layer corresponds to its own equalization factor, and the equalization factor of a layer acts on the model parameters of the layer. The first adjustment of the model parameters in the first substructure based on the equalization factor includes the first adjustment of the model parameters of the corresponding layer based on the equalization factors of each layer.
[0047] In a specific application scenario, the first substructure includes layer A, layer B, and layer C, and the equalization factors of the first substructure include a, b, and c, which are equalization factors of the three layers respectively. The weight of layer A is first adjusted using the equalization factor a, the weight of layer B is first adjusted using the equalization factor b, and the weight of layer C is first adjusted using the equalization factor c.
[0048] It can be understood that the forward calculation of the model can be considered as a calculation based on the input and the model parameters. The model parameters are first adjusted based on the equalization factor, and the input is second adjusted based on the equalization factor. The forward calculation of the model can be equivalent to the calculation of the input after the second adjustment and the model parameters after the first adjustment. The first adjustment and the second adjustment can be inverse operations.
[0049] In a specific application scenario, symmetric quantization is used as an example for explanation. Symmetric quantization can be expressed as the following formula:
[0050]
[0051] in, represents the model parameters of the kth layer, represents the quantization accuracy, Indicates the use Bit Quantization Pseudo-quantized value of the layer. Indicates the quantization step size.
[0052] In a specific application scenario, the equivalent way of forward calculation of the model for one layer can be written as:
[0053]
[0054] Where S represents the equalization factor of the kth layer. Indicates the quantization step size after equalization.
[0055] It is understandable that the equalization factor can be learned layer by layer or block by block.
[0056] The reasoning of the floating point model is X true value, where X represents the input (activation) of the kth layer. The loss function is used to learn the equalization factor of each channel. The specific loss function is expressed as:
[0057]
[0058] The output of this layer of the original model is taken as the true value and compared with the output of the first quantization structure.
[0059] In some implementation scenarios, the first substructure includes multiple layers, and the equalization factor of the first substructure includes the equalization factor of each layer. Performing a second adjustment on the input based on the equalization factor may include: performing a second adjustment on the input of the corresponding layer based on the equalization factor of each layer.
[0060] In a specific application scenario, after determining the model parameters of the current first quantization structure, the first calibration input is processed using the first quantization structure to obtain a first quantization processing result, which may include: for each layer, using the layer to process the corresponding first equalization input to obtain the first output of the layer, the first equalization input corresponding to each layer is obtained by making a second adjustment to the first original input of each layer based on the current equalization factor, wherein the first original input of the first layer in the first quantization structure is the first calibration input, the first original input of the non-first layer is the first output of the previous layer, and the first output of the last layer is the first quantization processing result.
[0061] Further, performing a second adjustment on the first original input of each layer based on the equalization factor to obtain the first equalized input corresponding to each layer may include performing a second adjustment on the first original input of each layer based on the equalization factor corresponding to the layer to obtain the first equalized input corresponding to the layer.
[0062] After the equalization factor is adjusted, the next time the first quantization structure is used for processing, the first original input of each layer is correspondingly also second-adjusted using the current equalization factor.
[0063] In some embodiments, the equalization factor corresponding to a layer is a one-dimensional vector, the number of elements contained in the one-dimensional vector is the same as the number of channels of the corresponding layer, and one element corresponds to one channel. If the first substructure includes multiple layers, the equalization factor of the first substructure includes a one-dimensional vector corresponding to each layer.
[0064] In some implementation scenarios, the elements contained in the equalization factor are between 0 and 1, so that the equalization factor can achieve the effect of narrowing the parameter distribution space.
[0065] In some embodiments, a model parameter of a layer may be divided into sub-parameters corresponding to each channel, and a first adjustment may be performed on the model parameter of a layer by multiplying the sub-parameter corresponding to each channel in the model parameter of the layer by an element corresponding to the channel.
[0066] In some embodiments, the original input may be divided into sub-features corresponding to each channel, and the second adjustment is to divide the sub-features corresponding to each channel in the original input by the elements corresponding to the channel.
[0067] In a specific application scenario, the first adjustment is to multiply the parameter value contained in the sub-parameter by the corresponding element, and the product is used as the new parameter value. The second adjustment is to divide the feature value contained in the sub-feature by the corresponding element, and the quotient is used as the new feature value.
[0068] In some embodiments, the modules in the original model can be divided into a first type of modules and a second type of modules, wherein the first type of modules have scaling invariance and the second type of modules do not have scaling invariance.
[0069] Among them, for the first type of modules, the model parameters therein can be preprocessed using a cross-layer balancing method.
[0070] For the second type of modules, the second type of modules can be divided to obtain a plurality of first substructures, and the model parameters are preprocessed using a learnable equalization factor.
[0071] In a specific application scenario, the common continuous convolution-activation-convolution-activation structure (such as Conv-ReLU-Conv-ReLU structure) in convolutional neural networks (CNN) is a first-class module, and the distribution of weights can be adjusted by cross-layer balancing. This method is very effective for the distribution balancing of grouped convolution and depth-separable convolution.
[0072] In a specific application scenario, the attention module can be used as the second type of module.
[0073] It is understandable that after adjusting the model parameters of the second type of modules in the original model, the model to be quantized can be obtained. Alternatively, in some cases, after adjusting the model parameters of the first type of modules and the second type of modules in the original model respectively, the model to be quantized can be obtained.
[0074] See also Figure 2 , Figure 2 : is a flow chart of an embodiment of the quantitative accuracy acquisition step 1 of the present application. Specifically, this step may include:
[0075] Step S210: Divide the original model into a number of second substructures.
[0076] The second substructure may include one or more layers. If the second substructure includes multiple layers, the multiple layers may use the same quantization accuracy.
[0077] In some embodiments, a layer in the original model serves as a second substructure.
[0078] Step S220: Determine the quantization accuracy corresponding to each second substructure based on the sensitivity of each second substructure to the model output.
[0079] It is understandable that the quantization of model parameters will affect the output of the model, and the impact of the second substructure on the model output can be characterized by sensitivity. The setting of quantization accuracy will affect the sensitivity of the second substructure to the model output, and thus affect the final model performance. For the same second substructure, the higher the sensitivity, the greater the impact of the corresponding quantization accuracy on the model output.
[0080] A plurality of corresponding candidate quantization precisions are set for each second substructure to form a candidate quantization precision set corresponding to each second substructure. The quantization precision finally used by each second substructure is one selected from the corresponding candidate quantization precision set.
[0081] In some embodiments, the candidate quantization precisions of all second substructures are the same. In some cases, different candidate quantization precisions may also be set for different second substructures.
[0082] In some embodiments, the setting of quantization accuracy will affect the properties of the model. In addition to selecting quantization accuracy based on sensitivity, the quantization accuracy can also be selected in combination with the properties of the model. The properties of the model may include properties related to the hardware device on which it is deployed, and may include, for example, at least one of the size of the model and the amount of computation of the model.
[0083] In some embodiments, based on the sensitivity of each second substructure to the model output, determining the quantization accuracy corresponding to each second substructure includes: constructing a candidate quantization accuracy set corresponding to each second substructure; taking the minimum sum of the sensitivities corresponding to all second substructures as the optimization target, searching in the candidate quantization accuracy set corresponding to each second substructure, and obtaining a target search result that meets the optimization target as the quantization accuracy corresponding to the original model. The target search result includes a candidate quantization accuracy selected from each candidate quantization accuracy set.
[0084] A candidate quantization precision is selected from a set as the quantization precision of the second substructure corresponding to the set.
[0085] Furthermore, the sensitivity of the second substructure to the model output can be obtained through the following steps: based on the selected candidate quantization precision, the second substructure is quantized to obtain a second quantization structure, and the second quantization structure is used to replace the second substructure in the original model to obtain a partial quantization model of the second substructure under the quantization precision.
[0086] The quantization process performed on the second substructure may be set according to actual application requirements, and may be, for example, symmetric quantization or asymmetric quantization.
[0087] It can be understood that the original model can be divided into several second substructures, and a second substructure can be quantized and replaced separately to obtain a partial quantization model corresponding to the second substructure, that is, the other parts of the partial quantization model except the second substructure are the same as the original model.
[0088] In a specific application scenario, a layer is used as a second substructure for example. The model has a total of K layers, of which the kth layer is the second substructure currently being processed. The parameters W of each layer of the original model are expressed as The optional candidate quantization precision set of this layer is expressed as . Indicates the use Bit Quantization Pseudo quantization value of the layer. Use The partial quantization model obtained by quantizing the kth layer can be expressed as .
[0089] Computer vision processing is performed on the second sample image using the partial quantization model and the original model respectively to obtain a second reference processing result and a second quantization processing result. Based on the difference between the second reference processing result and the second quantization processing result, the sensitivity of the second substructure under the candidate quantization accuracy is determined.
[0090] In some embodiments, the difference between the second reference processing result and the second quantization processing result may be directly used as the sensitivity of the second substructure under the candidate quantization precision.
[0091] In a specific application scenario, the Euclidean distance may be used to represent the difference between the second reference processing result and the second quantization processing result, and may be used as the sensitivity of the second substructure under a certain candidate quantization precision.
[0092] In some embodiments, a target loss may be determined based on the second quantization processing result and the true label of the second sample image, and a reference loss may be determined based on the second reference processing result and the true label of the second sample image. Based on the difference between the reference loss and the target loss, the sensitivity of the second substructure under the candidate quantization accuracy is determined.
[0093] In a specific application scenario, a single layer is used as an example. For a supervised learning, the loss function of the model can be expressed as:
[0094]
[0095] in, Represents each sample data and label respectively, and L represents the loss function. We use The bit configuration of the dequantized model , then the corresponding loss value calculation can be expressed as:
[0096]
[0097] in, Indicates the use Bit Quantization The pseudo-quantized value of the layer, Represents the quantization step size. Taking symmetric quantization as an example, it can be expressed as:
[0098]
[0099] but The set of mixed precision loss values for a layer can be expressed as:
[0100]
[0101] Define usage Bit Quantization The sensitivity of the layer is:
[0102]
[0103] but The mixed precision sensitivity set of a layer can be expressed as:
[0104]
[0105] The optimization objective function can be expressed as:
[0106]
[0107] In some embodiments, the number of the second sample images may be multiple. Further, the differences between the second reference processing results and the second quantization processing results of all the second sample images may be accumulated to represent the sensitivity. Alternatively, the difference values corresponding to all the second sample images may be counted, and the central trend statistical value, such as the average value, may be used to represent the sensitivity.
[0108] Similarly, the difference between the reference loss and the target loss of all the second sample images can be accumulated to represent the sensitivity. The difference values corresponding to all the second sample images can also be counted, and the central tendency statistical value, such as the average value, can be used to represent the sensitivity.
[0109] In some embodiments, the optimization target may further include hardware constraints. The hardware constraints may include at least one of a model size constraint and a model computation amount constraint. The model size and the model computation amount are related to the selected quantization accuracy.
[0110] In some implementation scenarios, the model size constraint may include that the size of the target model is less than or equal to a model size threshold.
[0111] In some implementation scenarios, the model computation amount constraint may include that the computation amount of the target model is less than or equal to a model computation amount threshold.
[0112] In a specific application scenario, the size of the target model can be expressed as:
[0113]
[0114] in, Indicates The number of bits used by the layer, Indicates The number of parameters for the layer.
[0115] In a specific application scenario, the computational complexity of the target model can be expressed as:
[0116]
[0117] in, Indicates the amount of computation at the kth layer, which is related to the quantization accuracy used. Quantization can reduce the amount of model computation, and different quantization accuracy can reduce the amount by different multiples. The amount of computation of the target model can be calculated based on the computation of the original model and the reduction multiple corresponding to the quantization accuracy used.
[0118] In some embodiments, target quantization processing is performed on model parameters of the model to be quantized based on quantization accuracy to obtain a target model, which may include: dividing the adjusted original model into several third substructures; for each third substructure, determining the truncation factor and the initial value of the quantization adjustment parameter corresponding to the third substructure; and performing the following preset processing on each parameter to be adjusted: adjusting the parameter to be adjusted based on the difference between the third reference processing result and the third quantization processing result.
[0119] The truncation factor is a parameter used to adjust the quantization step size. The quantization adjustment parameter can adjust the quantization result.
[0120] In a specific application scenario, the value of the truncation factor may be between 0 and 1.
[0121] Wherein, each third substructure is processed as above to obtain the final target model.
[0122] The parameters to be adjusted include at least one of a truncation factor, a quantization adjustment parameter and a model parameter of the third substructure.
[0123] The third reference processing result is obtained by processing the second calibration input using the fourth substructure. The fourth substructure is a structure in the original model corresponding to the third substructure.
[0124] The third quantization processing result is obtained by processing the second calibration input using the third quantization structure. The third quantization structure can be obtained by performing a target quantization operation on the model parameters of the current third substructure using the quantization accuracy corresponding to the third substructure, the current truncation factor and the current quantization adjustment parameter.
[0125] The adjustment of the parameter to be adjusted will affect the model parameters of the third quantization structure. After the parameter to be adjusted is adjusted, the third quantization structure can be updated accordingly.
[0126] After the parameters to be adjusted are adjusted and no longer change, the third quantization structure is fixed, and all the third quantization structures can form a target model.
[0127] The second calibration input may be the third sample image or a preprocessing result of the third sample image. For details, reference may be made to the relevant description about the first sample image in the above embodiment.
[0128] It can be understood that the fourth substructure and the third quantization substructure have the same input, and the outputs of the fourth substructure and the third quantization substructure are compared, so as to determine the difference between the third quantization substructure and the original original model.
[0129] In a specific application scenario, take one layer as an example to introduce a learnable truncation factor , to select the appropriate cutoff range, which can be expressed as follows:
[0130]
[0131] in, Represents the parameters of the model to be quantized, which are the results after being adjusted by the equalization factor.
[0132] In some embodiments, the third substructure includes one or more layers. The truncation factor includes the truncation factor of each layer, and the quantization adjustment parameter includes the quantization adjustment parameter of each layer. The truncation factors of different layers may be the same or different, and the quantization adjustment parameters of different layers may be the same or different.
[0133] And, each layer in the third substructure has its own quantization accuracy and equalization factor. It should be noted that when the model to be quantized is constructed based on the target model, the model parameters have been adjusted using the equalization factor. During the target quantization process, the model will process the input. At this time, it is necessary to use the predetermined equalization factor corresponding to the layer to make a second adjustment to the layer input.
[0134] In some embodiments, using the quantization accuracy corresponding to the third substructure, the current truncation factor and the current quantization adjustment parameter, the target quantization operation on the model parameters of the current third substructure may include: for each layer in the third substructure, the product of the truncation factor of the current layer and the quantization step length is used as the quantization step length of the new layer, wherein the quantization step length is determined based on the model parameters of the current layer and the corresponding quantization accuracy. The product of the rounded-down result of the target parameter and the quantization step length of the layer is used as the quantization parameter of the layer, thereby obtaining the third quantization structure.
[0135] The target parameter is the sum of the target mapping result and the target ratio. The target mapping result is obtained by target mapping the quantization adjustment parameter of the current layer. The target ratio is the ratio of the model parameter of the current layer to the quantization step size of the layer.
[0136] In a specific application scenario, taking one layer as an example, for adaptive weight quantization, we use the general method:
[0137]
[0138] in, represents the quantitative adjustment parameter, Yes and The learnable parameters of the same dimension can be initialized as .in, Represents the parameters of the model to be quantized, which are the results after being adjusted by the equalization factor.
[0139] In some embodiments, for the model parameters, based on the difference between the third reference processing result and the third quantization processing result, adjusting the parameter to be adjusted may include: using the fourth substructure to process the second calibration input to obtain the third reference processing result, and using the third quantization structure to process the second calibration input to obtain the third quantization processing result. Based on the difference between the third reference processing result and the third quantization processing result corresponding to all the third sample images, a quantization loss is obtained, and the model parameters are adjusted once based on the quantization loss.
[0140] It is understandable that there may be multiple third sample images. Each time a batch of images are selected for processing, the third sample images may be selected for processing multiple times. Whenever all the third sample images are processed once, the differences between the third reference processing results and the third quantization processing results corresponding to all the third sample images are accumulated to obtain the quantization loss, and the model parameters are adjusted once based on the accumulated quantization loss.
[0141] For the truncation factor and the quantization adjustment parameter, they can be set to be updated once after all the third sample images are processed once, or they may not be limited by the update condition. Exemplarily, they are updated once after each batch of third sample images is processed.
[0142] In some embodiments, when a stop condition is met, the adjustment of the parameter to be adjusted is stopped. The stop condition can be set according to actual application needs. For example, it can be set to stop after all third sample images are processed for a preset number of times.
[0143] In a specific application scenario, for a third substructure and a corresponding third quantization structure, it can be set that after all third sample images have been processed 20 times, the adjustment of the parameter to be adjusted is stopped.
[0144] In a specific application scenario, all third sample images can be divided into 5 batches. After all third sample images are processed once, the truncation factor, quantization adjustment parameter and model parameters of the third substructure are adjusted together. After each batch of images is processed, the gradient is not cleared.
[0145] In a specific application scenario, all third sample images can be divided into 5 batches. After each batch of images is processed, the truncation factor and the quantization adjustment parameter are adjusted once, and the model parameters of the third quantization structure are updated accordingly. After all third sample images are processed once, that is, after all 5 batches of images are processed, the model parameters of the third substructure are adjusted, and the model parameters of the third quantization structure are updated accordingly.
[0146] In a specific application scenario, the stop condition may be set to end the adjustment of the parameters to be adjusted after all the third sample images have been processed seven times.
[0147] Furthermore, the adjustment step of the model parameters can be controlled, so that the range in which the model parameters can be adjusted can be limited. In this way, it is possible to avoid the weight of the target model being too different from the weight of the original model.
[0148] In some implementation scenarios, the gradient of the quantization loss relative to the model parameters can be used to determine the adjustment direction of the model parameters, and the learning rate is set as the adjustment step of the model parameters; the model parameters are adjusted based on the adjustment step and the adjustment direction, thereby limiting the adjustment step.
[0149] In a specific application scenario, the preset processing steps of the model parameters are executed a preset number of times, and the product of the preset number of times and the learning rate is a preset value. On the one hand, the number of adjustments of the model parameters is limited, and on the other hand, the adjustment step size is limited, thereby limiting the range in which the model parameters can be adjusted.
[0150] In a specific application scenario, the optimizer SignSGD based on gradient sign learning is used, and the weight update form is:
[0151]
[0152] in, represents the learning rate, Represents the parameters of the model to be quantized, which is the result after being adjusted by the equalization factor. express Assume that the number of rounds of tuning for each layer is epochs, using the range of the Sign function , set the learning rate to:
[0153]
[0154] Therefore, the weights are updated once per epoch (the model loss gradient for each batch is not cleared), so the weights go through The maximum range of epoch tuning is:
[0155]
[0156] This method can optimize the rounding error and limit the change of model weights, ensuring the stability of the original model weights. At the same time, it only requires less parameter update cost and tuning iteration rounds.
[0157] It is understandable that in the target quantization process, not only the model parameters are quantized, but also the input of each layer of the model is quantized.
[0158] In some embodiments, the third quantization structure processes the second calibration input to obtain a third quantization processing result, which may include: for each layer in the third quantization structure, using the layer to process the corresponding quantization input to obtain the second output of the layer, wherein the quantization input of the layer is obtained by quantizing the second equalized input of the layer based on a preset quantization ratio, and the second equalized input is obtained by adjusting the second original input of the layer based on the equalization factor.
[0159] Among them, the second original input of the first layer in the third quantization structure is the second calibration input, the second original input of the non-first layer is the second output of the previous layer, and the second output of the last layer is the third quantization processing result.
[0160] In a specific application scenario, the designed tuning method adopts Loss function for learning:
[0161]
[0162] in, Indicates the quantization method of activation The output of the corresponding module in the original model is taken as the true value, and the output of the third quantization structure is compared with the true value.
[0163] in, Represents the output of the corresponding module of the original model, Represents the parameters in the original model, which have not been adjusted by the equalization factor. middle, Represents the parameters of the model to be quantized, which are the results after being adjusted by the equalization factor.
[0164] The quantization of the second equalization input may be symmetric quantization or asymmetric quantization, that is, the quantization process does not include the parameter to be adjusted. In some cases, the quantization of the second equalization input may be a target quantization operation, and the specific contents of the target quantization operation may be referred to above.
[0165] Specifically, the equalization factor is determined in advance and includes the equalization factor of each layer, and the equalization factor of the layer is used to perform a second adjustment on the second original input of the corresponding layer to obtain the second equalized input. For details, reference may be made to the relevant description of the second adjustment in the above embodiment.
[0166] The preset quantization ratio may be randomly generated. Each layer corresponds to a preset quantization ratio, and different layers may correspond to the same or different preset quantization ratios. The same layer may correspond to the same or different preset quantization ratios when processing different batches of images.
[0167] Specifically, a value to be quantized can be selected from the second equalized input, and the ratio of the value to be quantized to the total number of characteristic values in the second equalized input is equal to the quantization ratio. The value to be quantized is quantized to obtain a quantization result, and the value to be quantized in the second equalized input is replaced with the corresponding quantization result to obtain a quantized input.
[0168] In a specific application scenario, in order to alleviate the dependence of quantization tuning on the calibration data set (sample image), during the tuning process, the quantization of the activation values of each layer is randomly discarded, that is, some activations are quantized and some are not quantized. The random factor is a dynamic value, denoted as The final optimization objective function is:
[0169]
[0170] in , Indicates the proportion of randomly selected quantization .
[0171] It can be understood that the first sample image, the second sample image and the third sample image may be the same or different.
[0172] In some embodiments, the second substructure is a layer, and the first substructure may include one or more second substructures, and the fourth substructure may include one or more second substructures.
[0173] In some embodiments, the original model may be a pre-trained neural network model, and may include a large language model, etc.
[0174] This method guides the selection of mixed precision of the model based on sensitivity analysis and hardware-aware dynamic programming search, so as to identify the model's sensitive layers and high-latency layers and achieve lower bit compression. A learnable equivalent equalization strategy is designed to eliminate outliers that are not friendly to quantization. A block-by-block random tuning strategy is set to reduce the rounding error and truncation error of quantization while ensuring the generalization of the quantized model. This solution realizes the low-latency and lossless compression deployment of the model on the end-side chip, greatly reducing the resources and costs of model deployment.
[0175] See also Figure 3 , Figure 3 It is a schematic diagram of the framework of an embodiment of a computer vision processing device of the present application.
[0176] In this embodiment, the computer vision processing device 30 includes an acquisition module 31 and a computer vision processing module 32. The acquisition module 31 is used to acquire image data; the computer vision processing module 32 is used to perform computer vision processing on the image data using the stored target model to obtain a computer vision processing result; wherein the target model is obtained by the following processing: obtaining the quantization accuracy corresponding to the original model; obtaining a number of first substructures from the original model, and determining the equalization factor corresponding to each first substructure based on the first reference processing result of each first substructure on the first sample image; adjusting the model parameters of the corresponding first substructure in the original model based on each equalization factor to obtain the model to be quantized; the equalization factor is a parameter used to adjust the distribution interval of the model parameter; performing target quantization processing on the model parameters of the model to be quantized based on the quantization accuracy to obtain the target model.
[0177] See also Figure 4 , Figure 4 It is a schematic diagram of the framework of an embodiment of the electronic device of the present application.
[0178] The electronic device 40 includes a memory 41 and a processor 42, and the processor 42 is used to execute the program instructions stored in the memory 41 to implement the steps in any of the above computer vision processing method embodiments. In a specific implementation scenario, the electronic device 40 may include, but is not limited to: computer equipment, electrical equipment, microcomputers, desktop computers, servers, and in addition, the electronic device 40 may also include mobile devices such as laptops and tablet computers, which are not limited here.
[0179] Specifically, the processor 42 is used to control itself and the memory 41 to implement the steps in any of the above-mentioned computer vision processing method embodiments. The processor 42 can also be called a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 42 can be implemented by an integrated circuit chip.
[0180] See also Figure 5 , Figure 5It is a schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application.
[0181] The computer-readable storage medium 50 provided in this embodiment stores program instructions 51 that can be executed by a processor. When the program instructions 51 are executed by the processor, they are used to implement the steps in any of the above-mentioned computer vision processing method embodiments.
[0182] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.
[0183] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as units or components can be combined or integrated into another subsystem, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0184] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or part of the contribution to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (processor) to perform all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.
Claims
1. A computer vision processing method, characterized in that: The method comprises: The processing device acquires the image data; Performing computer vision processing on the image data using the target model stored in the processing device to obtain a computer vision processing result; The target model is obtained by the following processing: Get the quantization accuracy corresponding to the original model; Acquire a plurality of first substructures from the original model, and for each of the first substructures, determine an initial value of an equalization factor corresponding to the first substructure; process a first calibration input using the first substructure to obtain a first reference processing result; and process the first calibration input using a first quantization structure to obtain a first quantization processing result; and adjust the equalization factor corresponding to the first substructure based on a difference between the first reference processing result and the first quantization processing result; Adjusting the model parameters of the first substructure corresponding to the original model based on each of the equalization factors to obtain a model to be quantized; the equalization factor is a parameter used to adjust the distribution interval of the model parameters; Based on the quantization accuracy, target quantization processing is performed on the model parameters of the model to be quantized to obtain the target model.
2. The method according to claim 1, characterized in that The first calibration input is a first sample image or a preprocessing result of the first sample image; the first quantization structure is obtained by adjusting the model parameters in the first substructure based on the current equalization factor.
3. The method according to claim 2, characterized in that The first quantization structure is obtained by the following processing: Performing a first adjustment on the model parameters in the first substructure based on the current equalization factor to obtain an equalization substructure corresponding to the first substructure; performing quantization processing on the model parameters in the equalization substructure based on the quantization accuracy corresponding to the first substructure to obtain the first quantization structure; The first quantization structure includes several layers; and the first calibration input is processed by using the first quantization structure to obtain a first quantization processing result, including: For each of the layers, the corresponding first equalized input is processed using the layer to obtain the first output of the layer, and the first equalized input of each layer is obtained by performing a second adjustment on the first original input of each layer based on the current equalization factor; wherein, the first original input of the first layer in the first quantization structure is the first calibration input, the first original input of the non-first layer is the first output of the previous layer, and the first output of the last layer is the first quantization processing result.
4. The method according to claim 3, characterized in that: The first substructure includes one or more layers; the equalization factor corresponding to the first substructure includes the equalization factor of each layer; the equalization factor of the layer is a one-dimensional vector, and the number of elements contained in the equalization factor of the layer is the same as the number of channels of the corresponding layer; The first adjustment is to multiply the sub-parameter corresponding to each of the channels in the model parameters of the layer by the element corresponding to the channel; The second adjustment is to divide the sub-feature corresponding to each of the channels in the first original input by the element corresponding to the channel.
5. The method according to claim 1, characterized in that The original model includes a first type of module and a second type of module, the first type of module has scaling invariance, and the plurality of first substructures are obtained by dividing the second type of module.
6. The method according to claim 5, characterized in that The second type of module is the attention mechanism module.
7. The method according to claim 1, characterized in that The obtaining of the quantization accuracy corresponding to the original model includes: Dividing the original model into a plurality of second substructures; the second substructures include one or more layers; Based on the sensitivity of each of the second substructures to the model output, the quantization accuracy corresponding to each of the second substructures is determined.
8. The method according to claim 7, characterized in that The determining, based on the sensitivity of each of the second substructures to the model output, the quantization accuracy corresponding to each of the second substructures comprises: Constructing candidate quantization precision sets corresponding to each of the second substructures respectively; The minimization of the sum of the sensitivities corresponding to all the second substructures is taken as the optimization goal, and a search is performed in the candidate quantization precision set corresponding to each of the second substructures to obtain a target search result that meets the optimization goal as the quantization precision corresponding to the original model, and the target search result includes a candidate quantization precision selected from each of the candidate quantization precision sets.
9. The method according to claim 8, characterized in that For each of the second substructures, the sensitivity corresponding to the second substructure is obtained through the following steps: Based on the selected candidate quantization precision, quantize the second substructure to obtain a second quantization structure, and use the second quantization structure to replace the second substructure in the original model to obtain a partial quantization model of the second substructure under the candidate quantization precision; Performing computer vision processing on the second sample image using the partial quantization model and the original model respectively to obtain a second reference processing result and a second quantization processing result; Based on the difference between the second reference processing result and the second quantization processing result, the sensitivity of the second substructure under the candidate quantization precision is determined.
10. The method according to claim 8, characterized in that The optimization target also includes hardware constraints, and the hardware constraints include at least one of a model size constraint and a model calculation amount constraint.
11. The method according to claim 1, characterized in that The performing target quantization processing on the model parameters of the to-be-quantized model based on the quantization accuracy to obtain the target model comprises: Dividing the model to be quantized into a plurality of third substructures, and determining, for each of the third substructures, an initial value of a truncation factor and a quantization adjustment parameter corresponding to the third substructure; Perform the following preset processing on each parameter to be adjusted: Based on the difference between the third reference processing result and the third quantization processing result, the parameter to be adjusted is adjusted; wherein the parameter to be adjusted includes at least one of the truncation factor, the quantization adjustment parameter and the model parameter of the third substructure; the third reference processing result is obtained by processing the second calibration input using the fourth substructure, the third quantization processing result is obtained by processing the second calibration input using the third quantization structure, the fourth substructure is a structure corresponding to the third substructure in the original model, and the third quantization structure is obtained by performing a target quantization operation on the model parameters of the current third substructure using the quantization accuracy corresponding to the third substructure, the current truncation factor and the quantization adjustment parameter; the second calibration input is a third sample image or a preprocessing result of the third sample image; The target model is formed by using the third quantized structures after the preset processing.
12. The method according to claim 11, characterized in that The third substructure includes one or more layers; the truncation factor includes the truncation factor of each layer, and the quantization adjustment parameter includes the quantization adjustment parameter of each layer; And / or, using the quantization accuracy corresponding to the third substructure, the current truncation factor and the quantization adjustment parameter to perform a target quantization operation on the current model parameters of the third substructure includes: For each layer in the third substructure, the product of the truncation factor of the current layer and the quantization step size of the layer is used as a new quantization step size of the layer; the quantization step size of the layer is determined based on the model parameters of the current layer and the corresponding quantization precision; The product of the rounded-down result of the target parameter and the quantization step of the layer is used as the quantization parameter of the layer to obtain the third quantization structure; the target parameter is the sum of the target mapping result and the target ratio, the target mapping result is obtained by target mapping the quantization adjustment parameter of the current layer, and the target ratio is the ratio of the model parameter of the current layer to the quantization step of the layer.
13. The method according to claim 11 or 12, characterized in that: The adjusting the parameter to be adjusted based on the difference between the third reference processing result and the third quantization processing result comprises: Processing the second calibration input using the fourth substructure to obtain the third reference processing result, and processing the second calibration input using the third quantization structure to obtain the third quantization processing result; Obtaining a quantization loss based on a difference between the third reference processing result and the third quantization processing result corresponding to all the third sample images; Using the learning rate as the adjustment step size of the model parameter, determining the adjustment direction of the model parameter based on the gradient of the quantization loss relative to the model parameter, and adjusting the model parameter based on the adjustment step size and the adjustment direction; The preset processing step of the model parameter is executed a preset number of times, and the product of the preset number of times and the learning rate is a preset value.
14. The method according to claim 13, characterized in that The third quantization structure includes several layers; and the processing of the second calibration input by using the third quantization structure to obtain the third quantization processing result includes: For each of the layers, the corresponding quantized input is processed using the layer to obtain the second output of the layer, the quantized input of each layer is obtained by quantizing the second equalized input of the layer based on a preset quantization ratio, and the second equalized input is obtained by adjusting the second original input of the layer based on the equalization factor; wherein, the second original input of the first layer in the third quantization structure is the second calibration input, the second original output of the non-first layer is the second output of the previous layer, and the second output of the last layer is the third quantization processing result.
15. The method according to claim 14, characterized in that The quantizing the second equalized input of the layer based on the preset quantization ratio includes: Selecting a value to be quantized from the second equalized input, wherein the ratio of the value to be quantized to the total number of characteristic values in the second equalized input is equal to the preset quantization ratio; Quantizing the value to be quantized to obtain a quantization result; The to-be-quantized value in the second equalized input is replaced with the corresponding quantization result to obtain the quantized input.
16. A computer vision processing device, characterized in that: The device comprises: An acquisition module, used for acquiring image data; A computer vision processing module, used to perform computer vision processing on the image data using the stored target model to obtain a computer vision processing result; The target model is obtained by the following processing: Get the quantization accuracy corresponding to the original model; A plurality of first substructures are obtained from the original model, and for each of the first substructures, an initial value of an equalization factor corresponding to the first substructure is determined; a first calibration input is processed using the first substructure to obtain a first reference processing result; and a first quantization structure is used to process the first calibration input to obtain a first quantization processing result; based on the difference between the first reference processing result and the first quantization processing result, the equalization factor corresponding to the first substructure is adjusted; model parameters of the corresponding first substructure in the original model are adjusted based on each of the equalization factors to obtain a model to be quantized; the equalization factor is a parameter used to adjust the distribution interval of the model parameter; Based on the quantization accuracy, target quantization processing is performed on the model parameters of the model to be quantized to obtain the target model.
17. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores program instructions, and when the program instructions are executed by the processor, the method according to any one of claims 1 to 15 is implemented.
18. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the method according to any one of claims 1 to 15 is implemented.
Citation Information
Patent Citations
Convolutional neural network model compression method and device combining quantization and pruning
CN116384470A
Systems and Methods of Cross Layer Rescaling for Improved Quantization Performance
US20200302299A1