Image Processing Method and Related Methods, Devices, Electronic Devices, and Storage Media

By quantizing and cropping training of the image processing model, the cropping coefficients are dynamically adjusted to determine the numerical range of weight parameters, the accuracy loss problem during the deployment of neural network models in the prior art is solved, and higher processing task accuracy and inference speed are achieved.

CN119418175BActive Publication Date: 2025-06-17ANHUI LISTENAI CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510032189.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2024-12-04
Filing Date
2025-01-09
Publication Date
2025-06-17
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

The prior art has high accuracy losses when deploying neural network models, making it difficult to improve the accuracy of tasks such as image processing.

Method used

By extracting the image feature representation of the to-process image and quantizing the image processing model based on the quantization tool of the current device, the quantization model is obtained. The model is at least trimming training, and the cropping coefficient is decreasing to determine the numerical range of weight parameters, and then the final result is obtained through solution quantization processing.

Benefits of technology

This method can ensure the inference speed of neural network model while improving the accuracy of the model in processing tasks, and dynamically decreasing the crop coefficient to make the weight parameters close to uniform distribution, alleviating the accuracy loss caused by quantization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119418175B_ABST
    Figure CN119418175B_ABST
Patent Text Reader

Abstract

The present application discloses an image processing method and related methods, devices, electronic devices, and storage media. Among them, the image processing method includes: extracting an image feature representation of an image to be processed, and quantifying an image processing model deployed on a current device based on a quantization tool on the current device to obtain a first quantized model; wherein, the image processing model is obtained by at least pruning training, and during the pruning process, a pruning coefficient is decreased, and the pruning coefficient is used to determine a numerical range for pruning weight parameters in the image processing model; processing the image feature representation based on the first quantized model to obtain a temporary processing result; and dequantifying the temporary processing result based on the quantization tool to obtain a final processing result of the image to be processed. The above solution can ensure the inference speed of the neural network model as much as possible and improve the accuracy of the neural network model in its processing tasks.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of a Chinese patent application with the application number 2024117727591 and the application title "Image Processing Method and Related Methods, Devices, Electronic Devices, and Storage Media" filed with the National Intellectual Property Administration on December 04, 2024, the entire content of which is incorporated herein by reference. Technical Field

[0002] This application relates to the field of artificial intelligence technology, and in particular, to an image processing method and related methods, devices, electronic devices, and storage media. Background Art

[0003] In recent years, neural network models represented by tasks such as images, speech, and natural language have shown great application value. After the neural network model is trained, it is usually deployed on a server or an edge (embedded) device for inference.

[0004] Generally, the neural network model after floating-point training needs to be converted into a low-bit model through quantization technology to improve the model inference speed. However, when deploying neural network models, such as image processing models, through existing technologies, there is still a high accuracy loss, making it difficult to improve the accuracy of image processing and other tasks. In view of this, how to ensure the inference speed of the neural network model as much as possible and improve the accuracy of the neural network model in its processing tasks has become an urgent problem to be solved. Summary of the Invention

[0005] The main technical problem to be solved by this application is to provide an image processing method and related methods, devices, electronic devices, and storage media, which can ensure the inference speed of the neural network model as much as possible and improve the accuracy of the neural network model in its processing tasks.

[0006] To solve the above technical problem, in the first aspect of this application, an image processing method is provided, including: extracting an image feature representation of an image to be processed, and quantifying an image processing model deployed on the current device based on a quantization tool on the current device to obtain a first quantized model; wherein, the image processing model is obtained by at least pruning training, and during the pruning process, the pruning coefficient is decreased, and the pruning coefficient is used to determine the numerical range of pruning the weight parameters in the image processing model; processing the image feature representation based on the first quantized model to obtain a temporary processing result; and dequantifying the temporary processing result based on the quantization tool to obtain the final processing result of the image to be processed.

[0007] To solve the above technical problems, a second aspect of the present application provides a voice processing method, including: extracting a voice feature representation of the voice to be processed, and quantizing the voice processing model deployed on the current device based on a quantization tool on the current device to obtain a second quantized model; wherein, the voice processing model is obtained by at least pruning training, and during the pruning process, a pruning coefficient is decreased, and the pruning coefficient is used to determine the numerical range of pruning the weight parameters in the voice processing model; processing the voice feature representation based on the second quantized model to obtain a temporary processing result; and dequantizing the temporary processing result based on the quantization tool to obtain the final processing result of the voice to be processed.

[0008] To solve the above technical problems, a third aspect of the present application provides a natural language processing method, including: extracting a statement feature representation of the statement to be processed, and quantizing the natural language processing model deployed on the current device based on a quantization tool on the current device to obtain a third quantized model; wherein, the natural language processing model is obtained by at least pruning training, and during the pruning process, a pruning coefficient is decreased, and the pruning coefficient is used to determine the numerical range of pruning the weight parameters in the natural language processing model; processing the statement feature representation based on the third quantized model to obtain a temporary processing result; and dequantizing the temporary processing result based on the quantization tool to obtain the final processing result of the statement to be processed.

[0009] To solve the above technical problems, a fourth aspect of the present application provides an image processing device, including: an extraction quantization module, a model processing module, and a dequantization module. The extraction quantization module is configured to extract an image feature representation of the image to be processed, and quantize the image processing model deployed on the current device based on a quantization tool on the current device to obtain a first quantized model; wherein, the image processing model is obtained by at least pruning training, and during the pruning process, a pruning coefficient is decreased, and the pruning coefficient is used to determine the numerical range of pruning the weight parameters in the image processing model; the model processing module is configured to process the image feature representation based on the first quantized model to obtain a temporary processing result; and the dequantization module is configured to dequantize the temporary processing result based on the quantization tool to obtain the final processing result of the image to be processed.

[0010] To solve the above technical problems, a fifth aspect of the present application provides a voice processing device, including: an extraction and quantization module, a model processing module, and a dequantization module. The extraction and quantization module is configured to extract a voice feature representation of the voice to be processed, and quantize the voice processing model deployed on the current device based on the quantization tool on the current device to obtain a second quantized model; wherein, the voice processing model is obtained by at least pruning training, and during the pruning process, a pruning coefficient is decreased, and the pruning coefficient is used to determine the numerical range of pruning the weight parameters in the voice processing model. The model processing module is configured to process the voice feature representation based on the second quantized model to obtain a temporary processing result. The dequantization module is configured to dequantize the temporary processing result based on the quantization tool to obtain the final processing result of the voice to be processed.

[0011] To solve the above technical problems, a sixth aspect of the present application provides a natural language processing device, including: an extraction and quantization module, a model processing module, and a dequantization module. The extraction and quantization module is configured to extract a statement feature representation of the statement to be processed, and quantize the natural language processing model deployed on the current device based on the quantization tool on the current device to obtain a third quantized model; wherein, the natural language processing model is obtained by at least pruning training, and during the pruning process, a pruning coefficient is decreased, and the pruning coefficient is used to determine the numerical range of pruning the weight parameters in the natural language processing model. The model processing module is configured to process the statement feature representation based on the third quantized model to obtain a temporary processing result. The dequantization module is configured to dequantize the temporary processing result based on the quantization tool to obtain the final processing result of the statement to be processed.

[0012] To solve the above technical problems, a seventh aspect of the present application provides an electronic device, including at least a memory and a processor coupled to each other. The memory stores at least program instructions, and the processor is configured to execute the program instructions to implement the image processing method in the first aspect above, or implement the voice processing method in the second aspect above, or implement the natural language processing method in the third aspect above.

[0013] To solve the above technical problems, an eighth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor. The program instructions are used to implement the image processing method in the first aspect above, or implement the voice processing method in the second aspect above, or implement the natural language processing method in the third aspect above.

[0014] In the above solution, the image feature representation of the image to be processed is extracted, and the image processing model deployed on the current device is quantized based on the quantization tool on the current device to obtain a first quantized model. The image processing model is at least obtained through pruning training. During the pruning training process, the pruning coefficient is decreased. The pruning coefficient is used to determine the numerical range for pruning the weight parameters in the image processing model. Then, the image feature representation is processed based on the first quantized model to obtain a temporary processing result. Thus, the temporary processing result is dequantized based on the quantization tool to obtain the final processing result of the image to be processed. Furthermore, on the one hand, by quantizing the neural network model deployed on the current device and then applying it, the inference speed of the neural network model can be ensured as much as possible. On the other hand, the neural network model is at least obtained through pruning training, and since the pruning coefficient is dynamically decreased during the pruning training process, after the weight parameters are pruned within the numerical range determined by the pruning system during the training process, they gradually stabilize within the range determined by the final pruning coefficient, making the weight parameters as close to a uniform distribution as possible, which helps to alleviate the accuracy loss caused by model quantization as much as possible and can improve the accuracy of the neural network model in its processing tasks as much as possible. Therefore, the inference speed of the neural network model can be ensured as much as possible, and the accuracy of the neural network model in its processing tasks can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a schematic flowchart of an embodiment of the image processing method of the present application;

[0016] Figure 2 is a schematic flowchart of an embodiment of the voice processing method of the present application;

[0017] Figure 3 is a schematic flowchart of an embodiment of the natural language processing method of the present application;

[0018] Figure 4 is a schematic framework diagram of an embodiment of the image processing apparatus of the present application;

[0019] Figure 5 is a schematic framework diagram of an embodiment of the voice processing apparatus of the present application;

[0020] Figure 6 is a schematic framework diagram of an embodiment of the natural language processing apparatus of the present application;

[0021] Figure 7 is a schematic framework diagram of an embodiment of the electronic device of the present application;

[0022] Figure 8 is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The following will describe in detail the solution of the embodiment of the present application with reference to the accompanying drawings of the specification.

[0024] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0025] The terms "system" and "network" in this article are often used interchangeably in this article. The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the fragment " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "plurality" in this article means two or more than two.

[0026] Please refer to Figure 1 , Figure 1 , which is a schematic flowchart of an embodiment of the image processing method of the present application. Specifically, the following steps may be included:

[0027] Step S11: Extract the image feature representation of the image to be processed, and quantize the image processing model deployed on the current device based on the quantization tool on the current device to obtain the first quantization model.

[0028] In the embodiment of the present disclosure, the image processing model is at least obtained through pruning training. During the pruning process, the pruning coefficient is decreased. The pruning coefficient is used to determine the numerical range of the weight parameters in the pruned image processing model. Exemplarily, taking the numerical range represented as [min, max] as an example, where min represents the lower limit value of the numerical range and max represents the upper limit value of the numerical range. Then a possible example of pruning the weight parameters using the numerical range can be: if the weight parameter is greater than the upper limit value max, the weight parameter can be reset to the upper limit value max; if the weight parameter is less than the lower limit value min, the weight parameter can be reset to the lower limit value min; if the weight parameter is within the above numerical range, the weight parameter can remain unchanged. Of course, the above example is only a possible example of pruning the weight parameters using the numerical range, and does not limit other possible ways of pruning the weight parameters accordingly, and no further examples will be given here. In addition, it should be noted that during the pruning training process, the image processing model processes data with the weight parameters after pruning, and adjusts the weight parameters before pruning (i.e., the original weight parameters) based on the running results.

[0029] In an implementation scenario, an image processing model can be used to perform image processing tasks, which can include but are not limited to: object detection, object segmentation, object classification, etc. That is to say, the image processing model can be an object detection model (such as YOLO, etc.), or an object segmentation model (such as U-Net, etc.), or an object classification model. Here, the specific image processing tasks performed by the image processing model and the network model structure are not limited. It should be noted that when the image processing model is an object detection model, the running result of the image processing model is the object detection result (such as the rectangular box where the target object is located); when the image processing model is an object segmentation model, the running result of the image processing model is the object segmentation result (such as the connected domain composed of pixel points belonging to the target object); when the image processing model is for object classification, the running result of the image processing model is the object classification result (such as the specific category to which the target object belongs).

[0030] In an implementation scenario, as a possible example, during the pruning training process, the pruning coefficient can be initialized first, and then in each iteration of the image processing model: based on the current values of the weight parameters and the pruning coefficient in the image processing model, a first range is obtained, and based on the first range, the weight parameters in the image processing model are pruned, and based on the running result of the image processing model after pruning, the weight parameters in the image processing model are adjusted, thus completing one round of training. Furthermore, the pruning coefficient can be decreased, and the image processing model is returned to perform a new round of iteration. In the above manner, in each iteration of the pruning training, first, the current values of the weight parameters and the pruning coefficient are used to determine the first range, then the weight parameters are pruned using the first norm, and then based on the running result of the image processing model after pruning, the weight parameters before pruning in the image processing model are adjusted, and then the pruning coefficient is decreased to return to perform a new round of iteration. This can dynamically adjust the numerical range for pruning the weight parameters following the changes of the pruning coefficient and the weight parameters, and the weight parameters can gradually stabilize within the range determined by the final pruning coefficient, making the weight parameters as close as possible to a uniform distribution, which helps to alleviate the accuracy loss caused by model quantization as much as possible.

[0031] In a specific implementation scenario, before initializing the clipping coefficient, the target value of the clipping coefficient (i.e., to what value the clipping coefficient decreases) can be determined first, so as to initialize the clipping coefficient to a value higher than the target value. Exemplarily, in the embodiments of the present disclosure, considering that when quantifying a certain set of data, if it satisfies a uniform distribution, its quantization error is often smaller. Therefore, the embodiments of the present disclosure expect the distribution of the weight parameters to be closer to the uniform distribution. If a certain set of data X satisfies a uniform symmetric distribution in the interval [-T, T], its maximum value and average value need to satisfy: T = max(X) = 2 * mean(|X|). Where max() represents taking the maximum value, mean() represents taking the average value, and || represents taking the absolute value. That is to say, the target value can be set to 2. Of course, as another possible example, in order to avoid the model accuracy being more damaged due to the target value being extremely set to 2, the target value can be set to slightly greater than 2, such as 3, etc. The specific value of the target value is not limited here. Still taking the target value set to 2 as an example, the clipping coefficient can be initialized to a value higher than the target value, such as 4, 5, etc. The initial value of the clipping coefficient is not limited here.

[0032] In a specific implementation scenario, in each iteration, the average value can be obtained first based on the absolute value of the weight parameters in the image processing model, and then the upper limit value clip_data of the first range can be obtained based on the product of the average value and the current value:

[0033] ……(1)

[0034] In the above formula (1), represents the weight parameter, represents taking the absolute value, represents taking the average value, represents the current value of the clipping coefficient. In addition, the lower limit value of the first range and the upper limit value are opposite to each other. That is to say, after clipping the weight parameter using the first range, the weight parameter can be expressed as:

[0035] ……(2)

[0036] In the above formula (2), represents the lower limit value, represents the upper limit value, represents the weight parameter, Denotes a clipping function. For the specific implementation, reference can be made to the aforementioned clipping process, which will not be elaborated here. It should be noted that when clipping the weight parameters in the image processing model, the weight parameters of each network layer in the image processing model can be clipped layer by layer. Exemplarily, the weight parameters in the first network layer can be clipped first. At this time, the average value can be obtained based on the absolute values of the weight parameters in the first network layer, and then the upper limit value of the first range can be obtained based on the product of the average value and the current value of the clipping coefficient. The lower limit value of the first range is the opposite of the upper limit value. Then, the weight parameters in the first network layer are clipped based on this first range, and then the weight parameters in the second network are clipped. The specific process can refer to the clipping process of the first network layer above, which will not be elaborated here. Taking the linear layer "linear" as an example, if it is not clipped during forward inference, its "forward" function can be expressed as:

[0037] ……(3)

[0038] In the above formula (3), represents the "forward" function of the linear layer "linear", represents the original parameters, including the original weight parameters and the original bias parameters , represents the input tensor, represents the return content (i.e., the output result) of the linear layer "linear", that is, according to the original weight parameters and the original bias parameters process the input tensor to obtain the result: . Different from the aforementioned unclipped "forward" function, if the linear layer "linear" is clipped during forward inference, its "forward" function can be expressed as:

[0039]

[0040]

[0041] ……(4)

[0042] In the above formula (4), represents the original weight parameters, It represents the weight parameters after cropping. For the specific meanings of other parameters, please refer to Formulas (1), (2), and (3), which will not be elaborated here. On this basis, similar to the image to be processed, during the cropping training process, the image feature representation of the sample image can also be expressed in tensor form. After the image feature representation of the sample image is cropped, it is processed layer by layer through each network layer, and the operation result of the image processing model after cropping can be obtained. Based on this, the weight parameters before cropping can be adjusted. Exemplarily, taking the object detection task as an example, the sample image can be labeled with the true region of the sample object. Then, the operation result of the image processing model after cropping can include the predicted region of the sample object in the sample image. Therefore, based on the difference between the true region and the predicted region, at least the weight parameters before cropping in the image processing model can be adjusted, and the bias parameters can also be adjusted synchronously. Or, taking the object classification task as an example, the sample image can be labeled with the true category of the sample object. Then, the operation result of the image processing model after cropping can include the predicted category of the sample object in the sample image. Therefore, based on the difference between the true category and the predicted category, at least the weight parameters before cropping in the image processing model can be adjusted, and the bias parameters can also be adjusted synchronously. The above examples only take the object detection task and the object classification task as examples for the specific process of adjusting the weight parameters before cropping. The same can be applied to other image processing tasks and will not be elaborated one by one here.

[0043] In a specific implementation scenario, after each round of iteration, the cropping coefficient can be decreased. As a possible example, based on the current iteration round, the amplitude value for decreasing the cropping coefficient can be determined, and the current iteration round is negatively correlated with the amplitude value. That is to say, the closer the current iteration round is, the larger the amplitude value for decreasing the cropping coefficient can be; conversely, the later the current iteration round is, the smaller the amplitude value for decreasing the cropping coefficient can be. In other words, in the early stage of training, the cropping coefficient can be decreased significantly to accelerate the weight distribution towards a uniform distribution, that is, to pursue fast and efficient in the early stage of training. While in the later stage of training, the cropping coefficient can be decreased slightly to force the weight distribution to tend towards a uniform distribution more accurately, that is, to pursue fine and accurate in the later stage of training. Exemplarily, the amplitude value can be within the numerical range of 0.2 to 0.02. For example, when the first cropping training is performed, the amplitude value for decreasing the cropping coefficient can be 0.2; when the second cropping training is performed, the amplitude value for decreasing the cropping coefficient can be 0.19; when the third cropping training is performed, the amplitude value for decreasing the cropping coefficient can be 0.18, and so on, which will not be elaborated one by one here. Or, as another possible example, each time cropping training is performed, the cropping coefficient can also be decreased by a preset value, that is, the preset value remains constant during the training process of the image processing model. Exemplarily, the preset value can be set to 0.05, 0.04, etc., and the specific numerical value of the preset value is not limited here.

[0044] In an implementation scenario, before pruning training, the image processing model can be pre-trained. It should be noted that during the pre-training process, the weight parameters in the image processing model are not pruned. Exemplarily, taking object detection as an example, the image processing model can be used to process a sample image to obtain the predicted region of the sample object in the sample image, and based on the difference between the predicted region and the true region of the sample object in the sample image, the network parameters (such as weight parameters, bias parameters, etc.) of the image processing model can be adjusted; or, taking object classification as an example, the image processing model can be used to process a sample image to obtain the predicted category of the sample object in the sample image, and based on the difference between the predicted category and the true category of the sample object in the sample image, the network parameters (such as weight parameters, bias parameters, etc.) of the image processing model can be adjusted. Of course, the above examples are only several possible examples of pre-training the image processing model before pruning training, and other possible situations will not be exemplified one by one here.

[0045] In an implementation scenario, after pruning training, the image processing model can also be quantized. It should be noted that at the beginning of the quantization training, the weight parameters in the image processing model are the weight parameters after the pruning training is completed. Specifically, based on the final values of the weight parameters and the pruning coefficient in the image processing model after the pruning training is completed, a second range can be obtained, and based on the second range, the weight parameters in the image processing model can be pruned. On this basis, quantization training can be performed on the image processing model after pruning according to the second range to obtain an image processing model for deployment on the current device. It should be noted that as described above, when pruning the weight parameters in the image processing model, the weight parameters in each network layer of the image processing model can be pruned layer by layer. For specific details, reference can be made to the foregoing relevant description, which will not be elaborated here. In the above manner, the weight parameters in the image processing model are pruned after the pruning training of the image processing model is completed. At this time, the pruned weight parameters are as close as possible to a uniform distribution, thereby reducing as much as possible the accuracy loss caused by subsequent parameter quantization. Moreover, based on this, further quantization training is helpful to further reduce the accuracy loss caused by parameter quantization.

[0046] In a specific implementation scenario, for the specific process of determining the second range, reference can be made to the foregoing relevant description of determining the first range, which will not be elaborated here.

[0047] In a specific implementation scenario, quantization training can be achieved through technical details of quantization training such as TQT (Trained Quantization Thresholds) training based on W4A8 (i.e., weight quantization to 4 bits and activation quantization to 8 bits). Exemplarily, taking TQT training as an example, for the sake of description, the quantization factor of the weight parameter can be denoted as s_w, the quantization factor of the input (i.e., tensor) can be denoted as s_i, and the quantization factor of the output (i.e., activation) can be denoted as s_o. It should be noted that the quantization factor s_w of the weight parameter can be directly initialized according to the weight parameter weight of each network layer, and the quantization factors s_i of the input (i.e., tensor) and s_o of the output (i.e., activation) can be obtained by the image processing model through statistics on a small calibration set. The specific process can refer to technical details such as TQT training and will not be limited here. Then the quantization process can be described as:

[0048]

[0049] ……(5)

[0050] In the above formula (5), n is the quantization bit. Taking W4A8 as an example, for the weight parameter, its quantization bit n can be set to 4. In addition, q max = 2 n-1 - 1, q min = - 2 n-1 . To achieve automatic differentiation of the pseudo - quantization process during quantization training and for simplicity, it can be set that , where means no forward differentiation. In addition, for the specific meanings of other relevant parameters, refer to the technical details of pertensor quantization and will not be elaborated here. For example, the quantization factor s_w of the weight parameter is expressed as an exponential power of 2, which can facilitate multiplication or division calculations for each tensor (i.e., perform shift operations), helping to greatly improve the inference speed.

[0051] In a specific implementation scenario, different from the initialization method of the scaling factor (i.e., the aforementioned quantization factor) in conventional TQT training, as another possible example, as described above, before quantization training, each network layer in the image processing model is respectively initialized with a first scaling factor for quantizing input parameters / output parameters (e.g., they can be respectively expressed as the quantization factor s_i of the aforementioned input parameters and the quantization factor s_o of the output parameters). Then, for any network layer, its first scaling factor can be initialized through the following steps: First, statistical variables of the target parameters can be initialized for the network layer. It should be noted that when initializing the first scaling factor of the input parameters of the network layer, the target parameter is the input parameter, and when initializing the first scaling factor of the output parameters of the network layer, the target parameter is the output parameter. In addition, when initializing the first scaling factor of the input parameters or output parameters of each network layer, the statistical variables can all be initialized to 0. On this basis, the maximum absolute value of the target parameters of the network layer can be obtained when the image processing model processes the sample images in the calibration set. Then, the current value of the statistical variable and the above maximum absolute value can be weighted to obtain the new current value of the statistical variable. It should be noted that when weighting, the weight factors of the two can be set as needed. For example, when emphasizing the statistical variable more, the weight factor of the statistical variable can be set larger, and when emphasizing the maximum absolute value more, the weight factor of the maximum absolute value can be set larger. Exemplarily, taking the initialization of the first scaling factor of the input parameters as an example, the update process of the statistical variable can be expressed as:

[0052] ……(6)

[0053] In the above formula (6), the right side of the equal sign represents the current value of the statistical variable, and the left side of the equal sign represents the new current value of the statistical variable after update, represents the maximum absolute value of the input parameters, 0.9 represents the weight factor of the statistical variable, and 0.1 represents the weight factor of the maximum absolute value. Of course, the update process of the statistical variable shown in formula (6) is only one possible example when taking the initialization of the first scaling factor of the input parameters as an example, and other possible update methods are not exemplified one by one here. Similarly, taking the initialization of the first scaling factor of the output parameters as an example, the update process of the statistical variable can be expressed as:

[0054] ……(7)

[0055] In the above formula (7), the right side of the equal sign represents the current value of the statistical variable, and the left side of the equal sign represents the new current value of the statistical variable after update, Represents the maximum absolute value of the output parameter, 0.9 represents the weight factor of the statistical variable, and 0.1 represents the weight factor of the maximum absolute value. Of course, formula (7) is only a possible example of the update process of the statistical variable when the first scaling factor of the output parameter is initialized. Other possible update methods are not given examples here. Then, for the next sample image in the calibration set, the maximum absolute value of the target parameter of the network layer when the image processing model processes the sample image in the calibration set can be returned to execute the aforementioned acquisition. The maximum absolute value of the target parameter of the network layer when the image processing model processes the sample image in the calibration set can be repeated until all the sample images in the calibration set are processed. Finally, based on the latest current value of the statistical variable, the first scaling factor expressed as an exponential power of 2 can be obtained. As a possible example, after the latest current value of the input parameter or output parameter of any network layer with respect to the statistical variable, the following formula can be substituted to obtain the first scaling factor expressed as an exponential power of 2:

[0056] …… (8)

[0057] In the above formula (8), Indicates the latest current value of the statistical variable, n is the quantization bit, and round indicates rounding. In addition, when the target parameter is an input parameter, s represents the first scaling factor of the input parameter (such as the aforementioned s_i), and when the target parameter is an output parameter, s represents the first scaling factor of the output parameter (such as the aforementioned s_o). The above method initializes the statistical variable of the target parameter for the network layer, obtains the maximum absolute value of the target parameter of the network layer when the image processing model processes the sample image in the calibration set, and then weights it based on the current value and the maximum absolute value of the statistical variable to obtain the new current value of the statistical variable, and for the next sample image in the calibration set, returns to execute the maximum absolute value of the target parameter of the network layer when the image processing model processes the sample image in the calibration set, until all sample images in the calibration set are processed, so as to obtain the first scaling factor expressed as an exponential power of 2 based on the latest current value of the statistical variable, and then can weight and smooth the statistical variable. Compared with directly taking the maximum absolute value to obtain the scaling factor, it can avoid the scaling factor affected by the outliers that may appear in the target parameter as much as possible.

[0058] In a specific implementation scenario, different from the initialization method of the scaling factor (i.e., the aforementioned quantization factor) in conventional TQT training, as another possible example, as mentioned above, each network layer in the image processing model is initialized with a second scaling factor for quantizing the weight parameter (such as the aforementioned quantization factor s_w of the weight parameter) before quantization training, and for any network layer, the second scaling factor of its weight parameter can be initialized by the following steps: First, the value obtained by traversing within the preset numerical range can be used as the exponential power of 2 to obtain the candidate scaling factor of the weight parameter. It should be noted that the upper limit and lower limit of the preset numerical range can be set according to the maximum possible value and the minimum possible value of the weight parameter. Exemplarily, in general, the weight parameter is in the numerical range of -100~100, so the preset numerical range can be set to -8~8, then integers (such as -8, -7, -6, ..., 6, 7, 8, etc.) can be traversed within this preset numerical range, and the traversed values ​​are used as the exponential power of 2 to obtain the candidate scaling factor of the weight parameter. For ease of description, the traversed value can be recorded as factor, and the candidate scaling factor can be recorded as Next, the quantization deviation of the candidate scaling factor can be obtained based on the weight parameter and the weight parameter obtained by quantizing the weight parameter using the candidate scaling factor and then dequantizing it. , where W represents the weight parameter, Represents the use of candidate scaling factors Quantize the weight parameter W, Represents dequantization, and the quantization deviation can be obtained by L2 norm measurement. Finally, based on the quantization deviation of each candidate scaling factor, a candidate scaling factor can be selected as the second scaling factor. Exemplarily, the candidate scaling factor corresponding to the minimum quantization deviation can be selected as the second scaling factor. In the above method, the candidate scaling factor of the weight parameter is obtained by traversing the obtained value within the preset numerical range as the exponential power of 2, and the quantization deviation of the candidate scaling factor is obtained based on the weight parameter and the weight parameter quantized by the candidate scaling factor and then dequantized, and then the candidate scaling factor is selected as the second scaling factor based on the quantization deviation of each candidate scaling factor. Compared with directly initializing the scaling factor according to the maximum value, the quantization loss caused by outliers that may appear in the weight parameter can be minimized.

[0059] In a specific implementation scenario, after quantization, the image feature representation in the form of a tensor can be processed by the image processing model after quantization to obtain an operation result, and based on this, the weight parameters in the image processing model are fine-tuned so that the weight parameters after pruning adapt to the quantization process. Then, after the quantization training converges, the weight parameters can be pruned again with reference to the foregoing relevant descriptions, and after pruning, further quantization can be performed with reference to the foregoing relevant descriptions to obtain a first quantization model. It should be noted that the quantization tools may include at least one of the following: W4A8 quantization, symmetric quantization, linear quantization, pertensor quantization, and the scaling factor is a power of 2. In an actual application scenario, according to the calculation scheme of mapping from 32-bit data to low bits during quantization, it can be divided into linear quantization and non-linear quantization, symmetric quantization and asymmetric quantization; according to whether the object of the quantization scheme is tensor or channel (i.e., the channel), it can be divided into pertensor quantization and per-channel quantization; according to whether the number of bits after model quantization is exactly the same, mixed quantization can be introduced. In addition, when the quantization tools include at least one of W4A8 quantization, symmetric quantization, linear quantization, pertensor quantization, and the scaling factor is a power of 2, the specific process of model quantization can refer to the foregoing formula (5) and related descriptions, which will not be elaborated here.

[0060] In an implementation scenario, according to the application scenario of image processing, the image to be processed can be different images. Exemplarily, in a traffic management scenario, the image to be processed can be a road traffic image (such as, an intersection capture image, etc.); or, in an industrial scenario, the image to be processed can be a product capture image (such as, an image of the product to be inspected, etc.); or, in an educational scenario, the image to be processed can be an image of a test paper capture, etc. Of course, the above examples are only several possible examples of the image to be processed, and the specific content of the image to be processed is not limited here.

[0061] In an implementation scenario, feature extraction of the image to be processed can be achieved through network layers such as convolution, so that the image to be processed can be converted into an image feature representation in the form of a tensor through feature extraction. The specific process of feature extraction can refer to the technical details of network layers such as convolution, which will not be elaborated here.

[0062] Step S12: Process the image feature representation based on the first quantization model to obtain a temporary processing result.

[0063] It should be noted that the image feature representation is processed layer by layer through the first quantization model, and the temporary processing result can be obtained at the output layer of the first quantization model. Since the temporary processing result superimposes quantization factors, it cannot be directly output as the final processing result, but rather dequantization processing needs to be performed before output.

[0064] Step S13: Dequantize the temporary processing result based on the quantization tool to obtain the final processing result of the image to be processed.

[0065] Specifically, with reference to formula (5), the dequantization process can be achieved in combination with the quantization factor. As a possible example, the dequantization can be expressed as , of course, the above formula shows the dequantization of the model weight parameters, and the dequantization of the model output results can be deduced by analogy. For specific details, please refer to the relevant descriptions of dequantization in quantization training such as TQT, which will not be elaborated here. Exemplarily, in an object detection task, the final processing result may include, but is not limited to, the target region of the target object in the image to be processed; or, in an object segmentation task, the final processing result may include, but is not limited to, the connected domain formed by the pixel points belonging to the target object in the image to be processed; or, in an object classification task, the final processing result may include, but is not limited to, the target category of the target object in the image to be processed. Of course, the above examples are only several possible examples in the actual application process, and other possible situations will not be listed one by one here.

[0066] In the above solution, the image feature representation of the image to be processed is extracted, and the image processing model deployed on the current device is quantized based on the quantization tool on the current device to obtain the first quantization model. The image processing model is at least obtained through pruning training, and the pruning coefficient is decreased during the pruning training process. The pruning coefficient is used to determine the numerical range for pruning the weight parameters in the image processing model. Then, the image feature representation is processed based on the first quantization model to obtain the temporary processing result. Thus, the temporary processing result is dequantized based on the quantization tool to obtain the final processing result of the image to be processed. Furthermore, on the one hand, by quantizing the neural network model deployed on the current device and then applying it, the inference speed of the neural network model can be ensured as much as possible. On the other hand, the neural network model is at least obtained through pruning training, and since the pruning coefficient is dynamically decreased during the pruning training process, after the weight parameters are pruned within the numerical range determined by the pruning system during the training process, they gradually stabilize within the range determined by the final pruning coefficient, making the weight parameters as close to a uniform distribution as possible, which helps to alleviate the accuracy loss caused by model quantization as much as possible and can improve the accuracy of the neural network model in its processing task as much as possible. Therefore, the inference speed of the neural network model can be ensured as much as possible, and the accuracy of the neural network model in its processing task can be improved.

[0067] Please refer to Figure 2 , Figure 2It is a schematic flowchart of an embodiment of the voice processing method of this application. It should be noted that the embodiments of the present disclosure mainly illustrate the differences in the application scenarios of voice processing compared with the foregoing application scenarios of image processing. For the same or similar parts, reference can be made to the disclosed embodiments of the foregoing image processing method (for example, the image processing model in the disclosed embodiments of the foregoing image processing method can be replaced with a language processing model), which will not be elaborated here. Specifically, in the application scenario of voice processing, the embodiments of the present disclosure may include the following steps:

[0068] Step S21: Extract the voice feature representation of the voice to be processed, and quantize the voice processing model deployed on the current device based on the quantization tool on the current device to obtain a second quantization model.

[0069] In the embodiments of the present disclosure, the voice processing model is at least obtained through pruning training. During the pruning process, the pruning coefficient is decreased. The pruning coefficient is used to determine the numerical range of the weight parameters in the pruned voice processing model. It should be noted that the pruning training of the voice processing model can refer to the relevant description of the pruning training of the image processing model in the foregoing disclosed embodiments, which will not be elaborated here. In addition, the specific meaning of the pruning coefficient and the determination method of the numerical range can refer to the foregoing disclosed embodiments, which will not be elaborated here either. In addition, during the pruning training and quantization training of the voice processing model, different from using sample images in the foregoing disclosed embodiments, sample voices can be used in the embodiments of the present disclosure.

[0070] In one implementation scenario, similar to the image processing model in the foregoing disclosed embodiments, before the voice processing model performs pruning training, the voice processing model can be pre-trained; or, after the voice processing model performs pruning training, the voice processing model can be quantized trained. It should be noted that the specific processes of the pre-training and quantization training of the voice processing model can refer to the pre-training and quantization training of the image processing model in the foregoing disclosed embodiments, which will not be elaborated here.

[0071] In one implementation scenario, depending on the application scenario, the voice to be processed may also be different. Exemplarily, in the voice recognition scenario, the voice to be processed may be the voice to be recognized. For example, the voice to be recognized can be collected by a voice assistant (such as a car machine, etc.); or, in the voice translation scenario, the voice to be processed may be the voice to be translated. For example, the voice to be translated can be collected by a translation machine, etc. Of course, the above examples are only several possible examples of voice processing tasks, and other possible situations will not be exemplified one by one here. For example, voice processing tasks may also include but are not limited to voice simultaneous translation tasks, etc. Correspondingly, the voice processing model may specifically include but is not limited to: voice recognition model, voice translation model, voice simultaneous translation model, etc. The network structure of the voice processing model is not limited here.

[0072] In an implementation scenario, feature extraction is performed on the speech to be processed, and a speech feature representation in the form of a tensor can be obtained. For specific details, reference can be made to the relevant description of the image feature representation in the foregoing disclosed embodiments, which will not be elaborated herein. In addition, for the specific process of quantizing the speech processing model based on the quantization tool, reference can be made to the relevant description of quantizing the image processing model based on the quantization tool in the foregoing disclosed embodiments, which will not be elaborated herein either.

[0073] Step S22: Process the speech feature representation based on the second quantization model to obtain a temporary processing result.

[0074] For specific details, reference can be made to the relevant description of "processing the image feature representation based on the first quantization model to obtain a temporary processing result" in the foregoing disclosed embodiments, which will not be elaborated herein.

[0075] Step S23: Dequantize the temporary processing result based on the quantization tool to obtain the final processing result of the speech to be processed.

[0076] For specific details, reference can be made to the relevant description of "dequantizing the temporary processing result based on the quantization tool to obtain the final processing result of the image to be processed" in the foregoing disclosed embodiments, which will not be elaborated herein either. In addition, for the specific process of quantizing the speech processing model after pruning training, reference can be made to the quantization training of the image processing model in the foregoing embodiments of the image processing method, which will not be elaborated herein either.

[0077] In the above solution, the speech feature representation of the speech to be processed is extracted, and the speech processing model deployed on the current device is quantized based on the quantization tool on the current device to obtain a second quantization model. The speech processing model is at least obtained through pruning training, and the pruning coefficient is decreased during the pruning process. The pruning coefficient is used to determine the numerical range of pruning the weight parameters in the speech processing model. Then, the speech feature representation is processed based on the second quantization model to obtain a temporary processing result. Thus, the temporary processing result is dequantized based on the quantization tool to obtain the final processing result of the speech to be processed. Furthermore, on the one hand, by quantizing the neural network model deployed on the current device and then applying it, the inference speed of the neural network model can be ensured as much as possible. On the other hand, the neural network model is at least obtained through pruning training, and since the pruning coefficient is dynamically decreased during the pruning training process, after the weight parameters are pruned within the numerical range determined by the pruning system during the training process, they gradually stabilize within the range determined by the final pruning coefficient, making the weight parameters as close as possible to a uniform distribution, which helps to alleviate the accuracy loss caused by model quantization as much as possible and can improve the accuracy of the neural network model in its processing task as much as possible. Therefore, the inference speed of the neural network model can be ensured as much as possible, and the accuracy of the neural network model in its processing task can be improved.

[0078] Please refer toFigure 3 , Figure 3 is a schematic flowchart of an embodiment of the natural language processing method of this application. It should be noted that the embodiments of the present disclosure mainly illustrate the differences in the application scenarios of natural language processing compared to the aforementioned application scenarios of image processing. For the same or similar parts, reference can be made to the disclosed embodiments of the aforementioned image processing method (for example, the image processing model in the disclosed embodiments of the aforementioned image processing method can be replaced with a natural language processing model), which will not be elaborated here. Specifically, in the application scenario of natural language processing, the embodiments of the present disclosure may include the following steps:

[0079] Step S31: Extract the statement feature representation of the statement to be processed, and quantize the natural language processing model deployed on the current device based on the quantization tool on the current device to obtain a third quantization model.

[0080] In the embodiments of the present disclosure, the natural language processing model is at least obtained through pruning training. During the pruning process, the pruning coefficient is gradually decreased. The pruning coefficient is used to determine the numerical range for pruning the weight parameters in the natural language processing model. It should be noted that the pruning training of the natural language processing model can refer to the relevant descriptions of the pruning training of the image processing model in the aforementioned disclosed embodiments, which will not be elaborated here. In addition, the specific meaning of the pruning coefficient and the determination method of the numerical range can refer to the aforementioned disclosed embodiments, which will not be elaborated here either. In addition, during the pruning training and quantization training of the natural language processing model, different from using sample images in the aforementioned disclosed embodiments, sample texts can be used in the embodiments of the present disclosure.

[0081] In an implementation scenario, similar to the image processing model in the aforementioned disclosed embodiments, before the natural language processing model performs pruning training, the natural language processing model can be pre-trained; or, after the natural language processing model performs pruning training, the natural language processing model can be quantized for training. It should be noted that the specific processes of pre-training and quantization training of the natural language processing model can refer to the pre-training and quantization training of the image processing model in the aforementioned disclosed embodiments, which will not be elaborated here.

[0082] In an implementation scenario, depending on the application scenario, the statement to be processed may also vary. Exemplarily, in an intelligent dialogue scenario, the statement to be processed may be an input statement expecting the current device to make a response (e.g., "Please help me plan a travel itinerary from City A to City B"); or, in an assisted writing scenario, the statement to be processed may be the theme scope and reference materials for assisted writing (e.g., "Please write a speech for me with reference to the following materials"). Of course, the above examples are only several possible examples of natural language processing tasks, and other possible situations will not be exemplified one by one here. For example, natural language processing tasks may also include but are not limited to intelligent office work, etc. Correspondingly, the natural language processing model may specifically include but is not limited to: intelligent dialogue models (such as large language models, etc.), assisted collaboration models, etc. The network structure of the speech processing model is not limited here.

[0083] In an implementation scenario, feature extraction is performed on the statement to be processed, and a statement feature representation in tensor form can be obtained. For the specific details, reference can be made to the relevant description of image feature representation in the foregoing disclosed embodiments, which will not be elaborated here. In addition, for the specific process of quantifying the natural language processing model based on a quantization tool, reference can be made to the relevant description of quantifying the image processing model based on a quantization tool in the foregoing disclosed embodiments, which will not be elaborated here either.

[0084] Step S32: Process the statement feature representation based on the third quantization model to obtain a temporary processing result.

[0085] For the specific details, reference can be made to the relevant description of "processing the image feature representation based on the first quantization model to obtain a temporary processing result" in the foregoing disclosed embodiments, which will not be elaborated here.

[0086] Step S33: Dequantize the temporary processing result based on the quantization tool to obtain the final processing result of the statement to be processed.

[0087] For the specific details, reference can be made to the relevant description of "dequantizing the temporary processing result based on the quantization tool to obtain the final processing result of the image to be processed" in the foregoing disclosed embodiments, which will not be elaborated here either. In addition, for the specific process of quantizing the natural language processing model after pruning training, reference can be made to the quantization training of the image processing model in the foregoing embodiments of the image processing method, which will not be elaborated here either.

[0088] In the above solution, the statement feature representation of the statement to be processed is extracted, and the natural language processing model deployed on the current device is quantized based on the quantization tool on the current device to obtain a third quantized model. The natural language processing model is at least obtained through pruning training. During the pruning process, the pruning coefficient is decreased. The pruning coefficient is used to determine the numerical range for pruning the weight parameters in the natural language processing model. Then, the statement feature representation is processed based on the third quantized model to obtain a temporary processing result. Thus, the temporary processing result is dequantized based on the quantization tool to obtain the final processing result of the statement to be processed. On the one hand, by quantizing the neural network model deployed on the current device and then applying it, the inference speed of the neural network model can be ensured as much as possible. On the other hand, the neural network model is at least obtained through pruning training, and since the pruning coefficient is dynamically decreased during the pruning training process, after the weight parameters are pruned within the numerical range determined by the pruning system during the training process, they gradually stabilize within the range determined by the final pruning coefficient, making the weight parameters as close as possible to a uniform distribution, which helps to alleviate the accuracy loss caused by model quantization as much as possible and can improve the accuracy of the neural network model in its processing tasks as much as possible. Therefore, the inference speed of the neural network model can be ensured as much as possible, and the accuracy of the neural network model in its processing tasks can be improved.

[0089] Please refer to Figure 4 , Figure 4 FIG. is a schematic framework diagram of an embodiment of the image processing device of the present application. The image processing device 40 includes: an extraction and quantization module 41, a model processing module 42, and a dequantization module 43. The extraction and quantization module 41 is configured to extract the image feature representation of the image to be processed, and quantize the image processing model deployed on the current device based on the quantization tool on the current device to obtain a first quantized model. Among them, the image processing model is at least obtained through pruning training. During the pruning process, the pruning coefficient is decreased. The pruning coefficient is used to determine the numerical range for pruning the weight parameters in the image processing model. The model processing module 42 is configured to process the image feature representation based on the first quantized model to obtain a temporary processing result. The dequantization module 43 is configured to dequantize the temporary processing result based on the quantization tool to obtain the final processing result of the image to be processed.

[0090] In the above solution, the image processing device 40 extracts the image feature representation of the image to be processed, quantifies the image processing model deployed on the current device based on the quantization tool on the current device to obtain a first quantized model, and the image processing model is at least obtained through pruning training. During the pruning training process, the pruning coefficient is decreased. The pruning coefficient is used to determine the numerical range for pruning the weight parameters in the image processing model. Then, based on the first quantized model, the image feature representation is processed to obtain a temporary processing result. Thus, the quantization tool is used to dequantize the temporary processing result to obtain the final processing result of the image to be processed. Furthermore, on the one hand, by quantizing the neural network model deployed on the current device and then applying it, the inference speed of the neural network model can be ensured as much as possible. On the other hand, the neural network model is at least obtained through pruning training, and since the pruning coefficient is dynamically decreased during the pruning training process, after the weight parameters are pruned within the numerical range determined by the pruning system during the training process, they gradually stabilize within the range determined by the final pruning coefficient, making the weight parameters as close as possible to a uniform distribution, which helps to alleviate the accuracy loss caused by model quantization as much as possible and can improve the accuracy of the neural network model in its processing tasks as much as possible. Therefore, the inference speed of the neural network model can be ensured as much as possible, and the accuracy of the neural network model in its processing tasks can be improved.

[0091] In some disclosed embodiments, the image processing device 40 includes a coefficient initialization module for initializing the pruning coefficient; the image processing device 40 includes an iterative training module for, in each iteration of the image processing model: obtaining a first range based on the current values of the weight parameters and the pruning coefficient in the image processing model, pruning the weight parameters in the image processing model based on the first range, and adjusting the weight parameters before pruning based on the running result of the image processing model after pruning; the image processing device 40 includes a decreasing loop module for decreasing the pruning coefficient and returning to execute a new round of iteration for the image processing model.

[0092] In some disclosed embodiments, the decreasing loop module includes an amplitude determination sub-module for determining the amplitude value for decreasing the pruning coefficient based on the current iteration round; wherein, the current iteration round is negatively correlated with the amplitude value; the decreasing loop module includes a first decreasing sub-module for decreasing the pruning coefficient based on the amplitude value of the decreasing pruning coefficient.

[0093] In some disclosed embodiments, the decreasing loop module includes a second decreasing sub-module for decreasing the pruning coefficient by a preset value; wherein, the preset value remains constant during the training process of the image processing model.

[0094] In some disclosed embodiments, the iterative training module includes an average value calculation sub-module for obtaining an average value based on the absolute values of the weight parameters in the image processing model; the iterative training module includes an endpoint value determination sub-module for obtaining the upper limit value of a first range based on the product of the average value and the current value; wherein, the lower limit value and the upper limit value of the first range are opposite numbers to each other.

[0095] In some disclosed embodiments, the image processing apparatus 40 includes a range determination module for obtaining a second range based on the final values of the weight parameters and the clipping coefficient in the image processing model after the end of clipping training; the image processing apparatus 40 includes a parameter clipping module for clipping the weight parameters in the image processing model based on the second range; the image processing apparatus 40 includes a quantization training module for performing quantization training based on the image processing model after clipping according to the second range to obtain an image processing model for deployment on the current device.

[0096] In some disclosed embodiments, before quantization training, each network layer in the image processing model is respectively initialized with a first scaling factor for quantifying input parameters / output parameters. The image processing apparatus 40 includes an initialization module for initializing the statistical variables of the target parameters for the network layer; the image processing apparatus 40 includes an absolute value statistics module for obtaining the maximum absolute value of the target parameters of the network layer when the image processing model processes the sample images in the calibration set; wherein, when initializing the first scaling factor of the input parameters, the target parameter is the input parameter, and when initializing the first scaling factor of the output parameters, the target parameter is the output parameter; the image processing apparatus 40 includes a numerical weighting module for weighting based on the current value of the statistical variable and the maximum absolute value to obtain a new current value of the statistical variable; the image processing apparatus 40 includes a loop iteration module for, for the next sample image in the calibration set, returning to execute obtaining the maximum absolute value of the target parameters of the network layer when the image processing model processes the sample images in the calibration set until all the sample images in the calibration set are processed; the image processing apparatus 40 includes a factor determination module for obtaining a first scaling factor expressed as a power of 2 based on the latest current value of the statistical variable.

[0097] In some disclosed embodiments, before quantization training, each network layer in the image processing model is respectively initialized with a second scaling factor for quantifying the weight parameters. The image processing apparatus 40 includes a candidate factor module for obtaining candidate scaling factors for the weight parameters by taking the values obtained by traversing within a preset numerical range as powers of 2; the image processing apparatus 40 includes a deviation measurement module for obtaining the quantization deviation of the candidate scaling factor based on the weight parameters and the weight parameters after quantifying the weight parameters using the candidate scaling factor and then dequantifying them; the image processing apparatus 40 includes a factor selection module for selecting a candidate scaling factor as the second scaling factor based on the quantization deviations of the respective candidate scaling factors.

[0098] In some disclosed embodiments, the weight parameters in each network layer of the image processing model are pruned layer by layer.

[0099] In some disclosed embodiments, the quantization tools include at least one of the following: W4A8 quantization, symmetric quantization, linear quantization, pertensor quantization, and a scaling factor that is a power of 2.

[0100] Please refer to Figure 5 , Figure 5 which is a schematic framework diagram of an embodiment of the voice processing device of the present application. The voice processing device 50 includes: an extraction and quantization module 51, a model processing module 52, and a dequantization module 53. The extraction and quantization module 51 is configured to extract a voice feature representation of the voice to be processed, and quantize the voice processing model deployed on the current device based on the quantization tool on the current device to obtain a second quantized model; wherein, the voice processing model is obtained by at least pruning training, and during the pruning process, a pruning coefficient is decreased, and the pruning coefficient is used to determine the numerical range of the weight parameters to be pruned in the voice processing model; the model processing module 52 is configured to process the voice feature representation based on the second quantized model to obtain a temporary processing result; the dequantization module 53 is configured to dequantize the temporary processing result based on the quantization tool to obtain the final processing result of the voice to be processed.

[0101] In the above solution, the voice processing device 50 extracts the voice feature representation of the voice to be processed, quantizes the voice processing model deployed on the current device based on the quantization tool on the current device to obtain a second quantized model, and the voice processing model is obtained by at least pruning training, and during the pruning process, a pruning coefficient is decreased, and the pruning coefficient is used to determine the numerical range of the weight parameters to be pruned in the voice processing model. Then, the voice feature representation is processed based on the second quantized model to obtain a temporary processing result, and thus the temporary processing result is dequantized based on the quantization tool to obtain the final processing result of the voice to be processed. Furthermore, on the one hand, by quantizing the neural network model deployed on the current device and then applying it, the inference speed of the neural network model can be ensured as much as possible. On the other hand, the neural network model is also obtained by at least pruning training, and since the pruning coefficient is dynamically decreased during the pruning training process, after the weight parameters are pruned within the numerical range determined by the pruning system during the training process, they gradually stabilize within the range determined by the final pruning coefficient, making the weight parameters as close to a uniform distribution as possible, which helps to alleviate the accuracy loss caused by model quantization as much as possible and can improve the accuracy of the neural network model in its processing task as much as possible. Therefore, the inference speed of the neural network model can be ensured as much as possible, and the accuracy of the neural network model in its processing task can be improved.

[0102] Please refer to Figure 6 , Figure 6It is a schematic framework diagram of an embodiment of the natural speech processing device of the present application. The natural speech processing device 60 includes: an extraction and quantization module 61, a model processing module 62, and a dequantization module 63. The extraction and quantization module 61 is configured to extract the statement feature representation of the statement to be processed, and quantize the natural language processing model deployed on the current device based on the quantization tool on the current device to obtain a third quantized model; wherein, the natural language processing model is at least obtained through pruning training, and during the pruning process, the pruning coefficient is decreased, and the pruning coefficient is used to determine the numerical range of pruning the weight parameters in the natural language processing model; the model processing module 62 is configured to process the statement feature representation based on the third quantized model to obtain a temporary processing result; the dequantization module 63 is configured to dequantize the temporary processing result based on the quantization tool to obtain the final processing result of the statement to be processed.

[0103] In the above solution, the natural speech processing device 60 extracts the statement feature representation of the statement to be processed, quantizes the natural language processing model deployed on the current device based on the quantization tool on the current device to obtain a third quantized model, and the natural language processing model is at least obtained through pruning training. During the pruning process, the pruning coefficient is decreased, and the pruning coefficient is used to determine the numerical range of pruning the weight parameters in the natural language processing model. Then, the statement feature representation is processed based on the third quantized model to obtain a temporary processing result. Thus, the temporary processing result is dequantized based on the quantization tool to obtain the final processing result of the statement to be processed. Furthermore, on the one hand, by quantizing the neural network model deployed on the current device and then applying it, the inference speed of the neural network model can be ensured as much as possible. On the other hand, the neural network model is at least obtained through pruning training, and since the pruning coefficient is dynamically decreased during the pruning training process, after the weight parameters are pruned within the numerical range determined by the pruning system during the training process, they gradually stabilize within the range determined by the final pruning coefficient, making the weight parameters as close to a uniform distribution as possible, which helps to alleviate the accuracy loss caused by model quantization as much as possible and can improve the accuracy of the neural network model in its processing task as much as possible. Therefore, the inference speed of the neural network model can be ensured as much as possible, and the accuracy of the neural network model in its processing task can be improved.

[0104] Please refer to Figure 7 , Figure 7It is a schematic diagram of the framework of an embodiment of the electronic device 70 of the present application. The electronic device 70 includes a memory 71 and a processor 72 that are coupled to each other. At least program instructions are stored in the memory 71, and the processor 72 is configured to execute the program instructions to implement the steps in any of the above-described embodiments of the image processing method, or the steps in any of the above-described embodiments of the speech processing method, or the steps in any of the above-described embodiments of the natural language processing method. For details, reference may be made to the foregoing disclosed embodiments, which will not be elaborated herein. The electronic device 70 may include, but is not limited to, a smart phone, a tablet computer, a server, etc. The specific type of the electronic device 70 is not limited herein.

[0105] Specifically, the processor 72 is configured to control itself and the memory 71 to implement the steps in any of the above-described embodiments of the image processing method. The processor 72 may also be referred to as a CPU (Central Processing Unit). The processor 72 may be an integrated circuit chip with signal processing capabilities. The processor 72 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 72 may be implemented jointly by integrated circuit chips.

[0106] In the above solution, when the electronic device 70 implements the steps in any of the above-described embodiments of the image processing method, or the steps in any of the above-described embodiments of the speech processing method, or the steps in any of the above-described embodiments of the natural language processing method, on the one hand, by quantifying the neural network model deployed on the current device and then applying it, the inference speed of the neural network model can be ensured as much as possible. On the other hand, the neural network model is at least pruned and trained, and since the pruning coefficient is dynamically decreased during the pruning training process, after the weight parameters are pruned within the numerical range determined by the pruning system during the training process, they gradually stabilize within the range determined by the final pruning coefficient, making the weight parameters as close as possible to a uniform distribution, which helps to alleviate the accuracy loss caused by model quantization as much as possible and can improve the accuracy of the neural network model in its processing tasks as much as possible. Therefore, the inference speed of the neural network model can be ensured as much as possible, and the accuracy of the neural network model in its processing tasks can be improved.

[0107] Please refer to Figure 8 , Figure 8It is a schematic framework diagram of an embodiment of the computer-readable storage medium 80 of the present application. The computer-readable storage medium 80 stores program instructions 81 that can be run by a processor. The program instructions 81 are used to implement the steps in any of the above-described embodiments of the image processing method, or the steps in any of the above-described embodiments of the speech processing method, or the steps in any of the above-described embodiments of the natural language processing method.

[0108] In the above solution, the computer-readable storage medium 80 implements the steps in any of the above-described embodiments of the image processing method, or the steps in any of the above-described embodiments of the speech processing method, or the steps in any of the above-described embodiments of the natural language processing method. Therefore, on the one hand, by quantifying the neural network model deployed on the current device and then applying it, the inference speed of the neural network model can be ensured as much as possible. On the other hand, the neural network model has at least been pruned and trained. And because the pruning coefficient is dynamically decreased during the pruning training process, after the weight parameters are pruned within the numerical range determined by the pruning system during the training process, they gradually stabilize within the range determined by the final pruning coefficient, making the weight parameters as close as possible to a uniform distribution, which helps to alleviate the accuracy loss caused by model quantization as much as possible and can improve the accuracy of the neural network model in its processing tasks as much as possible. Therefore, the inference speed of the neural network model can be ensured as much as possible, and the accuracy of the neural network model in its processing tasks can be improved.

[0109] In some embodiments, the functions or modules included in the device provided by the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0110] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0111] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0112] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of these units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0113] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, may exist physically separately for each unit, or two or more units may be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0114] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods of each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0115] If the technical solution of the present application involves personal information, before the product applying the technical solution of the present application processes personal information, it has clearly informed the personal information processing rules and obtained the individual's independent consent. If the technical solution of the present application involves sensitive personal information, before the product applying the technical solution of the present application processes sensitive personal information, it has obtained the individual's separate consent and at the same time meets the requirement of "express consent". For example, at a personal information collection device such as a camera, a clear and prominent sign is set to inform that the personal information collection range has been entered and personal information will be collected. If an individual voluntarily enters the collection range, it is considered consent to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are informed by obvious signs / information, personal authorization is obtained through pop-up messages or by asking the individual to upload their personal information themselves; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

Claims

1. An image processing method, characterized in that: include: Extracting an image feature representation of the image to be processed, and quantizing an image processing model deployed on the current device based on a quantization tool on the current device to obtain a first quantized model; wherein the image processing model is obtained at least through cropping training, and the cropping coefficient is decreased during the cropping process, and the cropping coefficient is used to determine the numerical range of the weight parameter in the cropped image processing model; Processing the image feature representation based on the first quantization model to obtain a temporary processing result; De-quantizing the temporary processing result based on the quantization tool to obtain a final processing result of the image to be processed; The training steps of the image processing model include: Initializing the cropping factor; In each round of iteration of the image processing model: based on the current values ​​of the weight parameter and the cropping coefficient in the image processing model, a first range is obtained, and based on the first range, the weight parameter in the image processing model is cropped, and based on the running result of the image processing model after the cropping, the weight parameter before the cropping is adjusted; The crop factor is decremented and a new iteration is performed on the image processing model again.

2. The method according to claim 1, characterized in that The step of decreasing the crop factor comprises: Based on the current iteration round, determining a magnitude value of decreasing the clipping coefficient; wherein the current iteration round is negatively correlated with the magnitude value; The crop factor is decremented based on the magnitude value of the decremented crop factor.

3. The method according to claim 1, characterized in that The step of decreasing the crop factor comprises: The cropping factor is decreased by a preset value; wherein the preset value remains constant during the training process of the image processing model.

4. The method according to claim 1, characterized in that The obtaining of the first range based on the current value of the weight parameter and the cropping coefficient in the image processing model comprises: Calculating an average value based on the absolute value of the weight parameter in the image processing model; Based on the product of the average value and the current value, an upper limit value of the first range is obtained; wherein the lower limit value of the first range and the upper limit value are reciprocal numbers of each other.

5. The method according to claim 1, characterized in that After the clipping training is completed, the method further includes: Obtaining a second range based on final values ​​of the weight parameter and the cropping coefficient in the image processing model after the cropping training is completed; Based on the second range, cutting the weight parameter in the image processing model; Quantization training is performed based on the image processing model after being cropped according to the second range to obtain an image processing model for deployment on the current device.

6. The method according to claim 5, characterized in that Before the quantization training, each network layer in the image processing model is initialized with a first scaling factor for quantizing input parameters / output parameters, and for any of the network layers, the initialization step of the first scaling factor includes: Initialize the statistical variables of the target parameters for the network layer, and obtain the maximum absolute value of the target parameters of the network layer when the image processing model processes the sample images in the calibration set; wherein, when initializing the first scaling factor of the input parameter, the target parameter is the input parameter, and when initializing the first scaling factor of the output parameter, the target parameter is the output parameter; Performing weighting based on the current value and the maximum absolute value of the statistical variable to obtain a new current value of the statistical variable, and returning, for the next sample image in the calibration set, the maximum absolute value of the target parameter of the network layer when the image processing model is used to process the sample image in the calibration set, until all the sample images in the calibration set are processed; Based on the latest current value of the statistical variable, a first scaling factor expressed as an exponential power of 2 is obtained.

7. The method according to claim 5, characterized in that Before the quantization training, each network layer in the image processing model is initialized with a second scaling factor for quantizing weight parameters, and for any of the network layers, the initialization step of the second scaling factor includes: Taking the value obtained by traversing within the preset value range as the exponential power of 2, a candidate scaling factor of the weight parameter is obtained; Obtaining a quantization deviation of the candidate scaling factor based on the weight parameter and the weight parameter obtained by quantizing the weight parameter using the candidate scaling factor and then dequantizing the weight parameter; Based on the quantization deviations of the candidate scaling factors, the candidate scaling factors are selected as the second scaling factors.

8. The method according to claim 1, characterized in that The weight parameters in each network layer in the image processing model are trimmed layer by layer.

9. The method according to claim 1, characterized in that: The quantization tools include at least one of the following: W4A8 quantization, symmetric quantization, linear quantization, pertensor quantization, and a scaling factor of 2.

10. A speech processing method, characterized in that: include: Extracting speech feature representation of the speech to be processed, and quantizing the speech processing model deployed on the current device based on the quantization tool on the current device to obtain a second quantized model; wherein the speech processing model is obtained by at least trimming training, and the trimming coefficient is reduced during the trimming process, and the trimming coefficient is used to determine the numerical range of the weight parameter in the trimming of the speech processing model; Processing the speech feature representation based on the second quantization model to obtain a temporary processing result; Dequantizing the temporary processing result based on the quantization tool to obtain the final processing result of the speech to be processed; wherein the training step of the speech processing model includes: Initializing the cropping factor; In each round of iteration of the speech processing model: based on the current values ​​of the weight parameter and the clipping coefficient in the speech processing model, a first range is obtained, and based on the first range, the weight parameter in the speech processing model is clipped, and based on the running result of the speech processing model after clipping, the weight parameter before clipping is adjusted; The cropping factor is decremented and a new round of iteration is performed on the speech processing model again.

11. A natural language processing method, characterized in that: include: Extracting a sentence feature representation of the sentence to be processed, and quantizing the natural language processing model deployed on the current device based on a quantization tool on the current device to obtain a third quantized model; wherein the natural language processing model is obtained by at least trimming training, and the trimming coefficient is reduced during the trimming process, and the trimming coefficient is used to determine the numerical range of the weight parameter in the trimming of the natural language processing model; Processing the sentence feature representation based on the third quantization model to obtain a temporary processing result; De-quantizing the temporary processing result based on the quantization tool to obtain the final processing result of the sentence to be processed; wherein the training step of the natural language processing model includes: Initializing the cropping factor; In each round of iteration of the natural language processing model: based on the current values ​​of the weight parameter and the clipping coefficient in the natural language processing model, a first range is obtained, and based on the first range, the weight parameter in the natural language processing model is clipped, and based on the running result of the natural language processing model after clipping, the weight parameter before clipping is adjusted; The cropping factor is decremented, and a new round of iteration is performed on the natural language processing model.

12. An image processing device, characterized in that: include: An extraction and quantization module, used to extract the image feature representation of the image to be processed, and quantize the image processing model deployed on the current device based on the quantization tool on the current device to obtain a first quantization model; wherein the image processing model is obtained by at least cropping training, and the cropping coefficient is decreased during the cropping process, and the cropping coefficient is used to determine the numerical range of the weight parameter in the cropped image processing model; A model processing module, used for processing the image feature representation based on the first quantization model to obtain a temporary processing result; A dequantization module is used to dequantize the temporary processing result based on the quantization tool to obtain the final processing result of the image to be processed; wherein the training step of the image processing model includes: Initializing the cropping factor; In each round of iteration of the image processing model: based on the current values ​​of the weight parameter and the cropping coefficient in the image processing model, a first range is obtained, and based on the first range, the weight parameter in the image processing model is cropped, and based on the running result of the image processing model after the cropping, the weight parameter before the cropping is adjusted; The crop factor is decremented and a new iteration is performed on the image processing model again.

13. A speech processing device, characterized in that: include: An extraction and quantization module is used to extract the speech feature representation of the speech to be processed, and quantize the speech processing model deployed on the current device based on the quantization tool on the current device to obtain a second quantization model; wherein the speech processing model is obtained by at least trimming training, and the trimming coefficient is reduced during the trimming process, and the trimming coefficient is used to determine the numerical range of the weight parameter in the trimming speech processing model; A model processing module, used for processing the speech feature representation based on the second quantization model to obtain a temporary processing result; A dequantization module is used to dequantize the temporary processing result based on the quantization tool to obtain the final processing result of the speech to be processed; wherein the training step of the speech processing model includes: Initializing the cropping factor; In each round of iteration of the speech processing model: based on the current values ​​of the weight parameter and the clipping coefficient in the speech processing model, a first range is obtained, and based on the first range, the weight parameter in the speech processing model is clipped, and based on the running result of the speech processing model after clipping, the weight parameter before clipping is adjusted; The cropping factor is decremented and a new round of iteration is performed on the speech processing model again.

14. A natural language processing device, characterized in that: include: An extraction and quantization module is used to extract the sentence feature representation of the sentence to be processed, and quantize the natural language processing model deployed on the current device based on the quantization tool on the current device to obtain a third quantization model; wherein the natural language processing model is obtained by at least trimming training, and the trimming coefficient is reduced during the trimming process, and the trimming coefficient is used to determine the numerical range of the weight parameter in the trimming of the natural language processing model; A model processing module, used for processing the sentence feature representation based on the third quantitative model to obtain a temporary processing result; A dequantization module is used to dequantize the temporary processing result based on the quantization tool to obtain the final processing result of the sentence to be processed; wherein the training step of the natural language processing model includes: Initializing the cropping factor; In each round of iteration of the natural language processing model: based on the current values ​​of the weight parameter and the clipping coefficient in the natural language processing model, a first range is obtained, and based on the first range, the weight parameter in the natural language processing model is clipped, and based on the running result of the natural language processing model after clipping, the weight parameter before clipping is adjusted; The cropping factor is decremented, and a new round of iteration is performed on the natural language processing model.

15. An electronic device, characterized in that: The method comprises at least a memory and a processor coupled to each other, wherein the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the image processing method according to any one of claims 1 to 9, or the speech processing method according to claim 10, or the natural language processing method according to claim 11.

16. A computer-readable storage medium, characterized in that: Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the image processing method described in any one of claims 1 to 9, or the speech processing method described in claim 10, or the natural language processing method described in claim 11.

Citation Information

Patent Citations

  • Convolutional network full-integer quantization method and application method thereof

    CN110135580A

  • Neural network training and image processing method and device, equipment and medium

    CN110363297A