Data processing method, device and storage medium

By dynamically determining the quantization mapping coefficients of each layer of forward inference under each step length parameter of the Diffusion model and utilizing pre-calibrated quantization calibration parameters, the accuracy and speed issues in the Diffusion model quantization scheme are resolved, thus reducing storage space and improving accuracy without reducing inference efficiency.

CN118428482BActive Publication Date: 2025-09-12北京凌川科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410390115.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-01
Publication Date
2025-09-12
Estimated Expiration
2044-04-01

AI Technical Summary

Technical Problem

The existing Diffusion model quantization scheme cannot meet the needs of its iterative forward reasoning, resulting in high reasoning cost and insufficient accuracy. The quantization scheme in related technologies cannot reduce storage space and reasoning speed while ensuring accuracy.

Method used

By dynamically determining the quantization mapping coefficients used in each layer of forward inference under each step length parameter of the Diffusion model, the quantization mapping coefficients are calculated using pre-calibrated quantization calibration parameters to ensure quantization accuracy and speed.

Benefits of technology

Without sacrificing the inference efficiency of the Diffusion model, the model's running space is reduced and the inference accuracy is guaranteed, achieving good performance in quantization conditions such as INT8, INT4, and INT6.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118428482B_ABST
    Figure CN118428482B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a data processing method, device and storage medium. The above method includes obtaining a first step length parameter and a first basic data, wherein the above first basic data is characterization data of text content or image content; obtaining a first quantization calibration parameter corresponding to the above first step length parameter in the first network layer of the data processing model, wherein the above first quantization calibration parameter is a parameter obtained by performing quantization calibration on the above first network layer based on the calibration data and the above first step length parameter; calculating a first quantization mapping coefficient corresponding to the above first step length parameter and the above first network layer based on the above first basic data and the above first quantization calibration parameter; quantizing the above first basic data based on the above first quantization mapping coefficient, and inputting the quantization result into the above first network layer for data processing to obtain the first output data of the above first network layer. This method can ensure data processing accuracy and take into account data processing speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a data processing method, device, and storage medium. Background Art

[0002] The Diffusion model is a deep generative model. In a Diffusion model, data undergoes two phases: diffusion and inverse diffusion. In the diffusion phase, noise is continuously added to the original data, shifting the data distribution from its original distribution to a desired distribution. In the inverse diffusion phase, the data is restored to its original distribution. The Diffusion model relies on an iterative denoising process, resulting in significantly higher inference costs than conventional models. Because the Diffusion model performs iterative forward inference, with different inputs for each iteration, quantization schemes used in related technologies cannot meet the quantization requirements of the Diffusion model. Summary of the Invention

[0003] In order to solve at least one of the above technical problems, the present disclosure provides a data processing method, device, and storage medium. The technical solutions of the present disclosure are as follows:

[0004] According to a first aspect of an embodiment of the present disclosure, there is provided a data processing method, including:

[0005] Obtaining a first step length parameter and first basic data, where the first basic data is representation data of text content or image content;

[0006] Obtaining a first quantization calibration parameter corresponding to the first step length parameter at a first network layer of a data processing model, where the first quantization calibration parameter is a parameter obtained by performing quantization calibration on the first network layer based on calibration data and the first step length parameter;

[0007] Calculating, according to the first basic data and the first quantization calibration parameter, a first quantization mapping coefficient corresponding to the first step length parameter and the first network layer;

[0008] The first basic data is quantized based on the first quantization mapping coefficient, and the quantization result is input into the first network layer for data processing to obtain first output data of the first network layer.

[0009] In a possible embodiment, before obtaining the first step length parameter and the first basic data, the method includes:

[0010] Determining a quantization calibration loss corresponding to the first step length parameter, wherein the quantization calibration loss takes the calibration data and the calibration parameter as variables, and the quantization calibration loss indicates a quantization accuracy loss of the calibration data;

[0011] Get calibration data;

[0012] The first quantization calibration parameter is determined according to a relationship between the calibration data and the quantization calibration loss.

[0013] In a possible embodiment, the quantization calibration loss further includes an error control parameter, and determining the first quantization calibration parameter according to a relationship between the calibration data and the quantization calibration loss includes:

[0014] When the error control parameter is a first preset value, determining a corresponding target calibration parameter and a corresponding target quantization performance loss;

[0015] determining a corresponding target calibration parameter and a corresponding target quantization performance loss when the error control parameter is a second preset value, the first preset value being different from the second preset value;

[0016] Using the target calibration parameter corresponding to the minimum target quantization performance loss as the first quantization calibration parameter;

[0017] The target calibration parameter is the calibration parameter when the corresponding quantization calibration loss reaches the minimum value.

[0018] The target quantization performance loss indicates the accuracy loss generated in the output data of the data processing model after quantizing the calibration data using the corresponding target calibration parameters under the first step length parameter constraint, inputting the quantization result into the first network layer.

[0019] In a possible embodiment, after determining the first quantization calibration parameter based on the relationship between the calibration data and the quantization calibration loss, the method further includes:

[0020] Inputting a result obtained by quantizing the calibration data based on the first quantization calibration parameter into the first network layer, and using an output result of the first network layer as calibration data corresponding to a second network layer, where the second network layer is a network layer in the data processing model used to process the data output by the first network layer;

[0021] According to the relationship between the calibration data corresponding to the second network layer and the corresponding quantization calibration loss, the quantization calibration parameter corresponding to the first step length parameter in the second network layer is determined.

[0022] In a possible embodiment, after determining the first quantization calibration parameter based on the relationship between the calibration data and the quantization calibration loss, the method further includes:

[0023] When the first step length parameter is determined to correspond to the quantization calibration parameters of each network layer of the data processing model, a second step length parameter is obtained according to the first step length parameter;

[0024] Obtaining calibration data corresponding to the first network layer under the second step size parameter constraint according to output data of the data processing model under the first step size parameter constraint;

[0025] According to the relationship between the calibration data corresponding to the first network layer under the constraint of the second step size parameter and the corresponding quantization calibration loss, a second quantization calibration parameter is determined, where the second quantization calibration parameter is the quantization calibration parameter corresponding to the second step size parameter in the first network layer.

[0026] In a possible embodiment, determining the second quantization calibration parameter according to a relationship between the calibration data corresponding to the first network layer under the constraint of the second step size parameter and the corresponding quantization calibration loss includes:

[0027] Acquire a reference error control parameter, where the reference error control parameter is an error control parameter corresponding to the first quantization calibration parameter;

[0028] In a case where the error control parameter is the reference error control parameter, a calibration parameter when the corresponding quantization calibration loss obtains a minimum value is determined as the second quantization calibration parameter.

[0029] In a possible embodiment, when the error control parameter is a first preset value, determining a corresponding target calibration parameter and a corresponding target quantization performance loss includes:

[0030] quantizing the calibration data based on the corresponding target calibration parameters, and inputting the quantization result into the first network layer;

[0031] Obtaining a first data processing result output by the data processing model under the first step length parameter constraint;

[0032] directly inputting the calibration data into the first network layer to obtain a second data processing result output by the data processing model under the first step length parameter constraint;

[0033] The target quantization performance loss is determined according to a difference between the first data processing result and the second data processing result.

[0034] In a possible embodiment, after quantizing the first basic data based on the first quantization mapping coefficient and inputting the quantization result into the first network layer for data processing to obtain first output data of the first network layer, the method further includes:

[0035] using the first output data as second basic data of a second network layer, where the second network layer is a network layer in the data processing model used to process the first output data;

[0036] Obtaining a quantization calibration parameter corresponding to the first step length parameter at the second network layer;

[0037] Calculate, according to the second basic data and the quantization calibration parameter corresponding to the first step length parameter at the second network layer, a quantization mapping coefficient corresponding to the first step length parameter and the second network layer;

[0038] Based on the first step length parameter and the quantization mapping coefficient of the second network layer, the second basic data is quantized, and the quantization result is input into the second network layer for data processing to obtain output data of the second network layer.

[0039] In a possible embodiment, the method further includes:

[0040] When the target output data output by the data processing model under the constraint of the first step length parameter is obtained and the first step length parameter does not satisfy the step length parameter constraint condition, obtaining a second step length parameter according to the first step length parameter;

[0041] Obtaining a second quantization calibration parameter, where the second quantization calibration parameter is a quantization calibration parameter corresponding to the second step size parameter at the first network layer;

[0042] Calculating, according to the target output data and the second quantization calibration parameter, a second quantization mapping coefficient corresponding to the second step size parameter and the first network layer;

[0043] Quantization is performed based on the second quantization mapping coefficient and the target basic data, and the quantization result is input into the first network layer for data processing to obtain second output data of the first network layer.

[0044] In a possible embodiment, the method further includes:

[0045] When the target output data output by the data processing model under the first step length parameter constraint is obtained and the first step length parameter satisfies the step length parameter constraint condition, the target data is used as the final data processing result of the data processing model on the first basic data.

[0046] According to a second aspect of an embodiment of the present disclosure, there is provided a data processing apparatus, including:

[0047] A data acquisition module is configured to acquire a first step length parameter and first basic data, where the first basic data is representation data of text content or image content;

[0048] a quantization mapping coefficient determination module configured to obtain a first quantization calibration parameter corresponding to the first step length parameter at a first network layer of a data processing model, the first quantization calibration parameter being a parameter obtained by performing quantization calibration for the first network layer based on calibration data and the first step length parameter; and calculate a first quantization mapping coefficient corresponding to the first step length parameter and the first network layer based on the first basic data and the first quantization calibration parameter;

[0049] The data processing module is configured to quantize the first basic data based on the first quantization mapping coefficient, and input the quantization result into the first network layer for data processing to obtain first output data of the first network layer.

[0050] In one embodiment, the apparatus further comprises a pre-processing module, wherein the pre-processing module is configured to perform:

[0051] Determining a quantization calibration loss corresponding to the first step length parameter, wherein the quantization calibration loss takes the calibration data and the calibration parameter as variables, and the quantization calibration loss indicates a quantization accuracy loss of the calibration data;

[0052] Get calibration data;

[0053] The first quantization calibration parameter is determined according to a relationship between the calibration data and the quantization calibration loss.

[0054] In one embodiment, the quantization calibration loss further includes an error control parameter, and the pre-processing module is configured to perform:

[0055] When the error control parameter is a first preset value, determining a corresponding target calibration parameter and a corresponding target quantization performance loss;

[0056] determining a corresponding target calibration parameter and a corresponding target quantization performance loss when the error control parameter is a second preset value, the first preset value being different from the second preset value;

[0057] Using the target calibration parameter corresponding to the minimum target quantization performance loss as the first quantization calibration parameter;

[0058] The target calibration parameter is the calibration parameter when the corresponding quantization calibration loss reaches the minimum value.

[0059] The target quantization performance loss indicates the accuracy loss generated in the output data of the data processing model after quantizing the calibration data using the corresponding target calibration parameters under the first step length parameter constraint, inputting the quantization result into the first network layer.

[0060] In one embodiment, the pre-processing module is configured to perform:

[0061] Inputting a result obtained by quantizing the calibration data based on the first quantization calibration parameter into the first network layer, and using an output result of the first network layer as calibration data corresponding to a second network layer, where the second network layer is a network layer in the data processing model used to process the data output by the first network layer;

[0062] According to the relationship between the calibration data corresponding to the second network layer and the corresponding quantization calibration loss, the quantization calibration parameter corresponding to the first step length parameter in the second network layer is determined.

[0063] In one embodiment, the pre-processing module is configured to perform:

[0064] When the first step length parameter is determined to correspond to the quantization calibration parameters of each network layer of the data processing model, a second step length parameter is obtained according to the first step length parameter;

[0065] Obtaining calibration data corresponding to the first network layer under the second step size parameter constraint according to output data of the data processing model under the first step size parameter constraint;

[0066] According to the relationship between the calibration data corresponding to the first network layer under the constraint of the second step size parameter and the corresponding quantization calibration loss, a second quantization calibration parameter is determined, where the second quantization calibration parameter is the quantization calibration parameter corresponding to the second step size parameter in the first network layer.

[0067] In one embodiment, the pre-processing module is configured to perform:

[0068] Acquire a reference error control parameter, where the reference error control parameter is an error control parameter corresponding to the first quantization calibration parameter;

[0069] In a case where the error control parameter is the reference error control parameter, a calibration parameter when the corresponding quantization calibration loss obtains a minimum value is determined as the second quantization calibration parameter.

[0070] In one embodiment, the pre-processing module is configured to perform:

[0071] quantizing the calibration data based on the corresponding target calibration parameters, and inputting the quantization result into the first network layer;

[0072] Obtaining a first data processing result output by the data processing model under the first step length parameter constraint;

[0073] directly inputting the calibration data into the first network layer to obtain a second data processing result output by the data processing model under the first step length parameter constraint;

[0074] The target quantization performance loss is determined according to a difference between the first data processing result and the second data processing result.

[0075] In one embodiment, the data acquisition module is configured to execute: using the first output data as second basic data of a second network layer, where the second network layer is a network layer in the data processing model used to process the first output data;

[0076] The quantization mapping coefficient determination module is configured to execute: obtaining a quantization calibration parameter corresponding to the first step length parameter at the second network layer; and calculating a quantization mapping coefficient corresponding to the first step length parameter and the second network layer based on the second basic data and the quantization calibration parameter corresponding to the first step length parameter at the second network layer;

[0077] The data processing module is configured to execute: quantizing the second basic data based on the quantization mapping coefficient corresponding to the first step length parameter and the second network layer, and inputting the quantization result into the second network layer for data processing to obtain output data of the second network layer.

[0078] In one embodiment, the data acquisition module is configured to execute: when the target output data output by the data processing model under the first step length parameter constraint is acquired and the first step length parameter does not satisfy the step length parameter constraint condition, obtain the second step length parameter according to the first step length parameter;

[0079] The quantization mapping coefficient determination module is configured to: obtain a second quantization calibration parameter, where the second quantization calibration parameter is a quantization calibration parameter corresponding to the second step size parameter at the first network layer; and calculate a second quantization mapping coefficient corresponding to the second step size parameter and the first network layer based on the target output data and the second quantization calibration parameter;

[0080] The data processing module is configured to perform: quantization based on the second quantization mapping coefficient and the target basic data, and input the quantization result into the first network layer for data processing to obtain second output data of the first network layer.

[0081] In one embodiment, the data processing module is configured to execute: when the target output data output by the data processing model under the first step length parameter constraint is obtained and the first step length parameter satisfies the step length parameter constraint condition, the target data is used as the final data processing result of the data processing model on the first basic data.

[0082] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute instructions to implement a method as described in any one of the first aspects above.

[0083] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute any one of the methods in the first aspect of the embodiment of the present disclosure.

[0084] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, which includes a computer program, the computer program is stored in a readable storage medium, and at least one processor of a computer device reads and executes the computer program from the readable storage medium, so that the computer device performs any one of the methods in the first aspect of the embodiment of the present disclosure.

[0085] The technical solutions provided by the embodiments of the present disclosure bring at least the following beneficial effects:

[0086] The disclosed embodiment provides a data processing solution that essentially achieves quantization of any layer of forward reasoning under any step-size parameter by dynamically determining the quantization mapping coefficient used during forward reasoning at any layer under any step-size parameter of the Diffusion model. The quantization mapping coefficient is calculated based on the data to be processed (processing object) and the quantization calibration parameter of any layer under the arbitrary step-size parameter. The quantization calibration parameter is a fixed value obtained by pre-quantization calibration of the forward reasoning process of the layer under the step-size parameter based on the calibration data. Therefore, the fixed value is a known parameter before the specific quantization is executed, and will not reduce the calculation speed of the quantization mapping coefficient or the forward reasoning speed. However, it can significantly improve the quantization performance, reduce the space overhead during data processing, and control the data accuracy.

[0087] The present disclosure is applicable to data processing models that perform iterative forward reasoning, a typical example of which is the Diffusion model. Taking the Diffusion model as an example, the introduction of this fixed value is equivalent to pre-calibrating the iterative forward reasoning process at each layer under each step size parameter of the Diffusion model, thereby ensuring the reasoning accuracy of the quantized Diffusion model. Therefore, the data processing solution provided by the present disclosure can reduce the operating space occupied by the Diffusion model while ensuring the reasoning accuracy of the Diffusion model without substantially compromising the Diffusion model's reasoning efficiency, and can achieve good data processing performance for quantization scenarios such as INT8, INT4, and INT6.

[0088] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0090] Figure 1 is a schematic diagram of an implementation environment of a data processing method according to an exemplary embodiment;

[0091] Figure 2 is a flow chart showing a data processing method according to an exemplary embodiment;

[0092] Figure 3 The process of determining the quantitative calibration parameters according to an exemplary embodiment is shown as follows Figure 1 ;

[0093] Figure 4 The process of determining the quantitative calibration parameters according to an exemplary embodiment is shown as follows Figure 2 ;

[0094] Figure 5 is a flow chart showing a method for determining target quantization performance loss according to an exemplary embodiment;

[0095] Figure 6 is a comparison diagram of the implementation effects of the present disclosure according to an exemplary embodiment;

[0096] Figure 7 is a block diagram of a data processing device according to an exemplary embodiment;

[0097] Figure 8 A structural block diagram of a computer device according to an exemplary embodiment is shown. Figure 1;

[0098] Figure 9 A structural block diagram of a computer device according to an exemplary embodiment is shown. Figure 2 . DETAILED DESCRIPTION

[0099] In order to enable ordinary people in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0100] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar first objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0101] Before introducing the method embodiments provided in the present application, a brief introduction is first given to the relevant terms or nouns that may be involved in the method embodiments of the present application to facilitate understanding by those skilled in the art in the field of the present application.

[0102] Diffusion model: A diffusion model is a deep generative model. In a diffusion model, data undergoes two phases: diffusion and inverse diffusion. In the diffusion phase, noise is added to the original data, shifting the data distribution from its original distribution to a desired distribution. In the inverse diffusion phase, the data is restored to its original distribution.

[0103] Diffusion models are widely used in the field of image generation. When used for image generation, because the Diffusion model is a multi-step iterative model, it requires multiple iterations of forward reasoning to obtain the output result. Each forward reasoning requires the use of a corresponding step parameter (step). In addition, each forward reasoning also requires the use of noise image parameters. These noise image parameters serve as the reasoning basis for the forward reasoning of the Diffusion model. The Diffusion model reduces the noise in these noise image parameters layer by layer to obtain the image output result of this forward reasoning, that is, to update the noise image parameters. In the next iteration, the step is updated and forward reasoning continues using the updated noise image parameters until the final image generation result is obtained.

[0104] Int8 is a data type representing an 8-bit signed integer. Its value range is from -128 to 127, with -128 being the minimum and 127 being the maximum. In computer hardware, it requires only one byte of storage space. INT6 and INT4 are similar, representing 6-bit and 4-bit signed integers, respectively.

[0105] FP32 refers to 32-bit floating-point numbers, also known as single-precision floating-point numbers. It is widely used in graphics processing, scientific computing, machine learning, and deep learning. In computer hardware, it requires 4 bytes of storage space.

[0106] Quantization: Converting numbers that take up more storage space and have slower computation speed, such as FP32, to lower-precision representations, such as Int8, Int6, or Int4.

[0107] Before describing the embodiments of the present disclosure in detail, the relevant technical background related to the embodiments of the present disclosure is introduced to facilitate understanding by those skilled in the art in the art of this application.

[0108] As a high-performance deep generative model, the Diffusion model has been widely used in related technologies, especially in the field of image generation, where it has fully demonstrated its potential and superiority. However, because the Diffusion model relies on an iterative denoising process, its inference cost is much higher than that of ordinary models. The image generation time of common Diffusion models has reached the order of seconds, and the high computing power required has severely limited the widespread application of the Diffusion model.

[0109] Quantization has been proven to effectively accelerate Diffusion models, reducing the storage space occupied by the Diffusion model during runtime and increasing the speed of the Diffusion model's output results. For example, quantizing the input data of each forward inference of the Diffusion model to INT8 can significantly reduce the storage space occupied by the Diffusion model. Thanks to the hardware-friendly nature of INT8 multiplication and addition, it can also accelerate the model's inference speed. In the research of other neural network models, quantization has already developed mature solutions, but the Diffusion model is different from other neural network models because the Diffusion model performs an iterative forward inference process, and the input of each iteration is different. Therefore, the quantization solutions in related technologies cannot meet the quantization requirements of the Diffusion model.

[0110] The quantization of the Diffusion model can be performed based on the basic quantization formula. In [ ], x represents the quantized object, [] denotes the rounding operation, and clamp denotes the truncation operation. Numbers less than qmin are set to qmin, and numbers greater than qmax are set to qmax. s represents the quantization mapping coefficient. For INT8, qmin = -128 and qmax = 127. The setting of the quantization mapping coefficient in the basic quantization formula has a significant impact on the quantization results of the Diffusion model.

[0111] In a related art, a basic dynamic setting method of quantization mapping coefficients is provided. The basic dynamic setting method uses the formula s=abs(x)max() / (2 n-1 -1) to determine the quantization mapping coefficient s. This formula is the basic quantization mapping coefficient formula. In this basic quantization mapping coefficient formula, abs() represents the absolute value and max() represents the maximum value. That is, the quantization mapping coefficient is uniquely determined based on the quantized number with the largest absolute value among a batch of quantized numbers and the number of quantization bits n. However, this quantization method has been proven by many works to be not optimal and often cannot achieve optimal performance. Especially in the Diffusion model, if each forward inference uses this basic quantization mapping formula to quantize the input of this forward inference, then the error introduced by quantization will tend to grow significantly with multiple iterations, which will have a significant impact on the output accuracy of the Diffusion model.

[0112] Another related technique, a static quantization method, offers acceptable performance with INT8 quantization but dramatically degrades with INT6 or INT4. Furthermore, additional optimization methods are required to ensure the inference accuracy of the quantized Diffusion model, making it time-consuming.

[0113] In another related technology, an optimized dynamic quantization method is proposed, which dynamically calculates the quantization mapping coefficients for each forward inference by additionally setting the neural network. However, the neural network in this method requires additional training, and each dynamic calculation of the quantization mapping coefficients will introduce a large computational overhead, which reduces the inference speed of the Diffusion model.

[0114] It can be seen that no quantitative solution has been proposed in the relevant technologies that can guarantee the inference accuracy while reducing the inference overhead of the Diffusion model. The inference overhead includes the time resources required for inference and the space resources occupied. Specifically, the first related technology cannot guarantee the inference accuracy of the Diffusion model, the second related technology increases the inference overhead because it is more time-consuming, and the third related technology also increases the inference overhead because of the computational overhead.

[0115] In view of this, an embodiment of the present disclosure provides a data processing solution, which essentially realizes the quantization of each layer of forward reasoning under each step-size parameter by dynamically determining the quantization mapping coefficient used in each layer of forward reasoning under each step-size parameter of the Diffusion model. The quantization mapping coefficient is calculated based on the data to be processed (processing object) and the quantization calibration parameter of each layer under each step-size parameter. The quantization calibration parameter is a fixed value, which is obtained by pre-quantization calibration of the forward reasoning process of this layer under the step-size parameter based on the calibration data. Therefore, the fixed value is a known parameter before the specific quantization is executed, and will not reduce the calculation speed of the quantization mapping coefficient or the forward reasoning speed.

[0116] The present disclosure is applicable to data processing models that perform iterative forward reasoning, a typical example of which is the Diffusion model. Taking the Diffusion model as an example, the introduction of this fixed value is equivalent to pre-calibrating the iterative forward reasoning process at each layer under each step size parameter of the Diffusion model, thereby ensuring the reasoning accuracy of the quantized Diffusion model. Therefore, the data processing solution provided by the present disclosure can reduce the operating space occupied by the Diffusion model while ensuring the reasoning accuracy of the Diffusion model without substantially compromising the Diffusion model's reasoning efficiency, and can achieve good performance for quantization scenarios such as INT8, INT4, and INT6.

[0117] All data or information involved in this disclosure are authorized by the user or fully authorized by all parties.

[0118] Figure 1 FIG. 1 is a schematic diagram showing an implementation environment of a data processing method according to an exemplary embodiment. Taking the electronic device as a terminal as an example, see Figure 1 , the implementation environment specifically includes: a terminal 101 and a server 102.

[0119] Terminal 101 can be at least one of a smartphone, a smartwatch, a desktop computer, a laptop computer, and a portable computer. An application program that utilizes the data processing method disclosed herein can be installed and run on terminal 101. Users can log in to the application program through terminal 101 to access services related to the data processing method disclosed herein. Terminal 101 can generally refer to one of multiple terminals; this embodiment uses terminal 101 as an example. Those skilled in the art will appreciate that the number of terminals described above may be greater or lesser. For example, there may be only a few terminals, or there may be dozens, hundreds, or even more terminals. The embodiments of this disclosure do not limit the number of terminals or the types of devices.

[0120] Server 102 can be at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Server 102 can be connected to terminal 101 and other terminals via a wireless network or a wired network. Server 102 can implement the data processing method disclosed herein and transmit the data processing results to terminal 101, which can then use or display the data processing results. Of course, server 102 can also include other functional servers to provide more comprehensive and diverse services.

[0121] Figure 2 is a flow chart showing a data processing method according to an exemplary embodiment. Figure 2 As shown, the above method at least includes the following steps S210-S240.

[0122] In step S210, a first step length parameter and first basic data are acquired, where the first basic data is representation data of text content or image content.

[0123] The data processing method of the present disclosure may be applicable to a data processing model that completes a data processing process through iterative forward reasoning. The present disclosure uses a Diffusion model as an example to elaborate on the data processing method of the present disclosure.

[0124] The Diffusion model is an iterative forward reasoning model. The input data used in each forward reasoning includes two parts. One part is used to characterize the iterative information. This type of information can be understood as a step size parameter. The other part of the information is the direct processing object of the Diffusion model. This type of information is called basic data. For the Diffusion model, the input of each network layer in each iteration includes the corresponding step size parameter and the direct processing object of the network layer. The Diffusion model can process information in text or information in images. Therefore, the basic data is the representation data of the text content or image content. Through multiple iterations of forward reasoning, the Diffusion model can eventually output an image that fully matches the representation data. The present disclosure does not limit the representation data. For example, the representation data can be obtained by directly extracting features from the text content or image content, or the data output by any network layer after any iteration in the Diffusion model can be understood as a type of representation data of the text content or image content.

[0125] Step S210 represents the data processing process of any network layer in any iteration. Step S210 can obtain the corresponding step size parameter and basic data, which are the first step size parameter and the first basic data in step S210, and the corresponding network layer is the first network layer. Taking the first network layer in the first iteration as an example, the first step size parameter is the initial value of the step size parameter, which is a known parameter, and the first basic data is the initial data processing object of the Diffusion model. Taking the second network layer in the first iteration as an example, the first step size parameter is the initial value of the step size parameter, and the first basic data is the data output by the first network layer in the first iteration. Taking the first network layer in the second iteration as an example, the first step size parameter is the value of the initial value of the step size parameter after one iteration, and the first basic data is the data output by the last network layer in the first iteration.

[0126] In step S220, a first quantization calibration parameter corresponding to the first step length parameter in the first network layer of the data processing model is obtained, where the first quantization calibration parameter is a parameter obtained by performing quantization calibration on the first network layer based on the calibration data and the first step length parameter.

[0127] This disclosure pre-sets a quantization calibration parameter for each network layer at each iteration, thereby improving quantization calibration speed and accuracy. The quantization calibration parameter corresponding to the first step size parameter and the first network layer is referred to as the first quantization calibration parameter. The following text details the technical solution for setting the quantization calibration parameter corresponding to each network layer at each step size parameter based on calibration data, and is not further elaborated here.

[0128] In step S230, a first quantization mapping coefficient corresponding to the first step length parameter and the first network layer is calculated according to the first basic data and the first quantization calibration parameter.

[0129] The quantization of the Diffusion model can be performed based on the aforementioned basic quantization formula. In the related art, the basic quantization mapping coefficient formula s=abs(x)max() / (2 n-1 -1) to determine s, it can be seen that s is a parameter related to the quantized object. In order to solve the problem of directly using s=abs(x)max() / (2 n-1-1) to determine the problem of poor quantization performance of s. In the present disclosure, quantization calibration parameters are used to calibrate s. Therefore, in step S230, the first quantization mapping coefficient is calculated based on the first quantization calibration parameter and the first basic data. This method is essentially still a dynamic quantization method because the first quantization mapping coefficient is dynamically set based on the first quantization calibration parameter and the first basic data. The first quantization calibration parameter is pre-calibrated, and the first basic data is dynamic data closely related to the iteration situation and network layer.

[0130] In the present disclosure, based on the formula s=f(t)*abs(x)max() / (2 n-1 -1) determine the first quantization mapping coefficient, where s is the first quantization mapping coefficient, abs(x)max() / (2 n-1 -1) has the same meaning as above, x is the first basic data, f(t) is the first quantization calibration parameter, t represents the step size parameter, and in step S230, t corresponds to the first step size parameter. The value range of the first quantization calibration parameter is 0 to 1, which represents the pre-calibrated parameter obtained for different step size parameters or iterative steps. Unlike the related art, which requires the introduction of an additional neural network to calculate the quantization mapping coefficient, the present disclosure directly uses the pre-calculated quantization calibration parameter in combination with abs(x)max() / (2 n-1 -1) to determine the optimal quantization mapping coefficient. This can also achieve the goal of ensuring quantization accuracy and data processing performance, while saving computing resources and improving data processing speed.

[0131] In step S240, the first basic data is quantized based on the first quantization mapping coefficient, and the quantization result is input into the first network layer for data processing to obtain the first output data of the first network layer.

[0132] The quantization in step S240 can be performed based on the aforementioned basic quantization formula. In the example, x uses the first base data, [] indicates a rounding operation, clamp indicates a truncation operation, and s is the first quantization mapping coefficient used. For INT8, qmin = -128 and qmax = 127. This disclosure does not limit the data processing process of the first network layer; this process is determined by the Diffusion model. The output of this first network layer is the first output data. Similarly, this disclosure does not limit the data processing process of other network layers; all of these are related to the data processing model itself.

[0133] The first step size parameter and the first basic data represent the step size parameter and basic data used by any network layer of any iteration of the Diffusion model. The present disclosure essentially realizes the quantization of each layer of forward reasoning under each step size parameter by dynamically determining the quantization mapping coefficient used in the forward reasoning of each layer under each step size parameter of the Diffusion model. The quantization mapping coefficient is calculated based on the data to be processed (basic data) and the quantization calibration parameter of each layer under each step size parameter. The quantization calibration parameter is a fixed value. The fixed value is obtained by pre-quantization calibration of the forward reasoning process of the layer under the step size parameter based on the calibration data. Therefore, the fixed value is a known parameter before the specific quantization is executed, and will not reduce the calculation speed of the quantization mapping coefficient or the forward reasoning speed. Taking the Diffusion model as an example, the introduction of this fixed value is equivalent to pre-calibrating the forward inference process of each layer of iteration under each step parameter of the Diffusion model. Therefore, it can reduce the running space occupied by the Diffusion model while ensuring the inference accuracy of the Diffusion model without basically losing the inference efficiency of the Diffusion model, and can achieve good performance for quantization conditions such as INT8, INT4, and INT6.

[0134] In one embodiment, if the first network layer is not the last network layer of the Diffusion model, the first basic data can be quantized based on the first quantization mapping coefficient, and the quantization result can be input into the first network layer for data processing to obtain the first output data of the first network layer. Then, the first output data can be used as the second basic data of the second network layer, and the second network layer is the network layer used to process the first output data in the data processing model; the quantization calibration parameter corresponding to the first step length parameter in the second network layer is obtained; according to the second basic data and the quantization calibration parameter corresponding to the first step length parameter in the second network layer, the quantization mapping coefficient corresponding to the first step length parameter and the second network layer is calculated; based on the quantization mapping coefficient corresponding to the first step length parameter and the second network layer, the second basic data is quantized, and the quantization result is input into the second network layer for data processing to obtain the output data of the second network layer.

[0135] The quantization calibration parameters corresponding to the first step length parameter in the second network layer are also obtained through pre-calibration, and are based on the same inventive concept as the method for obtaining the first quantization calibration parameter described above. The method for determining the quantization mapping coefficients corresponding to the first step length parameter and the second network layer is also based on the same inventive concept as the method for determining the first quantization mapping coefficients described above. The specific method for quantizing the second basic data and inputting the quantization results into the second network layer for data processing to obtain the output data of the second network layer is based on the same inventive concept as step S240 described above and will not be repeated here.

[0136] Since the second network layer can be understood as the network layer of the data processing model (Diffusion model) after the first network layer, after the first network layer completes data processing under the first step length parameter constraint, the data processed by the first network layer can be used to continue to perform quantization-based data processing on the next second network layer, thereby further completing the quantization-based data processing of subsequent network layers under the iteration rounds constrained by the first step length parameter, until the final data processing result of the data processing model under the iteration rounds constrained by the first step length parameter is obtained. Compared with the data processing process of the data processing model under the iteration rounds constrained by the first step length parameter in the related technology, this can obtain better quantization performance and more accurate data processing results.

[0137] In one embodiment, if the first step length parameter is not the last step length parameter of the Diffusion model, the first step length parameter does not satisfy the step length parameter constraint. When the target output data finally output by the data processing model under the constraint of the first step length parameter is obtained, and the first step length parameter does not satisfy the step length parameter constraint, a second step length parameter is obtained according to the first step length parameter; a second quantization calibration parameter is obtained, and the second quantization calibration parameter is the quantization calibration parameter corresponding to the second step length parameter in the first network layer; based on the target output data and the second quantization calibration parameter, the second quantization mapping coefficient corresponding to the second step length parameter and the first network layer is calculated; quantization is performed based on the second quantization mapping coefficient and the target basic data, and the quantization result is input into the first network layer for data processing to obtain the second output data of the first network layer. The target output data can be understood as the final output result of a single round of the Diffusion model under the constraint of the first step length parameter, that is, the output image of this round.

[0138] The target output data is subjected to a new round of iterative processing. This process is identical to the inventive concept of the previous round of iterative processing. Therefore, based on the target output data, the basic data corresponding to the first network layer under the second step size parameter can be determined. Based on the basic data corresponding to the first network layer under the second step size parameter and the second quantization calibration parameter, the second quantization mapping coefficient corresponding to the second step size parameter and the first network layer can be calculated. Based on the second quantization mapping coefficient, the basic data corresponding to the first network layer under the second step size parameter is quantized and the quantization result is input into the first network layer for data processing, thereby obtaining the second output data of the first network layer.

[0139] The method for determining the second quantization calibration parameter and the second quantization mapping coefficient is based on the same inventive concept as the method for determining the first quantization calibration parameter and the first quantization mapping coefficient. The specific quantization method is also based on the same inventive concept as step S240 above, and will not be described in detail here.

[0140] Because the second step size parameter can be understood as the step size parameter for the next round after the first step size parameter, after the entire Diffusion model completes data processing under the constraints of the first step size parameter, the target output data can be used to start the next round of data processing, until all rounds of data processing are completed to obtain the final data processing result of the overall processing process. Compared with the data processing process of the data processing model under the same iterative rounds in related technologies, this can achieve better quantitative performance and more accurate data processing results.

[0141] Of course, once the target output data output by the data processing model under the constraints of the first step length parameter is obtained, and the first step length parameter satisfies the step length parameter constraint, the target data is used as the final data processing result of the data processing model on the first basic data. In other words, if the first step length parameter indicates the last iteration, then the first step length parameter satisfies the step length parameter constraint. The target output data output by the data processing model under the constraints of the first step length parameter can be used as the final data processing result of the model. In the case of the Diffusion model, a final processed image can be obtained. The generation speed and generation accuracy of this final image are significantly superior to those of related technologies due to the data processing solution disclosed herein.

[0142] Next, the present disclosure details the technical solution for performing quantization calibration on each network layer in each iteration Step to obtain the corresponding quantization calibration parameters. The calculation process of the quantization calibration parameters belongs to the preprocessing process and is located before step S210. Therefore, it will not reduce the calculation speed of the quantization mapping coefficients when applied. Please refer to Figure 3 , which is a flow chart of a method for determining a quantitative calibration parameter according to an exemplary embodiment. Figure 1The method comprises:

[0143] Step S310: Determine the quantization calibration loss corresponding to the first step length parameter, wherein the quantization calibration loss takes the calibration data and the calibration parameters as variables, and indicates the quantization accuracy loss of the calibration data.

[0144] The quantization calibration loss in the embodiments of the present disclosure is related to the calibration data and calibration parameters, indicating the quantization accuracy loss caused by directly quantizing the calibration data based on the calibration parameters. The present disclosure does not limit the expression of the quantization calibration loss. In one embodiment, the quantization calibration loss loss = ||Q(x)*sx|| p , where x is the calibration parameter, and s can be obtained by f(t)*abs(x)max() / (2 n-1 -1) calculation, when used with loss, x also refers to the calibration parameter, f(t) is the calibration parameter of step S310. ||.|| p represents the matrix norm, and p is the error control parameter.

[0145] Step S320: Obtain calibration data.

[0146] This disclosure does not limit the calibration data; it can be set according to actual circumstances. It is simply a processing object that can trigger the data processing model to perform data processing, and its setting method does not constitute an implementation limitation of this disclosure. Obviously, any data representing text content or image content can be used as the calibration data.

[0147] Step S330 : Determine the first quantization calibration parameter according to the relationship between the calibration data and the quantization calibration loss.

[0148] The present disclosure sets the quantization calibration loss and can determine the first quantization calibration parameter based on the relationship between the calibration data and the quantization calibration loss. The first quantization calibration parameter can be used for rapid quantization and ensuring quantization performance in the application stage.

[0149] Please refer to Figure 4 , which is a flow chart of a method for determining a quantitative calibration parameter according to an exemplary embodiment. Figure 2 The method comprises:

[0150] S410. When the above-mentioned error control parameter is a first preset value, determine the corresponding target calibration parameter and the corresponding target quantization performance loss, the above-mentioned target calibration parameter is the calibration parameter when the corresponding quantization calibration loss takes the minimum value, and the above-mentioned target quantization performance loss indicates the accuracy loss generated in the output data of the above-mentioned data processing model after quantizing the above-mentioned calibration data using the corresponding target calibration parameter under the constraint of the above-mentioned first step length parameter, and inputting the quantization result into the above-mentioned first network layer.

[0151] Based on quantization calibration loss loss=||Q(x)*sx|| p It can be seen that the calibration parameter f(t), the error control parameter p, and the calibration data x all affect the value of the quantization calibration loss. Therefore, when the error control parameter p is fixed, the calibration parameter that minimizes the quantization calibration loss can be obtained. Therefore, when the error control parameter is at the first preset value, the calibration parameter that minimizes the corresponding quantization calibration loss can be obtained, that is, the target calibration parameter corresponding to the first preset value of the error control parameter can be obtained.

[0152] Please refer to Figure 5 , which is a flow chart of a method for determining a target quantization performance loss according to an exemplary embodiment. When the error control parameter is a first preset value, determining the corresponding target calibration parameter and the corresponding target quantization performance loss includes:

[0153] S510. Quantize the calibration data based on the corresponding target calibration parameters, and input the quantization result into the first network layer.

[0154] Quantization is implemented based on the basic quantization formula. This process is based on the same inventive concept as the quantization in the previous article and will not be described in detail here.

[0155] S520. Obtain the first data processing result output by the above data processing model under the above first step length parameter constraint.

[0156] The data processing result output by the first network layer is further processed by the subsequent network layer of the data processing model to obtain the final data processing result under the iteration step indicated by the first step length parameter, that is, the first data processing result.

[0157] S530. Input the calibration data directly into the first network layer to obtain a second data processing result output by the data processing model under the first step length parameter constraint.

[0158] If we directly use the data processing model for data processing without quantization, we can obtain the final data processing result at the iteration step indicated by the first step length parameter, that is, the second data processing result. Taking the Diffusion model as an example, both the first data processing result and the second data processing result are images.

[0159] S540. Determine the target quantization performance loss based on the difference between the first data processing result and the second data processing result.

[0160] The target quantization performance loss is the loss of the difference between the first data processing result and the second data processing result, that is, the final precision loss incurred in a single iteration after quantization. This disclosure does not limit the specific calculation method of the target performance quantization loss, and it does not constitute an obstacle to the implementation of this disclosure.

[0161] In the present disclosure, when the error control parameters are determined, the corresponding target calibration parameters and target quantization performance loss can be calculated, so as to quantize the final accuracy loss generated by the target calibration parameters in a single iteration round when the error control parameters are determined. According to the loss, the quantization error control parameters and target calibration parameters with the smallest accuracy loss can be selected, and then reasonable quantization mapping coefficients can be obtained in the application stage to ensure the data processing accuracy in the application stage without reducing the data processing speed.

[0162] S420. When the error control parameter is a second preset value, determine a corresponding target calibration parameter and a corresponding target quantization performance loss, and the first preset value is different from the second preset value.

[0163] The error control parameter can be varied according to preset rules, and the selection of the error control parameter is not limited in this disclosure. For calibration data, if the error control parameter p is small, the target quantization performance loss obtained will more closely reflect the global error caused by quantization. If the error control parameter p is large, the target quantization performance loss will focus more on the precision error caused by the larger absolute values ​​of the quantized objects.

[0164] In one embodiment, the error control parameter p can be selected from several values ​​between 2 and 3.2, all of which result in corresponding target calibration parameters and corresponding target quantization performance losses. In step S420, the error control parameter p is selected from a different value than that in step S410, resulting in different target calibration parameters and corresponding target quantization performance losses. The first preset value and the second preset value can both be values ​​between 2 and 3.2. Obviously, several triplets can be obtained, each of which includes the error control parameter p, the corresponding target calibration parameter, and the corresponding target quantization performance loss.

[0165] S430. Use the target calibration parameter corresponding to the minimum target quantization performance loss as the first quantization calibration parameter.

[0166] The target calibration parameter in the triplet with the smallest target quantization performance loss is used as the first quantization calibration parameter. This disclosure provides a specific method for determining the first quantization calibration parameter. The obtained first quantization calibration parameter can be obtained in step S210 and used to quantize the first basic data, thereby improving the data processing speed of the data processing model through quantization, reducing the operating space, and controlling accuracy loss.

[0167] In the case of obtaining the first quantitative calibration parameter corresponding to the first network layer of the first step length parameter, it is also possible to continue to determine the quantitative calibration parameters corresponding to other network layers under the first step length parameter. If the first network layer is not the last network layer of the data processing model, then the result obtained after quantizing the calibration data based on the first quantitative calibration parameter can be input into the first network layer, and the output result of the first network layer can be used as the calibration data corresponding to the second network layer. The second network layer is the network layer used to process the data output by the first network layer in the data processing model; according to the relationship between the calibration data corresponding to the second network layer and the corresponding quantitative calibration loss, the quantitative calibration parameter corresponding to the first step length parameter in the second network layer is determined. Obviously, the inventive concept of determining the quantitative calibration parameter corresponding to the first step length parameter in the second network layer according to the relationship between the calibration data corresponding to the second network layer and the corresponding quantitative calibration loss is the same as the inventive concept of obtaining the first quantitative calibration parameter in the previous text, and will not be repeated here.

[0168] Since the second network layer can be understood as the network layer after the first network layer of the data processing model (Diffusion model), after the quantization calibration parameters (first quantization calibration parameters) corresponding to the first network layer are determined under the first step length parameter constraint, the quantization calibration parameters corresponding to the second network layer under the first step length parameter constraint can be further obtained, until the quantization calibration parameters of each network layer under the iteration round constrained by the first step length parameter are obtained. These quantization calibration parameters can be used to quickly obtain the corresponding quantization mapping coefficients in the application stage, thereby realizing high-performance quantization-based data processing.

[0169] When the quantization calibration parameters corresponding to each network layer under the constraint of the first step length parameter are obtained, if the first step length parameter is not the step length parameter of the last iteration round, the quantization calibration parameters of each network layer under the second step length parameter can be determined. Specifically, the second step length parameter can be obtained according to the first step length parameter after determining the quantization calibration parameters corresponding to each network layer of the data processing model respectively. The calibration data corresponding to the first network layer under the constraint of the second step length parameter can be obtained according to the output data of the data processing model under the constraint of the first step length parameter. In other words, the data processing object of the first network layer in the next iteration is obtained according to the final output of the forward reasoning of the previous iteration. This process is the same as the inventive concept of obtaining the second basic data in the previous text and will not be repeated.

[0170] Based on the relationship between the calibration data corresponding to the first network layer under the constraint of the second step size parameter and the corresponding quantization calibration loss, a second quantization calibration parameter is determined. The second quantization calibration parameter is the quantization calibration parameter corresponding to the second step size parameter in the first network layer. The process of deriving the second step size parameter based on the first step size parameter is the process of iterative forward reasoning itself and is also a characteristic of the Diffusion model. This disclosure does not elaborate on this process.

[0171] Because the second step size parameter can be understood as the step size parameter for the next round after the first step size parameter, after all network layers in the entire Diffusion model have obtained corresponding quantization calibration parameters under the constraints of the first step size parameter, the next round of data processing can be started to obtain the corresponding quantization calibration parameters for each network layer in the next round, and the quantization calibration parameters for each network layer in the next round are obtained. This continues until the quantization calibration parameters for each network layer in all rounds are obtained. This allows for rapid quantization of each network layer in each round during the application phase, fully ensuring quantization speed and data processing accuracy.

[0172] In one embodiment, the method for determining the second quantization calibration parameter can be based on the same inventive concept as the first quantization calibration parameter. In another embodiment, a technical solution is provided for improving the speed of determining the second quantization calibration parameter based on the quantization calibration process of the first network layer under the first step length parameter constraint. Specifically, a reference error control parameter can be obtained, where the reference error control parameter is the error control parameter corresponding to the first quantization calibration parameter. When the error control parameter is the reference error control parameter, the calibration parameter at which the corresponding quantization calibration loss reaches a minimum value is determined as the second quantization calibration parameter.

[0173] Having previously obtained several triplets, the error control parameter of the triplet containing the first quantization calibration parameter can be directly used as the error quantization parameter used in step S610, eliminating the need to regenerate several triplets. In other words, the error quantization parameter p used in calculating the second quantization calibration parameter is directly obtained, and the corresponding target calibration parameter can then be obtained, which is directly used as the second quantization calibration parameter. This embodiment fully utilizes the results of the previous round of iterative quantization calibration to improve the calculation speed of the quantization calibration parameter.

[0174] The data processing solution provided by the present disclosure is essentially a dynamic quantization solution combined with pre-calibrated quantization calibration parameters. The technical effect of this solution has significant advantages over the other related technologies mentioned above.

[0175] Please refer to Figure 6 , which is a comparison diagram of the implementation effects of the present disclosure according to an exemplary embodiment. Figure 6 The left figure compares the image generation results of the disclosed data processing solution (this solution) with the previously mentioned static quantization solution (Solution 1) and optimized dynamic quantization solution (Solution 2). The application scenario is the INT8 quantized DDIM model. DDIM (Denoising Diffusion Implicit Model) can be considered an optimization of DDPM (Denoising Diffusion Probabilistic Model), using data from the commonly used CIFAR-10 image classification dataset.

[0176] The FID score (Fréchet Inception Distance) is a metric used to evaluate image quality, particularly when evaluating images generated by generative adversarial networks. The FID score measures the similarity between the features of the real and generated images by calculating the distance between them. Lower FID scores indicate better quality generated images.

[0177] Figure 6 The left figure shows that the FID of the data processing scheme disclosed in the present invention is significantly lower than that of schemes 1 and 2, which fully demonstrates that the image generation effect of the present invention is better. When processing the same basic data, the present invention can generate images of better quality.

[0178] Figure 6 The figure on the right shows a comparison of inference latency with the static quantization scheme. The latency ratio achieved with this disclosure is significantly lower than that achieved with Scheme 2, and therefore lower than the latency achieved with the optimized dynamic quantization scheme. As a dynamic quantization scheme, this disclosure's latency performance is very close to that of Scheme 1, a static quantization scheme.

[0179] The present disclosure can utilize dynamic INT8 quantization to accelerate the data processing process of the Diffusion model, such as the image generation process. This process can bring an acceleration effect of 1.5 to 2 to the Diffusion model, while also significantly reducing storage requirements. And while ensuring the INT8 acceleration effect, it can obtain data processing performance far exceeding that of Solution 1, and compared to Solution 2, it achieves a faster reasoning speed. As a dynamic quantization solution, the present disclosure is a computationally friendly dynamic quantization method that introduces almost no additional computing overhead while ensuring quantization accuracy, and even basically achieves the reasoning efficiency of using static quantization. The reasoning speed of the present disclosure is already close to that of the static quantization solution, which has obviously achieved a very ideal reasoning speed improvement.

[0180] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0181] Figure 7 FIG. 1 is a block diagram of a data processing device according to an exemplary embodiment. Figure 7 , the device comprises:

[0182] The data acquisition module 710 is configured to acquire a first step length parameter and first basic data, where the first basic data is data representing text content or image content;

[0183] The quantization mapping coefficient determination module 720 is configured to obtain a first quantization calibration parameter corresponding to the first step length parameter at a first network layer of the data processing model, the first quantization calibration parameter being a parameter obtained by performing quantization calibration for the first network layer based on the calibration data and the first step length parameter; and calculate a first quantization mapping coefficient corresponding to the first step length parameter and the first network layer based on the first basic data and the first quantization calibration parameter.

[0184] The data processing module 730 is configured to quantize the first basic data based on the first quantization mapping coefficient, and input the quantization result into the first network layer for data processing to obtain the first output data of the first network layer.

[0185] In one embodiment, the apparatus further includes a pre-processing module 740, which is configured to execute:

[0186] Determine a quantization calibration loss corresponding to the first step length parameter, wherein the quantization calibration loss uses the calibration data and the calibration parameter as variables, and the quantization calibration loss indicates a loss in quantization accuracy of the calibration data;

[0187] Get calibration data;

[0188] The first quantization calibration parameter is determined according to the relationship between the calibration data and the quantization calibration loss.

[0189] In one embodiment, the quantization calibration loss further includes an error control parameter, and the pre-processing module 740 is configured to perform:

[0190] When the error control parameter is a first preset value, determining a corresponding target calibration parameter and a corresponding target quantization performance loss;

[0191] determining a corresponding target calibration parameter and a corresponding target quantization performance loss when the error control parameter is a second preset value, the first preset value being different from the second preset value;

[0192] Using the target calibration parameter corresponding to the minimum target quantization performance loss as the first quantization calibration parameter;

[0193] The target calibration parameters are the calibration parameters when the corresponding quantization calibration loss reaches the minimum value.

[0194] The target quantization performance loss indicates the accuracy loss in the output data of the data processing model after quantizing the calibration data using the corresponding target calibration parameters under the first step length parameter constraint, inputting the quantization result into the first network layer.

[0195] In one embodiment, the pre-processing module 740 is configured to perform:

[0196] Inputting a result obtained by quantizing the calibration data based on the first quantization calibration parameter into the first network layer, and using the output result of the first network layer as calibration data corresponding to the second network layer, the second network layer being the network layer in the data processing model used to process the data output by the first network layer;

[0197] According to the relationship between the calibration data corresponding to the second network layer and the corresponding quantization calibration loss, the quantization calibration parameter corresponding to the first step length parameter in the second network layer is determined.

[0198] In one embodiment, the pre-processing module 740 is configured to perform:

[0199] After determining the quantization calibration parameters corresponding to the first step length parameters at each network layer of the data processing model, obtaining the second step length parameters according to the first step length parameters;

[0200] Obtaining calibration data corresponding to the first network layer under the second step length parameter constraint based on the output data of the data processing model under the first step length parameter constraint;

[0201] According to the relationship between the calibration data corresponding to the above-mentioned first network layer under the above-mentioned second step size parameter constraint and the corresponding quantization calibration loss, the second quantization calibration parameter is determined, and the above-mentioned second quantization calibration parameter is the quantization calibration parameter corresponding to the above-mentioned second step size parameter in the above-mentioned first network layer.

[0202] In one embodiment, the pre-processing module 740 is configured to perform:

[0203] Obtaining a reference error control parameter, where the reference error control parameter is an error control parameter corresponding to the first quantization calibration parameter;

[0204] In a case where the error control parameter is the reference error control parameter, the calibration parameter when the corresponding quantization calibration loss takes a minimum value is determined as the second quantization calibration parameter.

[0205] In one embodiment, the pre-processing module 740 is configured to perform:

[0206] quantizing the calibration data based on the corresponding target calibration parameters, and inputting the quantization result into the first network layer;

[0207] Obtaining a first data processing result output by the data processing model under the first step length parameter constraint;

[0208] Directly inputting the calibration data into the first network layer to obtain a second data processing result output by the data processing model under the first step length parameter constraint;

[0209] The target quantization performance loss is determined based on a difference between the first data processing result and the second data processing result.

[0210] In one embodiment, the data acquisition module 710 is configured to: use the first output data as second basic data of a second network layer, where the second network layer is a network layer in the data processing model used to process the first output data;

[0211] The quantization mapping coefficient determination module 720 is configured to: obtain a quantization calibration parameter corresponding to the first step length parameter at the second network layer; and calculate a quantization mapping coefficient corresponding to the first step length parameter and the second network layer based on the second basic data and the quantization calibration parameter corresponding to the first step length parameter at the second network layer.

[0212] The above-mentioned data processing module 730 is configured to perform: quantizing the above-mentioned second basic data based on the above-mentioned first step length parameter and the quantization mapping coefficient of the above-mentioned second network layer, and inputting the quantization result into the above-mentioned second network layer for data processing to obtain the output data of the above-mentioned second network layer.

[0213] In one embodiment, the data acquisition module 710 is configured to: upon acquiring the target output data output by the data processing model under the first step length parameter constraint, and if the first step length parameter does not satisfy the step length parameter constraint, obtain a second step length parameter based on the first step length parameter;

[0214] The quantization mapping coefficient determination module 720 is configured to: obtain a second quantization calibration parameter, where the second quantization calibration parameter is a quantization calibration parameter corresponding to the second step size parameter at the first network layer; and calculate a second quantization mapping coefficient corresponding to the second step size parameter and the first network layer based on the target output data and the second quantization calibration parameter.

[0215] The data processing module 730 is configured to perform: quantization based on the second quantization mapping coefficient and the target basic data, and input the quantization result into the first network layer for data processing to obtain the second output data of the first network layer.

[0216] In one embodiment, the above-mentioned data processing module 730 is configured to execute: when the target output data output by the above-mentioned data processing model under the above-mentioned first step length parameter constraint is obtained, and the above-mentioned first step length parameter satisfies the step length parameter constraint condition, the above-mentioned target data is used as the final data processing result of the above-mentioned data processing model on the above-mentioned first basic data.

[0217] Regarding the apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments of the relevant methods and will not be elaborated on here.

[0218] Please refer to Figure 8 , which shows a structural block diagram of a computer device provided by an exemplary embodiment of the present disclosure Figure 1 The computer device may be a terminal. The computer device is used to implement the data processing method provided in the above embodiment. Specifically:

[0219] Typically, the computer device 800 includes a processor 801 and a memory 802 .

[0220] The processor 801 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 801 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), or PLA (Programmable Logic Array). The processor 801 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In an exemplary embodiment, the processor 801 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In an exemplary embodiment, the processor 801 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0221] The memory 802 may include one or more computer-readable storage media, which may be non-transitory. The memory 802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In an exemplary embodiment, the non-transitory computer-readable storage medium in the memory 802 is used to store at least one instruction, at least one program, code set, or instruction set, and is configured to be executed by one or more processors to implement the above-mentioned data processing method.

[0222] In an exemplary embodiment, computer device 800 may optionally include a peripheral device interface 803 and at least one peripheral device. Processor 801, memory 802, and peripheral device interface 803 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 803 via a bus, signal lines, or circuit boards. Specifically, the peripheral device includes at least one of a radio frequency circuit 804, a touchscreen display 805, a camera assembly 806, an audio circuit 807, a positioning assembly 808, and a power supply 809.

[0223] Those skilled in the art will understand that Figure 7The structure shown in the figure does not constitute a limitation on the computer device 800, and the computer device 800 may include more or fewer components than shown in the figure, or combine some components, or adopt a different component arrangement.

[0224] Please refer to Figure 9 , which shows a structural block diagram of a computer device provided by another exemplary embodiment of the present disclosure Figure 2 The computer device may be a server for executing the above data processing method. Specifically:

[0225] Computer device 900 includes a central processing unit (CPU) 901, a system memory 904 including a random access memory (RAM) 902 and a read-only memory (ROM) 903, and a system bus 905 connecting system memory 904 and CPU 901. Computer device 900 also includes a basic input / output system (I / O system) 906 that facilitates information transfer between various components within the computer, and a mass storage device 907 for storing an operating system 913, application programs 914, and other program modules 911.

[0226] The basic input / output system 906 includes a display 908 for displaying information and an input device 909, such as a mouse and keyboard, for user input. Both the display 908 and the input device 909 are connected to the central processing unit 901 via an input / output controller 910 connected to the system bus 905. The basic input / output system 906 may also include an input / output controller 910 for receiving and processing input from a variety of other devices, such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 910 also provides output to a display screen, printer, or other types of output devices.

[0227] The mass storage device 907 is connected to the central processing unit 901 via a mass storage controller (not shown) connected to the system bus 905. The mass storage device 907 and its associated computer-readable media provide non-volatile storage for the computer device 900. In other words, the mass storage device 907 may include a computer-readable medium (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.

[0228] Without loss of generality, computer-readable media may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include RAM, ROM, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc) or other optical storage, tape cassettes, magnetic tape, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that computer storage media are not limited to the above-mentioned ones. The above-mentioned system memory 904 and mass storage device 907 can be collectively referred to as memory.

[0229] According to various embodiments of the present disclosure, the computer device 900 may also be connected to a remote computer on a network such as the Internet for operation. That is, the computer device 900 may be connected to a network 912 via a network interface unit 911 connected to the system bus 905, or the network interface unit 911 may be used to connect to other types of networks or remote computer systems (not shown).

[0230] The memory further includes a computer program, which is stored in the memory and configured to be executed by one or more processors to implement the data processing method.

[0231] In an exemplary embodiment, a computer-readable storage medium is also provided, in which at least one instruction, at least one program, a code set or an instruction set is stored. When the at least one instruction, the at least one program, the code set or the instruction set is executed by a processor, the data processing method is implemented.

[0232] Optionally, the computer-readable storage medium may include: ROM (Read Only Memory), RAM (Random Access Memory), SSD (Solid State Drives), or an optical disk, etc. Among them, the random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).

[0233] In an exemplary embodiment, a computer-readable storage medium including program code is also provided, such as a memory including the program code, and the program code can be executed by a processor to perform the above-mentioned data processing method. Alternatively, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0234] In an exemplary embodiment, a computer program product is further provided, including a computer program, which implements the above data processing method when executed by a processor.

[0235] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0236] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A data processing method, characterized in that: include: Obtaining a first step length parameter and first basic data, where the first basic data is representation data of text content or image content; Obtaining a first quantization calibration parameter corresponding to the first step length parameter at a first network layer of a data processing model, where the first quantization calibration parameter is a parameter obtained by performing quantization calibration on the first network layer based on calibration data and the first step length parameter; Calculating, according to the first basic data and the first quantization calibration parameter, a first quantization mapping coefficient corresponding to the first step length parameter and the first network layer; The first basic data is quantized based on the first quantization mapping coefficient, and the quantization result is input into the first network layer for data processing to obtain first output data of the first network layer.

2. The data processing method according to claim 1, wherein: Before obtaining the first step length parameter and the first basic data, the method includes: Determining a quantization calibration loss corresponding to the first step length parameter, wherein the quantization calibration loss takes the calibration data and the calibration parameter as variables, and the quantization calibration loss indicates a quantization accuracy loss of the calibration data; Get calibration data; The first quantization calibration parameter is determined according to a relationship between the calibration data and the quantization calibration loss.

3. The data processing method according to claim 2, characterized in that: The quantization calibration loss also includes an error control parameter, and determining the first quantization calibration parameter according to a relationship between the calibration data and the quantization calibration loss includes: When the error control parameter is a first preset value, determining a corresponding target calibration parameter and a corresponding target quantization performance loss; determining a corresponding target calibration parameter and a corresponding target quantization performance loss when the error control parameter is a second preset value, the first preset value being different from the second preset value; Using the target calibration parameter corresponding to the minimum target quantization performance loss as the first quantization calibration parameter; The target calibration parameter is the calibration parameter when the corresponding quantization calibration loss reaches the minimum value. The target quantization performance loss indicates the accuracy loss generated in the output data of the data processing model after quantizing the calibration data using the corresponding target calibration parameters under the first step length parameter constraint, inputting the quantization result into the first network layer.

4. The data processing method according to claim 2 or 3, characterized in that: After determining the first quantization calibration parameter according to the relationship between the calibration data and the quantization calibration loss, the method further includes: Inputting a result obtained by quantizing the calibration data based on the first quantization calibration parameter into the first network layer, and using an output result of the first network layer as calibration data corresponding to a second network layer, where the second network layer is a network layer in the data processing model used to process the data output by the first network layer; According to the relationship between the calibration data corresponding to the second network layer and the corresponding quantization calibration loss, the quantization calibration parameter corresponding to the first step length parameter in the second network layer is determined.

5. The data processing method according to claim 4, characterized in that: After determining the first quantization calibration parameter according to the relationship between the calibration data and the quantization calibration loss, the method further includes: When the first step length parameter is determined to correspond to the quantization calibration parameters of each network layer of the data processing model, a second step length parameter is obtained according to the first step length parameter; Obtaining calibration data corresponding to the first network layer under the second step size parameter constraint according to output data of the data processing model under the first step size parameter constraint; According to the relationship between the calibration data corresponding to the first network layer under the constraint of the second step size parameter and the corresponding quantization calibration loss, a second quantization calibration parameter is determined, where the second quantization calibration parameter is the quantization calibration parameter corresponding to the second step size parameter in the first network layer.

6. The data processing method according to claim 5, characterized in that: The determining the second quantization calibration parameter according to a relationship between the calibration data corresponding to the first network layer under the constraint of the second step size parameter and the corresponding quantization calibration loss includes: Acquire a reference error control parameter, where the reference error control parameter is an error control parameter corresponding to the first quantization calibration parameter; In a case where the error control parameter is the reference error control parameter, a calibration parameter when the corresponding quantization calibration loss obtains a minimum value is determined as the second quantization calibration parameter.

7. The data processing method according to claim 3, characterized in that: The determining, when the error control parameter is a first preset value, a corresponding target calibration parameter and a corresponding target quantization performance loss includes: quantizing the calibration data based on the corresponding target calibration parameters, and inputting the quantization result into the first network layer; Obtaining a first data processing result output by the data processing model under the first step length parameter constraint; directly inputting the calibration data into the first network layer to obtain a second data processing result output by the data processing model under the first step length parameter constraint; The target quantization performance loss is determined according to a difference between the first data processing result and the second data processing result.

8. The data processing method according to claim 1, wherein: After quantizing the first basic data based on the first quantization mapping coefficient and inputting the quantization result into the first network layer for data processing to obtain first output data of the first network layer, the method further includes: using the first output data as second basic data of a second network layer, where the second network layer is a network layer in the data processing model used to process the first output data; Obtaining a quantization calibration parameter corresponding to the first step length parameter at the second network layer; Calculate, according to the second basic data and the quantization calibration parameter corresponding to the first step length parameter at the second network layer, a quantization mapping coefficient corresponding to the first step length parameter and the second network layer; Based on the first step length parameter and the quantization mapping coefficient of the second network layer, the second basic data is quantized, and the quantization result is input into the second network layer for data processing to obtain output data of the second network layer.

9. The data processing method according to claim 1 or 8, characterized in that: The method further comprises: When the target output data output by the data processing model under the constraint of the first step length parameter is obtained and the first step length parameter does not satisfy the step length parameter constraint condition, obtaining a second step length parameter according to the first step length parameter; Obtaining a second quantization calibration parameter, where the second quantization calibration parameter is a quantization calibration parameter corresponding to the second step size parameter at the first network layer; Calculating, according to the target output data and the second quantization calibration parameter, a second quantization mapping coefficient corresponding to the second step size parameter and the first network layer; Quantization is performed based on the second quantization mapping coefficient and the target basic data, and the quantization result is input into the first network layer for data processing to obtain second output data of the first network layer.

10. The data processing method according to claim 1 or 8, characterized in that: The method further comprises: When the target output data output by the data processing model under the first step length parameter constraint is obtained and the first step length parameter satisfies the step length parameter constraint condition, the target output data is used as the final data processing result of the data processing model on the first basic data.

11. A data processing device, characterized in that: include: A data acquisition module is configured to acquire a first step length parameter and first basic data, where the first basic data is representation data of text content or image content; a quantization mapping coefficient determination module, configured to execute obtaining a first quantization calibration parameter corresponding to the first step length parameter at a first network layer of a data processing model, the first quantization calibration parameter being a parameter obtained by performing quantization calibration for the first network layer based on calibration data and the first step length parameter; and, calculating, based on the first basic data and the first quantization calibration parameter, a first quantization mapping coefficient corresponding to the first step length parameter and the first network layer; The data processing module is configured to quantize the first basic data based on the first quantization mapping coefficient, and input the quantization result into the first network layer for data processing to obtain first output data of the first network layer.

12. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the data processing method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the data processing method according to any one of claims 1 to 10.

14. A computer program product, characterized in that The computer program product includes a computer program, which is stored in a readable storage medium. At least one processor of a computer device reads and executes the computer program from the readable storage medium, so that the computer device performs the data processing method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Neural network quantification method and device, electronic device and storage medium

    CN113159318A

  • Image generation method, device and equipment and computer readable storage medium

    CN117475038A