Model quantification method and device
By co-quantizing the weight parameters of the basic model, the problem of compressing the model while maintaining the model performance is solved, and efficient model quantization is achieved for edge-end devices.
Patent Information
- Application Number
- CN202411959559.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to maximize the compression of the model while maintaining the basic model feature extraction capabilities to suit edge devices with fewer computing resources.
By determining multiple accuracy to be quantized and performing joint quantization of the model weight parameters based on these accuracy, a second model with lower accuracy but less performance loss is obtained. The method includes multiple iterations of quantization processing, adjusting the accuracy and value of the weight parameters each iteration to meet the convergence conditions.
It realizes that the model is significantly compressed on the basis of maintaining model performance, thereby expanding the scope of use of the basic model and is suitable for edge devices with limited computing resources.
Smart Images

Figure CN119940427A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to, but is not limited to, the field of computer technology, and in particular to a model quantization method and device. Background Art
[0002] In recent years, deep learning models, especially basic models with large number of parameters, have shown superior performance in various fields. However, basic models have high requirements for computing resources and are not suitable for edge devices with less computing resources, which severely limits the scope of use of basic models.
[0003] Therefore, how to compress the model to the maximum extent while maintaining the feature extraction capability of the basic model to expand the scope of use of the basic model has become an urgent problem to be solved. Summary of the invention
[0004] In view of this, the present application at least provides a model quantization method and device.
[0005] The technical solution of this application is implemented as follows:
[0006] In one aspect, the present application provides a model quantization method, the method comprising:
[0007] Determine a plurality of quantized precisions corresponding to a first model, wherein a weight parameter of the first model has at least one first target precision;
[0008] Based on multiple precisions to be quantized and values of weight parameters, the weight parameters of the first model are jointly quantized to obtain a second model, and the weight parameters of the second model have at least one second target precision; the joint quantization processing represents multiple iterative quantization processing of the precision and value of the weight parameters of the first model; each iterative quantization processing includes adjusting the precision and value of the weight parameters of the first model respectively.
[0009] In some embodiments, based on the precision to be quantized and the value of the weight parameter, the weight parameter of the first model is jointly quantized to obtain the second model, including:
[0010] Based on the current iterative quantization processing, after adjusting the accuracy of the weight parameters of the first model for the current time, the accuracy of the weight parameters of the first model is kept unchanged, and the value of the weight parameters of the first model is adjusted based on the first loss until the adjusted first model meets the convergence condition to obtain the second model.
[0011] In some embodiments, the first loss is a difference between the first output information and the second output information;
[0012] The first output information is output information obtained by processing the input information based on the adjusted first model;
[0013] The second output information is output information obtained by processing the input information based on the full-precision model; or, the second output information is label information corresponding to the input information.
[0014] In some embodiments, based on multiple quantization precisions and values of weight parameters, the weight parameters of the first model are jointly quantized to obtain the second model, further comprising:
[0015] Based on the values of the weight parameters of the first model that meet the convergence condition in the previous iterative quantization process, determining a precision set from a plurality of precisions to be quantized; each precision in the precision set corresponds to a corresponding network layer in the first model;
[0016] Based on the precision set, the precision of the weight parameters of the first model in the current iterative quantization process is adjusted.
[0017] In some embodiments, based on the value of the weight parameter of the first model that satisfies the convergence condition in the previous iterative quantization process, determining the precision set from a plurality of precisions to be quantized includes:
[0018] Based on the second loss, the accuracy of the weight parameters of the corresponding network layer of the first model that meets the convergence condition in the previous iterative quantization process is adjusted until the adjusted first model meets the convergence condition, and the accuracy set of the weight parameters of the adjusted network layer is used as the accuracy set.
[0019] In some embodiments, the second loss is a difference between the third output information and the second output information;
[0020] The third output information is output information obtained by processing the input information based on the adjusted first model;
[0021] The second output information is output information obtained by processing the input information based on the full-precision model; or, the second output information is label information corresponding to the input information.
[0022] In some embodiments, before adjusting the accuracy of the weight parameters of the corresponding network layer of the first model that meets the convergence condition in the previous iterative quantization process based on the second loss, the method further includes:
[0023] Based on the weight parameters of the first model, the target inference delay is determined using the inference delay function; the inference delay function represents the relationship between the weight parameters fitted based on at least one hardware information and the inference delay; the at least one hardware information represents the information of at least one hardware used to deploy the second model; wherein the second loss includes the target inference delay.
[0024] In some embodiments, the inference delay function includes a plurality of first functions and second functions; each first function represents an inference delay function corresponding to a corresponding network layer of the first model; the second function represents an inference delay function corresponding to the first model;
[0025] Based on the weight parameters of the first model, using the inference delay function, determining the target inference delay, the method includes:
[0026] Based on each first function, determining a first inference delay corresponding to a corresponding network layer in the first model; and based on the second function, determining a second inference delay corresponding to the first model;
[0027] Based on the plurality of first inference latencies and the second inference latencies, a target inference latency is determined.
[0028] In some embodiments, before adjusting the precision of the weight parameters of the first model in the current iterative quantization process based on the precision set, the method further includes:
[0029] Acquire storage information of at least one hardware for deploying the second model;
[0030] Determining size constraint information based on the stored information and the weight parameters of the first model;
[0031] Among them, the second loss includes size constraint information.
[0032] On the other hand, the present application also provides a model quantization device, comprising:
[0033] A determination module, configured to determine a plurality of quantized precisions corresponding to a first model, wherein a weight parameter of the first model has at least one first target precision;
[0034] A quantization module is used to perform joint quantization processing on the weight parameters of a first model based on multiple precisions to be quantized and the values of weight parameters to obtain a second model, and the weight parameters of the second model have at least one second target precision; the joint quantization processing represents performing multiple iterative quantization processing on the precision and value of the weight parameters of the first model; each iterative quantization processing includes adjusting the precision and value of the weight parameters of the first model respectively.
[0035] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to illustrate the technical solution of the present application.
[0037] Figure 1A schematic diagram of the implementation process of a model quantization method provided in this application;
[0038] Figure 2 A schematic diagram of an embodiment of performing warm-up training on a full-precision model in the model quantization method provided in the present application;
[0039] Figure 3 A schematic diagram of an embodiment of performing joint quantization processing on a first model in the model quantization method provided by the present application;
[0040] Figure 4 A schematic diagram of the structure of a model quantization device provided in this application;
[0041] Figure 5 A hardware entity diagram of a computer device provided for this application. DETAILED DESCRIPTION
[0042] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions of the present application are further elaborated in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.
[0043] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0044] The terms "first / second / third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first / second / third" can be interchanged with a specific order or sequence where permitted so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing this application and are not intended to limit this application.
[0046] The weight parameters of the model are the parameters used to connect each neuron in the neural network. Through model training, the weight parameters can be continuously optimized to minimize the loss function and improve the prediction performance of the model.
[0047] The precision of the weight parameters refers to the bit width corresponding to the weight parameters, which has a significant impact on the model. Specifically, the impact of the precision of the weight parameters on the model is mainly reflected in the following aspects: First, the higher the precision of the model weight parameters, the higher the model performance, which can better fit the data in model training, thereby improving the accuracy of model reasoning; second, under the same hardware conditions, the higher the precision of the model weight parameters, the lower the training efficiency of the model; third, the higher the precision of the model weight parameters, the higher the memory occupied; fourth, the higher the precision of the model weight parameters, the higher the numerical stability of the model, which can reduce the problem of numerical instability caused by rounding errors.
[0048] It can be seen that when computing resources are sufficient, high-precision models can provide better model processing results. However, for most edge devices with limited computing resources, the running speed of high-precision models will be greatly limited or even unable to run normally.
[0049] To this end, the related art uses a quantization technology with a high compression ratio to compress the high-precision basic model to a lower precision so that the basic model can be applied to edge devices with less computing resources. However, the model quantization method in the related art still has the following problems:
[0050] First, computing resources are limited in edge scenarios, so the optimal quantization strategy usually includes extremely low-precision quantization requirements. However, the existing quantization methods that reduce the precision of weight parameters to 4 bits or less are not effective, especially the post-training quantization method. The quantized model suffers from severe performance loss in downstream tasks.
[0051] Secondly, the storage capacity of various edge devices varies greatly, and the use of unified high compression ratio quantization technology cannot meet the needs of different deployment objects in edge heterogeneous scenarios. For example, if the model size is large, the model cannot be loaded; if the model size is small, the hardware resources cannot be fully utilized and the model will suffer a large loss of accuracy. However, existing methods find it difficult to accurately balance model size and performance loss, resulting in the inability to configure the optimal quantization strategy on a specific device, resulting in a significant degradation of model performance;
[0052] Finally, hardware effects are not considered when quantizing the model, so the impact of parameter values and precision changes of model parameters on inference latency cannot be considered during the quantization process.
[0053] Based on this, the present application provides a model quantization method, which can be executed by a processor of a computer device, wherein the computer device may refer to a device with data processing capabilities such as a server, a laptop, a tablet computer, and a desktop computer.
[0054] Figure 1A schematic diagram of the implementation process of a model quantization method provided in this application, such as Figure 1 As shown, the method includes the following steps S101 to S102:
[0055] Step S101, determining a plurality of precisions to be quantized corresponding to a first model, wherein a weight parameter of the first model has at least one first target precision.
[0056] Here, the first model refers to a neural network system formed by extensive interconnection of multiple different types of neural networks.
[0057] In some embodiments, the first model is any model with artificial intelligence (AI) capabilities. For example, the first model can be a large model, where the large model refers to a machine learning model with large-scale parameters and complex computing structures. The large model can be a generative model, such as a raw text model or a raw image model.
[0058] The weight parameters of the first model have at least one first target accuracy, which means that the weight parameters of at least one network layer in the first model have a preset target accuracy.
[0059] In some embodiments, a model quantization process may be performed on any initial model so that a weight parameter of at least one network layer in the initial model has a first target accuracy.
[0060] In some embodiments, the full-precision model may be quantized according to a target precision gradient to obtain a first model.
[0061] The multiple quantization precisions refer to at least two target precisions that can be used to quantize multiple network layers in the first model when the first model is quantized.
[0062] In some embodiments, the plurality of quantization accuracies may include 16-bit floating point numbers, 8-bit floating point numbers, 4-bit floating point numbers, and 2-bit floating point numbers. Thus, when the first model is quantized, the target quantization accuracies may be determined from the plurality of quantization accuracies.
[0063] Step S102, based on the multiple precisions to be quantized and the values of the weight parameters, the weight parameters of the first model are jointly quantized to obtain a second model, and the weight parameters of the second model have at least one second target precision; the joint quantization processing represents performing multiple iterative quantization processing on the precision and value of the weight parameters of the first model; each iterative quantization processing includes adjusting the precision and value of the weight parameters of the first model respectively.
[0064] Here, after determining the first model and multiple precisions to be quantized, the precision used to quantize the weight parameters of multiple network layers in the first model is determined from the multiple precisions to be quantized, and the first model is quantized using the determined precision to obtain a second model with a second target precision.
[0065] In some embodiments, quantization is performed on each hidden layer in the first model. Here, the hidden layer refers to each layer of neural network between the input layer network and the output layer network in the first model.
[0066] The joint quantization process refers to optimizing the accuracy and value of the weight parameters of the first model respectively, so as to achieve the effect of joint quantization.
[0067] Specifically, the joint quantization process includes multiple iterative quantization processes, and each iterative quantization process includes adjusting the precision and value of the weight parameters of the first model respectively. For example, in some embodiments, in each iterative quantization process, the precision and value of the weight parameters are used as learnable objects for model training, and when the trained model meets the preset convergence condition, the iterative quantization process is terminated.
[0068] In this way, in the joint quantization process, first, the first model is used as the target model; then, in the first iterative quantization process, the accuracy and value of the weight parameters of the target model are adjusted respectively to obtain the adjusted model; thereafter, the adjusted model is updated to the target model to perform the second iterative quantization process based on the updated target model; in this way, when the number of iterations reaches a preset threshold, or the accuracy and value of the weight parameters of the target model meet the preset convergence conditions, the target model is used as the second model.
[0069] In some embodiments, in each iterative quantization process, the first model is quantized in an asynchronous quantization process. Here, the asynchronous quantization process means that each time the quantization process is performed, the multiple network layers of the first model are quantized using different quantization precisions.
[0070] In some embodiments, in each iterative quantization process, for each network layer in the first model, a corresponding quantization precision is determined from a plurality of precisions to be quantized.
[0071] In some embodiments, in each iterative quantization process, for the same type of network layer in the first model, a corresponding quantization precision is determined from a plurality of precisions to be quantized.
[0072] As can be seen from the above, in the model quantization method provided by the present application, first, multiple quantization accuracies corresponding to the first model are determined, wherein the weight parameters of the first model have at least one first target accuracy; then, based on the multiple quantization accuracies and the values of the weight parameters, the weight parameters of the first model are jointly quantized to obtain at least one second target accuracy, wherein the joint quantization process represents the execution of multiple iterative quantization processes on the accuracy and value of the weight parameters of the first model, and each iterative quantization process includes adjusting the accuracy and value of the weight parameters of the first model respectively. In this way, by jointly quantizing, that is, quantizing the first model to the second model by multiple iterative quantization processes, the model can gradually adapt to the error caused by quantization, thereby reducing the problem of reduced model generalization ability caused by truncated quantization; in addition, by adjusting the accuracy and value of the weight parameters in each iterative quantization process, the model quantization process can focus on the optimization of the model accuracy and value respectively, so that the quantized second model can more accurately express the pre-quantization model and improve the model quantization effect.
[0073] In some embodiments, based on the precision to be quantized and the value of the weight parameter, the weight parameter of the first model is jointly quantized to obtain the second model, that is, the above step S102 can be implemented as the following step S1021:
[0074] Step S1021, based on the current iterative quantization processing, after adjusting the accuracy of the weight parameters of the first model for the current time, keep the accuracy of the weight parameters of the first model unchanged, and adjust the value of the weight parameters of the first model based on the first loss until the adjusted first model meets the convergence condition to obtain the second model.
[0075] Here, in an iterative quantization process, the accuracy of the weight parameters of the first model is first optimized, and then the value of the weight parameters is optimized while keeping the accuracy of the weight parameters unchanged.
[0076] The first loss refers to the measure of the gap between the processing result obtained by the first model processing the input information and the actual result in the current iterative quantization process.
[0077] In some embodiments, in the current iterative quantization process, after adjusting the accuracy of the weight parameters of the first model, multiple iterative adjustments are performed on the values of the weight parameters of the first model based on the first loss until the adjusted first model meets the convergence condition at the current accuracy.
[0078] In some embodiments, the first loss is the difference between the first output information and the second output information; wherein the first output information is the output information obtained by processing the input information based on the adjusted first model; the second output information is the output information obtained by processing the input information based on the full-precision model; or, the second output information is the label information corresponding to the input information.
[0079] Here, the adjusted first model can be the first model after the accuracy of the weight parameters is adjusted in the current iterative quantization process; it can also be the first model in any state after the value of the weight parameters of the first model is adjusted based on the first loss, until the first model that meets the convergence conditions at the current accuracy.
[0080] In this way, the input information is input into the adjusted first model, and the first model is used to perform reasoning on the input information to obtain corresponding first output information.
[0081] When the second output information is output information obtained by processing the input information based on the full-precision model, by calculating the first loss between the first output information and the second output information, and using the first loss to adjust the first model, knowledge distillation of the full-precision model can be achieved, so that the adjusted first model can learn the behavior and performance of the full-precision model.
[0082] In the case where the second output information is the label information corresponding to the input information, by calculating the first loss between the first output information and the second output information and adjusting the first model using the first loss, the adjusted first model can learn the ability to map the input information to the corresponding label information. In some embodiments, the corresponding label information can be determined for the input information in the form of manual annotation; the input information and the corresponding label information can also be determined from an existing sample data set.
[0083] In this way, in the current iterative quantization process, the first loss is used to adjust the value of the weight parameter of the first model until the adjusted first model meets the convergence condition; thereafter, the first model that meets the convergence condition is used as the target model, and the next iterative quantization process is performed.
[0084] In some embodiments, based on the multiple precisions to be quantized and the values of the weight parameters, the weight parameters of the first model are jointly quantized to obtain the second model, that is, the above step S102 can also be implemented as the following steps S1022 to S1023:
[0085] Step S1022: determine a precision set from the multiple precisions to be quantized based on the values of the weight parameters of the first model that meet the convergence conditions in the previous iterative quantization process; each precision in the precision set corresponds to a corresponding network layer in the first model.
[0086] Here, after obtaining the first model that meets the convergence condition in the previous iterative quantization process, the first model that meets the convergence condition is used as the target model to perform the current iterative quantization process.
[0087] Based on the value of the weight parameter, the accuracy set is determined from multiple accuracy to be quantized, that is, based on the value of the weight parameter, the accuracy for quantizing each network layer in the first model is automatically searched within the range of multiple accuracy to be quantized, so as to determine the accuracy set for quantizing each network layer of the first model.
[0088] In some embodiments, the accuracy of the weight parameters in the first model can be used as a learnable parameter, and the first model can be trained using a sample data set to determine an accuracy set from multiple accuracies to be quantified.
[0089] In some embodiments, the precision set corresponding to each network layer of the first model includes multiple precisions determined from multiple precisions to be quantized, that is, in the optimized first model, the precisions of the weight parameters of at least some network layers are different, thereby realizing asynchronous quantization of each network layer in the first model.
[0090] Since the functions of the network layers in the first model are different and their impact on the model inference results is also different, the importance of the network layers in the first model is not the same. Therefore, by using the asynchronous quantization method to determine higher precision for the more important network layers and lower precision for the less important network layers, the overall weight parameter amount of the model can be reduced while maintaining the model inference accuracy.
[0091] In some embodiments, each precision in the precision set corresponds to a corresponding network layer in the first model, that is, multiple precisions in the precision set have a one-to-one correspondence with each network layer in the first model.
[0092] In some embodiments, each precision in the precision set corresponds to a network layer of a corresponding type in the first model, that is, multiple precisions in the precision set have a one-to-one correspondence with multiple types of network layers in the first model. For example, multiple precisions in the precision set correspond to an attention layer, a fully connected layer, a pooling layer, etc. in the first model.
[0093] Step S1023: Based on the precision set, adjust the precision of the weight parameters of the first model in the current iterative quantization process.
[0094] Here, in the current iterative quantization process, based on multiple precisions in the precision set, the precisions of the weight parameters of the corresponding network layers in the first model are adjusted respectively to achieve the effect of precision optimization.
[0095] In the above embodiment, based on the value of the weight parameter of the first model that meets the convergence condition in the previous iterative quantization process, a precision set is adaptively determined from multiple precisions to be quantized, and the precision set is used to adjust the precision of the weight parameters of the first model in the current iterative quantization process, thereby achieving the purpose of adaptively searching for the weight parameter precision, and then quantizing the weight parameters of the first model to the optimal precision in an asynchronous manner.
[0096] In some embodiments, the accuracy set is determined from the multiple precisons to be quantized based on the value of the weight parameter of the first model that meets the convergence condition in the previous iterative quantization process, that is, the above step S1022 can be implemented as the following step S1024:
[0097] Step S1024, adjusting the accuracy of the weight parameters of the corresponding network layer of the first model that meets the convergence condition in the previous iterative quantization process based on the second loss, until the adjusted first model meets the convergence condition, and taking the accuracy set of the weight parameters of the network layer after the adjustment as the accuracy set.
[0098] Here, the second loss refers to a measure of the gap between the processing result obtained by the first model processing the input information and the true result after adjusting the value of the weight parameter of the first model in the previous iterative quantization process.
[0099] In some embodiments, in this iterative quantization process, based on the second loss, the precision of the weight parameters of each network layer in the first model is iteratively adjusted multiple times until the adjusted first model meets the convergence condition. That is, the precision of the weight parameters of each network layer in the first model is used as a learnable parameter to perform model training, and the precision of the weight parameters is updated using the second loss during the model training process.
[0100] In some embodiments, the second loss is the difference between the third output information and the second output information; wherein the third output information is the output information obtained by processing the input information based on the adjusted first model; the second output information is the output information obtained by processing the input information based on the full-precision model; or, the second output information is the label information corresponding to the input information.
[0101] Here, the adjusted first model can be the first model after the value of the weight parameter is adjusted in the previous iterative quantization process; it can also be the first model in any state after the accuracy of the weight parameter of the first model is adjusted based on the second loss, until the first model that meets the convergence conditions is determined in the current accuracy optimization stage.
[0102] In this way, the input information is input into the adjusted first model, and the first model is used to perform reasoning on the input information to obtain corresponding third output information.
[0103] Similarly, when the second output information is output information obtained by processing the input information based on the full-precision model, by calculating the second loss between the third output information and the second output information, and using the second loss to adjust the first model, knowledge distillation of the full-precision model can be achieved, so that the adjusted first model can learn the behavior and performance of the full-precision model.
[0104] In the case where the second output information is the label information corresponding to the input information, by calculating the second loss between the third output information and the second output information, and adjusting the first model using the second loss, the adjusted first model can learn the ability to map the input information to the corresponding label information. In some embodiments, the corresponding label information can be determined for the input information in the form of manual annotation; the input information and the corresponding label information can also be determined from an existing sample data set.
[0105] In this way, in the current iterative quantization process, the second loss is used to adjust the accuracy of the weight parameters of the first model until the adjusted first model meets the convergence condition; thereafter, the value of the weight parameters of the first model is adjusted based on the first loss until the adjusted first model meets the convergence condition.
[0106] In some embodiments, before adjusting the accuracy of the weight parameters of the corresponding network layer of the first model that meets the convergence condition in the previous iterative quantization process based on the second loss, that is, the above step S1024, further includes the following step S1025:
[0107] Step S1025, based on the weight parameters of the first model, use the inference delay function to determine the target inference delay; the inference delay function represents the relationship between the weight parameters fitted based on at least one hardware information and the inference delay; the at least one hardware information represents the information of at least one hardware used to deploy the second model; wherein the second loss includes the target inference delay.
[0108] Here, at least one hardware refers to at least one hardware used to deploy the second model. For example, the at least one hardware may be an edge device, a server device, or other electronic device with data processing capability used to deploy the second model.
[0109] In some embodiments, at least one hardware may be a group of heterogeneous hardware devices for deploying the second model. Here, heterogeneous hardware devices refer to hardware devices with different processing capabilities, storage capacities, network bandwidths, and other characteristics that are mixed and used in the same system or cluster. For example, heterogeneous hardware devices composed of hardware such as a central processing unit (CPU), a graphics processing unit (GPU), and a field programmable gate array (FPGA).
[0110] At least one piece of hardware information refers to information about at least one piece of hardware used to deploy the second model. For example, the hardware information may be storage capacity information, available computing resource information, network resource information, etc. of the corresponding hardware. For example, the hardware information may be available memory information of the CPU, available video memory information of the GPU, available storage information of the FPGA, etc.
[0111] In some embodiments, given a set of hardware devices, inference is performed on the hardware devices using a model with specified weight parameter accuracy through offline sampling, thereby obtaining the relationship between the accuracy of the weight parameters and the inference delay, that is, determining the inference time function.
[0112] In some embodiments, a multilayer perceptron (MLP) network may be used to fit the inference delay function.
[0113] In this way, when determining the second loss function, based on the weight parameters of the first model corresponding to the second loss function, the target inference delay corresponding to the first model can be determined using the inference delay function, and then the target inference delay can be used as part of the second loss function.
[0114] In some embodiments, when using an MLP network to fit the inference delay function, the weight parameters of the first model corresponding to the second loss function can be used as the input of the MLP network, and the output of the MLP network can be used as the target inference delay.
[0115] In some embodiments, the inference delay function includes multiple first functions and second functions; each of the first functions represents the inference delay function corresponding to the corresponding network layer of the first model; the second function represents the inference delay function corresponding to the first model.
[0116] Here, when determining the relationship between weight parameters and inference delay through offline sampling, the inference delay corresponding to different network layers at different weight parameter accuracies can be sampled on a given hardware device, thereby obtaining the first function corresponding to the network layer; at the same time, the overall delay corresponding to the model at different weight parameter accuracies can be sampled on a given hardware device, thereby obtaining the second function corresponding to the model.
[0117] In some embodiments, the first functions correspond to the same type of network layers, or the same type of network layers correspond to the same first function.
[0118] In this way, the target inference delay is determined by using the inference delay function based on the weight parameters of the first model, that is, the above step S1025 can be implemented as the following steps S1026 to S1027:
[0119] Step S1026: Based on each of the first functions, determine the first inference delay corresponding to the corresponding network layer in the first model; and based on the second function, determine the second inference delay corresponding to the first model.
[0120] Here, based on the value of the weight parameter of each network layer in the first model and the corresponding first function, the first inference delay corresponding to each network layer is determined; at the same time, the value of the weight parameter of the first model and the second function are used to determine the second inference delay corresponding to the first model, that is, the overall inference delay of the first model.
[0121] Step S1027: determine the target inference delay based on multiple first inference delays and the second inference delays.
[0122] In some embodiments, a target inference delay is determined based on a plurality of first inference delays and their corresponding weight coefficients, a second inference delay and their corresponding weight coefficients.
[0123] In the above embodiment, the inference delay function (i.e., multiple first functions and second functions) is fitted by offline sampling the inference delay caused by different network layers under different weight parameter accuracies on a specified hardware device, and the target inference delay calculated based on the inference delay function is used as part of the second loss in the accuracy optimization stage, so that in the process of searching for the optimal accuracy, inference delay constraints can be introduced for network layers of different structures in the first model, so that the entire model quantization process can perceive the real hardware characteristics to minimize the inference delay while ensuring the model inference accuracy; in addition, using the MLP network to fit the inference delay function (including multiple first functions and second functions) has the advantage of being easy to implement.
[0124] In some embodiments, before adjusting the precision of the weight parameters of the first model in the current iterative quantization process based on the precision set, that is, before the above step S1023, the method further includes the following steps S1028 to S1029:
[0125] Step S1028, obtaining storage information of at least one hardware for deploying the second model;
[0126] Step S1029: determining size constraint information based on the storage information and the weight parameters of the first model; wherein the second loss includes the size constraint information.
[0127] As described above, the storage information of at least one hardware may be memory information, video memory information, or other storage information of at least one hardware used to deploy the second model.
[0128] In this way, based on the storage information of at least one hardware and the weight parameters of the current first model, the size constraint information of the first model is calculated, and the size constraint information is used as part of the second loss to limit the size of the quantized second model to near the storage limit of at least one hardware, so that the second model can better adapt to the hardware conditions.
[0129] In some embodiments, first, based on the weight parameters of the first model, the regularization term corresponding to the first model is determined; then, based on the storage information of the at least one hardware, the hardware constraint information is determined; finally, based on the regularization term information and the hardware constraint information, the size constraint information is determined.
[0130] In some embodiments, the regularization term corresponding to the first model is a group Lasso regularizer.
[0131] In some embodiments, before determining a plurality of quantization precisions corresponding to the first model, that is, before the above step S101, the method provided by the present application may further include the following step S103:
[0132] Step S103, quantizing the full-precision model according to the target precision gradient to obtain a first model.
[0133] Here, the full-precision model refers to a model whose weight parameters have higher precision.
[0134] In some embodiments, the full-precision model refers to a model whose weight parameters are 32-bit floating point numbers.
[0135] In some embodiments, the target accuracy gradient refers to a plurality of preset accuracy levels that decrease in a step-by-step manner.
[0136] For example, when the full-precision model is a 32-bit floating-point model, the target precision gradient may be a precision gradient composed of 16-bit floating-point numbers, 8-bit floating-point numbers, and 4-bit floating-point numbers. In this way, the full-precision model is quantized according to the preset target precision gradient, that is, the 32-bit full-precision model is quantized step by step to the first model with a weight parameter precision of 4 bits according to the precision gradients of 16-bit floating-point numbers, 8-bit floating-point numbers, and 4-bit floating-point numbers.
[0137] In some embodiments, the full-precision model may be quantized using an RTN (Round-to-Nearest) quantization algorithm.
[0138] Here, the RTN quantization algorithm refers to quantizing the value of the model weight parameter to the closest target precision based on the rounding principle, thereby realizing the approximate representation of the model weight parameter. For example, when the current model precision is 32 bits and the target precision is 16 bits, the RTN quantization algorithm can be used to round the 32-bit floating point value of the weight parameter to the closest 16-bit floating point value, thereby realizing model quantization.
[0139] In this way, when the full-precision model is quantized according to the preset target accuracy gradient, first, the RTN quantization algorithm is used to quantize the accuracy of the weight parameters of the full-precision model from full precision to the first gradient accuracy; then, while keeping the first gradient accuracy unchanged, the sample data set and the preset loss function are used to adjust the value of the weight parameters of the quantized model until the adjusted model meets the preset convergence conditions, and a model with the first gradient accuracy is obtained; thereafter, according to the above steps, the model with the first gradient accuracy is quantized step by step to the model with the last level of gradient accuracy, that is, the first model with the first target accuracy is obtained.
[0140] Next, combine Figure 2 , an embodiment of quantizing the full-precision model according to the target precision gradient to obtain a first model, that is, preheating training the full-precision model is described.
[0141] like Figure 2 As shown, the full-precision model 210 is a 32-bit floating-point model;
[0142] The first module 220 is a neural network with a weight parameter quantization function; the first module 220 includes N network layers, that is, the first network layer 221 (where W r,L1 represents the weight parameter before quantization in the first network layer), the second network layer 222 (where W r,L2 represents the weight parameter before quantization in the second network layer), ... the nth network layer 223 (where W r,Ln represents the weight parameter before quantization in the nth network layer) until the Nth network layer 224 (where W r,LNrepresents the weight parameter before quantization in the Nth network layer); wherein the N network layers represent the N hidden layers in the first module 220;
[0143] Each network layer in the first module 221 includes a pseudo-quantized node. The quantization process of the pseudo-quantized node is shown in the second network layer 222. Figure 2 As shown, in the pseudo-quantization node, first, the full-precision weight parameter W rL2 2221 quantized to 16-bit precision weight parameter W* qL2 2222; Then, the 16-bit precision weight parameter W* rL2 The precision of 2222 is restored to 32-bit floating point precision to obtain the weight parameter W qL2 2223; Afterwards, the weight parameter W is used during model training qL2 2223 processes the input data to obtain the output result of the second network layer 222; after the model converges, the 16-bit precision weight parameter W* rL2 2222 is the quantization result corresponding to 16-bit precision.
[0144] The sample data 230 is a predetermined sample data set;
[0145] The target accuracy gradient is {16-bit, 8-bit, 4-bit}.
[0146] In this way, the full-precision model 210 is quantized to a first model of 4-bit precision according to the target precision gradient, that is, the process of preheating training the full-precision model is as follows:
[0147] First, the weight parameters of each network layer in the first module 220 are updated using the weight parameters of each network layer in the full-precision model 210;
[0148] Then, the sample data 230 is input into the full-precision model 210 and the first module 220 respectively, so as to obtain the fourth output information 240 corresponding to the full-precision model 210 and the fifth output information 250 corresponding to the first module 220;
[0149] Afterwards, a third loss 260 between the fourth output information 240 and the fifth output information 250 is calculated;
[0150] Here, the formula (1) for calculating the third loss 260 is as follows:
[0151]
[0152] Among them, L distillrepresents the third loss; when a block-wise training method is adopted for each network layer in the first module 210, n represents the number of blocks; when a layer-wise training method is adopted for each network layer in the first module 210, n represents the number of layers; Yr Fourth output information 240 representing the full-precision model 210; Yq represents fifth output information 250 of the first module 220;
[0153] in, Yr The calculation formula (2) is as follows:
[0154] Yr=W r X+b r (2);
[0155] Among them, W r represents the weight parameter of the full-precision model 210; X represents the sample data input into the full-precision model 210; b r represents the network bias of the full precision model 210.
[0156] Y q The calculation formula (3) is as follows:
[0157] Yq=W q X+b q (3);
[0158] Among them, W q is the weight parameter obtained by performing precision restoration on the quantized weight parameter using the pseudo quantization node in the first module 220; X represents the sample data input into the first module 220; b q represents the network bias of the first module 220;
[0159] In the first module 220, the pseudo-quantization nodes in each network layer are used to calculate W q The formula (4) is as follows:
[0160] W q =Q fake (Θ b32~16 ,W r ) (4);
[0161] Among them, Θ b32~16 Indicates that the current quantization target is to quantize from 32-bit precision to 16-bit precision; Q fake represents the pseudo-quantization function, and Q fake =ε「W r / ε", where ε represents the scaling factor in the RTN quantization algorithm, and Among them, Θ represents the target accuracy of the current quantization.
[0162] Then, the first module 220 is updated using the third loss 260, and the quantization process is continued based on the updated first module 220 until the updated first module 220 converges at 4-bit precision, thereby obtaining a first model.
[0163] In this way, the total number of warm-up training rounds N of the first model that quantizes the full-precision model to 4-bit precision according to the target precision gradient 预热训练 It can be expressed as the following formula (5):
[0164] N 预热训练 =N b16 +N b8 +N b4 (5);
[0165] Among them, N b16 N represents the number of training rounds required to quantize from 32-bit precision to 16-bit precision; b8 N represents the number of training rounds required to quantize from 16-bit precision to 8-bit precision; b4 Indicates the number of training rounds required to quantize from 8-bit precision to 4-bit precision.
[0166] It can be seen from the above embodiment that before the first model is subjected to joint quantization processing, the full-precision model is subjected to step-by-step warm-up training according to the target precision gradient, and the precision of the weight parameters is gradually reduced from 32 bits to 16 bits, 8 bits, and 4 bits, until it converges to 4 bits of precision. In this way, the quantization loss in the model quantization process is used as a gradually increasing training noise, so that the model can gradually adapt to the error caused by quantization, avoiding the problem of reduced model generalization performance caused by truncated reduction of model precision, and solving the problem of serious performance loss caused by extremely low bit precision.
[0167] Next, combine Figure 3 , an embodiment of jointly quantizing the first model based on multiple precisions to be quantized is described.
[0168] like Figure 3 As shown, the full precision model 310 is a 32-bit precision model;
[0169] Each network layer in the first module 320 has a pseudo-quantized node, and the value of the weight parameter in the first module 320 is a learnable parameter;
[0170] The precision of the weight parameters of each network layer in the second module 330 is a learnable parameter, and the second module 320 has a precision search function, which can search for the precision that meets the model deployment requirements from the precision set to be quantized {16 bits, 8 bits, 4 bits, 2 bits};
[0171] The MLP network 340 is used to fit the inference delay function for each hidden layer of the quantized model and the overall model;
[0172] A size calculation module 350, a functional module for calculating size constraint information for the second module 330 based on the weight parameter of the second module 330 and the storage information of the hardware 360;
[0173] Hardware 360 is a plurality of heterogeneous hardware devices used to deploy the quantized second model.
[0174] Next, combine Figure 3 , through the following steps S301 to S305, an iterative quantization process in the joint quantization process is explained.
[0175] Step S301, updating the second module 330 using the weight parameters of each network layer in the first module 320 that have converged in the last iterative quantization process;
[0176] like Figure 3 As shown, in the last iterative quantization, the weight parameters of each network layer in the first module 320 after convergence are changed from W r,Ln Quantize to W * q,Ln ,like Figure 3 As shown in the second network layer in , the pseudo-quantization node is used to change the value of the weight parameter from W r,L2 Quantize to W * q,L2 In this way, the weight parameter value W of each network layer in the first module 320 after convergence is used * q,Ln Update each network layer in the second module 330 to obtain the values of the weight parameters of the first network layer to the Nth network layer in the second module 330, which are W * q,L1 , W * q,L2 …W * q,Ln1 , W * q,LN .
[0177] Here, when the current iterative quantization process is the first iterative quantization process in the joint quantization process, the weight parameters of each network layer in the first model with 4-bit precision obtained after preheating training of the full-precision module 310 are used to update the first module 320 and the second module 330;
[0178] Step S302, input the sample data 370 into the full-precision model 310 and the second module 330 respectively, obtain the fourth output information 311 of the full-precision model 310 and the sixth output information 331 of the second module 330; calculate the fourth loss L between the fourth output information 311 and the sixth output information 331 4; Input the weight parameters of the second module 330 into the MLP network 340 to calculate the target reasoning delay information L 延时 The weight parameters of the second module 330 and the storage information of the hardware 360 are input into the size calculation module 350 to calculate the size constraint information L 尺寸 ; Based on the fourth loss L 4 , target reasoning delay information L 延时 and size constraint information L 尺寸 The fifth loss was determined to be 380;
[0179] Here, the fifth loss 380 can be determined using the following formula (6):
[0180] L 5 =L 4 +αL 尺寸 +L 延时 (6);
[0181] Among them, L 5 represents the value of the fifth loss 380; α represents the size constraint information L 尺寸 The penalty coefficient is α∈[0,1].
[0182] Among them, the fourth loss L 4 The error between the fourth output information 311 and the sixth output information 331 is calculated by using the mean square error (MSE) loss row number;
[0183] Target reasoning delay information L 延时 It can be determined based on the following formula (7):
[0184]
[0185] in, T represents the inference delay corresponding to the i-th layer network in the second module 330; ε (M) represents the overall reasoning delay corresponding to the second module 330; λ 1 and λ 2 Respectively and T ε (M) corresponding penalty coefficient, and λ 1 and λ 2 ∈[0,1];
[0186] Among them, the inference delay corresponding to the i-th layer is It can be determined using the following formula (8):
[0187]
[0188] in, represents the value of the weight parameter of the i-th layer network in the second module 330; and They represent the weight parameters and bias of the MLP network that fits the inference delay of the i-th layer network.
[0189] The overall inference delay T corresponding to the second module 330 ε (M) can be determined using the following formula (9):
[0190]
[0191] Among them, W q Represents the values of weight parameters of all network layers in the second module 330; and They respectively represent the weight parameters and bias of the MLP network that fits the overall inference delay of the second module 330.
[0192] Dimensional constraint information L 尺寸 It can be determined using the following formula (10):
[0193]
[0194] Wherein, M represents the Lasso regularization term information corresponding to the second module 330; ε 存储 represents the storage usage limit available for deploying the model on Hardware 360; δ represents the relaxation factor used to limit the model size to ε 存储 Nearby, when δ is 0, L 尺寸 Occupancy limits will be strictly enforced 存储 .
[0195] Among them, the regularization term information M can be determined using the following formula (11):
[0196]
[0197] Wherein, all weight parameters k on the i-th layer of the second module 330 are divided into j groups for updating, Represents the value of the weight parameter of a group on the i-th layer; Indicates the precision corresponding to the weight parameter.
[0198] In this way, the expectation of the regularization term information M It can be determined using the following formula (12):
[0199]
[0200] Step S303, inputting the fifth loss 380 into the second module 330, so as to update the accuracy of each network layer in the second module 330 based on the fifth loss 380, until the second module 330 meets the convergence condition;
[0201] Here, the second module 330 adaptively determines the corresponding target accuracy from the set of accuracy to be quantized for each network layer based on the fifth loss 380, and applies the target accuracy to the corresponding network layer.
[0202] Step S304, for the second module 330 that meets the convergence condition, the accuracy of each network layer in the second module 330 is determined as the accuracy in the accuracy set, and quantization processing is performed on each network layer in the first module 320 based on the accuracy set to obtain the quantized first module 320;
[0203] Step S305, input the sample data 370 into the full-precision model 310 and the quantized first module 320 respectively, and obtain the fourth output information 311 corresponding to the full-precision model 310 and the seventh output information 321 corresponding to the first module 320; based on the fourth output information 311 and the first module 320, calculate the sixth loss 390 using the above formula (1); based on the sixth loss 390, update the first module 320 until the first module 320 meets the convergence condition;
[0204] Step S305: for the first module 320 that meets the convergence condition in this iterative quantization process, the weight parameters of each network layer in the second module are updated using the weight parameters of each network layer in the first module 320 to perform the next iterative quantization process; at the same time, when the number of iterations meets the preset number threshold or meets other convergence conditions, the iterative quantization is stopped, and the final quantized model, that is, the second model, is derived based on the following formula (13):
[0205] Θ * =min Θ L 5 (W q * ,Θ) (13);
[0206] Among them, W q * represents the value of the weight parameter corresponding to the first module 320 after convergence; Θ represents the accuracy of the weight parameter corresponding to the first module 320 after convergence; Θ * Represents the optimal accuracy hyperparameter determined from multiple sets of weight parameter values and accuracies.
[0207] It can be seen from the above embodiments that the model quantization method provided by the present application, when jointly quantizing the first model, adjusts the memory share of each network layer of the model by the accuracy of the weight parameters of each network layer, and uses the output information of the full-precision model processing sample output and the storage limitation information of the given hardware to constrain the quantization process of the model, thereby realizing the optimal quantization strategy on the edge heterogeneous device; at the same time, the second module can adaptively search for the optimal precision parameters based on the precision set to be quantized, so that the model can be quantized to the optimal precision in an asynchronous manner, which can effectively suppress the significant attenuation of model performance on the basis of maximizing the use of hardware resources; in addition, by introducing inference delay constraint information, the entire quantization process can perceive the real hardware special effects, while ensuring the inference accuracy of the model and minimizing the inference delay.
[0208] Based on the foregoing embodiments, the present application provides a model quantization device, which includes the units included and the modules included in the units, which can be implemented by a processor in a computer device; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.
[0209] Figure 4 A schematic diagram of the structure of a model quantization device provided in this application, such as Figure 4 As shown, the model quantization device 400 includes: a determination module 410 and a quantization module 420, wherein:
[0210] A determination module 410 is used to determine a plurality of quantization precisions corresponding to a first model, wherein a weight parameter of the first model has at least one first target precision;
[0211] The quantization module 420 is used to perform a joint quantization process on the weight parameters of the first model based on the multiple precisions to be quantized and the values of the weight parameters to obtain a second model, wherein the weight parameters of the second model have at least one second target precision; the joint quantization process represents multiple iterative quantization processes performed on the precision and value of the weight parameters of the first model; each iterative quantization process includes adjusting the precision and value of the weight parameters of the first model respectively.
[0212] In some embodiments, the quantization module 420 includes a first quantization module 430;
[0213] The first quantization module 430 is used to maintain the accuracy of the weight parameters of the first model unchanged after adjusting the accuracy of the weight parameters of the first model for the current iterative quantization processing, and adjust the value of the weight parameters of the first model based on the first loss until the adjusted first model meets the convergence condition to obtain the second model.
[0214] In some embodiments, the first loss is the difference between the first output information and the second output information;
[0215] The first output information is output information obtained by processing input information based on the adjusted first model;
[0216] The second output information is output information obtained by processing the input information based on a full-precision model; or, the second output information is label information corresponding to the input information.
[0217] In some embodiments, the quantization module 420 includes a second quantization module 440; the second quantization module 440 is configured to:
[0218] Based on the values of the weight parameters of the first model that meet the convergence condition in the previous iterative quantization process, determine a precision set from the multiple precisions to be quantized; each precision in the precision set corresponds to a corresponding network layer in the first model;
[0219] Based on the precision set, the precision of the weight parameters of the first model in the current iterative quantization process is adjusted.
[0220] In some embodiments, the second quantization module 440 is used to adjust the accuracy of the weight parameters of the corresponding network layer of the first model that meets the convergence condition in the previous iterative quantization process based on the second loss, until the adjusted first model meets the convergence condition, and the accuracy set of the adjusted weight parameters of the network layer is used as the accuracy set.
[0221] In some embodiments, the second loss is the difference between the third output information and the second output information;
[0222] The third output information is output information obtained by processing the input information based on the adjusted first model;
[0223] The second output information is output information obtained by processing the input information based on a full-precision model; or, the second output information is label information corresponding to the input information.
[0224] In some embodiments, the apparatus 400 further includes a delay calculation module 450;
[0225] The delay calculation module 450 is used to determine the target inference delay based on the weight parameters of the first model using the inference delay function; the inference delay function represents the relationship between the weight parameters fitted based on at least one hardware information and the inference delay; the at least one hardware information represents the information of at least one hardware used to deploy the second model; wherein the second loss includes the target inference delay.
[0226] In some embodiments, the inference delay function includes a plurality of first functions and second functions; each of the first functions represents an inference delay function corresponding to a corresponding network layer of the first model; the second function represents an inference delay function corresponding to the first model;
[0227] The delay calculation module 450 is used to:
[0228] Based on each of the first functions, determining a first inference delay corresponding to a corresponding network layer in the first model; and based on the second function, determining a second inference delay corresponding to the first model;
[0229] The target inference latency is determined based on a plurality of the first inference latency and the second inference latency.
[0230] In some embodiments, the apparatus 400 further includes a size calculation module 460; the size calculation module 460 is configured to:
[0231] Acquire storage information of at least one hardware for deploying the second model;
[0232] Determining size constraint information based on the stored information and a weight parameter of the first model;
[0233] Wherein, the second loss includes the size constraint information.
[0234] The description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.
[0235] It should be noted that in the embodiment of the present application, if the above-mentioned model quantization method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific hardware, software or firmware, or any combination of hardware, software, and firmware.
[0236] An embodiment of the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.
[0237] The embodiment of the present application provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, some or all of the steps in the above method are implemented. The computer-readable storage medium can be transient or non-transient.
[0238] An embodiment of the present application provides a computer program, including a computer-readable code. When the computer-readable code is run in a computer device, a processor in the computer device executes some or all of the steps for implementing the above method.
[0239] The embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be implemented specifically by hardware, software or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium, and in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK) and the like.
[0240] It should be noted here that the description of the various embodiments above tends to emphasize the differences between the various embodiments, and the same or similar aspects can be referenced to each other. The description of the above device, storage medium, computer program and computer program product embodiments is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the embodiments of the device, storage medium, computer program and computer program product of this application, please refer to the description of the method embodiment of this application for understanding.
[0241] It should be noted that Figure 5 A schematic diagram of a hardware entity of an electronic device in this application, such as Figure 5 As shown, the hardware entity of the electronic device 500 includes: a processor 501, a communication interface 502 and a memory 503, wherein:
[0242] The processor 501 generally controls the overall operation of the electronic device 500 .
[0243] The communication interface 502 enables the electronic device to communicate with other terminals or servers through a network.
[0244] The memory 503 is configured to store instructions and applications executable by the processor 501, and can also cache data to be processed or processed by the processor 501 and each module in the electronic device 500 (for example, image data, audio data, voice communication data, and video communication data), which can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM). Data transmission between the processor 501, the communication interface 502 and the memory 503 can be carried out through the bus 504.
[0245] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the serial number of each step / process mentioned above does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. The serial numbers of the embodiments of the present application mentioned above are for description only and do not represent the advantages and disadvantages of the embodiments.
[0246] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0247] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0248] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0249] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0250] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, etc., various media that can store program codes.
[0251] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0252] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.
Claims
1. A model quantization method, the method comprising: Determine a plurality of quantized precisions corresponding to a first model, wherein a weight parameter of the first model has at least one first target precision; Based on the multiple precisions to be quantized and the values of the weight parameters, jointly quantize the weight parameters of the first model to obtain a second model, wherein the weight parameters of the second model have at least one second target precision; The joint quantization process represents performing multiple iterative quantization processes on the accuracy and value of the weight parameters of the first model; Each iteration of the quantization process includes adjusting the precision and value of the weight parameters of the first model respectively.
2. The method according to claim 1, wherein the weight parameters of the first model are jointly quantized based on the precision to be quantized and the value of the weight parameter to obtain the second model, comprising: Based on the current iterative quantization processing, after adjusting the accuracy of the weight parameters of the first model for the current time, the accuracy of the weight parameters of the first model is kept unchanged, and the value of the weight parameters of the first model is adjusted based on the first loss until the adjusted first model meets the convergence condition to obtain the second model.
3. The method according to claim 2, wherein: The first loss is the difference between the first output information and the second output information; The first output information is output information obtained by processing input information based on the adjusted first model; The second output information is output information obtained by processing the input information based on a full-precision model; or, the second output information is label information corresponding to the input information.
4. The method according to claim 2, wherein the weight parameters of the first model are jointly quantized based on the multiple quantization precisions and the values of the weight parameters to obtain the second model, further comprising: Determine a precision set from the multiple precisions to be quantized based on the values of the weight parameters of the first model that meet the convergence condition in the previous iterative quantization process; Each precision in the precision set corresponds to a corresponding network layer in the first model; Based on the precision set, the precision of the weight parameters of the first model in the current iterative quantization process is adjusted.
5. The method according to claim 4, wherein the step of determining the precision set from the plurality of precisions to be quantized based on the value of the weight parameter of the first model that satisfies the convergence condition in the previous iterative quantization process comprises: Based on the second loss, the accuracy of the weight parameters of the corresponding network layer of the first model that meets the convergence condition in the previous iterative quantization process is adjusted until the adjusted first model meets the convergence condition, and the accuracy set of the adjusted weight parameters of the network layer is used as the accuracy set.
6. The method according to claim 5, wherein the second loss is a difference between the third output information and the second output information; The third output information is output information obtained by processing the input information based on the adjusted first model; The second output information is output information obtained by processing the input information based on a full-precision model; or, the second output information is label information corresponding to the input information.
7. The method according to claim 5 or 6, before adjusting the accuracy of the weight parameters of the corresponding network layer of the first model that meets the convergence condition in the previous iterative quantization process based on the second loss, further comprising: Determine a target inference delay using an inference delay function based on the weight parameters of the first model; The inference delay function represents the relationship between the weight parameter fitted based on at least one hardware information and the inference delay; The at least one hardware information represents information of at least one hardware used to deploy the second model; wherein the second loss includes the target inference latency.
8. According to the method of claim 7, the inference delay function comprises a plurality of first functions and second functions; each of the first functions represents the inference delay function corresponding to the corresponding network layer of the first model; The second function represents the inference delay function corresponding to the first model; The determining the target inference delay based on the weight parameter of the first model and using the inference delay function includes: Based on each of the first functions, determining a first inference delay corresponding to a corresponding network layer in the first model; and based on the second function, determining a second inference delay corresponding to the first model; The target inference latency is determined based on a plurality of the first inference latency and the second inference latency.
9. The method according to claim 4, before adjusting the precision of the weight parameters of the first model in the current iterative quantization process based on the precision set, further comprising: Acquire storage information of at least one hardware for deploying the second model; Determining size constraint information based on the stored information and a weight parameter of the first model; Wherein, the second loss includes the size constraint information.
10. A model quantization device, comprising: A determination module, configured to determine a plurality of quantized precisions corresponding to a first model, wherein a weight parameter of the first model has at least one first target precision; A quantization module, configured to perform a joint quantization process on the weight parameters of the first model based on the multiple precisions to be quantized and the values of the weight parameters to obtain a second model, wherein the weight parameters of the second model have at least one second target precision; The joint quantization process represents performing multiple iterative quantization processes on the accuracy and value of the weight parameters of the first model; Each iteration of the quantization process includes adjusting the precision and value of the weight parameters of the first model respectively.