Model quantification method and computing device

By quantifying the shared weight parameters of the pre-trained model in advance and combining the service-specific weight parameter training, the problem of imbalance between services in the process of quantization of neural network models is solved, efficient model updates and deployment are achieved, data transmission is reduced, and the quantitative effect of each service is improved.

CN120471116APending Publication Date: 2025-08-12HONOR DEVICE CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202411398783.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

When dealing with multiple services, it is difficult for the prior art to balance the quantitative effects of each service in the quantization process of neural network models. As a result, the good quantitative effect of one service may affect the quantitative effect of another service, and a large amount of data is required to be transmitted during the deployment of the model, which is very costly to train.

Method used

By quantifying the shared weight parameters of the pre-trained model in advance, and keeping them unchanged during subsequent training, and quantization perception training is carried out in combination with the business-specific weight parameters, the respective target business models are obtained, and the quantization and training process of each business are decoupled.

Benefits of technology

When business changes or new services are added, reduce the transmission amount of model update data, improve model update efficiency, avoid negative impacts on quantification and training between each service, and improve the quantitative effect of each service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471116A_ABST
    Figure CN120471116A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a model quantification method and computing equipment, and the method comprises the steps: training equipment quantifies a pre-training model, and obtains a pre-training quantification model; the training device builds a first original business model based on the pre-training quantification model; the training device performs quantitative perception training on the first original business model based on the N pieces of business data to obtain N target business models; the t-th target service model in the N target service models comprises a shared quantization weight parameter and a t-th service quantization weight parameter in the N target service models; one business corresponds to one target business model; t is an integer from 1 to N. According to the embodiment of the invention, the quantification effect of each service can be improved, and the deployment process of training during model change is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of terminal technology, and in particular to a model quantization method and computing device. Background Art

[0002] When a neural network model needs to handle multiple businesses, it is generally necessary to use dedicated neural network models for each business to ensure that different business requirements can be met. For multiple similar businesses, the neural network models for these businesses have similar feature extraction capabilities. Using different neural network models to execute related businesses requires the device to store a large amount of neural network model data, and a large amount of model data needs to be transmitted during model deployment, resulting in high training costs. Pre-training and fine-tuning methods can be used to train models for different businesses. Pre-training can pre-select models on large amounts of data to learn common features, and fine-tune the pre-trained model on the dataset for the feature task to optimize performance and adapt to the characteristics of each business. The above process can greatly reduce the cost of training and model deployment.

[0003] However, model deployment requires quantification processing. During the model quantification process, it is difficult to balance the quantification effects of various businesses. A good quantification effect for one business may lead to a poor quantification effect for another business. Summary of the Invention

[0004] The embodiments of the present application disclose a model quantization method and computing device, which can improve the quantization effect of various businesses and optimize the training deployment process when the model is changed.

[0005] In the first aspect, the present application provides a model quantization method, which is applied to a training device, and the method includes: the training device quantizes a pre-trained model to obtain a pre-trained quantized model; the pre-trained quantized model includes a shared quantization weight parameter; the training device builds a first original business model based on the pre-trained quantized model; the first original business model includes the shared quantization weight parameter and the initial business weight parameter; the training device quantizes the perception training of the first original business model based on N business data to obtain N target business models; the tth target business model in the N target business models includes the shared quantization weight parameter and the tth business quantization weight parameter in the N target business models; one business corresponds to one target business model; and t is an integer from 1 to N in sequence.

[0006] In the embodiment of the present application, since the pre-trained shared weight parameters are quantized in advance, the quantized shared weight parameters are kept unchanged during the subsequent training and quantization of the weights of each business, which can decouple the mutual influence of the parameter training and quantization processes of each business. In this way, when a certain business changes or a new business is added, it does not affect the model parameters corresponding to the business that has been trained and quantized. Therefore, in the case of subsequent business changes, the transmission of model update data can be reduced, the amount of data downloaded by the device is greatly reduced, and the efficiency of model update is improved. In addition, after the above-mentioned training and quantization process of each business, the negative impact of each other's quantization and training is avoided, and the quantization effect of each business is improved.

[0007] In a possible implementation, the training device quantizes the perception training of the first original business model based on N business data to obtain N target business models, including: the training device trains the initial business weight parameters of the first original business model based on N business data to obtain N business network models; the i-th business network model in the N business network models includes the shared quantization weight parameter and the i-th business weight parameter in the N business weight parameters; one business corresponds to one business network model; N is an integer greater than or equal to 2; i is an integer from 1 to N in sequence; the training device quantizes the N business weight parameters in the N business network models according to the quantization data of N businesses to obtain N target business models; the t-th target business model in the N target business models includes the shared quantization weight parameter and the t-th business quantization weight parameter in the N target business models; one business corresponds to one target business model; t is an integer from 1 to N in sequence. In this way, the pre-trained shared weight parameters can be quantized in advance, and in the subsequent training and quantization process of each business's weights, the quantized shared weight parameters are kept unchanged, which can decouple the mutual influence of each business's parameter training and quantization process. In this way, when a certain business changes or a new business is added, it will not affect the model parameters corresponding to the business that has been trained and quantized. Therefore, in the case of subsequent business changes, the transmission of model update data can be reduced, the amount of data downloaded by the device is greatly reduced, and the efficiency of model updates is improved. In addition, after the above-mentioned training and quantization process of each business, the negative impact of each business on the quantization and training is avoided, and the quantization effect of each business is improved.

[0008] The service weight parameter is a trained but unquantized weight parameter, and the service quantized weight parameter is a trained and quantized weight parameter.

[0009] In one possible implementation, the training device performs quantitative perceptual training on the initial service weight parameters of the first original service model based on N service data to obtain N service network models. This includes: the training device performs quantitative perceptual training on the initial service weight parameters of the first original service model based on the N service data to obtain N target service models. This simplifies the training process and reduces training costs.

[0010] The quantization-aware training includes inserting a pseudo-quantization node into the first original business model, then determining the first quantization parameter, inputting the business data of the first business, adjusting the quantization parameter according to the output result, and determining the target quantization parameter. For details, please refer to Figure 11 The relevant content in will not be repeated here.

[0011] In a possible implementation manner, the quantitative data is business data.

[0012] In a possible implementation, after obtaining N target business models, the method further includes: the training device trains the first original business model based on the newly added business data to obtain the N+1th business network model; the N+1th business network model includes the shared quantization weight parameter and the N+1th business weight parameter; the training device quantizes the N+1th business weight parameter in the N+1th business network model according to the quantization data of the newly added business to obtain the N+1th target business model; the N+1th target business model includes the shared quantization weight parameter and the N+1th business quantization weight parameter; the training device sends the first model update data to the using device, and the first model update data includes the N+1th business quantization weight parameter. In this way, when adding a new business, the training device can directly train and quantize the initial business weight parameters of the new business, and continue to use the existing shared quantization weight parameters, so as not to affect the parameters of other original business models, thereby improving the training and quantization efficiency and reducing the amount of data required to download and the installation time during the deployment process.

[0013] In one possible implementation, after obtaining N target business models, the method further includes: the training device trains the first original business model based on the business data of the Mth business to obtain the Mth business network model; the Mth business network model includes the shared quantization weight parameter and the Mth business weight parameter; the Mth business is any one of the Nth businesses; the training device quantizes the Mth business weight parameter in the Mth business network model according to the quantization data of the Mth business to obtain the Mth target business model; the Mth target business model includes the shared quantization weight parameter and the Mth business quantization weight parameter; the training device sends second model update data to the using device, and the second model update data includes the Mth business quantization weight parameter. In this way, when adjusting an existing business, the training device can directly train and quantize the initial business weight parameters of the business, or continue training using the existing shared quantization weight parameters, so as not to affect the parameters of other original business models, thereby improving training and quantization efficiency and reducing the amount of data required to download and the installation time during deployment.

[0014] In one possible implementation, the N business data are selected from the group consisting of text data, voice data, video data, and image data. Different neural network types can process different types of data, and the N business data need to be unified into one type of data to ensure the feasibility and effectiveness of the model.

[0015] In one possible implementation, the training device quantizes N service weight parameters in the N service network models according to the quantized data of N services to obtain N target service models, including: the training device quantizes the tth service weight parameter in the tth service network model according to the quantized parameters of the tth service to obtain the tth target service model. In this way, the data of each service quantizes its own service weight parameter, decoupling the quantization process of each service and reducing quantization interference.

[0016] In a possible implementation, the training device quantizes N service weight parameters in the N service network models according to the quantization data of N services respectively to obtain N target service models, including: the training device determines a first quantization parameter of the t-th service; the training device executes a first quantization target calculation process based on the first quantization parameter; the first quantization target calculation process includes: the training device quantizes the t-th service network model according to the first quantization parameter to obtain a quantized t-th quantized service model; the training device inputs the quantization data of the t-th service into the t-th service network model to obtain an unquantized output result; the quantization data of the t-th service is input into the t-th service network model to obtain an unquantized output result; According to the input of the t-th quantized business model, a quantized output result is obtained, and a first target value is calculated based on the unquantized output result and the quantized output result; after the training device executes the first quantized target calculation process, it determines whether quantization adjustment is required; if quantization adjustment is required, the training device adjusts the first quantization parameter of the t-th business and re-executes the first quantization target calculation process based on the adjusted first quantization parameter; if quantization adjustment is not required, the first quantization parameter corresponding to the minimum value of the first target value is determined as the t-th target quantization parameter, and the t-th business network model is quantized according to the t-th target quantization parameter to obtain the quantized t-th target business model. In this way, during the quantization process, the training device can adjust the quantization parameter to ensure that the quantization loss is smaller and the quantization effect is improved.

[0017] In one possible implementation, determining whether quantization adjustment is required includes: determining whether the first quantization loss calculation process has been executed based on all preset quantization parameters, and if the first quantization loss calculation process has been executed for all preset quantization parameters, determining that quantization adjustment is not required; if there is a preset quantization parameter for which the first quantization loss calculation process has not been executed, determining that quantization adjustment is required; or determining whether the loss of the first quantization parameter has converged based on a gradient descent method, and if the loss of the first quantization parameter has converged, determining that quantization adjustment is not required; if the loss of the first quantization parameter has not converged, determining that quantization adjustment is required. In this way, the quantization parameter can be adjusted to reduce the quantization loss.

[0018] In one possible implementation, the first quantization parameter includes one or more of an integer quantization scale, a zero point zero, an upper limit UperTh, and a lower limit LowTh. The training device quantizes the t-th service network model according to the first quantization parameter to obtain the quantized t-th quantized service model, including: the training device calculates the corresponding weight element Q in the t-th quantized service model based on the weight element x of the t-th service network model, the first quantization parameter, and the quantization mapping formula.

[0019] In one possible implementation, the training device builds a first original business model based on the pre-trained quantization model, including: according to the LoRA fine-tuning method, taking a single weight matrix in the pre-trained quantization model as a basic weight W, adding two low-rank matrices A and B, and obtaining the first original business model; or, according to the LoRA fine-tuning method, taking a plurality of consecutive neurons in the pre-trained quantization model as a basic weight W, adding two low-rank matrices A and B, and obtaining the first original business model; or, according to the adapter fine-tuning method, taking the entire hidden layer of the pre-trained quantization model as the mainboard model W, adding an adapter component, and obtaining the first original business model. In this way, for different fine-tuning methods, the training device can effectively reduce the number of updates to the weights in the model, reduce the impact of services on each other, and reduce computing costs and storage costs.

[0020] In a possible implementation, before the training device quantizes the pre-trained model to obtain the pre-trained quantized model, the method further includes: the training device pre-trains the original network model to obtain the pre-trained model.

[0021] In a second aspect, the present application provides a model quantization method, which is applied to a training device and includes:

[0022] The training device builds a second original business model based on the pre-trained quantization model; the second original business model includes an initial shared weight parameter and an initial business weight parameter; the training device trains the second original business model based on N business data to obtain N business network models; the i-th business network model in the N business network models includes the shared weight parameter and the i-th business weight parameter in the N business weight parameters; one business corresponds to one business network model; N is an integer greater than or equal to 2; i is an integer from 1 to N in sequence; the training device selects a target quantization parameter based on the target value of all businesses in the N businesses to determine N target business models; the target quantization parameter is used to quantize the shared weight parameter and business weight parameter in the N business network models, and the i-th target business model in the N target business models includes the shared quantization weight parameter and the t-th business quantization weight parameter; t is an integer from 1 to N in sequence.

[0023] In the embodiment of the present application, through the training and quantization process of each of the above services, the negative impact of quantization and training on each other among the services is avoided, and the quantization effect of each service can be improved.

[0024] In one possible implementation, the training device selects a target quantization parameter based on the loss values of all services among N services and determines N target service models, including: the training device determines a first quantization parameter; the training device executes a target calculation process based on the first quantization parameter; the target calculation process includes: the training device quantizes the N service network models according to the first quantization parameter to obtain N quantized service models; the training device inputs the quantized data of the N services into the N service network models respectively, and obtains N unquantized output results accordingly; the training device inputs the quantized data of the N services into the N quantized service models respectively, and obtains N quantized output results accordingly, and calculates target values based on the N unquantized output results and the N quantized output results; after the training device executes the target calculation process, it determines whether quantization adjustment is required; if quantization adjustment is required, the training device adjusts the first quantization parameter and re-executes the target calculation process based on the adjusted first quantization parameter; if quantization adjustment is not required, the first quantization parameter corresponding to the minimum value of the target value is determined as the target quantization parameter, and the N service network models are quantized according to the target quantization parameter to obtain N target service models. In this way, during the quantization process, the training device can adjust the quantization parameters to ensure smaller quantization loss and improve the quantization effect.

[0025] In a third aspect, the present application provides a computing device comprising: one or more processors and one or more memories; the one or more processors are coupled to the one or more memories, the one or more memories being used to store computer program code, the computer program code comprising computer instructions, and when the one or more processors execute the computer instructions, the computing device executes a model quantization method as in any possible implementation of the first aspect or the second aspect.

[0026] In a fourth aspect, the present application provides a computing device, comprising: one or more functional modules. The one or more functional modules are configured to execute the model quantization method in any possible implementation of the first aspect or the second aspect.

[0027] In a fifth aspect, an embodiment of the present application provides a computer storage medium comprising computer instructions, which, when executed on a computing device, enables a communication device to execute a model quantization method in any possible implementation of the first or second aspect.

[0028] In a sixth aspect, an embodiment of the present application provides a computer program product, which, when running on a computer, enables the computer to execute the model quantization method in any possible implementation of the first aspect or the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 This is a schematic diagram of an architecture for network model training and use provided by an embodiment of the present application;

[0030] Figure 2A This is a schematic diagram of the LoRA model structure provided in an embodiment of the present application;

[0031] Figure 2B This is a schematic diagram of the Adaptor model structure provided in an embodiment of the present application;

[0032] Figure 3 This is a schematic diagram of a quantization mapping provided in an embodiment of the present application;

[0033] Figure 4 This is a schematic diagram of a training model deployment process provided by an embodiment of the present application;

[0034] Figure 5 This is a flow chart of a model quantization method provided in an embodiment of the present application;

[0035] Figure 6A This is a schematic diagram of the original business model of LoRA provided in an embodiment of the present application;

[0036] Figure 6B This is a schematic diagram of another LoRA original business model provided in an embodiment of the present application;

[0037] Figure 6C This is a schematic diagram of an original business model of an adapter provided in an embodiment of the present application;

[0038] Figure 7 This is a schematic diagram of weight matrix segmentation in a neural network model provided in an embodiment of the present application;

[0039] Figure 8 This is a schematic diagram of a multi-task model training process provided by an embodiment of the present application;

[0040] Figure 9 This is a flow chart of another model quantization method provided in an embodiment of the present application;

[0041] Figure 10 This is a schematic diagram of another multi-task model training process provided in an embodiment of the present application;

[0042] Figure 11 This is another flow chart of a model quantization method provided in an embodiment of the present application;

[0043] Figure 12 This is another schematic diagram of the processing process of multi-task model training provided in an embodiment of the present application. DETAILED DESCRIPTION

[0044] The embodiments of the present application relate to a model quantization method and computing device, which are used to reduce the user's selection operations for picture size and save processing resources and energy consumption.

[0045] In order to facilitate understanding of the technical solutions in the embodiments of the present application, the relevant concepts involved in the embodiments of the present application are first introduced below.

[0046] 1. Multi-task neural network model:

[0047] Fields such as artificial intelligence, computer vision (CV), and natural language processing (NLP) typically involve techniques such as image processing and language processing. These processes are typically implemented using neural networks. Image processing can include tasks such as image classification, image recognition, semantic segmentation, image generation, and object detection. NLP is widely used in fields such as speech recognition, text generation, information extraction, text classification, and information recommendation.

[0048] Within the processing domain of the same type of neural network model, the same or similar feature extraction capabilities of the neural network model can be used for different processing services, providing the prerequisite for a single neural network model to handle multiple similar services. Using a separate neural network model for each processing service consumes a significant amount of device storage space, requires significant training and computational effort, and information cannot be shared between models. This creates a real need to use neural network models with smaller data volumes to handle more tasks (different services of the same type). Therefore, a single neural network model can be used to complete multiple different services, integrating features from different tasks to improve model performance and generalization capabilities. For example, the target task of recognizing user speech, i.e., extracting text information; and the task of communicating with the user, i.e., generating the next paragraph of text information, are all accomplished through the same neural network model.

[0049] For example, Honor phones, through features like Smart Video Editing and the YOYO Voice Assistant, require neural network models to provide different functional services. In the Smart Video Editing service, the phone receives a user's voice message, generates a video editing task based on this voice message, and then edits the video according to the task instructions. The user inputs "Edit the kitten in the video into a one-minute video." The electronic device extracts information from the audio: "Video object: kitten; Video duration: 1 minute." After analyzing and extracting the text information, it obtains the editing task instructions. The YOYO Voice Assistant can engage in conversation or respond based on user input. For example, if a user inputs "Hello, YOYO," the YOYO Voice Assistant responds: "Hi, hello." The YOYO Voice Assistant can also generate task instructions based on the user's voice message and switch between the first and second tasks based on the task instructions. For example, if the user inputs "Open camera," the YOYO Voice Assistant generates a command to launch the camera, launches the camera application, and the electronic device displays the camera application interface. The processing and analysis of text information in these multiple services can all be accomplished using a single neural network model.

[0050] It should be noted that the business data input into the network model in this application can be one of the types such as text, image, audio, video, etc. This application does not specifically limit the type of input data.

[0051] 2. Training architecture of neural network model:

[0052] In the actual use of multi-task neural network models, model training and execution are not performed on the same device. The following further explains the architecture and devices related to training and executing neural network models.

[0053] The training and use of neural network models typically occur on the cloud side and on the device side, respectively. The neural network model is trained on the cloud side and, after training is complete, deployed on the device side for use.

[0054] Figure 1 This is a schematic diagram of an architecture for network model training and use, which is exemplified by an embodiment of the present application. The architecture includes a training device 101 and one or more use devices 102. The training device 101 and the use device 102 communicate with each other via wired or wireless means, so the training device 101 can send the trained model (or network) to the use device 102; accordingly, the use device 102 receives the model and uses the model to complete a specific task.

[0055] Optionally, the using device 102 can feed back the results based on the model to the above-mentioned training device 101, so that the training device 101 can further train the model based on the results of the using device 102; the retrained model can be sent to the using device 102 to update the original model.

[0056] The training device 101 can be a device with strong computing capabilities, such as a server or a server cluster consisting of multiple servers. The training device 101 can include a first neural network. The first neural network can be used to process multiple services. These multiple services have different tasks but belong to the same category.

[0057] The device 102 is a device that needs to recognize (or detect) images, such as a handheld device (e.g., a mobile phone, a tablet computer, a PDA, etc.), a vehicle-mounted device (e.g., a car, a bicycle, an electric car, an airplane, a ship, etc.), a wearable device (e.g., a smart watch (such as iWatch, etc.), a smart bracelet, a pedometer, etc.), a smart home device (e.g., a refrigerator, a television, an air conditioner, an electric meter, etc.), an intelligent robot, a workshop equipment, etc.

[0058] 3. Training phase of the multi-task neural network model:

[0059] The training of a multi-task neural network model can be divided into multiple training stages, which can include pre-training and fine-tuning.

[0060] Pretraining refers to the process of pretraining a model or preselecting a training mode. This process results in a pretrained model. Pretraining involves setting up a network model to perform a specific task. After initializing the network model's parameters, the network is trained, continuously adjusting the parameters to minimize losses. If the network model pretraining meets certain requirements, the parameters are saved and designated as a pretrained model. The pretraining process allows the network model to capture a wide range of features by learning from a large amount of general data before processing task-specific data, improving its performance and versatility on the target task.

[0061] Fine-tuning is the process of applying a pre-trained model to a specific business dataset and adapting its parameters to that dataset. After receiving a specific task, the saved pre-trained model can be used as the initial parameters, with the parameters continuously adjusted during the training process. Further training the pre-trained model on a small amount of labeled data for the new task allows the model to learn the specific features and patterns relevant to the current task, thereby ensuring the unique characteristics of the new task.

[0062] 4. Pre-training process of different types of neural network models:

[0063] Different types of neural network models require different pre-training tasks. The following describes two different pre-training models: the generative language model llama and the representation pre-training model RoBerta, and their corresponding pre-training processes:

[0064] Example 1: The generative language model llama is a pretrained model. The original network model is pretrained to generate the generative language model llama. The pretrained task is the LanguageModel task, which predicts the next word given a preceding text. This data can be directly sourced from open-source data or customized, ensuring quality while closely matching the data distribution of the pretraining process. After pretraining, if you wish to fine-tune the llama model, you can use an open-source dataset.

[0065] Example 2: RoBerta, a pre-trained representation model, is a pre-trained model. The original network model is pre-trained to obtain RoBerta. The pre-training task is the MaskedLanguageModel task, which masks some input words and predicts the masked content from the following text. Data can also be sourced from open-source corpora or self-built.

[0066] 5. Parameter fine-tuning method:

[0067] The neural network model involved in the embodiments of the present application requires fine-tuning of the pre-trained model. Several fine-tuning methods are described below.

[0068] 1) Low-Rank Adaptation (LoRA)

[0069] LoRA is a widely used technique for fine-tuning large language models. It provides a training method that reduces the number of parameters required for training. LoRA reduces the number of parameters by inserting low-rank matrices into specific layers of a pre-trained model, making fine-tuning more efficient. This allows fine-tuning by adding a small number of parameters without changing the original model weights W.

[0070] Figure 2A This is a schematic diagram of a LoRA model structure disclosed exemplarily in an embodiment of the present application.

[0071] like Figure 2A As shown, LoRA adds two low-rank matrices A and B to a specific layer in the pre-trained model, and the original weight matrix W is adjusted to W+AB. When the input data is X, the original weight matrix is W, and the two low-rank matrices A and B are used, the forward propagation formula is:

[0072] Y=(W+AB)X

[0073] Among them, if the original size of the parameter W is d*d, the sizes of the matrices A*B are d*r and r*d respectively, and Y is the output data.

[0074] The above-mentioned LoRA reduces the number of updates to the number of layers by introducing a low-rank matrix, which can reduce computing costs and storage requirements.

[0075] 2) Series adapter

[0076] During parameter fine-tuning with an adapter, the adapter component is added to the pre-trained backbone model W. The adapter's network structure can be designed based on the characteristics of the task and typically includes several fully connected or convolutional layers. Initialization is performed using the parameters of the pre-trained backbone model W, which are then frozen. Only the adapter parameters are fine-tuned. Because the adapter has fewer parameters, the fine-tuning process is more efficient.

[0077] Figure 2B This is a schematic diagram of an Adaptor model structure disclosed exemplarily in an embodiment of the present application.

[0078] like Figure 2B As shown in the figure, the SeriesAdapter adds an adapter component C to the backbone model W in pre-training mode, and the original weight matrix is adjusted from W to CW. When the input data is X, the original weight matrix is W, and the adapter component is A, the forward propagation formula is:

[0079] Y=CWX

[0080] By reusing some parameters of the backbone model, the adapter reduces the amount of learning required during fine-tuning and improves training efficiency. Gradient isolation ensures that adapter training does not interfere with the backbone model, ensuring the stability of the fine-tuning process. Furthermore, the adapter possesses a certain degree of generalization capability, allowing it to be applied to different tasks and model architectures.

[0081] It should be noted that the above parameter fine-tuning methods are designed to illustrate two methods in this application. This application may also include other fine-tuning methods without limitation. In addition, the above is only fine-tuning of the operator. It can also be fine-tuning of a part of the network model (multiple consecutive neurons are W). For example, after pre-training, the local part of the network model is L(X), and after fine-tuning, it is: Y = L(X) + ABX.

[0082] Operators (OPs) are computing units in deep learning algorithms. In neural network models, operators correspond to the computational logic within a layer.

[0083] 6. Quantification:

[0084] Quantization is a technique for compressing deep learning model parameters to reduce computational complexity. Neural networks are approximate computational processes, and absolute precision is not required for each calculation. Therefore, in some cases, model parameters that require more bits of storage can be converted to fewer bits without compromising model accuracy. Quantization typically converts a floating-point model (the Tensor data type in common neural networks is typically float32) into a quantized model (Tensor data types such as int4). Quantization converts floating-point real numbers into fixed-point integers, essentially approximating continuous values to a finite number of discrete values. For example, the value range of a uint8 is 0 to 255, while the value range of a float32 is 0.0 to 1.0. Converting a float32 value to a uint8 value can be considered quantization. Quantization essentially rescales the value range and can be roughly understood as a linear mapping (or nonlinear mapping). Quantization can be used as an information compression method to effectively reduce the large amount of memory consumed by neural network models, addressing storage priorities and reducing the number of data reads.

[0085] Figure 3 This is a schematic diagram of a quantitative mapping disclosed in an exemplary embodiment of the present application. Figure 3 As shown, r represents a floating-point real number, q represents a quantized fixed-point integer, and the conversion process between floating-point real numbers and fixed-point integers can be:

[0086] r=S(qZ)

[0087] q=round(r / S+Z)

[0088] Where S is scale, which represents the proportional relationship between real numbers and integers; Z is zero point, which represents the integer corresponding to the real number after 0 is measured, S = (rmax-rmin) / (qmax-qmin), Z = round(qmax-rmax / S). rmax is the maximum value of r; rmin is the minimum value of r; qmax is the maximum value of q; qmin is the minimum value of q. The round function is a rounding function used to calculate and return a specified number of digits for rounding. It should be noted that the above quantization process is only an exemplary statement and is not limited in this application. Figure 3 This is just an example, and the present application may also use other quantification methods without limitation.

[0089] Several different quantization algorithms are described below:

[0090] Quantization can be generally divided into two types: post-training quantization (PTQ) and quantization-aware training (QAT). The following describes these two types of quantization:

[0091] Post-training quantization:

[0092] PTQ is a post-training quantization process. After model training is complete, a high-precision training model is obtained. PTQ requires statistical analysis of training model parameters, such as the maximum and minimum weights. Based on this statistical information, the parameters are then quantized from high-precision to low-precision. The advantage of PTQ is the simplicity of the quantization process, but the disadvantage is relatively poor quantization results.

[0093] Quantization-aware training:

[0094] QAT simulates the quantization during the inference phase during training. Simulating the effects of quantization during training allows the model to adapt to the information loss caused by quantization, thereby reducing the accuracy loss caused by quantization and ensuring higher performance.

[0095] One possible approach is to directly insert a pseudo-quantization node, perform training and quantization from scratch, optimize the parameters based on the gradient backpropagation, and directly obtain the trained and quantized model.

[0096] Another possible approach is to insert fake quantization nodes into the model after floating-point model training is complete, and then perform quantization training. The fake quantization nodes simulate quantization of data from float to int types, primarily affecting network weights and activations. After model training is complete, the fake quantization nodes are replaced with real quantization operations to generate the final quantized model. Fake quantization is actually a combination of quantization and dequantization, simulating the error caused by quantization rounding. While the input and output of fake quantization operations appear unchanged, they simulate rounding operations. This error is treated as training noise, and QAT fine-tuning adapts to this noise, thereby reducing accuracy loss during the final quantization to INT4.

[0097] Multi-task learning (MTL) is a machine learning method used to simultaneously learn multiple related tasks through shared representations. Its core idea is to leverage the correlation between different tasks and improve the model's generalization ability and learning efficiency by sharing model parameters or feature representations.

[0098] After multi-task learning training, a multi-task neural network model is obtained. The neural network model is then deployed to the user device, which can use the multi-task neural network model according to the needs of different tasks. However, in order to continuously improve the capabilities and effects of the neural network model, the neural network model needs to be continuously evolved and updated, which means that the trained network model needs to be retrained and deployed. The following examples illustrate two possible scenarios:

[0099] For example, a user device already has a neural network model deployed that can perform two tasks: voice information recognition and voice conversation. To provide users with a third task (e.g., user voice input "open camera"), the manufacturer needs to retrain the neural network model using the data from the three tasks and redeploy the retrained neural network model to the user device. This means that the user device needs to re-download a complete neural network model for use.

[0100] For example, a neural network model has been deployed on the device, which can complete Task 1, Task 2, and Task 3. After Task 3 was put into use, the user's performance was not good. The manufacturer retrained the neural network model and tested Task 3 to ensure that it met the standards before re-using it. The above process requires retraining the neural network model on the training device side according to the data of the above three tasks, and redeploying the retrained neural network model to the device, that is, the device needs to re-download a complete neural network model for use.

[0101] In order to ensure the effectiveness of neural network models or add and adjust different business functions, manufacturers need to continuously evolve neural network models, as well as retrain and redeploy neural network models before putting them into use.

[0102] Due to the addition of new types of tasks (business) or the readjustment of existing tasks, if the deployed neural network model is further trained, the execution effect of the existing types of tasks needs to be tested after training. The test results will inevitably have a negative effect on the previously existing training tasks. If the neural network model is completely retrained according to business needs, the training cycle will be long and the processing resources of the training equipment will be seriously occupied. In addition, after the training is completed, the complete data needs to be redeployed to the user device. Each time the user device updates the neural network model, the complete neural network model data needs to be downloaded. The amount of downloaded data is large, and the user-side user device update and installation time is long. For example, the mobile phone system version update includes a complete neural network model, and the version update time is long.

[0103] In response to the above problems, the present application proposes a model training method and a network model. The neural network model includes a set of shared weight parameters and N sets of business weight parameters. The neural network model can handle multiple tasks, and each task corresponds to a set of business weight parameters. During the neural network model training phase, the training device can first pre-train the original network model to obtain a pre-trained model, and determine the parameters of the pre-trained model as shared weight parameters. Then, the original business network model is built based on the pre-trained model. The original business network model includes the shared weight parameters and original business weight parameters in the pre-trained model. N sets of business data are respectively input into the original business network model to train N sets of business weight parameters. When training the i-th set of business weight parameters corresponding to the i-th business, the i-th set of business data is output to the original business network model for training to obtain the i-th set of business weight parameters. Then, the next set of business weight parameters is trained until the business weight parameters of all businesses are trained. Wherein, N is an integer greater than or equal to 2, and i is a positive integer from 1 to N. After the above-mentioned neural network model training is completed, the neural network model is deployed to the user device. During the neural network model usage phase, when processing the kth service, the electronic device first determines the kth service network model based on the shared weight parameters and the kth service weight parameter among the N groups of service weight parameters, inputs the kth service data into the kth service network model, and obtains an output result. Here, k is any integer from 1 to N.

[0104] In the above network training process, high-precision parameters are used, while in data deployment, low-precision parameters are needed to ensure that the network transmits less data, increase transmission speed, and reduce the pressure on device data storage. For example, Figure 4 This is a schematic diagram of a training model deployment process disclosed in an embodiment of the present application. Figure 4 As shown, the multi-task high-precision model is the weight element of FP16 or BE16, and the weight elements of FP16 or BE16 are quantized to obtain a multi-task low-precision model, which is the weight element of INT4.

[0105] Switching from high-precision parameters to low-precision parameters requires pre-quantization of the network model. After the low-precision model has been deployed to a user device, if a service needs to be adjusted (including updating existing services or adding new services), the training device must re-quantize the entire model after adjusting the service. After quantization is complete, since the weight parameters of all services have been adjusted, the training device must re-send the complete network model to the user device. Therefore, during the service adjustment process, many model parameters are re-sent, and the user device needs to download a large amount of data. For example, when an electronic device (user device) updates a multi-service neural network model, the quantization process causes all model parameters to change, and the electronic device needs to download the entire network model data, that is, the network model for the services that have not been adjusted also needs to be downloaded and installed. In addition, during the service adjustment process of the multi-service neural network model, the quantization process for the adjusted services requires adjusting the quantization method of the entire model. This may have a negative impact on the non-adjusted services compared to the already deployed network, resulting in the output results of some tasks in the multi-task neural network model being worse than before.

[0106] Based on the above problems, the present application proposes a model quantization method and training device, including: a training device quantizes a pre-trained model to obtain a pre-trained quantized model; the pre-trained quantized model includes a shared quantization weight parameter; the training device builds a first original business model based on the pre-trained quantized model; the first original business model includes the shared quantization weight parameter and the initial business weight parameter; the training device quantizes and perceives the first original business model based on N business data to obtain N target business models; the tth target business model in the N target business models includes the shared quantization weight parameter and the tth business quantization weight parameter in the N target business models; one business corresponds to one target business model; and t is an integer from 1 to N in sequence.

[0107] In the above implementation, since the pre-trained shared weight parameters are quantized in advance, the quantized shared weight parameters are kept unchanged during the fine-tuning process and the quantization process of the business parameters, which can decouple the influence of the parameter training and quantization process of each business. In this way, when a certain business changes or a new business is added, it will not affect the model parameters corresponding to the business that has been trained and quantized. Therefore, in the case of subsequent business changes, the transmission of model update data can be reduced, the amount of data downloaded by the device is greatly reduced, and the model update efficiency is improved. In addition, after the above-mentioned training and quantization process of each business, the negative impact of each business on the quantization and training is avoided, and the quantization effect of each business is improved.

[0108] Combining the training and quantization process of the above multi-task neural network model, the following describes the multi-service model training process in detail:

[0109] Figure 5 This is a flow chart of a model quantization method disclosed in an exemplary embodiment of this application. Figure 5 As shown, the electronic device may perform but is not limited to the following steps:

[0110] S501: The training device pre-trains the original network model to obtain a pre-trained model.

[0111] After the training device completes the construction of the original network model, it can pre-train the network model to obtain a pre-trained model. The type of pre-trained model is not limited, and it can process text, images, voice, and video, etc. The original network model built for different business types should be different. For example, the original network model is a model for processing language text; the original network model is a model for image clarity processing; the original network model is a model for identifying information in videos; and the original network model is a model for converting audio into text. The above is only an example of the original network model and is not limited. Before and after the original network model is trained to obtain the pre-trained model, the model processes the same business.

[0112] The following uses a large generative language model and a representation pre-training model as examples to illustrate the pre-training process:

[0113] For example, when the pre-trained model is a generative language model (llama), the pre-training can be performed by training the original network model through a language model task to obtain the pre-trained model. For example, input the previous paragraph of text and predict the next paragraph of text.

[0114] Exemplarily, when the pre-training model is a representation-type pre-training model, the pre-training may adopt a masked language model (MLM) task, that is, masking part of the words in the original network model input data and predicting the content of the masked part.

[0115] The pre-training process of the above two examples can directly use open source data or build a self-built database, which is not limited in this application.

[0116] In the embodiments of the present application, N business neural network models can process N businesses, with strong correlations between the N businesses. These businesses may have some common feature extraction capabilities, but also some different feature extraction capabilities. For example, in a multi-business neural network model that processes language text, the N businesses may require the common capability of extracting and analyzing text data, while some of the businesses may require capabilities such as generating the next consecutive sentence of the text, determining the processing task corresponding to the text, and responding to the output text with other text.

[0117] The weight parameters in the pre-trained model can be used as shared parameters in subsequent service network models. That is, the pre-trained model only includes the unquantized shared weight parameters and does not include the service weight parameters. Therefore, the pre-trained model can extract common features across different services. After the pre-trained model weight parameters are trained, subsequent training does not adjust these weight parameters. This allows the same feature extraction method to be used to extract common features across different services, minimizing model changes and ensuring model versatility.

[0118] Optionally, the pre-trained model in S501 may be a model that has already been trained, and there is no need to perform a pre-training process. This application may not perform a pre-training processing process.

[0119] In S501 , the pre-training process is completed, and then the pre-training model is fine-tuned according to the business. The fine-tuning process is described in S502 and S503 .

[0120] S502: The training device builds an original business model based on the pre-trained model.

[0121] During multi-task learning, tasks are highly correlated, yet distinct from one another. The weight parameters of the pre-trained model are used as shared weight parameters to build the original business model. The original business model includes shared weight parameters and initial business weight parameters. Initial business weight parameters are untrained and unquantized business weight parameters. Shared weight parameters are trained, unquantized weight parameters.

[0122] Based on the weight parameters of the pre-trained model, the initial business weight parameters are added to form the original business model. Different training methods should correspond to different original business models. The following describes the model building situations in different fine-tuning methods:

[0123] First, build the original business model according to the LoRA model.

[0124] The method of fine-tuning the LoRA model can be found in Figure 2A The basic idea of LoRA is to add low-rank matrices A and B to the pre-trained model. The key to LoRA fine-tuning is to determine the basic weight W in the pre-trained model, so as to accurately determine the location and size of the added low-rank matrices A and B. The following describes the corresponding LoRA fine-tuning situations for different W in the pre-trained model:

[0125] Case 1: Take the single weight matrix in the pre-trained model as W and add low-rank matrices A and B.

[0126] The training device can use a single operator of the pre-trained model as the original parameter W in LoRA and introduce low-rank matrices A and B. The following uses a pre-trained model as an example to illustrate the construction process:

[0127] Figure 6A This is a schematic diagram of an original business model of LoRA shown in an embodiment of the present application. Figure 6A As shown in , the pre-trained model can include an input layer, an output layer, and a hidden layer. The number of hidden layers is multiple, and information can be transmitted and processed between multiple layers. Neurons can receive result inputs from neurons in the previous layer. Figure 6A As shown, the hidden layer can include operators such as E1, W1, W2, W3, W4, W5, and E2, W6, among which W1, W2, W3, W4, W5, and W6 are operators for feature extraction; the neural network model can also include non-functional operators: E1, E2, etc. E1 and E2 can be activation functions that map the input of the neuron to the output end, and LayerNorm layer normalization operators for normalization processing. Among them, operators W1, W2, W3, W4, W5, and W6 can be used as the original weight W of LoRA, and E1 and E2 cannot be used as the original weight W of LoRA alone.

[0128] The following uses W1 as the basic weight W of LoRA as an example to illustrate the LoRA processing process. Figure 6A As shown, based on W1, two low-rank matrices A1 and B1 are introduced, so that W1X is replaced by (W1+A1B1)X. Among them, the matrix size of A1B1 can be equal to the matrix size of W1 (wherein, the matrix sizes of A1 and B1 may not be equal to the matrix size of W1, which is not limited in this application). A1 and B1 can be initialized in the original business model. For example, the initialization matrix A is a random number with a mean of 0, and the initialization matrix B is a completely zero matrix.

[0129] Optionally, Figure 6A In the example above, only one operator W1 is used as W to illustrate the addition of low-rank matrices A and B. This application can select one or more operators as W. For example, all operators in the hidden layer of the pre-trained model that can be used as W can be subjected to LoRA fine-tuning. This application does not limit the selected operator position and pre-trained model.

[0130] It should be noted that the above is only an example of the basic LoRA fine-tuning method. It can be a modified method of LoRA, such as LoRA+, VeRA, LoRA-fa, LoRA-drop, AdaLoRA, DoRA and Delta-LoRA fine-tuning methods, which are not limited in this application.

[0131] Case 2: Take multiple consecutive neurons in the pre-trained model as a W and add low-rank matrices A and B.

[0132] The training device can treat multiple consecutive neurons in the pre-trained model as a W and introduce low-rank matrices A and B. The following uses a pre-trained model as an example to illustrate the construction process:

[0133] Figure 6B This is a schematic diagram of another LoRA original business model construction exemplarily shown in an embodiment of the present application. Figure 6B The pre-trained model in can refer to the above Figure 6A Related description in . Figure 6B As shown, the entire hidden layer can be regarded as an original parameter W in LoRA. Assuming that the processing process of the hidden layer is L(X), the network is built in LoRA mode, and the processing process of the number of hidden layers is L(X)+ABX.

[0134] Optionally, Figure 6B The entire hidden layer is used as an original parameter W in LoRA, or a continuous part of the hidden layer network is used as the original weight W, and A and B are introduced to build the original network model. This application does not limit the range of W.

[0135] Optionally, in the process of building an original business model, low-rank matrices A and B can be added by combining the above-mentioned case 1 and case 2.

[0136] Second, build the original business model according to the adapter model.

[0137] Figure 6C This is a schematic diagram of an original business model of an adapter exemplarily shown in an embodiment of the present application. Figure 6C The pre-trained model in can refer to the above Figure 6A Related description in . Figure 6C As shown, the entire hidden layer can be used as an adapter backbone model W, and the adapter component C is added to obtain the original business model CWX.

[0138] For example, the last two hidden layers of the backbone model include the activation function activate and the normalization layer layernorm, etc., and the offset bias(B1) is added. The backbone model W can be expressed as LayerNorm(activate(W1X+B1)). After introducing the adapter component C, the output layer can be expressed as C(LayerNorm(activate(W1X+B1)))+B2, where B2 is also an offset. The above is merely an example and is not limited to this application.

[0139] The above two examples are only illustrated by LoRA and adapter fine-tuning. The fine-tuning process can include LoRA deformation fine-tuning, prefix tuning, prompt tuning, P-Tuning and P-Tuning v2 and other methods, which are not limited in this application.

[0140] S503: The training device trains the original service models based on the input data of the N services respectively to obtain N service network models.

[0141] The training data for the fine-tuning process includes N service data sets, where the N service data sets include the first service data set, the second service data set, ..., and the Nth service data set. The N service data sets correspond to N types of services, where N is an integer greater than or equal to 2. The training device can sequentially input each service data set into the original service model to train the service weight parameters in the service model.

[0142] Specifically, the training device inputs the first business data set into the original business model, the shared weight parameters in the original business model remain unchanged, the initial business weight parameters are trained, the first original business model corresponding to the first business is trained, and the first business network model is obtained... The training device inputs the i-th business data set into the original business model, the shared weight parameters in the original business model remain unchanged, the initial business weight parameters are trained, the i-th original business model corresponding to the i-th business is trained, and the i-th business network model is obtained... The training device inputs the N-th business data set into the original business model, the shared weight parameters in the original business model remain unchanged, the initial business weight parameters are trained, the N-th original business model corresponding to the N-th business is trained, and the N-th business network model is obtained. The i-th business network model among the N business network models includes the shared weight parameters and the i-th business weight parameter among the N business weight parameters; i is an integer from 1 to N in sequence, and N business network models are obtained through the above training process.

[0143] During the training process, the training device can sequentially train the service weight parameters in the original service model using the training data of different services. After the training is completed, the shared weight parameters and N service weight parameters corresponding to the N services are obtained. The training device can store the shared weight parameters and the N service weight parameters corresponding to the N services, as well as related data structure information. The data structure information can indicate the position of the shared weight parameters and the service weight parameters in the neural network model.

[0144] After S503, the training device can select target quantization parameters based on the loss values of all services in the N services, and determine N target service models based on the target quantization parameters. The target quantization parameters are used to quantize the shared weight parameters and service weight parameters in the N service network models. The i-th target service model in the N target service models includes the shared quantization weight parameters and the t-th service quantization weight parameters; t is an integer from 1 to N. This is explained in detail below through S504 to S509.

[0145] S504: The training device determines a first quantization parameter.

[0146] The quantization process establishes a data mapping relationship between fixed-point and floating-point data, approximating the continuous value of the signal to a finite number of discrete values, achieving good results at a minimal loss of precision. The first quantization parameter is an adjustment parameter that affects the quantization results of the N business network models.

[0147] Several possible implementations of the first quantization parameter are described below:

[0148] Implementation 1: The above mapping process can be expressed by the following quantitative mapping formula:

[0149] Q=round(scale*clip(x, UpperTh, LowTh))+zero

[0150] Among them, the first quantization parameter is an adjustment parameter for N business network models according to a specific quantization method. The first quantization parameter can be scale, zero, clip, UperTh and LowTh. Among them, Q represents the fixed-point element in the fixed-point weight after quantization, and x represents the floating-point element in the floating-point weight before quantization. Scale is the quantization scale, which is determined by the maximum and minimum values of the elements in the fixed-point weight after quantization, and can represent the proportional relationship between floating-point numbers and integers. For example, if linear uniform quantization is used, scale = (2N-1) / (Xmax-Xmin), Xmax is the maximum value after the clip operation; Xmin is the minimum value after the clip operation. Round is a rounding operation; the clip operation is a slicing operation, that is, selecting the range of the quantized object. UperTh is the upper limit of the clip operation. LowTh is the lower limit of the clip operation. At this time, the first quantization parameter can be a parameter for adjusting one or more of scale, zero, UperTh upper limit and LowTh lower limit, so that the quantization result can be adjusted.

[0151] Implementation method 2: The first quantization parameter is the scaling scale a, and a can be adjusted to determine the parameter of the quantization scale.

[0152] The first quantization parameter may include scale. For example, when the quantization mapping formula is Q=x / scale, scale=a(Xmax-Xmin) / (2N-1). When the quantization mapping formula is Q=round(scale*x)+zero, scale=a(2N-1) / (Xmax-Xmin).

[0153] Where a is the scaling factor, and N is the number of quantization bits. The electronic device can adjust a to adjust the first quantization parameter, scale. If the quantized value is int4, N is 4, and scale = a(Xmax - Xmin) / 15; if the quantized value is int8, N is 8, and scale = a(Xmax - Xmin) / 255. Xmax - Xmin is the difference between the maximum and minimum values of the weight elements in the current weight matrix.

[0154] The training device may have different scaling factors a, where a may include a1, a2, a3, a4, ..., aK. The scale size is calculated for each different a. For example, a1, a2, a3, a4, ..., aK are traversed. Scale1, scale2, scale3, ..., scaleK are obtained in sequence. Each weight element Q in the weight matrix is calculated for each K scale, thereby obtaining a quantized weight matrix. Where K is an integer greater than 2.

[0155] Implementation 3: The first quantization parameter is the scaling scale a and the slice size s*t. a and s*t can be adjusted to determine the quantization scale, the maximum value Xmax, and the minimum value Xmin.

[0156] The scaling factor a can refer to the content of a in Implementation 1. The slice size s*t is the matrix size of the weight elements in each weight matrix of the neural network. After the above segmentation, the maximum value Xmax and the minimum value Xmin are determined for the s*t weight elements in each segmentation unit.

[0157] The training device can adjust the scaling factor a to determine the quantization scale and the slice size s*t to determine the Xmax and Xmin in each slice. The quantization mapping formula can then be determined. Figure 7 Schematic diagram of weight matrix segmentation in a neural network model disclosed in an embodiment of the present application. Figure 7As shown, the weight matrix W4 in the pre-trained network model is segmented to obtain multiple groups of 16*16 (s*t) weight elements. The maximum value in each weight element group is Xmax, and the minimum value is Xmin. Thus, the Q corresponding to each weight element matrix can be calculated. All segmented weight element matrices contain s*t. The electronic device can change the size of s*t, thereby adjusting the first quantization parameter.

[0158] Implementation method 4: nonlinear quantization processing.

[0159] The training device can use nonlinear quantization to establish a quantization mapping relationship. For example, the quantization mapping formulas may be one or more of the following: Q = round((ex-1) / scale) + zero, Q = clip(x + zero, upper, lower), or Q = ax + zero. The above nonlinear quantization formulas are merely illustrative and not limiting.

[0160] Optionally, the quantization process can be segmented. For example, when x is greater than 0, Q = round(lg(x+1)); when x is less than 0, Q = round(-lg(-x+1)); when x is equal to 0, Q = 0. Where lg is a logarithmic function. This application does not limit the specific quantization process.

[0161] After determining the first quantization parameter, the training device performs a target loss calculation process based on the first quantization parameter. The target loss calculation process can be specifically referred to the contents of S505 and S506, which are described in detail below:

[0162] S505: The training device quantizes the N service network models according to the first quantization parameter to obtain N quantized service network models.

[0163] After the first quantization parameter is determined, a quantization formula can be determined based on the first quantization parameter. The quantized fixed-point element Q corresponding to each floating-point element x can then be calculated using the quantization mapping formula. By quantizing the service network models corresponding to the N services, N quantized service network models can be obtained.

[0164] S506: The training device calculates the target values of the N service network models before and after quantization respectively.

[0165] The target value can be either a quantitative loss value or a business loss value. The following describes each case separately:

[0166] Case 1: The target value is the quantized loss value.

[0167] After obtaining the N quantized business network models, the training device can calculate the loss values output by each business network model before and after quantization, and calculate the target loss based on the loss value corresponding to each business. The training device can input the quantized training data into the two business network models before and after quantization of the i-th business, respectively, to obtain the i-th unquantized output result and the i-th quantized output result. The training device can then calculate the loss value between the unquantized output result and the quantized output result according to the target loss function. Wherein, i is an integer from 1 to N in sequence. The training device can calculate N loss values corresponding to N businesses: loss(1), loss(2), loss(3), ..., loss(i), ..., loss(N). The quantized loss value is the loss before and after quantization.

[0168] The training device can establish a loss function and calculate the loss value based on the loss function. Among them, the loss function can be absolute value loss, square loss, cross entropy loss, mean square error, log-likelihood loss, KL loss, etc., which is not limited in this application.

[0169] The above-mentioned training device inputs the quantized data of N businesses into N business network models respectively, and obtains N unquantized output results accordingly; inputs the quantized data of N businesses into N quantized business models respectively, and obtains N quantized output results accordingly, and calculates N loss values based on the N unquantized output results and the N quantized output results. Afterwards, the training device can also calculate the target loss value based on the N loss values. After calculating the loss values of the N business network models before and after quantization, the training device can also calculate the quantized loss value based on the N loss values. That is, the training device can be provided with a target loss function, and the dependent variable of the target loss function includes the N loss values corresponding to the N businesses. For example, the target loss function: quantized loss value Loss = loss(1)+loss(2)+loss(3)+…+loss(i)+…+loss(N). The loss weights of the above-mentioned different businesses are the same or may be different, which is not limited in this application. The training device can calculate the target loss value Loss based on the target loss function.

[0170] Case 2: The target value is the business loss value.

[0171] The training device can establish a business loss function, which is the loss between the expected business result and the actual output result. The expected business result is the ideal output result (i.e., the correct result) of the quantitative business model. The actual output result is the output result of the quantified quantitative business model. After determining the expected business result and the actual output result, the training device can calculate the business loss value based on the business loss function. Among them, the business loss function can be absolute value loss, square loss, cross entropy loss, mean square error, log-likelihood loss, KL loss, etc., which is not limited in this application. Among them, different quantitative business models have different ways of calculating their business expectations and actual output results, which is not limited in this application.

[0172] S507: The training device determines whether quantitative adjustment is required. If quantitative adjustment is required, execute S509; if not, execute S508.

[0173] After the target loss calculation process of S505 and S506 is executed, the training device may need quantitative adjustment, that is, start to execute S507.

[0174] The training device can determine whether quantization adjustment is needed by determining whether the traversal is completed or by determining whether the loss function has converged. The following describes two different implementation methods:

[0175] In one possible implementation, the training device determines whether N service network models have been quantized according to all parameters in the first quantization parameter to obtain quantization loss values before and after quantization. If it is determined that all first quantization parameters have been traversed, it is determined that quantization adjustment is not required; if there are first quantization parameters that have not been traversed, it is determined that quantization adjustment is required.

[0176] In another possible implementation, based on the target loss calculated multiple times in S506, it is determined whether convergence has occurred. When the quantization loss has converged, it is determined that no quantization adjustment is required; when the quantization loss has not converged, it is determined that quantization adjustment is required. Specifically, after calculating the quantization loss value Loss, the training device can determine whether the quantization loss has converged based on multiple consecutive quantization loss values. That is, when the quantization loss value continues to decrease and tends to be stable (extraction descent), the training device can determine that the quantization loss has converged; otherwise, it has not converged. The above-mentioned process of establishing the calculation objective function can establish the objective function based on discrete values, derive the objective function, and obtain the result of whether it has converged.

[0177] For example, if the difference between the quantization loss values Loss calculated multiple times is less than a threshold, the quantization loss is determined to have converged; otherwise, it is determined not to have converged. After executing S506 for the first time, S507 is executed, and the electronic device can determine that a quantization loss value is currently obtained and that the quantization loss value Loss cannot be determined, i.e., it is determined not to have converged.

[0178] In another possible implementation, if the service loss values of N services of the training device in S506 meet preset requirements, it is determined that no quantitative adjustment is required; if the service loss values do not meet the preset requirements, it is determined that quantitative adjustment is required. The preset requirements may include the difference between the service loss value and the expected value being less than a preset value or converging. Optionally, whether the service loss value converges can refer to the processing process of whether the quantified loss converges, which will not be described in detail.

[0179] S508: The training device adjusts the first quantization parameter.

[0180] When the training device determines that quantization adjustment is required, the training device needs to adjust the first quantization parameter. After the first quantization parameter is adjusted, S505 and S506 are re-executed. Different quantization parameters may correspond to different adjustment methods. The following describes two different adjustment situations.

[0181] In one possible implementation, a fixed selection range is set for the quantization parameter. Different first quantization parameters are sequentially traversed according to the selection range to execute S505 and S506. For example, in implementation 1 of S504, the first quantization parameter is scaling a, and scaling a can be adjusted to a scaling scale that has not yet been quantized. Similarly, other implementations of S504 can also be adjusted in the same manner as described above, and are not further described.

[0182] In one possible implementation, the first quantization parameter may be adjusted using a gradient descent method, i.e., the first quantization parameter is adjusted in a direction in which the quantization loss value or the service loss value decreases. If a computational relationship between the first quantization parameter and the loss function can be established, the first quantization parameter may be updated using a gradient descent method.

[0183] It should be noted that the above two methods for adjusting the first quantization parameter are merely exemplary descriptions and are not specifically limited.

[0184] S509: The training device determines the first quantization parameter corresponding to the minimum value of the target value as the target quantization parameter, and quantizes the N service network models according to the target quantization parameter to obtain N quantized target service models.

[0185] In one possible scenario, when the target value is a quantization loss value, the minimum value of the quantization loss value is determined as the minimum value of the target value. When the target loss converges, the first quantization parameter corresponding to the minimum value of the quantization loss of the training device is determined as the target quantization parameter, and a quantization mapping formula is determined according to the target quantization parameter. N service network models are quantized according to the corresponding quantization mapping formula to obtain N quantized target service models.

[0186] In another possible scenario, when the target value is a service loss value, the minimum value of the difference between the service loss value and the service expectation value is determined as the minimum value of the target value. The training device determines the first quantization parameter corresponding to the minimum value of the service loss value as the target quantization parameter, and quantizes N service network models according to the target quantization parameter to obtain N quantized target service models.

[0187] After obtaining N quantized target business models, the training device may store the shared weight parameters and N business weight parameters of the N target business models, and store the structural data corresponding to the weight matrix.

[0188] above Figure 5 In the implementation method, the target loss function can be jointly determined by the losses of all businesses to adjust the quantization parameters so that the quantization results are beneficial to all businesses, thereby achieving quantization balance and ensuring that the quantization results will not have a negative impact on some businesses, thereby ensuring the output effect of the business network model after quantization.

[0189] Combine Figure 5 The multi-task model training process in this application is exemplified below by taking the execution of LoRA training method by some parameters in the neural network model as an example. Figure 8 This is a schematic diagram of a multi-task model training process disclosed in an embodiment of the present application. Figure 8 As shown in the figure, the combing results of different stages are explained in order of processing time:

[0190] like Figure 8As shown, after pre-training, the shared weight parameters of the pre-trained model (partial parameters W of the shared weight parameters) are obtained, and the pre-trained model is fine-tuned according to different business data, and the business weight data A and B are trained to obtain N corresponding business weight parameters: A1 and B1, A2 and B2, ..., AN and BN. Among them, LoRA includes the shared weight parameters W and business weight parameters A and B in the pre-trained model. The training model is built by combining the N business weight parameters with W in the pre-trained model to obtain N business network models, namely business model 1, business model 2, ..., business model N. The training model then quantizes the N business network models according to the quantization parameters to obtain N quantized business network models, namely quantized business model 1, quantized business model 2, ..., quantized business model N. Among them, after quantization, W becomes the weight matrix w, after quantization, weight matrix A becomes the weight matrix a; after quantization, weight matrix B becomes the weight matrix b. The training model calculates quantization loss values loss(1), loss(2), ..., loss(N) based on N business network models before and after quantization. After calculating the quantization losses of the N business network models, the quantization loss value Loss is calculated based on the N quantization losses. After calculating the quantization loss values, it is determined whether to adjust the quantization parameters based on the quantization loss values. If quantization adjustment is not required, N target quantization business models can be determined. If quantization adjustment is required, the process of quantizing the N LoRAs and calculating the target values is re-executed according to the adjusted quantization parameters to determine whether the target values have converged.

[0191] It should be noted that Figure 8 The W and w shown in the figure may refer to the entire network model, or may refer to a part of the network model or one or more weight matrices, which is not limited in this application.

[0192] Figure 5 and Figure 8 During the multi-task model training process, since the dependent variable of the objective loss function is the quantized loss value of the business network model for all businesses, there is a balance between the quantized losses of all businesses. It is possible that the loss of one business will decrease while the loss of another business will increase. This has improved to a certain extent, but it still cannot fully take into account the quantization effect of all businesses. In addition, the above quantization process means that if the neural network model adds new businesses or adjusts existing businesses in the future, the entire network model needs to be completely requantized, and the quantization process is long. In addition, after quantization is completed, the entire model needs to be redeployed to the device where it is used, which requires a large amount of data to download and a long deployment and installation time.

[0193] Based on the above training process of the task model, Figure 5 and Figure 8Based on this, a multi-task model training method is further proposed. The training device can first quantize the pre-trained model, then fine-tune the quantized pre-trained model to obtain N incompletely quantized business network models. The business weight parameters in each business network model are quantized according to each business, and N fully quantized business network models are obtained. This can avoid the situation where the same quantization process has a trade-off between different businesses, avoid quantization interference between businesses, and improve the quantization effect. Furthermore, when a business adjustment occurs, the training and quantization process only trains and quantizes the data of the adjusted business, and the data of other businesses does not change. Therefore, the quantization time required is short, the data download amount during deployment is small, and the installation time is even shorter.

[0194] Figure 9 This is a flow chart of another model quantization method proposed exemplarily in the embodiment of this application. Figure 9 As shown, the electronic device may perform but is not limited to the following steps:

[0195] S901: The training device pre-trains the original network model to obtain a pre-trained model.

[0196] The execution process of S901 may refer to the relevant content of S501 and will not be described in detail.

[0197] S902: The training device quantizes the pre-trained model to obtain a pre-trained quantized model.

[0198] Among them, after the training device obtains the pre-trained model, it can quantize the pre-trained model. Specifically, the training device can quantize the pre-trained model using quantization-aware training or post-training quantization methods to obtain a pre-trained quantized model (S901 to S902 are completed by QAT), wherein the processing of QAT and PTQ can refer to the above description and will not be repeated. The pre-trained quantized model includes shared quantization weight parameters, which are weight parameters that have been trained and quantized.

[0199] Optionally, during the pre-training and quantization process of the original network model, quantization-aware training can be performed directly, that is, S901 and S902 are performed together.

[0200] Among them, in the processing process of S902, it is necessary to ensure the versatility of the quantization results. The quantized model extracts the common features of multiple businesses. After that, the shared weight parameters of the quantized pre-trained model are no longer adjusted, which can provide an accurate basis for subsequent quantization.

[0201] S903: The training device builds an original business model based on the pre-trained quantization model.

[0202] The training device builds the original business model based on the pre-trained quantization model. The training device can fine-tune the partial structure or one or more weight matrices in the pre-trained quantization model according to different fine-tuning methods based on the pre-trained quantization model. Figures 6A to 6C Build the original business model by fine-tuning the method in Figures 6A to 6C The W in the figure should be replaced by the pre-trained model that has been quantized in S902. Figures 6A to 6C The difference is that the original business model includes the shared quantized weight parameters and the initial business weight parameters. The shared quantized weight parameters in the original business model are trained and quantized parameters and remain unchanged; the initial business weight parameters are untrained and unquantized parameters and need to be trained and quantized later.

[0203] In S903, after the electronic device obtains the original service model, it can perform quantization-aware training on the original service model based on N service data to obtain N target service models. The target service model includes shared quantization weight parameters and service quantization weight parameters. The quantization-aware training process from S904 to S914 first trains the floating-point weights of the service parameters, then inserts pseudo-quantization nodes for quantization-aware training to obtain N target service models.

[0204] S904: The training device trains the original service models based on the N service data to obtain N service network models.

[0205] After the training device trains the original service model based on the N service data, the training device can train the original service model based on the N service data. The specific training process of S904 can refer to the relevant description of S503 and will not be repeated here.

[0206] After the training is completed, the shared quantization weight parameters in the business network model are quantized parameters, and the business weight parameters are trained but unquantized parameters. Since the business weight parameters have not yet been quantized, the business weight parameters of each business are quantized separately with the goal of minimizing the loss of each business. The training device quantizes the N business weight parameters in the N business network models according to the quantization data of the N businesses, and obtains N target business models; the t-th target business model in the N target business models includes the shared quantization weight parameters and the t-th business quantization weight parameters in the N target business models; one business corresponds to one target business model; t is an integer from 1 to N. The following is an explanation through the processing process of S905 to S914:

[0207] S905: The training device determines a first service network model for a first service from N service network models.

[0208] The first service is the tth service, and the tth service is the service from the 1st service to the Nth service. After executing S904, the training device can obtain the service weight parameters corresponding to the N services, that is, obtain N service weight parameters. When the service network model of the first service is first determined, the first service network model is a service network model formed by combining the first service weight parameter of the first service and the shared quantization weight parameter.

[0209] The training device determines a first service network model for the first service from the N service network models, wherein the service weight parameters in the first service network model have not yet been quantified. This application does not limit the order of determining the first service network model.

[0210] S906: The training device determines a first quantization parameter of the first service.

[0211] In S905 or S913, after the first service network model is determined, the service weight parameters in the first service network model are quantized. First, a first quantization parameter of the first service needs to be determined.

[0212] The method for determining the first quantization parameter may refer to the relevant content of S504 and will not be described in detail. It should be noted that the quantization parameters used for each service may be different or the same, and this application does not limit this.

[0213] After determining the first quantization parameter in S906 or S910 , the training device performs a first quantization target calculation process based on the first quantization parameter; the first quantization target calculation process includes the processing procedures of S907 and S908 .

[0214] S907: The training device quantizes the first service network model according to the first quantization parameter to obtain a quantized first quantized service model.

[0215] After the training model determines the first quantization parameter, a quantization mapping formula can be determined based on the first quantization parameter. The weight element of the service weight parameter in the first service network model is calculated based on the quantization mapping formula, thereby obtaining the quantized service weight parameter. The quantized service quantization weight parameter is then substituted for the pre-quantized service weight parameter in the first service network model to obtain the quantized first quantized service model.

[0216] S908: The training device calculates a first target value of the first service network model.

[0217] After the training device obtains the two first business network models before and after quantization, it can input the quantized training data into the business network model respectively to obtain an unquantized output result and a quantized output result. That is, the training device inputs the quantized data of the first business into the first business network model to obtain an unquantized output result; and inputs the quantized data of the first business into the first quantized business model to obtain a quantized output result. The training device can calculate the quantization loss value between the unquantized output result and the quantized output result or the business loss value of the quantization perception training to obtain a first target value. Among them, the first target value can refer to the relevant description of the target value in S506 and will not be repeated.

[0218] The specific description of S908 may refer to the related description of calculating the target value in S506, without limitation.

[0219] S909: The training device determines whether quantitative adjustment is required. If quantitative adjustment is required, execute S911; if not, execute S910.

[0220] The processing of S909 may refer to the relevant content of S507 and will not be described in detail.

[0221] S910: The training device adjusts a first quantization parameter of a first service.

[0222] The processing process of S910 may refer to the relevant content of S508 and will not be described in detail.

[0223] S911: The training device determines a first quantization parameter corresponding to the minimum value of the first target value as a target quantization parameter, quantizes the first service network model according to the target quantization parameter, and obtains a quantized first target service model.

[0224] The first target business model is a quantified business network model corresponding to the first business among the N businesses, rather than a business network model for all businesses.

[0225] Among them, S911 can refer to the relevant description of S509 and will not be repeated here.

[0226] S912: The training device determines whether all service network models have been quantized. If all service network models have been quantized, S914 is executed; if some service network models have not been quantized, S913 is executed.

[0227] S913: The training device determines a service network model that has not been quantized among the N service network models as a first service network model.

[0228] If it is determined in S912 that some service network models have not yet been quantized, the quantized service network models among the N service network models are no longer quantized, and the one service network model among the N service network models that has not yet been quantized is determined as the first service network model. Specifically, the training device may determine the next service network model that has not yet been quantized as the first service network model.

[0229] For example, among N services, the service weight parameters of the 1st service to the tth service have been quantized, and the service weight parameters of the t+1th service to the Nth service have not yet been quantified. The t+1th service network model corresponding to the t+1th service can be determined as the first service network model in sequence.

[0230] The processing of S906, S913 and S912 can ensure that all services are traversed and the service weight parameters of all services are quantified, thereby ensuring the integrity of the quantification.

[0231] S914: The training device determines a target model based on the quantized target business models of all businesses.

[0232] When all service network models are quantized in S912 , the training device may determine a target model based on the quantized target service models of all services.

[0233] Among them, the target model is the model after the entire multi-business neural network model is trained and quantized. After obtaining the target model, it can be deployed to the device.

[0234] Combine Figure 9 The multi-task model training process in this application is exemplified below by taking the neural network model partial weight matrix to execute the LoRA training method as an example. Figure 10 This is a schematic diagram of another multi-task model training process disclosed in an embodiment of the present application. Figure 10 As shown in the figure, the combing results of different stages are explained in order of processing time:

[0235] like Figure 10As shown, after pre-training, the shared weight parameters of the pre-trained model are obtained. The shared weight parameters W of the pre-trained model are quantized to obtain a pre-trained quantized model. The shared weight parameters w in the pre-trained quantized model are fine-tuned according to different service data. The service weight data A and B are trained to obtain N corresponding service weight parameters: A1 and B1, A2 and B2, ..., AN and BN. LoRA includes the quantized shared weight parameters w in the pre-trained model and the unquantized service weight parameters A and B. The training model is constructed by combining the N service weight parameters with W in the pre-trained model to obtain N service network models, namely, service model 1, service model 2, ..., service model N. The training model then quantizes the unquantized service weight parameters A and B in the N service network models according to different services, obtaining N quantized service network models, namely, quantized service model 1, quantized service model 2, ..., quantized service model N. The quantized weight matrix A is weight matrix a, and the quantized weight matrix B is weight matrix b. In the process of quantizing the business weight matrix of the first business, the training model calculates the quantization loss value loss(1) (target value) based on the first business network model before and after quantization, and judges whether loss(1) converges. If loss(1) converges, the quantization method of the business weight matrix in the first business network model can be determined, and then the target quantized business model 1 after quantization is determined; if loss(1) does not converge, the current quantization parameter is adjusted, LoRA1 is re-executed for quantization, and the quantization loss is calculated to judge whether loss(1) converges.

[0236] above Figure 9 and Figure 10 In the implementation method, the training device can first quantize the pre-trained model, and the quantized result can be directly used in the fine-tuning and quantization process of the business weight matrix, keeping the quantization result of the pre-trained model unchanged, so that the fine-tuning and quantization of the business weight matrix only involve the relevant training and quantization process of the business, and do not involve the adjustment of the shared weight parameters, thereby avoiding the influence of the subsequent parameter training process of each business part on each business, improving the model effect after training and quantization, and improving the model output effect. Furthermore, if a business change occurs after the model is deployed, the training and quantization process will not affect the model parameters of other businesses. The redeployment process after adjustment is simpler, the amount of data downloaded is smaller, and the installation time is also shorter.

[0237] It should be noted that the business data and quantitative data of the N businesses may be the same data, corresponding to the businesses respectively.

[0238] Figure 11 This is another flow chart of a model quantization method proposed exemplarily in the embodiment of this application. Figure 9As shown, the electronic device may perform but is not limited to the following steps:

[0239] S1101: The training device pre-trains the original network model to obtain a pre-trained model.

[0240] S1102: The training device quantizes the pre-trained model to obtain a pre-trained quantized model.

[0241] S1103: The training device builds an original business model based on the pre-trained quantization model.

[0242] Among them, S1101 to S1103 can refer to the relevant description of S901 to S903 and will not be repeated here.

[0243] In S1103, after the electronic device obtains the original business model, it can perform quantization perception training on the original business model based on N business data to obtain N target business models. The target business model includes shared quantization weight parameters and business quantization weight parameters. Among them, the quantization perception training process from S1104 to S1112 directly inserts the original business model into the pseudo-quantization node, and performs quantization perception training on the initial business weight parameters in the N original business models to obtain N target business models.

[0244] S1104: The training device determines a first quantization parameter of the first service.

[0245] In S1103 or S1111, after the first service network model is determined, the service weight parameters in the first service network model are quantized. First, a first quantization parameter of the first service needs to be determined.

[0246] The method for determining the first quantization parameter may refer to the relevant contents of S504 and S906 and will not be described in detail. It should be noted that the quantization parameters used for each service may be different or the same, and this application does not limit this.

[0247] S1105: Perform quantization perception training on the initial service weight parameters of the original service model according to the first quantization parameter to obtain a first service quantization model.

[0248] After the training model determines the first quantization parameter, a quantization mapping formula can be determined based on the first quantization parameter. The weight element of the initial service weight parameter in the original service model is calculated based on the quantization mapping formula to obtain the service quantization weight parameter. The quantized service quantization weight parameter is then used to replace the pre-quantized service weight parameter in the first service network model to obtain the quantized first service quantization model.

[0249] S1106: The training device calculates a first target value for quantitative perception training of the first service network model.

[0250] After the training device obtains the first service network model, it can calculate the first target value of the quantization-aware training. The target value can refer to the relevant description in S506 and will not be repeated here.

[0251] S1107: The training device determines whether quantitative adjustment is required.

[0252] The processing process of S1107 may refer to the relevant content of S507 and will not be described in detail.

[0253] S1108: The training device adjusts the first quantization parameter of the first service.

[0254] The processing of S1108 may refer to the relevant contents of S508 and will not be described in detail.

[0255] S1109: The training device determines the first quantization parameter corresponding to the minimum value of the first target value as the target quantization parameter, quantizes the first service network model according to the target quantization parameter, and obtains a quantized first target service model.

[0256] Among them, S1105 to S1108 can refer to the processing process of S907 to S910, without limitation.

[0257] S1110: The training device determines whether all service network models have been quantized. If all service network models have been quantized, S1112 is executed; if some service network models have not been quantized, S1111 is executed.

[0258] S1111: The training device determines a service network model that has not been quantized among the N service network models as a first service network model.

[0259] S1112: The training device determines a target model based on the quantized target business models of all businesses.

[0260] Among them, S1110 to S1112 can refer to the processing process of S912 to S914, without limitation.

[0261] Optionally, when the business data of each business performs quantitative perception training on the original business data, you can use Figure 9 The quantization-aware training method in Figure 11 The quantization perception training method in is not limited in this application.

[0262] Combine Figure 11 The multi-task model training process in this application is exemplified below by taking the neural network model partial weight matrix to execute the LoRA training method as an example. Figure 12This is another schematic diagram of a multi-task model training process disclosed in an embodiment of the present application. Figure 12 As shown in the figure, the combing results of different stages are explained in order of processing time:

[0263] like Figure 12 As shown, after pre-training, the shared weight parameters of the pre-trained model are obtained, and the shared weight parameters W of the pre-trained model are quantized to obtain a pre-trained quantized model. The shared weight parameters w in the pre-trained quantized model are fine-tuned according to different service data. The service weight data A and B are trained to obtain N corresponding service quantized weight parameters: a1 and b1, a2 and b2, ..., aN and bN. LoRA includes the quantized shared weight parameters w in the pre-trained model and the trained and quantized service weight parameters a and b. During the quantization of the service weight matrix of the tth service, the training model determines whether it has converged based on the target value loss(t) of the quantization-aware training. If loss(t) converges, the target quantization parameters of the service weight matrix in the tth service network model can be determined, and the quantized target quantized service model t can be determined. If loss(t) does not converge, the current quantization parameters are adjusted, and LoRA1 is re-executed to perform QAT fine-tuning and calculate the quantization-aware loss to determine whether loss(t) has converged. t is an integer from 1 to N.

[0264] above Figure 11 and Figure 12 In the embodiment of Figure 9 and Figure 10 Based on this, the training equipment can perform training and quantization at the same time, reducing the model training and quantization process and speeding up the training and quantization process.

[0265] Combine Figure 9 and Figure 11 In the multi-task model training and quantization process, after the multi-task neural network model training is completed, the training device can store a shared weight parameter and N business weight matrices, as well as the corresponding data structure parameter information. If the business changes (adding business or adjusting business) in the future, training can be performed in the following ways. The following describes the training and quantization methods for different situations:

[0266] Case 1: Adding new business

[0267] exist Figure 9After completing the training and quantization process and deploying the target model on the device, if a new service needs to be added, the training device can input the newly added service data into the original service model based on the original service model known in S903 to obtain the N+1 service network model. Then, the first quantization parameter of the N+1 service network model is determined, and the service weight parameters in the N+1 service network model are quantized according to the first quantization parameter. Specifically, the training device trains the first original service model based on the newly added service data to obtain the N+1 service network model; the N+1 service network model includes a shared quantization weight parameter and an N+1 service weight parameter; the training device quantizes the N+1 service weight parameter in the N+1 service network model according to the quantization data of the newly added service to obtain the N+1 target service model; the N+1 target service model includes a shared quantization weight parameter and an N+1 service quantization weight parameter. The quantized N+1 target service model is obtained. Then, the quantized service weight parameters in the N+1 target service model are deployed.

[0268] In the above process, when adding new services, the training equipment can directly train and quantize the initial service weight parameters of the new services, and continue to use the existing shared quantization weight parameters, so as not to affect the parameters of other original service models. This can improve the training and quantization efficiency and reduce the amount of data required to download and the installation time during the deployment process.

[0269] Scenario 2: Adjusting business

[0270] exist Figure 9 The training and quantization process is completed in S903. After the target model is deployed on the device, if the existing business needs to be adjusted, the training device can train the first original business model based on the business data of the Mth business on the basis of the original business model known in S903 to obtain the Mth business network model; the Mth business network model includes shared quantization weight parameters and Mth business weight parameters; the Mth business is any business among the Nth business; the training device quantizes the Mth business weight parameters in the Mth business network model according to the quantization data of the Mth business to obtain the Mth target business model; the Mth target business model includes shared quantization weight parameters and Mth business quantization weight parameters; the training device sends the second model update data to the device, and the second model update data includes the Mth business quantization weight parameters. The quantized business weight parameters in the Mth quantized business model are then deployed.

[0271] Optionally, the adjusted training data or quantized data of the Mth service is input into the deployed Mth target service model, and the Mth quantized service model is continuously trained and quantized to obtain a new Mth target service model. The new Mth target service model is then deployed to the user device.

[0272] After training is completed in the above two cases, the training device can send the changed model parameters to the user device. The following describes the deployment process of the model parameters in the above two cases:

[0273] Case 1: Adding new business

[0274] After obtaining the quantized service weight parameters, the user device can download the network model update data from the training device. The network model update data includes the quantized service weight parameters and model structure parameters for the newly added service (N+1). After the user device downloads the network model update data, the service weight parameters and shared weight parameters for the (N+1) service are included in the network model for the (N+1) service.

[0275] Scenario 2: Adjusting business

[0276] After obtaining the quantized service weight parameters, the user device can download network model update data from the training device. In this case, the network model update data includes the quantized Kth service weight parameters and model structure parameters corresponding to the adjusted service. After the user device downloads the network model update data, the downloaded Kth service weight parameters are used to replace the original service weight parameters for the service. The network model used for the Kth service includes the Kth service weight parameters and shared weight parameters.

[0277] As used in the above embodiments, the term “when…” may be interpreted to mean “if…” or “after…” or “in response to determining…” or “in response to detecting…”, depending on the context. Similarly, the phrases “upon determining…” or “if (stated condition or event) is detected” may be interpreted to mean “if determining…” or “in response to determining…” or “upon detecting (stated condition or event)” or “in response to detecting (stated condition or event)”, depending on the context.

[0278] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk).

[0279] Those skilled in the art will appreciate that all or part of the process steps in the above-described method embodiments can be implemented by a computer program instructing the relevant hardware. The program can be stored in a computer-readable storage medium, and when executed, the program can include the process steps in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A model quantization method, characterized in that: The method is applied to a training device, and the method comprises: The training device quantizes the pre-trained model to obtain a pre-trained quantized model; the pre-trained quantized model includes shared quantization weight parameters; The training device builds a first original business model based on the pre-trained quantization model; the first original business model includes the shared quantization weight parameter and the initial business weight parameter; The training device quantizes and perceives the first original business model based on N business data to obtain N target business models; the tth target business model among the N target business models includes the shared quantization weight parameter and the tth business quantization weight parameter among the N target business models; one business corresponds to one target business model; and t is an integer from 1 to N respectively.

2. The method according to claim 1, characterized in that The training device quantizes and perceives the first original service model based on the N service data to obtain N target service models, including: The training device trains the initial service weight parameters of the first original service model based on N service data to obtain N service network models; the i-th service network model among the N service network models includes the shared quantization weight parameter and the i-th service weight parameter among the N service weight parameters; one service corresponds to one service network model; N is an integer greater than or equal to 2; and i is an integer from 1 to N in sequence; The training device quantizes the N business weight parameters in the N business network models according to the quantization data of N businesses to obtain N target business models; the tth target business model in the N target business models includes the shared quantization weight parameter and the tth business quantization weight parameter in the N target business models; one business corresponds to one target business model; and t is an integer from 1 to N respectively.

3. The method according to claim 1, characterized in that The training device quantifies and perceives the initial service weight parameters of the first original service model based on the N service data, and obtains N service network models respectively, including: The training device performs quantitative perception training on the initial business weight parameters of the first original business model based on N business data, and obtains N target business models respectively.

4. The method according to claim 1 or 2, characterized in that After obtaining N target business models, the method further includes: The training device trains the first original service model based on the newly added service data to obtain an N+1th service network model; the N+1th service network model includes the shared quantization weight parameter and the N+1th service weight parameter; The training device quantizes the N+1th service weight parameter in the N+1th service network model according to the quantized data of the newly added service to obtain the N+1th target service model; the N+1th target service model includes the shared quantized weight parameter and the N+1th service quantized weight parameter; The training device sends first model update data to the using device, where the first model update data includes the N+1th service quantization weight parameter.

5. The method according to claim 1 or 2, characterized in that After obtaining N target business models, the method further includes: The training device trains the first original service model based on the service data of the Mth service to obtain an Mth service network model; the Mth service network model includes the shared quantization weight parameter and the Mth service weight parameter; the Mth service is any one of the Nth services; The training device quantizes the Mth service weight parameter in the Mth service network model according to the quantized data of the Mth service to obtain the Mth target service model; the Mth target service model includes the shared quantized weight parameter and the Mth service quantized weight parameter; The training device sends second model update data to the using device, where the second model update data includes the Mth service quantization weight parameter.

6. The method according to any one of claims 1 to 5, characterized in that The N business data are one of text data, voice data, video data, and image data.

7. The method according to claim 2, characterized in that The training device quantizes N service weight parameters in the N service network models according to the quantized data of the N services, to obtain N target service models, including: The training device determines a first quantization parameter of the tth service; The training device executes a first quantization target calculation process based on the first quantization parameter; the first quantization target calculation process includes: the training device quantizes the t-th service network model according to the first quantization parameter to obtain a quantized t-th quantized service model; the training device inputs quantized data of the t-th service into the t-th service network model to obtain an unquantized output result; the training device inputs quantized data of the t-th service into the t-th quantized service model to obtain a quantized output result, and calculates a first target value based on the unquantized output result and the quantized output result; After the training device executes the first quantitative target calculation process, it determines whether quantitative adjustment is required; if quantitative adjustment is required, the training device adjusts the first quantitative parameter of the tth business and re-executes the first quantitative target calculation process based on the adjusted first quantitative parameter; if quantitative adjustment is not required, the first quantitative parameter corresponding to the minimum value of the first target value is determined as the tth target quantitative parameter, and the tth business network model is quantized according to the tth target quantization parameter to obtain the quantized tth target business model.

8. The method according to claim 7, characterized in that The first quantization parameter includes one or more of an integer quantization scale, a zero point zero, an upper limit UperTh, and a lower limit LowTh. The training device quantizes the t-th service network model according to the first quantization parameter to obtain a quantized t-th quantized service model, including: The training device calculates a corresponding weight element Q in a tth quantized service model based on the weight element x of the tth service network model, a first quantization parameter, and a quantization mapping formula.

9. The method according to any one of claims 1 to 8, characterized in that The training device builds a first original business model based on the pre-trained quantization model, including: According to the LoRA fine-tuning method, the single weight matrix in the pre-trained quantization model is used as a basic weight W, and low-rank matrices A and B are added to obtain a first original service model; or, According to the LoRA fine-tuning method, multiple consecutive neurons in the pre-trained quantization model are used as a basic weight W, and low-rank matrices A and B are added to obtain a first original business model; or, According to the adapter fine-tuning method, the hidden layer of the pre-trained quantization model is used as the main board model W, and the adapter component is added to obtain the first original business model.

10. A model quantization method, characterized in that: The method is applied to a training device, and the method comprises: The training device builds a second original business model based on the pre-trained quantization model; the second original business model includes an initial shared weight parameter and an initial business weight parameter; The training device trains the second original service model based on N service data to obtain N service network models; the i-th service network model among the N service network models includes the shared weight parameter and the i-th service weight parameter among the N service weight parameters; one service corresponds to one service network model; N is an integer greater than or equal to 2; and i is an integer from 1 to N in sequence; The training device selects target quantization parameters based on the target values of all services in N services and determines N target business models; the target quantization parameters are used to quantize the shared weight parameters and business weight parameters in the N business network models, and the i-th target business model in the N target business models includes the shared quantization weight parameters and the t-th business quantization weight parameters; the t is an integer from 1 to N in sequence.

11. The method according to claim 10, characterized in that The training device selects target quantization parameters based on the loss values of all services in N services and determines N target service models, including: The training device determines a first quantization parameter; The training device executes a target calculation process based on the first quantization parameter; the target calculation process includes: the training device quantizes the N service network models according to the first quantization parameter to obtain N quantized service models; the training device inputs the quantized data of the N services into the N service network models respectively, and obtains N unquantized output results accordingly; the quantized data of the N services are input into the N quantized service models respectively, and obtains N quantized output results accordingly, and calculates a target value based on the N unquantized output results and the N quantized output results; After the training device executes the target calculation process, it determines whether quantitative adjustment is required; if quantitative adjustment is required, the training device adjusts the first quantization parameter and re-executes the target calculation process based on the adjusted first quantization parameter; if quantitative adjustment is not required, the first quantization parameter corresponding to the minimum value of the target value is determined as the target quantization parameter, and N business network models are quantized according to the target quantization parameter to obtain N target business models.

12. A computing device, characterized in that include: One or more processors and one or more memories; the one or more processors are coupled to the one or more memories, the one or more memories are used to store computer program code, the computer program code including computer instructions, when the one or more processors execute the computer instructions, causing the computing device to perform the method according to any one of claims 1 to 11.

13. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on a computing device, the computing device is caused to perform the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Multi-task large model fine tuning method based on adapters and low-rank adaptation

    CN116822611A

  • Multi-task training method and device, storage medium and electronic equipment

    CN116957070A

  • Model fine tuning method and device, electronic equipment and storage medium

    CN117493879A

  • Large model-based medical text information governance method and system

    CN118114718A

  • Model quantification method and device, electronic equipment, vehicle and storage medium

    CN118673996A