Model training method and device

By introducing a loss function related to the penalty term and shift multiplication in neural network model training, quantizing the model parameters as the target weight, solving the problems of high computational complexity and large resource occupancy in the existing technology, and achieving efficient data processing and precision retention.

CN119940482APending Publication Date: 2025-05-06SMARTER SILICON (SHANGHAI) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311459635.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In practical applications, existing neural network models have high computational complexity, occupy a lot of computing resources and energy, and require complex multiplication operations.

Method used

A model training method is adopted to calculate the loss function value through the preset loss function, and update the model parameters according to the loss function value until the training target is met. This loss function includes the basic loss function term and the penalty term. The penalty term is related to the shift multiplication of the model parameters and is used to quantify the weight parameter as the target weight.

Benefits of technology

By quantizing the weight parameters as the target weight, it can ensure processing accuracy while improving the efficiency of the neural network model for data processing, reducing multiplication operations, and saving circuit area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940482A_ABST
    Figure CN119940482A_ABST
Patent Text Reader

Abstract

The invention provides a model training method and device, and the method comprises the steps: executing training corresponding to a neural network model, employing a preset loss function to calculate a loss function value in the training process, updating model parameters according to the loss function value until a training target is satisfied, and obtaining the model parameters; wherein the loss function comprises a basic loss function term and a penalty term, and the penalty term is related to a shift multiplication corresponding to a model parameter of the neural network; weight parameters in the model parameters can be quantized into target weights, the target weights are binary sequences comprising at least one target numerical value, and errors between the weight parameters and the numerical value of the target weights meet requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of deep learning technology, and in particular to a model training method and device. Background Art

[0002] In the application process of neural network model, data is usually processed by accelerating multiplication operation through multiplication-addition calculation array. Therefore, in the process of processing the neural network model, usually only the error between the output result of the neural network model and the real result needs to be considered.

[0003] At present, in order to reduce the computational complexity of neural network models in practical applications, model parameters are usually quantized into data types such as INT8, FP16 or BF16. Although this can reduce the amount of calculation, complex multiplication operations are still required on the data, which will also occupy most of the system's computing resources and energy. Summary of the invention

[0004] In view of this, the present application provides a model training method, comprising:

[0005] Execute training corresponding to the neural network model, during the training process, use a preset loss function to calculate a loss function value, and update the model parameters according to the loss function value until the training objectives are met, thereby obtaining the model parameters;

[0006] Among them, the loss function includes: a basic loss function term and a penalty term, the penalty term is related to the shift multiplication corresponding to the model parameters of the neural network; the weight parameters in the model parameters can be quantized into target weights, the target weights are binary sequences including at least one target value, and the error between the weight parameters and the values ​​of the target weights meets the requirements.

[0007] The above method performs the training corresponding to the neural network model including:

[0008] Execute the first stage training corresponding to the neural network model, wherein a preset initial loss function is used to calculate a first loss function value during the first stage training, until the first stage training is completed, and obtain the first model parameters;

[0009] Update the penalty term of the initial loss function, and re-execute the second stage training corresponding to the neural network model, continue to update the model parameters based on the first model parameters until the second stage training is completed, and obtain the second model parameters. The weight parameters in the second model parameters can be quantified into target weights.

[0010] In the above method, the updating of the penalty term of the initial loss function includes:

[0011] During the second stage training process, based on the initial loss function, the penalty term is gradually changed to obtain several second loss functions. The second loss function values ​​calculated based on the several second loss functions are used to update the model parameters based on the first model parameters until the training objectives are met and the above-mentioned second stage training is completed.

[0012] In the above method, the penalty term is used to characterize at least the number of digits of the target value in the target weight and the difference between the value of the target weight and its nearest power of 2.

[0013] In the above method, the penalty term is the product of the penalty term function and the penalty term coefficient. During the first stage of training, the penalty term coefficient is a preset fixed value.

[0014] In the above method, the updating of the penalty term of the initial loss function includes:

[0015] The penalty term coefficient of the penalty term is updated.

[0016] In the above method, the penalty term at least includes the sum of a first penalty term and a second penalty term, the first penalty term is the product of a first penalty term function and a second penalty term coefficient, and the second penalty term is the product of a second penalty term function and a second penalty term coefficient;

[0017] The first penalty item is used to characterize the number of digits of the target value in the target weight; the second penalty item is used to characterize the difference between the value of the target weight and its nearest power of 2; the first penalty item coefficient is used to change the number of digits of the target value in the weight represented by the first penalty item, and the second penalty item coefficient is used to change the difference between the value of the target weight represented by the second penalty item and its nearest power of 2.

[0018] In the above method, the updating of the penalty item coefficient of the penalty item includes:

[0019] Update the first penalty item coefficient and / or the second penalty item coefficient.

[0020] In the above method, the updating of the penalty item coefficient of the penalty item includes:

[0021] The penalty term coefficient is increased based on the penalty term coefficient of the initial loss function.

[0022] The above method, after completing the training of the neural network model, further includes:

[0023] Applying verification data to verify the neural network model to obtain a verification result;

[0024] If the verification result does not meet the preset verification conditions, the training process corresponding to the neural network model is re-executed until the neural network model meets the target requirements.

[0025] A model training device, comprising:

[0026] A training unit, used to perform training corresponding to the neural network model, during which a preset loss function is used to calculate a loss function value, and the model parameters are updated according to the loss function value until the training objectives are met, thereby obtaining the model parameters;

[0027] Among them, the loss function includes: a basic loss function term and a penalty term, the penalty term is related to the shift multiplication corresponding to the model parameters of the neural network; the weight parameters in the model parameters can be quantized into target weights, the target weights are binary sequences including at least one target value, and the error between the weight parameters and the values ​​of the target weights meets the requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0029] Figure 1 A method flow chart of a model training method provided in an embodiment of the present application;

[0030] Figure 2 A flowchart of another method of a model training method provided in an embodiment of the present application;

[0031] Figure 3 A flowchart of another method of a model training method provided in an embodiment of the present application;

[0032] Figure 4 A device structure diagram of a model training device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0033] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0034] In this application, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of more restrictions, the elements defined by the sentence "comprise one..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0035] The present application can be used in many general or special computing device environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multi-processor devices, distributed computing environments including any of the above devices or equipment, etc.

[0036] The embodiment of the present application provides a model training method, which can be applied to a variety of system platforms, and its execution subject can be a computer terminal or a processor of various mobile devices. The method flow chart of the method is as follows: Figure 1 As shown, specifically including:

[0037] S101: Execute training corresponding to the neural network model. During the training process, a preset loss function is used to calculate a loss function value.

[0038] The loss function includes a basic loss function and a penalty term. The basic loss function is related to the predicted value and the true value output by the neural network model during the training process, and the penalty term is related to the shift multiplication corresponding to the model parameters of the neural network model. The specific expression of the loss function is:

[0039] L(w,b)=l(w,b)+P

[0040] Among them, l(w,b) is the basic loss function, w and b are the weight parameter and bias parameter in the model parameters respectively, P is the penalty term, and the penalty term includes one or more functions, which is related to the weight parameter in the model parameters of the neural network model.

[0041] It should be noted that the penalty term is used to characterize at least the number of digits of the target value in the target weight after the weight parameter in the model parameter is quantized to the target weight and the difference between the target weight value and the power of 2 that is closest to it. Appropriate penalty terms can be configured according to the requirements for the target weight.

[0042] It should also be noted that the shift multiplication corresponding to the model parameters of the neural network model refers to the number of bits of the data input into the neural network model determined based on the weight parameters in the model parameters. The number of bits of the shift is related to the number and number of bits of the target data in the weight parameters.

[0043] In this application, the loss function is used to calculate the loss function value, which is to determine the error between the predicted value and the true value output by the neural network model according to the basic loss function in the loss function. The basic loss function can be a mean square error loss function, a cross entropy loss function, a KL divergence loss function or an exponential loss function, etc. Taking the mean square error loss function as an example, its specific expression is:

[0044]

[0045] Where n is the number of training data, y1 is the true value, and y2 is the predicted value.

[0046] S102: Determine whether the training target is met.

[0047] If the training target is not met, S103 is executed; if the training target is met, S104 is executed.

[0048] It should be noted that the training goal can be that the weight parameters in the model parameters of the neural network model can be quantified as the target weights, the loss function value calculated by the loss function is less than the first threshold, or the loss function value converges, or the number of training times in the training process reaches the target number.

[0049] S103: Update model parameters according to the loss function value.

[0050] It is understandable that when the neural network model training does not meet the training objectives, the model parameters of the neural network model are updated according to the loss function value, and the training corresponding to the neural network model as described in S101 above is continued, so that the neural network model can meet the training objectives during the training process.

[0051] It should be noted that, in the process of updating the model parameters, the weights in the model parameters are updated with reference to the relevant requirements of the target weights represented by the penalty terms.

[0052] In one embodiment, when updating the model parameters, the update can be performed according to a power of 2 close to the weight parameter in the model parameters. For example, before updating the model parameters, the weight parameter in the model parameters is 90=1011010, and the weight parameter is directly updated to 64=1000000 or 128=10000000, etc., which are powers of 2, until the loss function converges to obtain the final updated weight.

[0053] In another embodiment, when updating the model parameters, the weight parameters can be updated according to the law of gradual increase or gradual decrease based on the weight parameters in the original model parameters, so that the updated weight parameters are close to the power of 2 until the loss function of the neural network model converges. Among them, the weight parameters can be gradually increased or gradually decreased according to the change of the loss function value during the training process. For example, during the training process of the neural network model, the weight parameters are first updated in a gradually increasing manner. If the error between the predicted value and the true value represented by the loss function obtained during the training process after the weight parameters are updated increases, the weight parameters are updated in a gradually decreasing manner until the loss function converges and the error between the predicted value and the true value represented by the loss function is less than a preset value.

[0054] Optionally, when updating the model parameters, while updating according to the rule of gradual increase or gradual decrease, the weight parameters can also be updated with the goal of reducing the number of bits of the target value in the binary sequence corresponding to the weight parameters. For example, before updating the model parameters, the weight parameters in the model parameters are 90=1011010, and the weight parameters contain a 4-bit target value "1". If the update is performed according to the rule of gradual decrease, the weight parameters obtained after gradually updating the weight parameters of the model parameters can be 88=1011000, 84=1010100, 82=1010010, 80=1010000, etc., until the loss function converges and the final updated weight is obtained.

[0055] S104: Obtain model parameters.

[0056] It is understandable that during the training process, when the training can meet the goal, the current model parameters of the neural network model are obtained. The weight parameters in the model parameters can be quantified into target weights, and the target weights are binary sequences including at least one target value, and the error between the weight parameters and the target weight values ​​meets the requirements. Among them, the target value in the binary sequence is "1".

[0057] In the present application, the training corresponding to the neural network model is performed. During the training process, a preset loss function is used to calculate the loss function value, and the model parameters are updated according to the loss function value until the training objectives are met and the model parameters are obtained.

[0058] Among them, the loss function includes: a basic loss function term and a penalty term, the penalty term is related to the shift multiplication corresponding to the model parameters of the neural network; the weight parameters in the model parameters can be quantized into target weights, the target weights include at least one binary sequence of target values, and the error between the weight parameters and the target weight values ​​meets the requirements.

[0059] The present application updates the model parameters in the neural network model by applying a loss function, so that the weight parameters in the model parameters of the neural network model can be quantified into target weights. After the weight parameters are quantified into target weights, the target weights can be applied to shift the data input into the neural network model to replace the complex multiplication and addition algorithm in the original neural network model. This can improve the efficiency of data processing by the neural network model while ensuring the processing accuracy of the neural network model.

[0060] For the logic circuit corresponding to the neural network model, the original neural network model needs to deploy multiple adders to realize complex multiplication and addition operations. The more bits of the binary value corresponding to the weight parameter in the model parameter, the more adders need to be deployed. After the neural network model is trained, the weight parameters in the model parameters can be quantified into target weights, and the original multiplication and addition operations can be replaced by shifting. Therefore, the original multiple adders can be replaced by shifters in the logic circuit, which can improve the efficiency of data processing and save circuit area.

[0061] In this application, the training process corresponding to the neural network model is implemented as follows:

[0062] Execute the first stage training corresponding to the neural network model, use the preset initial loss function to calculate the first loss function value during the first stage training, until the first stage training is completed, and obtain the first model parameters;

[0063] Update the penalty term of the initial loss function, and re-execute the second stage training corresponding to the neural network model, continue to update the model parameters based on the first model parameters, until the second stage training is completed, obtain the second model parameters, and the weight parameters in the second model parameters can be quantified into target weights.

[0064] It should be noted that the penalty term in the initial loss function used in the first stage of training is in a fixed function form and is related to the weight parameters in the model parameters. During the first stage of training, the model parameters in the neural network model are adjusted according to the first loss function value calculated by the initial loss function until the first model parameters are obtained after the first stage of training is completed. Among them, the condition for adjusting the model parameters according to the first loss function value calculated by the initial loss function is that the model parameters are updated when the first loss function value does not reach the first convergence condition of the first stage of training. If the first loss function value calculated according to the initial loss function reaches the first convergence condition, the second stage of training is entered. In the second stage of training, the penalty term will change according to the change of the first loss function value calculated by the initial loss function. While updating the penalty term, the model parameters of the neural network model continue to be updated on the basis of the first model parameters in accordance with the method of updating the model parameters during the first stage of training.

[0065] Specifically, the specific implementation method of updating the penalty term of the initial loss function in the second stage is:

[0066] During the second stage of training, based on the initial loss function, the penalty term is gradually changed to obtain several second loss functions. The model parameters are updated based on the first model parameters based on the second loss function values ​​calculated based on the several second loss functions until the training objectives are met and the above-mentioned second stage of training is completed.

[0067] In the present application, the first stage training process of the neural network model is: obtaining training data and the true value corresponding to the training data; applying the neural network to process the training data to obtain the predicted value corresponding to the training data; obtaining the initial loss function according to the true value, the predicted value, the penalty term and the current weight parameter of the neural network model, wherein the penalty term in the initial loss function is obtained according to the current weight parameter. When the initial loss function does not converge or is not less than a preset threshold, the model parameters of the neural network model are adjusted according to the error between the true value and the predicted value, and the neural network is reapplied to process the training data until the initial loss function obtained after processing the training data converges, completing the first stage training.

[0068] The second stage training process of the neural network model is: based on the initial loss function, the penalty term is gradually changed. After each change of the penalty term, the training process corresponding to the first stage training is performed until the second loss function value calculated by the second loss function corresponding to the currently changed penalty term reaches the convergence condition, and the second stage training is completed.

[0069] Based on the method provided in the above embodiment, refer to Figure 2 , the training process for executing the first stage of training corresponding to the neural network model includes the following implementation steps:

[0070] S201: Obtain training data and true values ​​corresponding to the training data.

[0071] The training data may be data such as images, sounds or text. The true value of the training data is the true result obtained after processing the training data. For example, if the training data is an image, and it is necessary to annotate the person in the image, then the true person annotation information corresponding to the image is the true value of the image.

[0072] S202: Apply the neural network model to process the training data to obtain a prediction value corresponding to the processing result.

[0073] Before training, the network structure of the neural network model is initialized, and the hidden layers, number of neurons, activation functions, etc. in the neural network model are set. At the same time, after the initial neural network model is established, the initial model parameters of the neural network model are configured, where the initial model parameters include initial weight values ​​and initial bias values, etc.

[0074] It should also be noted that the specific method of applying the neural network model to process the training data is to process the training data based on the model parameters in the neural network model. Before processing the training data based on the model parameters, the initial weight value and the initial bias value may be pseudo-quantized, and the initial weight value may be a negative number.

[0075] Among them, the prediction value is the prediction result obtained after the neural network model processes the training data.

[0076] S203: Obtain a first loss function value of an initial loss function based on an error between the true value and the predicted value.

[0077] S204: Determine whether the first loss function value converges.

[0078] If yes, execute S206; if no, return to execute S205.

[0079] S205: Update model parameters and return to execute S202.

[0080] Specifically, the method of updating the model parameters may be to update the weight values ​​and bias values ​​in the model parameters according to the gradient descent method.

[0081] Optionally, after updating the model parameters, the updated weight values ​​and bias values ​​may be pseudo-quantized again to process the data based on the quantized weight values ​​and bias values.

[0082] S206: Obtain first model parameters.

[0083] It should be noted that during the training process of the first stage of training, the penalty term in the initial loss function is related to the weight parameters in the neural network model.

[0084] During the first stage of training, the first loss function value obtained by calculating the initial loss function is mainly affected by the model parameters in the neural network model. The basic loss function is used to ensure the gap between the predicted value and the true value. By updating the model parameters, the predicted value output by the neural network model after processing the training data can be close to the true value of the training data, so as to ensure the accuracy of the neural network model in processing the data.

[0085] Based on the method provided in the above embodiment, refer to Figure 3, the training process of executing the second stage training corresponding to the neural network model includes:

[0086] S301: Update the penalty item.

[0087] It should be noted that, in some embodiments, the penalty term is the product of the penalty term function and the penalty term coefficient. During the first stage of training, the penalty term coefficient in the penalty term is a preset fixed value, and the penalty term coefficient is usually set as small as possible, and can even be set to 0; during the second stage of training, the penalty term is gradually changed when the second stage of training is not completed. During the second stage of training, the updated penalty term is updated based on the penalty term of the initial loss function during the first stage of training.

[0088] Among them, the specific way to update the penalty item is to update the penalty item coefficient in the penalty item, specifically, to increase the penalty item coefficient on the basis of the penalty item coefficient in the initial loss function.

[0089] It should also be noted that the penalty item can be used to characterize the number of digits of the target value in the target weight and / or the difference between the value of the target weight and its nearest power of 2. Therefore, the penalty item can be a single penalty item or the sum of multiple penalty items. The number of penalty items can be set according to the specific weight requirements. For example, if the penalty item only characterizes the number of digits of the target value in the target weight, the penalty item can be the first penalty item; if the penalty item only characterizes the difference between the value of the target weight and its nearest power of 2, the penalty item can be the second penalty item; if the penalty item is used to characterize the number of digits of the target value in the target weight and the difference between the value of the target weight and its nearest power of 2, the penalty item is the sum of the first penalty item and the second penalty item. In addition, the penalty item may also include other penalty items for specifying requirements related to the target weight.

[0090] Similarly, the first penalty item is the product of the first penalty item function and the first penalty item coefficient, and the second penalty item is the product of the second penalty item function and the second penalty item coefficient. The first penalty item coefficient is used to change the number of digits of the target value in the weight represented by the first penalty item, and the second penalty item coefficient is used to change the difference between the value of the target weight represented by the second penalty item and its closest power of 2. Therefore, updating the penalty item coefficient of the penalty item specifically updates the first penalty item coefficient and / or the second penalty item coefficient.

[0091] Specifically, the penalty term is the sum of the first penalty term and the second penalty term. Therefore, the specific expression of the loss function is:

[0092] L(w,b)=l(w,b)+λ1P1(w)+λ2P2(w)

[0093] Among them, λ1 is the first penalty term coefficient, and P1(w) is the first penalty term function. λ2 is the second penalty term coefficient, and P2(w) is the second penalty term function. In the first stage of training, the weight parameters in the neural network model are mainly updated, so the penalty term coefficients λ1 and λ2 can be set to preset minimum values. For example, during the first stage of training, λ1 and λ2 are both 0 or other parameters less than 1. In the second stage of training, the method of updating the penalty term is specifically to update the penalty term coefficient. If the loss function contains multiple penalty terms, the penalty term coefficient of each penalty term can be adjusted separately, or the penalty term coefficients can be adjusted at the same time. In the second stage of training, the adjustment process of the penalty term can be adjusted according to the specific needs of the neural network model, which will not be repeated here.

[0094] P1(w) is a penalty function for shift multiplication, which is used to measure the number of bits of the target value required for the target weight (such as popcount function and other functions used to identify the target value in the weight). The specific form of the first penalty function is:

[0095] P1(w)=∑popcount(|Quant(w_i)|)

[0096] Among them, Quant is a function that converts floating-point numbers into fixed-point numbers, W_i is the i-th bit of the weight, and popcount is used to identify the target value "1" in each bit of each weight to obtain the number of bits with the target value "1".

[0097] λ1 is the first penalty term coefficient and is a non-negative hyperparameter used to control the influence of the first penalty term function. The larger λ1 is, the fewer the number of digits representing the target value in the target weight.

[0098] P2(w) is a penalty function for shift multiplication, which is used to calculate the difference between the weight w and its closest power of 2. The specific form of the first penalty function is:

[0099] P2(w)=∑|w_i-2 round(log2(abs(w_i))) | / scale

[0100] Among them, round means rounding to the nearest integer, log2 is the logarithm with base 2, abs means absolute value, and ∑ is the summation operation. All bit differences in the weight are summed to get the total difference. The quantization scale is related to the data type of the weight and is also a power of 2. For example, if the weight data type is int8, then scale = 2^8.

[0101] λ2 is the second penalty term coefficient and is a non-negative hyperparameter used to control the influence of the second penalty term function. The larger λ2 is, the smaller the difference between the value representing the target weight and its closest power of 2 is.

[0102] S302: Apply the neural network model to process the training data to obtain a prediction value corresponding to the processing result.

[0103] It should be noted that during the second stage of training, the model parameters in the neural network model may be pseudo-quantized in advance.

[0104] S303: Obtain a second loss function value of the second loss function based on the error between the true value and the predicted value.

[0105] It should be noted that the second loss function is the initial loss function after the penalty term is changed. Due to the change of the penalty term, the second loss function value calculated according to the second loss function also changes.

[0106] S304: Determine whether the second loss function value converges.

[0107] If yes, execute S306; if no, return to execute S305.

[0108] S305: Update the model parameters based on the model parameters updated last time, and return to execute S301.

[0109] It should be noted that the method of updating the model parameters may be to update the weight values ​​and bias values ​​in the model parameters according to the gradient descent method.

[0110] After the model parameters are updated, the updated weight values ​​and bias values ​​may be pseudo-quantized again to process the data based on the quantized weight values ​​and bias values.

[0111] S306: Obtain second model parameters.

[0112] Among them, the weight parameters in the second model parameters can be quantified into target weights.

[0113] It should also be noted that the model parameters of the neural network model may include bias parameters and activation functions in addition to weight parameters. The neural network model processes the input data through multiple network layers based on the model parameters and outputs the processing results corresponding to the data. Each network layer sets the corresponding model parameters.

[0114] In the embodiment of the present application, the model parameters of the neural network model are mainly adjusted during the first stage of training, so that when the neural network model processes data, the error between the output result and the true result is within the minimum error range. When entering the second stage of training, the penalty term is adjusted so that the result of data processing during the model training process is close to the true result, and the weight parameters of the neural network model can be quantified into target weights. When the neural network model processes data based on the target weight, it can process the data by shifting the data, further shortening the process of processing the data, ensuring the processing accuracy of the neural network model while improving the data processing efficiency of the neural network.

[0115] Based on the method provided in the above embodiment, after completing the second stage of training, the neural network model can be verified to further evaluate the performance of the neural network model. The specific verification process is as follows:

[0116] Apply the verification data to verify the neural network model and obtain the verification results;

[0117] If the verification result does not meet the preset verification conditions, the training process corresponding to the neural network model is re-executed until the neural network model meets the target requirements.

[0118] It should be noted that the verification data can be data such as images, text or audio. The verification data is input into the neural network model that has completed the second stage of training to obtain the verification result output by the neural network model. If the error between the verification process and the actual result corresponding to the verification data is greater than the verification threshold, it indicates that the verification result does not meet the verification condition. The training process corresponding to the neural network model is re-executed.

[0119] When the training process corresponding to the neural network model is re-executed, reference is made to the first stage training and the second stage training corresponding to the above S201 to S206 and S301 to S306 respectively.

[0120] Optionally, when re-executing the training process corresponding to the neural network model, the verification data can be used as training data for re-training the neural network model.

[0121] It should also be noted that, during the re-executing of the training process corresponding to the neural network model, if the number of training times of the second stage training reaches the preset maximum number, or the training process of the second stage training ends, the training of the neural network model ends.

[0122] This application uses a progressive quantization method to divide the quantization process into multiple stages, gradually reducing the quantization accuracy. During the first stage of training, the weight parameters in the model parameters are not quantized, and the original floating point numbers are maintained. After quantization, the weight values ​​close to the power of 2 are not considered first, and then the weight values ​​close to the power of 2 are considered. This can ensure that the model has sufficient expressive power in the early stage of training, thereby accelerating the convergence speed.

[0123] For the Mean Squared Error (MSE) loss function, we can adjust it to include a basic loss function and a penalty term. During the model training process, the basic loss function l(w,b) and the penalty terms P1(w) and P2(w) are optimized at the same time. In this way, the weights can be made more suitable for calculation using shift multiplication while maintaining the model performance.

[0124] The specific implementation processes and derivative methods of the above-mentioned embodiments are all within the protection scope of this application.

[0125] and Figure 1 Corresponding to the method described above, the embodiment of the present application also provides a model training device for Figure 1 The specific implementation of the method in the embodiment of the present application, the model training device provided in the embodiment of the present application can be applied to a computer terminal or various mobile devices, and its structural diagram is as follows Figure 4 As shown, specifically including:

[0126] The training unit 401 is used to perform training corresponding to the neural network model. During the training process, a preset loss function is used to calculate a loss function value, and the model parameters are updated according to the loss function value until the training objectives are met, thereby obtaining the model parameters.

[0127] Among them, the loss function includes: a basic loss function term and a penalty term, the penalty term is related to the shift multiplication corresponding to the model parameters of the neural network; the weight parameters in the model parameters can be quantized into target weights, the target weights are binary sequences including at least one target value, and the error between the weight parameters and the values ​​of the target weights meets the requirements.

[0128] In the embodiment of the present application, the training unit 401 includes:

[0129] A first training subunit is used to perform a first stage of training corresponding to the neural network model, wherein a preset initial loss function is used to calculate a first loss function value during the first stage of training until the first stage of training is completed, thereby obtaining a first model parameter;

[0130] The second training subunit is used to update the penalty term of the initial loss function and re-execute the second stage training corresponding to the neural network model, continue to update the model parameters based on the first model parameters until the second stage training is completed, and obtain the second model parameters. The weight parameters in the second model parameters can be quantized into target weights.

[0131] In the embodiment of the present application, the second training subunit updates the penalty term of the initial loss function, specifically for:

[0132] During the second stage training process, based on the initial loss function, the penalty term is gradually changed to obtain several second loss functions. The second loss function values ​​calculated based on the several second loss functions are used to update the model parameters based on the first model parameters until the training objectives are met and the above-mentioned second stage training is completed.

[0133] In the embodiment of the present application, the penalty term is used at least to characterize the number of digits of the target value in the target weight and the difference between the value of the target weight and its closest power of 2.

[0134] In the embodiment of the present application, the penalty term is the product of the penalty term function and the penalty term coefficient. During the first stage of training, the penalty term coefficient is a preset fixed value.

[0135] In the embodiment of the present application, the training unit 401 updates the penalty term of the initial loss function, specifically for:

[0136] The penalty term coefficient of the penalty term is updated.

[0137] In the embodiment of the present application, the penalty term at least includes the sum of a first penalty term and a second penalty term, the first penalty term is the product of a first penalty term function and a second penalty term coefficient, and the second penalty term is the product of a second penalty term function and a second penalty term coefficient;

[0138] The first penalty item is used to characterize the number of digits of the target value in the target weight; the second penalty item is used to characterize the difference between the value of the target weight and its nearest power of 2; the first penalty item coefficient is used to change the number of digits of the target value in the weight represented by the first penalty item, and the second penalty item coefficient is used to change the difference between the value of the target weight represented by the second penalty item and its nearest power of 2.

[0139] In the embodiment of the present application, the training unit 401 updates the penalty item coefficient of the penalty item, specifically for:

[0140] Update the first penalty item coefficient and / or the second penalty item coefficient.

[0141] In the embodiment of the present application, the training unit 401 updates the penalty item coefficient of the penalty item, specifically for:

[0142] The penalty term coefficient is increased based on the penalty term coefficient of the initial loss function.

[0143] In the embodiment of the present application, after the training of the neural network model is completed, the method further includes:

[0144] A verification unit is used to verify the neural network model using verification data to obtain a verification result; if the verification result does not meet a preset verification condition, re-execute the training process corresponding to the neural network model until the neural network model meets the target requirements.

[0145] For the specific working process of each unit and sub-unit in the model training device disclosed in the above embodiment of the present application, please refer to the corresponding content in the model training method disclosed in the above embodiment of the present application, which will not be repeated here.

[0146] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.

[0147] Those skilled in the art may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein may be implemented by electronic hardware, computer software, or a combination of both.

[0148] In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0149] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A model training method, comprising: Execute training corresponding to the neural network model, during the training process, use a preset loss function to calculate a loss function value, and update the model parameters according to the loss function value until the training objectives are met, thereby obtaining the model parameters; Among them, the loss function includes: a basic loss function term and a penalty term, the penalty term is related to the shift multiplication corresponding to the model parameters of the neural network; the weight parameters in the model parameters can be quantized into target weights, the target weights are binary sequences including at least one target value, and the error between the weight parameters and the values ​​of the target weights meets the requirements.

2. According to the method of claim 1, performing training corresponding to the neural network model comprises: Execute the first stage training corresponding to the neural network model, wherein a preset initial loss function is used to calculate a first loss function value during the first stage training, until the first stage training is completed, and obtain the first model parameters; Update the penalty term of the initial loss function, and re-execute the second stage training corresponding to the neural network model, continue to update the model parameters based on the first model parameters until the second stage training is completed, and obtain the second model parameters. The weight parameters in the second model parameters can be quantified into target weights.

3. According to the method of claim 1 or 2, the updating of the penalty term of the initial loss function comprises: During the second stage training process, based on the initial loss function, the penalty term is gradually changed to obtain several second loss functions. The second loss function values ​​calculated based on the several second loss functions are used to update the model parameters based on the first model parameters until the training objectives are met and the above-mentioned second stage training is completed.

4. According to the method of claim 1, the penalty term is used to characterize at least the number of digits of the target value in the target weight and the difference between the value of the target weight and its closest power of 2.

5. According to the method of claim 2, the penalty term is the product of the penalty term function and the penalty term coefficient, and during the first stage of training, the penalty term coefficient is a preset fixed value; The updating of the penalty term of the initial loss function comprises: The penalty term coefficient of the penalty term is updated.

6. The method according to claim 5, wherein the penalty term at least comprises the sum of a first penalty term and a second penalty term, wherein the first penalty term is the product of a first penalty term function and a second penalty term coefficient, and the second penalty term is the product of a second penalty term function and a second penalty term coefficient; The first penalty term is used to represent the number of digits of the target value in the target weight; The second penalty term is used to represent the difference between the value of the target weight and its closest power of 2; The first penalty item coefficient is used to change the number of digits of the target value in the weight represented by the first penalty item, and the second penalty item coefficient is used to change the difference between the value of the target weight represented by the second penalty item and its closest power of 2.

7. The method according to claim 6, wherein updating the penalty item coefficient of the penalty item comprises: Update the first penalty item coefficient and / or the second penalty item coefficient.

8. The method according to claim 6, wherein updating the penalty item coefficient of the penalty item comprises: The penalty term coefficient is increased based on the penalty term coefficient of the initial loss function.

9. The method according to claim 1 or 5, after completing the training of the neural network model, further comprising: Applying verification data to verify the neural network model to obtain a verification result; If the verification result does not meet the preset verification conditions, the training process corresponding to the neural network model is re-executed until the neural network model meets the target requirements.

10. A model training device, comprising: A training unit, used to perform training corresponding to the neural network model, during which a preset loss function is used to calculate a loss function value, and the model parameters are updated according to the loss function value until the training objectives are met, thereby obtaining the model parameters; Among them, the loss function includes: a basic loss function term and a penalty term, the penalty term is related to the shift multiplication corresponding to the model parameters of the neural network; the weight parameters in the model parameters can be quantized into target weights, the target weights are binary sequences including at least one target value, and the error between the weight parameters and the values ​​of the target weights meets the requirements.