A multi-category model training method, medium and device based on gradient balance

By smoothing and attenuating the sample weights according to the loss function gradient distribution in the multi-category image classification task, the problem that the existing technology is difficult to take into account both simple samples and difficult samples, and more efficient model training and better model accuracy are achieved.

CN112633359BActive Publication Date: 2025-05-06CHENGDU AITNENG ELECTRIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011509570.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-18
Publication Date
2025-05-06
Estimated Expiration
2040-12-18

AI Technical Summary

Technical Problem

The prior art is difficult to learn both simple and difficult samples in multi-category image classification tasks, resulting in poor model training results.

Method used

During the training process, the weights allocated by the sample are smoothed and attenuated according to the statistical results of the gradient distribution of the loss function corresponding to the input sample, and the impact of samples with different degrees of difficulty on the model is balanced.

Benefits of technology

This method can better explore the ability of data and models, shorten the training time of model, improve model accuracy and generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112633359B_ABST
    Figure CN112633359B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-category model training method, medium and device based on gradient balance, wherein the model training method comprises obtaining a loss function by inputting training sample data into a selected neural network model for training; statistically analyzing the gradient distribution of the loss function when training the model; allocating sample weights according to the distribution results; weight smoothing; weight attenuation; obtaining new gradients to update the network model; the present invention adjusts the weights of sample allocation according to the distribution of the loss function gradient obtained by inputting training sample data during the training process, balances the influence of samples of different degrees of difficulty on the model, shortens the model training time and improves the accuracy of the model at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence image processing technology, and in particular to a multi-category model training method, medium and equipment based on gradient balance. Background Art

[0002] With the advent of the big data era and the continuous improvement of computing power, deep learning technology driven by big data has developed rapidly. Deep learning has been widely used in many application fields such as image classification, target detection (typical applications such as face recognition, pedestrian recognition, vehicle recognition, etc.), image segmentation, etc. Behind the application of these fields are many excellent neural network structures that have emerged in recent years.

[0003] It can be said that a mature application model = a large amount of data + an excellent network structure + a suitable training method. Through a suitable training method and a large amount of data, the network model can automatically learn useful features and knowledge from the data, thereby completing the corresponding task objectives.

[0004] In recent years, a large number of public datasets have been established. For large-scale image multi-classification tasks, typical datasets include ImageNet, COCO, VOC, and OpenImage. A large number of excellent network structures have been designed and have excellent performance on these datasets.

[0005] At present, the training method of most networks is still based on the training method of stochastic gradient descent. This training method is very effective, but its potential has not been fully utilized. One of the important reasons is that samples of different difficulty levels are not used. For some tasks, too many simple samples will cover the impact of a small number of difficult samples on the model, making it difficult for the model to learn difficult samples; similarly, paying too much attention to difficult samples will also reduce the model's ability to learn simple samples. Therefore, there is currently no training strategy that can take into account the learning of simple samples and difficult samples at the same time, which also indirectly shows that the current mainstream training method needs to be further improved. Summary of the invention

[0006] In order to solve the above technical problems, the present invention provides a multi-category model training method, medium and device based on gradient balance, which inputs data into the model in batches for training. During the training process, according to the statistical results of the gradient distribution of the loss function corresponding to the input samples, the weights assigned to the samples are smoothed and attenuated to balance the influence of samples of different degrees of difficulty on the model. This method can better tap the capabilities of data and models, and under the same training data and network model, it can significantly shorten the time of model training and improve the accuracy of the trained model.

[0007] The present invention provides a multi-category model training method based on gradient balance, and the specific technical scheme is as follows:

[0008] S1: Obtain the loss function by inputting the training sample data into the selected neural network model for training;

[0009] S2: The gradient of the loss function when statistically training the model;

[0010] Train the constructed network model through forward propagation, calculate the loss function, and perform statistics on the gradient of the loss function;

[0011] S3: assign sample weights;

[0012] According to the obtained gradient distribution statistics, a corresponding weight is assigned to each sample;

[0013] S4: weight smoothing;

[0014] Smooth the weights corresponding to the obtained samples to reduce excessive weights;

[0015] S5: Attenuation processing of weight after smoothing;

[0016] During the training process, the smoothing term is attenuated so that it decays to 0 after the training.

[0017] S6: Get new gradients to update the network model;

[0018] The smoothed and attenuated weight matrix is ​​multiplied by the original loss function gradient to obtain a new gradient, and the parameters of the network model are updated through back propagation.

[0019] Furthermore, the neural network model in step S1 is a machine learning model based on the chain rule and stochastic gradient descent training method, and the training sample data is divided into several batches and input into the model.

[0020] Furthermore, in step S2, the gradient of the loss function is statistically analyzed to obtain statistical sample data, and the statistical method is to independently analyze the gradient after each model iteration.

[0021] Furthermore, interval partitioning statistics can be used for the gradient statistics of the loss function, and the gradient is evenly divided into several intervals.

[0022] Furthermore, in step S3, the weight of the sample is calculated for the obtained gradient distribution statistics, the inverse of the gradient distribution statistics is taken to obtain a distribution matrix, and the obtained distribution matrix is ​​normalized to obtain a weight matrix of the sample.

[0023] Furthermore, in step S4, the obtained weight matrix is ​​smoothed by setting the weights to a power less than 1.

[0024] Furthermore, in step S5, during the training process, the smoothing term may be decayed to 0 using an exponential decay method.

[0025] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned multi-category model training method based on gradient balance.

[0026] The present invention also provides an electronic device, comprising:

[0027] Memory, storing computer programs with algorithms;

[0028] A processor is connected to the memory data and executes the above-mentioned multi-category model training method based on gradient balance when calling the computer program.

[0029] A display is connected to the processor and the memory data, and displays an operation interaction interface related to the multi-category model training method based on gradient balance.

[0030] The beneficial effects of the present invention are as follows:

[0031] 1. This method inputs the training sample data in batches to obtain the gradient distribution of the loss function, assigns sample weights according to the gradient distribution, and reduces the weight differences between samples of different difficulty levels by smoothing and attenuating the weights, thus avoiding the influence of the increase in the number of difficult samples on the training effect. This method makes full use of the training sample data and greatly improves the generalization ability and accuracy of the trained model under the same training data set and network model, while significantly shortening the training time of the model and improving the accuracy of the model.

[0032] 2. This method achieves self-balancing of gradients and can be applied to the training of any machine learning model based on the chain rule and stochastic gradient descent training method, such as convolutional neural network models and recurrent neural network models. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a schematic flow chart of the method of the present invention;

[0034] Figure 2 It is a schematic diagram of the structure of an electronic device of the present invention. DETAILED DESCRIPTION

[0035] The following description clearly and completely describes the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0036] Example 1

[0037] An embodiment of the present invention provides a multi-category model training method based on gradient balance, which can be widely used in computer devices, such as personal computers, cloud servers, computing clusters, or other electronic devices that support executable model training.

[0038] like Figure 1 As shown, the training method comprises the following steps:

[0039] S1: The training sample data is divided into several batches and input into the selected neural network model for training to obtain the loss function;

[0040] Among them, this embodiment uses the OpenImage dataset with multiple categories and large data volume as a training sample, and the neural network model is a machine learning model based on the chain rule and stochastic gradient descent training method, such as VGG, Inception, ResNet and other convolutional neural network models.

[0041] S2: The gradient of the loss function when statistically training the model;

[0042] The constructed network model is trained by forward propagation, the loss function is calculated, and the gradient of the loss function is counted; the statistical methods of the gradient include independent statistics of the gradient after each iteration, cumulative statistics of the gradient after each iteration, and statistics of the gradient of each epoch. In this embodiment, the gradient after each model iteration is independently counted.

[0043] A variety of methods can be used to obtain the gradient distribution statistics. Different gradient distribution statistics methods are selected according to different problems. In this embodiment, a multi-label classification problem is taken as an example, so the loss function is the cross entropy loss function. The gradient is divided into multiple intervals according to size, and the distribution of the gradient in a certain interval during the training process is counted. The gradient statistics are performed separately for each category. In this way, the influence of difficult and easy samples in each category can be more specifically balanced during training. The interval division methods include average interval division, cluster interval division, segment division, and other methods.

[0044] The specific steps are as follows:

[0045] S201: After the training data set is input, the loss function is calculated and derived. The absolute value range of the input sample derivative is [0, U], where U is the upper boundary of the absolute value range of the sample derivative. The absolute value range of the derivative is evenly divided into m segments, that is, divided into interval segment.

[0046] S202: Each batch of training data samples input will obtain corresponding derivatives, and the derivative distribution will be statistically analyzed for each category separately. According to the number of derivatives in each interval, a distribution matrix K of size m×n is established for statistics. Each time a training data sample is input into the model, the distribution matrix K is updated.

[0047] S3: assign sample weights;

[0048] According to the obtained gradient distribution statistics, a corresponding weight is assigned to each sample. For category k, its distribution is:

[0049] a 1k , a 2k , …, a mk

[0050] where a ik Represents the number of derivatives of category k in the i-th interval.

[0051] According to the distribution of category k, the row vector of category k in the weight matrix W is calculated as:

[0052] [w 1k , w 2k , ..., w mk ]

[0053] where w ik represents the weight of the category k in the corresponding i, and the w ik The calculation formula is as follows:

[0054] w ik =(a 1k +a 2k +…+a mk ) / (a ik × m)

[0055] S4: weight smoothing;

[0056] The weights corresponding to the obtained samples are smoothed, and the excessive weights are reduced to make the network training smoother. The weights can be smoothed by setting a maximum value for the weights. The weights exceeding the maximum value will be forced to be set to the maximum value; the weights can also be smoothed by setting the weights to a power less than 1. In this implementation, the weight matrix obtained above is smoothed by setting the weights to a power less than 1, where the power exponent is set to 0.5;

[0057] S5: Attenuation processing of weight after smoothing;

[0058] During the training process, the smoothing term is attenuated so that the smoothing term decays to 0 after the training. i The attenuation formula is:

[0059] S i =0.5-0.5×i / Z

[0060] Where Z is the total number of training steps set for model training.

[0061] S6: Get new gradients to update the network model;

[0062] The smoothed and attenuated weight matrix Multiply the original loss function gradient, i.e. the original derivative G, to obtain a new gradient And update the parameters of the network model through back propagation.

[0063] Example 2

[0064] Based on the above-mentioned model training method, embodiment 2 of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the multi-category model training method based on gradient balance.

[0065] Example 3

[0066] Based on the above model training method, embodiment 3 of the present invention provides an electronic device, such as Figure 2 As shown, the device includes a memory, a processor and a display;

[0067] The memory stores a computer program having an algorithm;

[0068] The processor is connected to the memory data, and executes the above-mentioned multi-category model training method based on gradient balance when calling the computer program.

[0069] The display is data-connected to the processor and the memory, and displays an operation interaction interface related to the multi-category model training method based on gradient balance.

[0070] The present invention is not limited to the above-mentioned specific embodiments, but extends to any new features or any new combination disclosed in this specification, as well as any new method or process steps or any new combination disclosed.

Claims

1. A multi-category model training method based on gradient balance, characterized in that: The method comprises the following steps: S1: Obtain the OpenImage dataset as a training sample, input the training sample data into the selected neural network model for training to obtain the loss function; S2: The gradient distribution of the loss function when statistically training the model; Train the constructed network model through forward propagation, calculate the loss function, and perform statistics on the gradient of the loss function; The specific steps are as follows: S201: After the training data set is input, the loss function is calculated and derived. The absolute value range of the input sample derivative is [0, U], where U is the upper boundary of the absolute value range of the sample derivative. The absolute value range of the derivative is evenly divided into m segments, that is, divided into The interval segment; S202: Each batch of input training data samples will obtain corresponding derivatives, and the derivative distribution is separately counted for each category. According to the number of derivatives in each interval, a distribution matrix K of size m×n is established for statistics. Each time a training data sample is input into the model, the distribution matrix K is updated; S3: assign sample weights; According to the obtained gradient distribution statistics, a corresponding weight is assigned to each sample. For category k, its distribution is: a 1k ,a 2k ,…,a mk where a ik represents the number of derivatives of category k in the i-th interval; According to the distribution of category k, the row vector of category k in the weight matrix W is calculated as: [In 1k ,In 2k ,…,In mk ] where w ik represents the weight of the category k in the corresponding i, and the w ik The calculation formula is as follows: w ik =(a 1k +a 2k +…+a mk ) / (a ik ×m) S4: weight smoothing; Smooth the weights corresponding to the obtained samples to reduce excessive weights; S5: Attenuation processing of weight after smoothing; During the training process, the smoothing term is attenuated so that the smoothing term decays to 0 after the training. i The attenuation formula is: S i =0.5-0.5×i / Z Where Z is the total number of training steps set for model training; S6: Get new gradients to update the network model; The smoothed and attenuated weight matrix is ​​multiplied by the original loss function gradient to obtain a new gradient, and the parameters of the network model are updated through back propagation.

2. The multi-category model training method based on gradient balance according to claim 1, characterized in that: The neural network model described in step S1 is a machine learning model based on the chain rule and stochastic gradient descent training method, and the training sample data is divided into several batches and input into the model.

3. The multi-category model training method based on gradient balance according to claim 1, characterized in that: In step S2, the gradient of the loss function is statistically analyzed to obtain statistical sample data, and the statistical method is to independently analyze the gradient after each model iteration.

4. The multi-category model training method based on gradient balance according to claim 2, characterized in that: The statistical method that can be used for the gradient statistics of the loss function is interval division statistics, which divides the gradient into several intervals on average.

5. The multi-category model training method based on gradient balance according to claim 1, characterized in that: In step S3, the weight of the sample is calculated for the obtained gradient distribution statistics, the inverse of the gradient distribution statistics is taken to obtain a distribution matrix, and the obtained distribution matrix is ​​normalized to obtain a weight matrix of the sample.

6. The multi-category model training method based on gradient balance according to claim 1, characterized in that: In step S4, the obtained weight matrix is ​​smoothed by setting the weight to a maximum value or setting the weight to a power less than 1.

7. The multi-category model training method based on gradient balance according to claim 1, characterized in that: In step S5, during the training process, the smoothing term may be decayed to 0 by uniform decay, exponential decay or cosine decay.

8. A computer-readable storage medium storing a computer program, characterized in that: When the program is executed by a processor, the multi-category model training method based on gradient balance described in any one of claims 1 to 7 is implemented.

9. An electronic device, characterized in that: The electronic device comprises: Memory, storing computer programs with algorithms; A processor, connected to the memory data, and executing the multi-category model training method based on gradient balance according to any one of claims 1 to 7 when calling the computer program; A display is connected to the processor and the memory data, and displays an operation interaction interface related to the multi-category model training method based on gradient balance.

Citation Information

Patent Citations

  • Machine learning model training method and device

    CN107784312A

  • Live-action riding training method based on mixed reality technology

    CN110490978A