Method for effectively improving precision of neural network model
By introducing a backup model in neural network training and using sliding average update parameters, the parameter jitter problem caused by stochastic gradient descent method is solved, a more stable and efficient training process is achieved, and the accuracy and robustness of the neural network model are improved.
Patent Information
- Application Number
- CN202311778114.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, the stochastic gradient descent method adds random factors when calculating the gradient, resulting in jitter in parameter updates, which easily falls into local optimality, and it is difficult to improve the accuracy of the neural network model.
Based on the original training process, a backup model Model2 is added, and the parameters of Model2 are updated in a sliding average manner. Model2 is used to smooth the model update process and avoid local optimal traps.
By sliding average update of model parameters, a smoother update path is achieved, avoiding the model from falling into local minimum values, improving the efficiency and stability of training, and improving the robustness and accuracy of the neural network model.
Smart Images

Figure CN120197667A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural networks, and particularly relates to a method for effectively improving the accuracy of neural network models. Background Art
[0002] With the development of artificial intelligence technology, neural network models increasingly require higher accuracy support. Especially for image recognition and classification tasks, such as human / object detection, etc., the requirements for neural network training are also getting higher and higher. In the prior art, the process of neural network training can be divided into three steps:
[0003] 1. Define the structure of the neural network and the output result of forward propagation;
[0004] 2. Define the loss function and select the algorithm for backpropagation optimization;
[0005] 3. Generate a session and repeatedly run the backpropagation optimization algorithm on the training data;
[0006] The backpropagation algorithm implements an iterative process. At the beginning of each iteration, a part of the training data is taken first, and the prediction result of the neural network is obtained through the forward propagation algorithm. Since the training data all have correct answers, the gap between the prediction result and the correct answer can be calculated. Based on this gap, the backpropagation algorithm will correspondingly update the values of the neural network parameters to make them closer to the real answer.
[0007] As Figure 1 shown, the method specifically includes: initializing variables, number of training times = 0; selecting a part of the training data; obtaining the predicted value through forward propagation; updating the variables through backpropagation; determining whether the training target is reached? If so, end the training; if not, determine whether the number of training times is reached? If so, end the training; if not, the number of training times + 1,; further return to the step of selecting a part of the training data and repeat.
[0008] However, the defect of the prior art lies in that: the stochastic gradient descent method adds a random factor when calculating the gradient, there is jitter in the update of the parameters, and it is easy to fall into the local optimum. The stochastic gradient descent method (SGD) is a variant of the gradient descent algorithm, used to find the optimal value of the model parameters to minimize the loss function. Different from the traditional gradient descent algorithm, the stochastic gradient descent method does not calculate the gradient of the loss function using the entire dataset at each step. Instead, it randomly selects a small batch of data samples at each step to calculate the gradient, and uses the gradient of this small batch to update the model parameters. This process is called randomness because the selection of each small batch is random.
[0009] In addition, commonly used terms in the prior art include:
[0010] 1. Deep learning: Most deep learning methods use neural network architectures, which is why deep learning models are usually referred to as deep neural networks. The term "deep" generally refers to the number of hidden layers in a neural network. Traditional neural networks only contain 2 to 3 hidden layers, while deep networks may contain up to 150 hidden layers.
[0011] 2. Neural network model training, the learning of a neural network, that is, the training process, refers to the process in which the neurons in the input layer receive input information, transmit it to the neurons in the middle layer, and finally transmit it to the neurons in the output layer, and the output layer outputs the processing result of the information. Deep learning models are trained by using a large amount of labeled data, while neural network architectures directly learn features from the data without manual feature extraction.
[0012] 3. Training set: It is a set of samples used for training, mainly used to train the parameters in a neural network. Validation set: A set of samples used to verify the model performance. After different neural networks are trained on the training set, the validation set is used to compare and judge the performance of each model. Here, different models mainly refer to neural networks corresponding to different hyperparameters, or neural networks with completely different structures. Test set: For a trained neural network, the test set is used to objectively evaluate the performance of the neural network. Extended information: Time series data is data collected at different times and is used to describe the situation of the phenomenon changing over time.
[0013] 4. Epoch traversal: The process of passing through all the data samples in the dataset through a neural network once and returning once is called one training, that is, one epoch. Because passing through the complete dataset once in a neural network is not enough, and the complete dataset needs to be passed through the same neural network multiple times. Therefore, multiple epochs need to be set.
[0014] 5. Batch: Translated into Chinese as batch (the batch in batches). When training a neural network model, for example, if there are 1000 samples and these samples are divided into 10 batches, it is 10 batches. The size of each batch is 100, that is, batchsize = 100. Each time the model is trained and the weights are updated, a batch of samples is taken to update the weights. Summary of the Invention
[0015] To solve the above problems, the purpose of this application is to: by adding a backup model2 on the basis of the original training process, and updating this model2 in the way of moving average while training model. In the practical process, model2 is mostly better than model, and there are also significant differences in individual networks, showing better performance.
[0016] Specifically, the present invention provides a method for effectively improving the accuracy of a neural network model, and the method includes the following steps:
[0017] S1. Data preparation, that is, preparation of the data set required for the training task:
[0018] The data set includes a training set, a validation set and a test set.
[0019] The training set is a sample set for training.
[0020] The validation set is a sample set for validating the model performance.
[0021] After different neural networks are trained on the training set, the validation set is used to compare and judge the performance of each model.
[0022] For the trained neural network, the test set is used to objectively evaluate the performance of the neural network.
[0023] S2. Data preprocessing before training. The specific method of neural network data preprocessing depends on the specific requirements of the training data, task and model, including cropping, scaling, color transformation, data augmentation.
[0024] S3. Define the network structure Model and set it to the train mode.
[0025] S4. Copy out a backup network Model2 at the same time.
[0026] Model2 = Model.clone(), because the model is to be trained on the graphics card, and the directly copied Model2 is also on the graphics card.
[0027] S5. Traverse from 0 for epochs:
[0028] An epoch is a process in which the entire training data set is completely input into the neural network and undergoes one forward propagation and one backward propagation; a batch refers to a small part of the training data used for parameter update during each training process; in each epoch, the training data set is divided into multiple batches, and the neural network performs forward propagation and backward propagation on each batch, calculates the gradients and updates the model parameters; this process is repeated multiple times until the entire data set is traversed.
[0029] S6. Traverse the batch starting from 0:
[0030] As shown in the following step S7, one sample of a batch is run in one round, and the next round is carried out after the end; S7. The Model performs forward network calculation, which is set to the train mode. After one round (i.e., one batch) ends, gradient backpropagation will be carried out and the parameters of the Model will be updated; The forward network calculation belongs to forward propagation, which is a key step in the convolutional neural network CNN. It is used to pass the input data into the network. After a series of calculations and activation function processing, the output is finally generated; The parameters of the CNN include the weights and biases of the convolutional kernels;
[0031] S8. At the same time, the parameters of Model2 are updated according to the following formula:
[0032] Model2 = Model2 * alpha + Model * (1 - alpha),
[0033] That is, Model2 is updated in the way of exponential moving average. alpha is reserved for history and (1 - alpha) is used currently; The above formula is the exponential moving average, which is the same as that of Model. If it is CPU training, it is on the CPU. If it is GPU training, it is on the GPU;
[0034] S9. Both Model and Model2 are set to the eval mode to test the validation set, and the validation results acc1 and acc2 are obtained respectively;
[0035] The eval mode means that the current network enters the validation mode, which is opposite to the previous train mode;
[0036] The process of obtaining the validation results is to perform the forward propagation operation of the model using the validation dataset to get the predicted output of the model. These predicted outputs will be compared with the true labels to evaluate the performance of the model;
[0037] S10. Determine whether the epoch has ended? If not, return to step S6, that is, repeat steps S6 - S9 for each batch until the current epoch ends and the next round starts; If so, after all epochs end, compare the larger value of the highest accuracy of Model and Model2 with the highest value of the accuracy of the previous epoch, and keep the larger one each time. The finally left one is the largest value. If acc1 > acc2, then keep the parameters of Model as the final best.pt, otherwise keep the parameters of Model2 as the final best.pt.
[0038] In step S2, it also includes image cropping, image edge padding, image resizing, color gamut conversion, channel data exchange, mean subtraction, and coefficient multiplication.
[0039] In step S3, the network structure of the defined neural network model is implemented by writing a class of the network model. This class defines the hierarchical structure of the neural network, including the input layer, hidden layer, and output layer, as well as the parameters and activation functions of each layer; the network structure is defined according to the task and the problem to be solved; different types of neural networks (such as convolutional neural networks, recurrent neural networks, Transformers, etc.) require different structures; in addition, if there is a pre-trained model, it is loaded after the network structure Model is defined. In a neural network, loading a pre-trained model means using the learned features and weights, and then fine-tuning on the basis of this model or using it as the basis for a specific task to accelerate and improve the performance of the model.
[0040] In step S4, it can also run on a CPU without a graphics card, and in this case, it is placed on the CPU.
[0041] In step S8, the historical retention alpha is adjustable, and the empirical value is 99%.
[0042] Therefore, the advantages of this application are as follows: the method is simple. Using the method of the present invention presents a smooth transformation method. A smoother update means that it will not deviate far from the optimal point. A smoother update path helps to avoid the model falling into local minima. The final network model has more or less a connection with the past, improving the training efficiency and stability, being beneficial to enhancing the robustness of the neural network model, and to a certain extent improving the accuracy of the model. Thus, it further improves the efficiency of the neural network model deployed on a chip or system, especially for image recognition and classification tasks, such as human / object detection and other projects. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention.
[0044] Figure 1 It is a schematic diagram of the neural network training process in the prior art.
[0045] Figure 2 It is a schematic flow diagram of this method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] In order to more clearly understand the technical content and advantages of the present invention, the present invention will be further described in detail below with reference to the drawings.
[0047] In theory, any method related to neural network training can be applied here. The tasks I have done personally in this paper mainly include image recognition and classification tasks, such as face detection, face recognition, etc.
[0048] This application proposes a method for quickly improving the training efficiency of big data, such as Figure 2 shown
[0049] including the following steps:
[0050] S1. Data preparation, that is, preparing the dataset required for the training task:
[0051] The dataset includes a training set, a validation set, and a test set.
[0052] The training set is a set of samples used for training.
[0053] The validation set is a set of samples used to verify the model performance.
[0054] After different neural networks are trained on the training set, the validation set is used to compare and judge the performance of each model.
[0055] For the trained neural network, the test set is used to objectively evaluate the performance of the neural network.
[0056] S2. Data preprocessing before training. The specific method of neural network data preprocessing depends on the specific requirements of the training data, task, and model, including cropping, scaling, color transformation, data augmentation; it also includes image cropping (crop), image edge padding (padding), image resizing (resize), color gamut conversion, channel data exchange, mean subtraction, and coefficient multiplication.
[0057] S3. Define the network structure Model and set it to the train mode.
[0058] The network structure of the defined neural network model is implemented by writing a class for the network model. This class defines the hierarchical structure of the neural network, including the input layer, hidden layer, and output layer, as well as the parameters and activation functions of each layer; the network structure is defined according to the task and the problem to be solved; different types of neural networks (such as convolutional neural networks, recurrent neural networks, Transformers, etc.) require different structures; in addition, if there is a pre-trained model, it is loaded after the network structure Model is defined. In a neural network, loading a pre-trained model means using the learned features and weights, and then fine-tuning on the basis of this model or using it as the basis for a specific task to accelerate and improve the performance of the model.
[0059] S4. Copy out a backup network Model2 at the same time.
[0060] Model2 = Model.clone(). Since the model is trained on the graphics card, the directly copied Model2 is also on the graphics card; it can also run on a CPU without a graphics card, in which case it is placed on the CPU.
[0061] S5. Traverse from 0 for epochs:
[0062] An epoch is the process in which the entire training dataset is completely input into the neural network and undergoes one forward propagation and one backward propagation; a batch refers to a small part of the training data used for parameter update during each training process; in each epoch, the training dataset is divided into multiple batches, and the neural network performs forward propagation and backward propagation on each batch, calculates the gradients, and updates the model parameters; this process is repeated multiple times until the entire dataset is traversed.
[0063] S6. Traverse from 0 for batches:
[0064] As shown in the following step S7, one sample of a batch is run in one round, and then the next round is carried out; S7. The Model performs forward network calculation, set to the train mode. After one round, that is, after one batch, gradient backpropagation will be carried out and the parameters of the model Model will be updated; the forward network calculation belongs to forward propagation, which is a key step in the convolutional neural network CNN. It is used to pass the input data into the network, and after a series of calculations and activation function processing, the output is finally generated; the parameters of the CNN include the weights and biases of the convolutional kernels.
[0065] S8. At the same time, update the parameters of Model2 according to the following formula:
[0066] Model2 = Model2 * alpha + Model * (1 - alpha),
[0067] That is, the parameters of Model2 are updated in a moving average manner, with the historical value retaining alpha and the current value using (1 - alpha); the above formula is the moving average, which is the same as Model. If it is trained on the CPU, it is on the CPU, and if it is trained on the GPU, it is on the GPU; the historical retention alpha is adjustable, and the empirical value is 99%.
[0068] S9. Both Model and Model2 are set to the eval mode to test the validation set, and the validation results acc1 and acc2 are obtained respectively.
[0069] The eval mode means that the current network enters the validation mode, which is opposite to the previous train mode.
[0070] The process of obtaining the verification result is to perform the forward propagation operation of the model using the verification dataset to obtain the predicted output of the model, and these predicted outputs will be compared with the true labels to evaluate the performance of the model;
[0071] S10. Determine whether the epoch has ended? If not, return to step S6, that is, repeat steps S6 - S9 for each batch until the current epoch ends and the next round begins; if so, after all epochs end, compare the larger value of the highest accuracy of Model and Model2 with the highest value of the accuracy of the previous epoch, and retain the larger one each time. The final remaining value is the maximum value. If acc1 > acc2, then retain the parameters of Model as the final best.pt; otherwise, retain the parameters of Model2 as the final best.pt.
[0072] Among them, the description is as follows: Figure 2 The left dotted box 1 in the figure is the original neural network training process, and the right dotted box 2 is the added process based on the left to obtain a better network effect.
[0073] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for effectively improving the accuracy of a neural network model, characterized in that, The method includes the following steps: S1. Data preparation, i.e., preparing the dataset required for the training task: The dataset includes a training set, a validation set, and a test set. The training set is a set of samples for training. The validation set is a set of samples for validating the model performance. After different neural networks finish training on the training set, the validation set is used to compare and judge the performance of each model. For the trained neural network, the test set is used to objectively evaluate the performance of the neural network. S2. Data preprocessing before training. The specific method of neural network data preprocessing depends on the specific requirements of the training data, task, and model, including cropping, scaling, color transformation, and data augmentation. S3. Define the network structure Model and set it to the train mode. S4. Copy out a backup network Model2 at the same time. Model2 = Model.clone(), because the model is to be trained on the graphics card, and the directly copied Model2 is also on the graphics card. S5. Traverse from 0 for epochs: An epoch is a process in which the entire training dataset is completely input into the neural network and undergoes one forward propagation and one backward propagation. A batch refers to a small part of the training data used for parameter update during each training process. In each epoch, the training dataset is divided into multiple batches, and the neural network performs forward propagation and backward propagation on each batch, calculates the gradients, and updates the model parameters. This process is repeated multiple times until the entire dataset is traversed. S6. Traverse from 0 for batches: As shown in step S7 below, one sample of a batch is run in one round, and then the next round is carried out after it ends. S7. Model performs forward network calculation and is set to the train mode. After one round (i.e., one batch) ends, gradient backpropagation will be carried out and the parameters of model Model will be updated. The forward network calculation belongs to forward propagation, which is a key step in the convolutional neural network CNN. It is used to pass the input data into the network, and after a series of calculations and activation function processing, the output is finally generated. The parameters of the CNN include the weights and biases of the convolutional kernels. S8. At the same time, update the parameters of Model2 according to the following formula: Model2 = Model2 * alpha + Model * (1 - alpha), That is, Model2 is updated in the way of exponential moving average. The history retains alpha, and the current uses (1 - alpha). The above formula is the exponential moving average, and it is the same as Model. If it is CPU training, it is on the CPU. If it is GPU training, it is on the GPU. S9. Both Model and Model2 are set to the eval mode to test the validation set, and the validation results acc1 and acc2 are obtained respectively. The eval mode means that the current network enters the validation mode, which is opposite to the previous train mode. The process of obtaining the validation results is to perform the forward propagation operation of the model using the validation dataset to get the predicted output of the model, and these predicted outputs will be compared with the true labels to evaluate the performance of the model. S10. Determine whether the epoch has ended? If not, return to step S6, that is, repeat steps S6 - S9 for each batch until the current epoch ends and the next round begins; if so, after all epochs end, compare the larger value of the highest accuracy of Model and Model2 with the highest value of the accuracy of the previous epoch, and retain the larger one each time. The final remaining value is the maximum value. If acc1 > acc2, then retain the parameters of Model as the final best.pt, otherwise retain the parameters of Model2 as the final best.pt.
2. A method for effectively improving the accuracy of a neural network model according to claim 1, characterized in that, In the step S2, it also includes image cropping (crop), image edge padding (padding), image resizing (resize), color gamut conversion, channel data exchange, mean subtraction, and coefficient multiplication.
3. A method for effectively improving the accuracy of a neural network model according to claim 1, characterized in that, In the step S3, the network structure of the defined neural network model is implemented by writing a class for the network model. This class defines the hierarchical structure of the neural network, including the input layer, hidden layer, and output layer, as well as the parameters and activation functions of each layer; define the network structure according to the task and the problem to be solved; different types of neural networks require different structures; in addition, if there is a pre - trained model, it is loaded after the network structure Model is defined. In a neural network, loading a pre - trained model means using the learned features and weights, and then fine - tuning on the basis of this model or using it as the basis for a specific task to accelerate and improve the performance of the model.
4. A method for effectively improving the accuracy of a neural network model according to claim 1, characterized in that In the step S4, it can also run on a CPU without a graphics card, and in this case, it is placed on the CPU.
5. A method for effectively improving the accuracy of a neural network model according to claim 1, characterized in that In the step S8, the historical retention alpha is adjustable, and the empirical value is 99%.