Method for rapidly improving big data training efficiency
By performing alternate training of parity batches on super-large data sets, the problem of long training time for super-large data sets by neural network supervised learning is solved, and the effect of reducing training time and improving training efficiency is achieved.
Patent Information
- Application Number
- CN202311778710.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, neural network supervised learning has a long training time for super-large data sets, which results in long training of models and high trial cost.
By performing alternate training of the super-large data sets with parity batches, it will not reduce the overall data volume, but also reduce the training time, reduce the training speed, and improve the training speed.
Without reducing data diversity, the time-consuming and time-consuming of neural network training is significantly reduced, training efficiency is improved, and trial cost is reduced.
Smart Images

Figure CN120198767A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural networks, and particularly relates to a method for rapidly improving the training efficiency of big data. Background Art
[0002] In the prior art, the idea of neural network supervised learning is that on a labeled data set with known answers, the results given by the model should be as close as possible to the real answers. By adjusting the parameters in the neural network to fit the training data, the model can provide prediction capabilities for unknown samples. Under the condition of ensuring the quality and distribution balance of the sample data, the scale of the sample data determines the accuracy of the neural network training results. The larger the sample data volume, the higher the accuracy. Especially for image recognition and classification tasks, such as human / object detection, etc.
[0003] However, the defect of the prior art is that if the data set is too large, the training time will inevitably become longer, and a model may take a week or even longer to run, so the cost of the experiment is relatively high.
[0004] In addition, the commonly used terms in the prior art include:
[0005] 1. Deep learning: Most deep learning methods use neural network architectures, which is why deep learning models are usually called deep neural networks. The term "deep" usually refers to the number of hidden layers in the neural network. Traditional neural networks only contain 2 to 3 hidden layers, while deep networks may contain up to 150 hidden layers.
[0006] 2. The number of training times in a neural network is the number of times that 1 batch of training images passes through the network for training once (one forward propagation + one backward propagation) during training, and the weights are updated once for each iteration; during testing, it is the number of times that 1 batch of test images passes through the network once (one forward propagation).
[0007] 3. For all training sets, after training for one epoch (of course, it can also be set by oneself), the validation set is used to test the training effect of the model. Due to the non-intersection of the training set and the validation set, the results on the validation set are of reference significance.
[0008] 4. The differences among epoch, iteration, and batchsize, which are often seen in deep learning:
[0009] (1) batchsize: Batch size. In deep learning, generally SGD training is adopted, that is, each time during training, batchsize samples are taken from the training set for training;
[0010] (2) batch: 1 batch is equal to training once with batchsize samples;
[0011] (3)Epoch: One epoch is equal to training once using all the samples in the training set.
[0012] 5. Training set: The dataset used to train the parameters in the model;
[0013] Validation set: Used to check the state and convergence of the model during training;
[0014] Test set: The test set is used to evaluate the generalization ability of the model. That is, after the model determines the hyperparameters using the validation set, the data division can refer to three principles:
[0015] 1). For a small-scale sample set (in the order of tens of thousands), the commonly used allocation ratio is 60% for the training set, 20% for the validation set, and 20% for the test set;
[0016] 2). For a large-scale sample set (more than one million), as long as the number of the validation set and the test set is sufficient. For example, if there are 1 million pieces of data, then leave 10,000 for the validation set and 10,000 for the test set. For 10 million pieces of data, also leave 10,000 for the validation set and 10,000 for the test set;
[0017] 3). The fewer the hyperparameters, or the easier it is to adjust the hyperparameters, then the proportion of the validation set can be reduced and more can be allocated to the training set.
[0018] 6. Dataset preprocessing: Preprocessing is a very useful step for training a model. It can effectively help the training process converge faster and reduce incorrect learned features. Especially for pictures, preprocessing before training is very necessary.
[0019] 7. Training mode and validation mode:
[0020] When training a model in Pytorch, add: model.train() at the front;
[0021] When testing the model, use: model.eval() at the front. Because operators such as BatchNormalize and Dropout are different in training and testing, if set incorrectly, there will be a deviation in the results. Summary of the Invention
[0022] To solve the above problems, the purpose of this application is: By performing odd-even batch alternating training on an ultra-large dataset, it can neither reduce the overall data volume nor increase the training duration of the network. Without reducing data diversity (i.e., without affecting the model accuracy), it reduces the training time consumption and improves the training speed.
[0023] Specifically, the present invention provides a method for rapidly improving the efficiency of big data training, and the method comprises the following steps:
[0024] S1, data preparation, that is, training data preparation:
[0025] This method is for training of ultra-large datasets. The ultra-large datasets include millions to tens of millions of images. In the case of taking a large ratio of the training set and the validation set, that is, the vast majority of them are used as the training set and a small part is used as the validation set, and the test set is the data collected in the actual scenario;
[0026] Among them, the dataset includes a training set, a validation set, and a test set. The training set refers to the sample set for training, and the validation set is the sample set for validating the model performance; different tasks or networks require different datasets, which can be determined according to the actual situation;
[0027] S2, preprocessing of training data: including dimensionality transformation, cropping, scaling, mean removal, and variance normalization;
[0028] S3, defining the network structure of the Model network and parameter initialization:
[0029] Defining the network structure of the neural network model is achieved by writing a class of the network model. This class defines the hierarchical structure of the neural network, including the input layer, hidden layer, and output layer, as well as the parameters and activation functions of each layer; the network structure is defined according to the task and the problem to be solved; different types of neural networks require different structures; in addition, other layers and operations can be added to meet specific requirements;
[0030] Parameter initialization can help the model converge to the global minimum more easily, accelerate the training process, and improve the generalization ability of the model;
[0031] Defining the loss function loss, defining the optimizer, setting the training parameters related to the learning rate, epoch, and batchsize;
[0032] S4, traversing the epoch, assuming the current is i;
[0033] Epoch refers to the process in which the entire training dataset is completely input into the neural network and undergoes one forward propagation and one backward propagation;
[0034] Batch refers to a small part of the training data used for parameter update during each training process. Usually, in each epoch, the training dataset is divided into multiple batches, and the neural network performs forward propagation and backward propagation on each batch, calculates the gradients, and updates the model parameters; this process is repeated multiple times until the entire dataset is traversed; there is no specific way, and this is the standard way of neural network training;
[0035] S5, Determine the even or odd epoch: If it is an even epoch, perform training on the even batches; if it is an odd epoch, perform training on the odd batches, and cross - loop; specifically, it includes:
[0036] Traverse the batches. Each time, pack according to the batch size and prepare to enter the network for training. One group is one batch size. Assume the current is the j - th batch;
[0037] If i % 2 == j % 2, start training the current batch size dataset; otherwise, jump out and perform training on the next batch size, j = j + 1;
[0038] S6, train training: Specify Model to be in the training mode model.train(); Using model.train() sets the model to the train mode but does not start the actual training;
[0039] S7, eval verification: After traversing all batches, set the model to the verification mode model.eval(), perform a forward operation using the validation set to check the training accuracy of this epoch, and also perform the operation according to the batch size;
[0040] S8, Determine whether the epoch has ended? If not, return to step S4, set i = i + 1, and start training for the next epoch; if so, end;
[0041] S9, After training all epochs, end the training of the current model, save the network with the highest training accuracy this time, and perform actual testing using the test set.
[0042] In the above - mentioned step S1, the case of the large ratio includes: 100w:1w.
[0043] In the above - mentioned step S2, during the training mode of the pre - processing, there is also data augmentation: random flipping, chromaticity transformation, and brightness transformation.
[0044] In the above - mentioned step S3, if there is a pre - trained model or fine - tuning parameters or continue to interrupt the model training, after defining the network model, it is necessary to load the pre - trained model to be used.
[0045] The above - mentioned step S5 further includes:
[0046] S5.0, Determine whether it is an even epoch, expressed as determining whether epoch % 2 == 0? If so, perform step S5.1; if not, perform step S5.3;
[0047] S5.1, Traverse the batches;
[0048] A batch refers to a small portion of training data used for parameter update during each training process. Generally, the entire training dataset is divided into multiple batches, and each batch contains a certain number of training samples; S5.2, determine whether it is an even batch. batchsize can be expressed as determining whether batchsize % 2 == 0? If not, return to step S5.1; if so, proceed to step S6; S5.3, traverse the batch, using the same method as in step S5.1;
[0049] S5.4, determine whether it is an odd batch of batchsize, expressed as determining whether batchsize % 2 != 0? If not, return to step S5.1; if so, proceed to step S6.
[0050] Therefore, the advantages of this application are as follows: This method is simple and can solve the problems of slow training speed and long training time caused by extremely large datasets (millions or even tens of millions of images) during the training of deep learning models. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention.
[0052] Figure 1 It is a flowchart of this method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] In order to more clearly understand the technical content and advantages of the present invention, the present invention will be further described in detail below with reference to the drawings.
[0054] In theory, this method can be used for anything related to neural network training. The tasks demonstrated in the embodiments of this application are mainly image recognition and classification tasks, such as human / object detection, etc.
[0055] This application proposes a method for quickly improving the training efficiency of big data, such as Figure 1 shown, including:
[0056] S1, data preparation, that is, training data preparation:
[0057] This method is mainly for the training of extremely large datasets, which include millions or even tens of millions of images. If the dataset is very small and the training time is not considered, it can be ignored. Take the case where the ratio of the training set to the validation set is extremely large, that is, the vast majority is used as the training set and a small part is used as the validation set (for example, 100w:1w). The test set is preferably data collected from the actual scene;
[0058] Among them, the dataset includes a training set, a validation set, and a test set. The training set refers to the sample set used for training, and the validation set is the sample set used to verify the model performance. Different tasks or networks require different datasets, which can be determined according to the actual situation.
[0059] S2. Preprocess the training data, including dimensional transformation, cropping / scaling, mean subtraction, and variance normalization. In the training mode, data augmentation (such as random flipping, chromaticity transformation, brightness transformation, etc.) may also be included.
[0060] S3. Define the network structure of the Model network and initialize the parameters:
[0061] Defining the network structure of a neural network model is achieved by writing a class for the network model. This class defines the hierarchical structure of the neural network, including the input layer, hidden layers, and output layer, as well as the parameters and activation functions of each layer. The network structure is defined according to the task and the problem to be solved. Different types of neural networks (such as convolutional neural networks, recurrent neural networks, Transformers, etc.) require different structures. In addition, other layers and operations can be added to meet specific requirements, such as batch normalization, pooling layers, Dropout, etc. Parameter initialization can help the model converge to the global minimum more easily, accelerate the training process, and improve the generalization ability of the model.
[0062] Define the loss function loss, define the optimizer, and set the training parameters related to the learning rate, epoch, and batch size.
[0063] If there is a pre-trained model or fine-tuning parameters or continue to interrupt the training of the model, after defining the network model, it is necessary to load the pre-trained model to be used.
[0064] S4. Iterate through the epochs. Assume the current one is i.
[0065] An epoch refers to the process in which the entire training dataset is completely input into the neural network and undergoes one forward propagation and one backward propagation.
[0066] A batch refers to a small part of the training data used for parameter update during each training process. Usually, in each epoch, the training dataset is divided into multiple batches, and the neural network performs forward propagation and backward propagation on each batch, calculates the gradients, and updates the model parameters. This process is repeated multiple times until the entire dataset is traversed. There is no specific way, and this is the standard way of neural network training.
[0067] S5. Determine the even and odd epochs: If it is an even epoch, perform the training of the even batches; if it is an odd batch, perform the training of the odd batches, with cross-circulation. Among them, it includes:
[0068] Traverse the batch. Each time, pack according to the batch size and prepare to enter network training. One group is one batch size. Assume the current is the j-th one;
[0069] If i % 2 == j % 2, start training the current batch size dataset; otherwise, jump out and proceed to the next batch size training, j = j + 1;
[0070] The step S5 further includes:
[0071] S5.0, judge whether it is an even-numbered epoch, which can be expressed as judging whether epoch % 2 == 0? If so, proceed to step S5.1; if not, proceed to step S5.3;
[0072] S5.1, traverse the batch;
[0073] Batch refers to a small part of the training data used for parameter update during each training process. Usually, the entire training dataset is divided into multiple batches, and each batch contains a certain number of training samples; S5.2, judge whether it is the even-numbered batch size, which can be expressed as judging whether batch size % 2 == 0? If not, return to step S5.1; if so, proceed to step S6; S5.3, traverse the batch, in the same way as step S5.1;
[0074] S5.4, judge whether it is the odd-numbered batch size, which can be expressed as judging whether batch size % 2!= 0? If not, return to step S5.1; if so, proceed to step S6;
[0075] S6, train training: Specify Model as the training mode model.train(); Using model.train() sets the model to the train mode, but does not start the actual training;
[0076] S7, eval verification: After traversing all batches, set the model to the verification mode model.eval(), perform a forward operation with the verification set, and check the training accuracy of this epoch (also calculated according to the batch size);
[0077] Typically, in each epoch, the training dataset is divided into multiple batches, and the neural network performs forward propagation and backward propagation on each batch, calculates the gradients, and updates the model parameters. This process is repeated multiple times until the entire dataset is traversed. At the end of each epoch, the performance of the model is usually evaluated to understand the improvement of the model. The specific method depends on the task type (classification, regression, etc.) and the dataset.
[0078] S8, determine whether the epoch is over? If not, return to step S4 (go back to step S4, perform i = i + 1, and start training for the next epoch); if so, end.
[0079] S9, after all epochs are trained, end the training of the current model, save the network with the highest accuracy in this training, and perform actual testing with the test set.
[0080] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for quickly improving the efficiency of big data training, characterized in that, The method includes the following steps: S1, Data preparation, i.e., training data preparation: This method is for training on ultra-large datasets, which include millions to tens of millions of images. In the case of a large ratio of the training set to the validation set, i.e., the vast majority is used as the training set and a small part as the validation set, and the test set is data collected from the actual scenario. S2, Preprocessing of training data: including dimensional transformation, cropping, scaling, mean removal, and variance normalization. S3, Define the network Model network structure and parameter initialization: Define the loss function loss, define the optimizer, and set training parameters such as the learning rate, epoch, and batchsize. S4, Traverse epoch, assuming the current is i; Epoch refers to the process where the entire training dataset is completely input into the neural network and undergoes one forward propagation and one backward propagation. S5, Determine even or odd epochs: If it is an even epoch, perform training on even batches; if it is an odd epoch, perform training on odd batches, with cross-cycling. Specifically, it includes: Traverse batches. Each time, pack according to batchsize and prepare to enter the network for training. One batchsize forms a group. Assume the current is the jth. If i % 2 == j % 2, start training the current batchsize dataset; otherwise, skip to the next batchsize training, and j = j + 1. S6, Train: Specify Model to be in training mode model.train(); Using model.train() sets the model to the train mode. S7, Evaluate: After traversing all batches, set the model to evaluation mode model.eval(), perform one forward operation using the validation set, and check the training accuracy of this epoch, also operating according to batchsize. S8, Determine if the epoch has ended? If not, return to step S4, set i = i + 1, and start training for the next epoch; if so, end. S9, After training all epochs, end the training of the current model, save the network with the highest training accuracy this time, and perform actual testing using the test set.
2. A method for quickly improving the efficiency of big data training according to claim 1, characterized in that, In step S1, the case of a large ratio includes: 100w:1w.
3. A method for rapidly improving the efficiency of big data training according to claim 1, characterized in that, In step S2, during the preprocessing in the training mode, there is also data augmentation: random flipping, chromaticity transformation, and brightness transformation.
4. A method for rapidly improving the efficiency of big data training according to claim 1, characterized in that In step S3, if there is a pre-trained model or fine-tuning parameters or continue interrupted model training, after defining the network model, it is necessary to load the pre-trained model to be used.
5. A method for quickly improving the efficiency of big data training according to claim 1, characterized in that, Step S5 further includes: S5.0, Determine if it is an even epoch, represented as determining if epoch % 2 == 0? If so, proceed to step S5.1; if not, proceed to step S5.
3. S5.1, Traverse batches; A batch refers to a small portion of training data used for parameter updates during each training process. Usually, the entire training dataset is divided into multiple batches, and each batch contains a certain number of training samples; S5.2, Determine whether it is an even batch size, which can be expressed as determining whether batchsize % 2 == 0? If not, return to step S5.1; if so, proceed to step S6; S5.3, Traverse the batch, using the same method as in step S5.1; S5.4, Determine whether it is an odd batch size, expressed as determining whether batchsize % 2 != 0? If not, return to step S5.1; if so, proceed to step S6.