Training method of Transform based on jump-out local minimum
By constructing and optimizing the point θ2 near the local minimum value points in the parameter space of the deep neural network, the problem of easily falling into local minimum values in deep neural network training is solved, and a higher image classification accuracy is achieved.
Patent Information
- Application Number
- CN202510142010.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
Deep neural networks are prone to fall into local minimum values during training, which affects the final training effect, especially in the field of image classification. How to design effective methods to jump out of local minimum values to improve classification accuracy is an important issue.
Through the method of neuronal division, a point θ1 near the local minimum value point is constructed in the parameter space, and another point θ2 with equal training losses is constructed in θ1, and finally θ2 is further optimized to jump out of the local minimum value point.
The process of jumping out of local minimum points is realized, lower training losses and higher classification accuracy are obtained, and the performance of Transformer in image classification tasks is improved.
Smart Images

Figure CN120071083A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence, in the direction of deep learning, and has invented a training method for Transformer based on jumping out of local minima, which is applied in the field of image classification. Background Art
[0002] In recent years, deep learning has increasingly become an important branch and main force in the research of machine learning and artificial intelligence. Deep learning is widely used in daily life, promoting the rapid development of fields such as image classification.
[0003] In deep learning, as an important network architecture of deep neural networks, how to optimize the parameter values of the Transformer network and improve the classification accuracy of images is an important issue currently studied in the field of deep learning. The present invention can be applied in the field of image classification, aiming to improve the classification accuracy of images.
[0004] During the training process of Transformer, methods of gradient descent such as SGD, Adam and other optimizers are usually selected to optimize and train the objective function. The gradient descent method uses the gradient of the loss of the objective function with respect to the parameters to find the direction in which the loss drops fastest as the direction for optimizing the network parameters, making the network step by step towards a local minimum. For convex optimization problems, the gradient descent method can ensure that it will eventually converge to the global minimum. However, when facing highly non-convex optimization problems such as deep neural networks, the gradient descent method is very likely to make the training fall into a local minimum, affecting the final training effect. How to design an effective method to jump out of the local minimum so that a Transformer with a smaller loss value and more accurate classification can be obtained is a topic that should be deeply studied.
[0005] The present invention proposes an effective training method for Transformer based on jumping out of local minima. By the method of neuron splitting, the model jumps out of the local minimum, enabling the neural network to continue training and improving the training accuracy of the neural network. Finally, a Transformer with a higher classification accuracy is obtained, which can classify images more accurately. Summary of the Invention
[0006] The problem to be solved by the present invention is to design a training method for Transformer based on jumping out of local minima. In response to this problem, the present invention first uses the Adam optimizer widely used in deep learning to train a randomly initialized Transformer. When the training loss no longer decreases and reaches the convergence state, the obtained Transformer at this time is a local minimum point in the parameter space, and this local minimum point is denoted as θ *. Then construct the local minimum point θ in the parameter space * for a point θ near it 1 , θ 1 is close to the point θ of the trained Transformer * but the parameter values have changed somewhat. Then, using θ 1 as a basis, construct another point θ 1 that has the same training loss as θ 2 , θ 2 has the same structure as the point θ in the Transformer 1 , although the parameter values have changed but it has the same training loss value. Continuing to train θ 2 can reduce the training loss of θ 2 to a level lower than the training loss of the local minimum point θ * in the previously obtained Transformer. Eventually, the training loss no longer decreases and converges to which represents successfully jumping out of the local minimum point θ * , and obtaining a Transformer with better classification performance than θ * .
[0007] The present invention is generally divided into four major parts:
[0008] (1) First, train the randomly initialized Transformer. When the training loss converges, obtain a local minimum point θ * of the Transformer in the parameter space.
[0009] (2) Construct a point near the local minimum point θ * in the parameter space. This point has the same structure as the point θ * but the parameter values have changed somewhat. Denote this point as θ 1 .
[0010] (3) Then construct another point in the parameter space that has the same training loss as the constructed point θ 1 , that is, construct another point θ 1 whose training loss value does not change although the parameter values have changed somewhat compared to θ 2 .
[0011] (4) Finally, further train and optimize θ 2 to achieve jumping out of the local minimum point θ * , and obtain a Transformer with a lower training loss and better classification performance compared to θ * .
[0012] The specific technical solution of the method proposed by the present invention is as follows:
[0013] 1. Use the Adam optimizer widely used in deep learning to train the randomly initialized Transformer. When the training loss converges, that is, the training loss no longer decreases, the obtained Transformer at this time is a local minimum point in the parameter space, and this local minimum point is denoted as θ * .
[0014] 2. Select the Feed-Forward fully connected module in the last layer module in the optimized θ * to change its weight parameters to construct a point near the parameter space of θ * , denoted as θ 1 .
[0015] 3. Then continue to change the weight parameters in the Feed-Forward fully connected module in the last layer of θ 1 to construct another point in the parameter space with the same training loss as the constructed point θ 1 , that is, to construct another point whose training loss value does not change although the parameter values have changed compared with θ. This point is denoted as θ 2 .
[0016] 4. Finally, further train and optimize θ 2 so that the training loss can be reduced to a lower level than the training loss of the local minimum point θ * in the Transformer obtained from the previous training, indicating that it has successfully jumped out of the local minimum point θ * , and a Transformer with a better classification effect than θ * is obtained, denoted as
[0017] The present invention designs a training method for Transformer based on jumping out of local minima, which helps to improve the situation of getting stuck in local minima in the process of training Transformer caused by gradient descent optimization algorithms such as SGD and Adam, and a Transformer with smaller loss and higher classification accuracy can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a schematic diagram of the Feed-Forward fully connected module in the Transformer
[0019] Figure 2 is a training loss landscape diagram near θ * in the parameter space
[0020] Figure 3 is the structure θ 1 schematic diagram
[0021] Figure 4 is the structure θ 2 schematic diagram
[0022] Figure 5 is from θ * to θ 2 visualization curve of training loss in the direction Specific implementation manner
[0023] This method is mainly divided into four parts. First, use the Adam optimizer widely used in deep learning to train the randomly initialized Transformer. When the training loss no longer decreases, a local minimum point of the Transformer in the parameter space is obtained at this time, and this local minimum point is denoted as θ * . Then, construct a point near the local minimum point θ * in the parameter space, denoted as θ 1 . Next, construct another point θ 1 in the parameter space that has the same training loss as θ 2 , θ 2 and θ 1 have the same structure but the parameter values are changed and the training loss values are equal. Finally, continue to optimize θ 2 , which can reduce the training loss to a lower level than that of θ * , indicating successful jumping out of the local minimum and obtaining a better network denoted as
[0024] The specific process is as follows:
[0025] Step 1: Dataset preparation
[0026] The CIFAR-10 public dataset is used in the experiment. The image data comes from the real world, including 60,000 color images with a size of 32x32. There are 10 categories in total, namely airplane, car, bird, cat, deer, dog, frog, horse, ship, and truck, and each category has 6,000 pictures. There are 50,000 images in the dataset as the training set and 10,000 images as the test set.
[0027] Step 2: Train the initial network until it converges to the local minimum
[0028] The experiment uses the VisionTransformer-Tiny (ViT-Ti) network structure. VisionTransformer is a model structure in which the Transformer is applied to the field of computer vision, and its basic structure is the same as that of the original Transformer. The number of layers L of the adopted ViT-Ti network is 12, and each layer includes a self-attention module and a Feed-Forward module. Figure 1 is a schematic diagram of the Feed-Forward fully connected module in the Transformer, and the encoding width d of the network model = 192, the hidden layer width d of the Feed-Forward hidden = 512, the number of attention heads head = 3. The ReLU activation function is used after the Feed-Forward hidden layer, and finally includes a fully connected layer classification head, and the number of output neurons is set to 10, the number of categories in the CIFAR10 dataset.
[0029] The randomly initialized Transformer is trained using the Adam optimizer widely used in deep learning. During training, the number of warm-up epochs warmup_epoch = 20 is set, and the epoch is used to represent the current training epoch. Then, the calculation formula for the learning rate lr can be expressed as:
[0030]
[0031] The training loss of the Transformer converges after 11,500 training epochs, and a local minimum point θ is obtained. * . At this time, the training loss R(θ * ) of θ is obtained. * ) = 0.00552, the classification accuracy rate on the training set is 91.166, and the classification accuracy rate on the test set is 78.01.
[0032] Randomly select 2 direction vectors δ and ξ from the network parameters, and use two-dimensional contour lines to draw the training loss landscape near θ in the parameter space. The drawing function is f(α, β) = R(θ * + αδ + βξ), where α and β are two scaling parameter values, and the parameters are uniformly taken from -1.0 to 1.0. * + αδ + βξ), α and β are two scaling parameter values, and the parameters are uniformly taken from -1.0 to 1.0. Figure 2 is the training loss landscape near θ * . The point (0, 0) in the figure represents θ * , and the corresponding loss is R(θ * ) = 0.00552. The experimental results show that the training loss values of the parameter models near θ in the parameter space are all higher than R(θ * ), which can indicate that θ * ) *is a local minimum point.
[0033] Step 3: Construct the point θ 1
[0034] At the local minimum point θ of the Transformer * In the fully connected module of the last layer, randomly select two neurons in the hidden layer, change the weight parameters of these two neurons, and obtain θ 1 , assuming the selected neurons are the i-th neuron and the j-th neuron respectively, then θ 1 The specific settings of the weight parameters changed in
[0035]
[0036] where represents the weight vector of the i-th neuron in the hidden layer of the fully connected module corresponding to the input layer, represents the weight vector of the j-th neuron in the hidden layer of the fully connected module corresponding to the input layer, represents the bias of the i-th neuron in the hidden layer, represents the bias of the j-th neuron in the hidden layer, represents the weight vector of the i-th neuron in the hidden layer corresponding to the output layer, represents the weight vector of the j-th neuron in the hidden layer corresponding to the output layer.
[0037] Figure 3 is the schematic diagram of the parameter construction of the fully connected module of θ 1 . The output vector of the output layer of θ 1 is denoted as o, d model represents the size of the model encoding dimension and is also the width of the input layer of the fully connected module, d hidden represents the dimension size of the hidden layer of the fully connected module, x represents the input vector of the input layer, assuming the neuron selected for the next operation is the k-th neuron, p ∈ {1, 2,..., d hidden}\{i, j, k}, represents the neurons in the hidden layer other than i, j, k, b 2 represents the bias vector of the output layer, represents the weight vector of the k-th neuron in the hidden layer corresponding to the output layer, represents the weight vector of the k-th neuron in the hidden layer corresponding to the input layer, represents the bias of the k-th neuron in the hidden layer, represents the weight vector of the p-th neuron in the hidden layer corresponding to the output layer, represents the weight vector of the p-th neuron in the hidden layer corresponding to the input layer, Denote the bias of the $p$-th neuron in the hidden layer, then the output $o$ of the output layer is:
[0038] When the $k$-th neuron in the hidden layer is activated:
[0039]
[0040] When then:
[0041]
[0042] where $\sigma(x)=\max(0,x)$ is the ReLU activation function. The training loss of the constructed point $\theta$ 1 is $R(\theta$ 1 ) = 0.00553.
[0043] Step 4: Construct another point $\theta$ 1 in the parameter space that has the same training loss as $\theta$ 2
[0044] Denote the output of the $k$-th neuron in the middle hidden layer of the fully connected module in the last layer of the $i$-th training sample $X$ i after passing through the ReLU activation function in the Transformer as Since the Transformer network itself usually has a relatively deep network depth and complex network parameters, and at the same time the training samples are diverse, for different training samples $X$ i and $X$ j the outputs of the $k$-th neuron in the middle hidden layer of the fully connected module in the last layer after passing through the ReLU activation function and will not be the same, that is $i\neq j$.
[0045] For any $k$-th neuron in the middle hidden layer of the fully connected module in the last layer of the constructed Transformer network $\theta$ 1 for a training dataset containing $N$ training samples, the outputs after passing through the ReLU activation function can always form a list After sorting this list by the value, a threshold $\eta$ can always be selected to divide the sorted list into two parts, one part is greater than or equal to the threshold $\eta$, and the other part is less than or equal to the threshold $\eta$.
[0046] After selecting the threshold $\eta$, construct another Transformer network $\theta$ 1 that has the same training loss as $\theta$ 2 In $\theta$ 1Arbitrarily select a neuron in the hidden layer of the fully connected module of the last layer, and change the weight parameters of this neuron and the weight parameters of the two neurons selected in step 2 to obtain θ 2 , assuming that the neurons selected in step 2 are the i-th neuron and the j-th neuron respectively, and the neuron selected in this step is the k-th neuron, then the specific settings of the weight parameters in θ 2 are as follows:
[0047]
[0048] where represents the weight vector from all neurons in the input layer to the i-th neuron in the hidden layer of the fully connected module in the last layer of θ 2 , represents the weight vector from all neurons in the input layer to the k-th neuron in the hidden layer of the fully connected module in the last layer of θ 1 , represents the weight vector from all neurons in the input layer to the j-th neuron in the hidden layer of the fully connected module in the last layer of θ 2 , represents the weight vector from all neurons in the hidden layer to the i-th neuron in the output layer of the fully connected module in the last layer of θ 2 , represents the weight vector from all neurons in the hidden layer to the j-th neuron in the output layer of the fully connected module in the last layer of θ 2 , represents the weight vector from all neurons in the hidden layer to the k-th neuron in the output layer of the fully connected module in the last layer of θ 1 . represents the bias of the k-th neuron in the hidden layer of the fully connected module in the last layer of θ 2 , represents the bias of the 2nd neuron in the hidden layer of the fully connected module in the last layer of θ 2 , represents the bias of the k-th neuron in the hidden layer of the fully connected module in the last layer of θ 2 , represents the bias of the k-th neuron in the hidden layer of the fully connected module in the last layer of θ 1 , b′ 2 represents the bias vector of the output layer of the fully connected module in the last layer of θ 2 , b 2 represents the bias vector of the output layer of the fully connected module in the last layer of θ 1 . represents θ 1The weight vector corresponding to the k-th neuron in the hidden layer of the fully connected module in the last layer with respect to the output layer, and η is the threshold selected according to the previous method.
[0049] Figure 4 is θ 2 Schematic diagram of the parameter structure of the fully connected module. Let the output of the j-th neuron in the output layer of θ 2 be denoted as o′, p ∈ {1, 2,..., d hidden}\{i, j, k}, then the output o′ of the output layer in θ 2 is:
[0050] When , the k-th neuron is activated:
[0051]
[0052] When :
[0053]
[0054] When :
[0055]
[0056] The output o′ of θ 2 constructed in this way is identically equal to o, so there is always a training loss R(θ 2 ) = R(θ 1 ). The loss of the experimentally obtained θ 2 is R(θ 2 ) = 0.00553, and R(θ 2 ) = R(θ 1 ). Figure 5 Visualizes the loss values in the parameter space in the direction from θ * to θ 2 , with a scaling range from -1.5 to 1.5. The experimental results prove that θ * and θ 2 are in different regions in the parameter space.
[0057] Step 5: Further optimize θ 2 to reduce the training loss to a level lower than that of θ *
[0058] Continue to train θ 2 using the Adam optimizer. The training loss will continue to decrease during the optimization process, and finally converge to a better model whose training loss successfully jumps out of the local minimum. The classification accuracy rates on both the training set and the test set are higher than θ * , and it is a Transformer with better classification effect than θ * . Finally, the loss of obtained from the experiment is . At the same time, compared with θ * , the classification accuracy rate on the test set has increased from 78.01% to 78.44%, proving that the method designed in the present invention can successfully jump out of the local minimum and obtain a Transformer with a smaller training loss, and its classification accuracy rate on the test set is higher.
Claims
1. A training method based on a Transformer that jumps out of a local minimum, characterized in that: It includes the following steps: (1) Initialize the Transformer randomly and train it with the Adam optimizer until the training loss stops decreasing and convergence is reached. A local minimum point in the parameter space of the Transformer is obtained, which is denoted as θ * ; (2) Construct another network with the same structure as θ* but with different parameter values, denoted as θ1; (3) Construct another point θ2 in the parameter space with the same training loss as θ1 but with different parameter values; (4) Further optimize θ2 to reduce the network training loss to less than θ * The training loss is lower than that of * , we get a point with better classification effect, recorded as 2. The training method based on Transformer escaping local minima according to claim 1, characterized in that: The step 1 specifically includes: Construct a Transformer with L layers. Each layer in the Transformer includes a self-attention module and a fully connected module. The network parameters of the Transformer are denoted by θ. Each fully connected module in the Transformer is a fully connected network containing a single hidden layer. The two-dimensional weight matrix and one-dimensional bias vector from the input layer to the hidden layer in the fully connected module of the lth layer of the Transformer are denoted by W. l,1 and b l,1 , the two-dimensional weight matrix and one-dimensional bias vector from the hidden layer to the output layer are denoted as W l,2 and b l,2 , the input vector of the fully connected module of the lth layer is represented as x l , the output vector is represented as o l , σ(x) = max(0, x) represents the ReLU activation function, then the output o of the fully connected module of the lth layer l for: o l =W l,2 σ(W l,1 x l +b l,1 )+b l,2 The predicted output of network θ is recorded as The true label is recorded as y. For the training task with N classification categories, the cross entropy loss R(θ) of the network θ is: where yi represents the i-th component of vector y, Represents the i-th component in the network prediction output, which is also the probability of being predicted as the i-th class; The L-layer Transformer is randomly initialized and trained with the Adam optimizer until the loss stops decreasing and converges to a local minimum. The local minimum point θ is obtained. * , the training loss obtained at this time is expressed as R(θ * ).
3. The training method based on Transformer escaping local minima according to claim 1, characterized in that: The step 2 specifically includes: In the fully connected module of the last layer in the local minimum point θ* of Transformer, two neurons in the hidden layer are randomly selected, and the weight parameters of these two neurons are changed to obtain θ1. Assuming that the selected neurons are the i-th neuron and the j-th neuron, the specific settings of the weight parameters changed in θ1 are as follows: in represents the weight vector of the input layer corresponding to the i-th neuron in the hidden layer of the fully connected module, represents the weight vector of the jth neuron in the hidden layer of the fully connected module corresponding to the input layer, represents the bias of the i-th neuron in the hidden layer, represents the bias of the jth neuron in the hidden layer, represents the weight vector of the output layer corresponding to the i-th neuron in the hidden layer, Represents the weight vector of the output layer corresponding to the jth neuron in the hidden layer.
4. The training method based on Transformer escaping local minima according to claim 1, characterized in that: The step 3 specifically includes: For any i-th training sample X i The output of any k-th neuron in the middle hidden layer of the fully connected module in the last layer of Transformer after the ReLU activation function is expressed as Since the Transformer network itself is usually deep and has complex network parameters, and the training samples are diverse, different training samples X i and X j The output of the kth neuron in the middle hidden layer of the fully connected module in the last layer after the ReLU activation function and will not be the same, that is For any k-th neuron in the middle hidden layer of the fully connected module in the constructed Transformer network θ1, for a training data set containing N training samples, there is always an output after the ReLU activation function that can form a list After sorting the list according to the value, a threshold η can always be selected to divide the sorted list into two parts, one part is greater than or equal to the threshold η, and the other part is less than or equal to the threshold η. The median of the list is selected as the selected threshold η; After selecting the threshold η, construct another Transformer network θ2 with the same training loss as θ1. Select any neuron in the hidden layer of the last fully connected module in θ1, change the weight parameter of this neuron and the weight parameters of the two neurons selected in step 2, and obtain θ2. Assuming that the i-th neuron and the j-th neuron are selected in step 2, and the k-th neuron is selected in this step, the specific settings of the weight parameters in θ2 are as follows: in represents the weight vector from all neurons in the input layer to the i-th neuron in the hidden layer in the fully connected module of the last layer in θ2, represents the weight vector from all neurons in the input layer to the kth neuron in the hidden layer in the fully connected module of the last layer in θ1, represents the weight vector from all neurons in the input layer to the jth neuron in the hidden layer in the fully connected module of the last layer in θ2, represents the weight vector from all neurons in the hidden layer to the i-th neuron in the output layer in the fully connected module of the last layer in θ2, represents the weight vector from all neurons in the hidden layer to the jth neuron in the output layer in the fully connected module of the last layer in θ2, represents the weight vector from all neurons in the hidden layer to the kth neuron in the output layer in the fully connected module of the last layer in θ1; represents the bias of the i-th neuron in the hidden layer of the last fully connected module in θ2, represents the bias of the second neuron in the hidden layer of the last fully connected module in θ2, represents the bias of the kth neuron in the hidden layer of the last fully connected module in θ2, represents the bias of the kth neuron in the hidden layer of the last fully connected module in θ1, b′ 2 represents the bias vector of the output layer in the fully connected module of the last layer in θ2, b 2 represents the bias vector of the output layer in the fully connected module of the last layer in θ1; represents the weight vector of the output layer corresponding to the kth neuron in the hidden layer of the last layer of the fully connected module in θ1, and η is the threshold selected according to the previous method; The training loss of θ2 constructed in this way will have R(θ2)=R(θ1).
5. The training method based on Transformer escaping local minima according to claim 1, characterized in that: The step 4 specifically includes: Continue to train the network θ2 through the Adam optimizer, and the training loss can continue to decrease. When the loss stops decreasing, it converges to a better model. Its loss Successfully jumped out of the local minimum.