A method for classifying cultural resource images using hierarchical training combined with label smoothing
Through hierarchical training combined with label smoothing methods, the ResNet-18 network is optimized, the information transmission problem is solved, the generalization ability and accuracy of the object detection model are improved, and the effective fusion of high-resolution low-level features and high-level semantic information is achieved.
Patent Information
- Application Number
- CN202210937374.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-05
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-08-05
AI Technical Summary
The existing image classification and object detection models have problems such as information loss, gradient disappearance or explosion during the information transmission process, which makes it difficult to train deep networks and poor generalization capabilities, and the design space of feature pyramid network architecture is huge and difficult to optimize.
The method of layered training combined with label smoothing is adopted. By adding a fully connected layer and classifier to the ResNet-18 network, the cross-entropy loss function of label smoothing is optimized. The low-level uses label smoothing to weaken the supervision intensity, and the deep tightening and smoothing is improved to improve the supervision intensity, achieving bottom-up backpropagation.
It improves the generalization ability of the model, can effectively integrate different levels of features, and improves the accuracy and generalization ability of object detection.
Smart Images

Figure CN115311494B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multi-scale object detection and classification, and particularly relates to a method for classifying cultural resource images by using hierarchical training combined with label smoothing. Background Art
[0002] In recent years, ResNet and its variants have achieved great success. ResNet proposed a residual block with shortcut connections, solved the degradation problem of deep networks, and accelerated network training. Its core is the residual learning unit, and the main idea is to add a direct connection channel in the network, that is, the idea of shortcut. The previous image network structure performed a non-linear transformation on the performance input, while the Highway Network allowed a certain proportion of the output of the previous network layer to be retained. The idea of Resnet is also very similar to that of the Highway Network, allowing the original input information to be directly passed to the subsequent layers. A neural network layer can learn not the entire output, but the residual of the output of the previous network. As an extremely deep network, it can obtain shallow information from the significantly increased depth. Due to the superiority of ResNet, variants combining ResNet with other network technologies have emerged, such as SENet, SKNet, SIMNet, etc. The success of ResNet and its variants in various industries has attracted the interest of researchers, and they have studied the working mechanism of ResNet. A common feature of ResNet and its variants is that "shortcut connections" are applied to route data across network layers in the network, and its working mechanism proposes the idea of residual learning. Traditional convolutional networks or fully connected networks will more or less have problems such as information loss and loss during information transmission, and at the same time, it will also lead to gradient disappearance or gradient explosion, resulting in the inability to train very deep networks. Resnet solves this problem to a certain extent by directly routing the input information to the output, protecting the integrity of the information. The entire network only needs to learn the part of the difference between the input and the output, simplifying the learning objective and difficulty.
[0003] Learning visual feature representation is a fundamental problem in computer vision. In the past few years, great progress has been made in designing the model architectures of deep convolutional networks for image classification and object detection. Different from image classification that predicts the class probabilities of images, object detection has its own challenges, namely detecting and localizing multiple objects at a wide range of scales and positions. To solve this problem, many modern object detectors usually use the pyramid feature representation method, which represents an image with multi-scale feature layers. The Feature Pyramid Network (FPN) is one of the representative model architectures for generating pyramid feature representations for object detection. It adopts a typical backbone model designed for image classification and builds a feature pyramid by successively combining two adjacent layers in the feature layers of the backbone model through top-down and lateral connections. High-level features are semantically strong but have low resolution. They are upsampled and combined with high-resolution features to produce high-resolution and strong-semantic feature representations. Although FPN is simple and effective, it may not be the best architecture design. Recently, PANet has shown that adding an additional bottom-up path to FPN features can improve the feature representation of low-resolution features. And many recent works have proposed various cross-scale connections or operations to combine features to produce pyramid-style feature representations. The challenge in designing the feature pyramid structure lies in its huge design space. As the number of layers increases, the possible number of connections for combining different-scale features grows exponentially. Summary of the Invention
[0004] To overcome the above-mentioned deficiencies of the prior art, the object of the present invention is to propose a method that combines a hierarchical training method with label smoothing to weaken the supervision intensity. Different from traditional hierarchical learning methods, the optimization method technology proposed by the present invention is based on the existing FPN method. It proposes hierarchical learning combined with preventing the model from predicting labels too confidently during training, and uses label smoothing, which can improve the poor generalization ability, to weaken the supervision intensity. That is, during hierarchical training of the network, for the shallow network, label smoothing is used to weaken the supervision intensity, and for the deep network, the smoothing is tightened to increase the supervision intensity, so that the model has better generalization ability. At the same time, using the network structure of resnet-18, each residual block is regarded as a layer, and a fully connected layer and a classifier are added in the middle of each residual block connecting to the next residual block, and the loss of each layer is calculated by cross-entropy with the label after label smoothing. The final loss is obtained by adding the losses of each layer. During training, the lower-level layer first optimizes the layer parameters through the gradient of this additional loss. As the training deepens, the top-level parameters are gradually optimized, that is, bottom-up backpropagation is adopted. By using smoothed labels, the supervision degree of the model is reduced, giving a certain ambiguity to the labels, achieving an effect similar to multi-labels.
[0005] To achieve the above object, the technical solution adopted by the present invention is:
[0006] A method for weakening the supervision intensity by combining a hierarchical training method with label smoothing, comprising the following steps:
[0007] Step 1, collect open-source image data from the public cultural cloud;
[0008] Step 2, assume the use of the image model Resnet, determine that the algorithm idea is to perform predictions independently at different feature layers, utilize the hierarchical idea of FPN, and assume the image pyramid algorithm idea adopts a bottom-up idea, that is, the forward process of the network. During the forward process, the size of the feature map will change after passing through some layers, while it will not change after passing through other some layers. The layers that do not change the size of the feature map are grouped into a stage. Therefore, the features extracted each time are the outputs of the last layer of each stage, and in this way, a feature pyramid can be formed; for the hierarchically organized feature maps, calculate the loss by pre-calculating the loss of one level during the training of the lower layers for backpropagation, adjusting the hierarchical parameters of the lower front layers in advance, adding the losses of each layer to finally obtain a total loss, using this loss for backpropagation, and at the same time, for the layer features of each stage, add the same regularization methods as L1, L2, and dropout, that is, label smoothing;
[0009] Step 3, utilize the determined model and algorithm idea to perform image processing, perform data augmentation such as flipping and blurring on the data, send the processed data into the model, and standardize the data; adopt an image size of 3*3*3, then its number of channels is 128. During training, use 16, 32, 64, and 128 channels as a stage respectively, which can also be regarded as a layer. Utilize the network structure of resnet-18, regard each block as a layer, add a fully connected layer and classification in the middle of connecting each block to the next block, and calculate the hierarchical loss with the label after label smoothing through cross-entropy loss, that is, consider the data of one layer as processed, send the processed layer data into the next layer, perform the same operation on the layer data, and finally combine the feature information of the previous channel layers when passing through 128 channels. After global average pooling and full connection, perform classification, and the loss obtained at this time is the sum of the losses of the previous layers. Use the total loss for backpropagation, and continuously optimize the hierarchical parameters of the lower layers first through this additional loss gradient during training. As the training deepens, gradually optimize the top layer parameters;
[0010] The loss function of the model uses cross-entropy. For each sample i, its loss function is:
[0011]
[0012] Note: y i is the sample label, which is 1 or 0; x i is the sample data; P is the probability function; is all sample data for yi Probability;
[0013] After data randomization, the probability of a new label being the same as y i is 1 - ε, and the probability of being different is ε (i.e., 1 - y i ). Therefore, when using label randomization for the training data, with probability 1 - ε, its loss function is the same as the above formula, and with probability ε, it is:
[0014]
[0015] Note: y i is the sample label, either 1 or 0; x i is the sample data; P is the probability function; is the probability that all sample data is y i ; 1 - ε is the probability of a new label being the same as y i ; ε (i.e., 1 - y i ) is the probability of being different;
[0016] By taking the weighted average of the above two formulas according to the probability, we can obtain:
[0017]
[0018] Note: y i is the sample label, either 1 or 0; x i is the sample data; P is the probability function; is the probability that all sample data is y i ; 1 - ε is the probability of a new label being the same as y i ; ε (i.e., 1 - y i ) is the probability of being different;
[0019] Simplify the above formula. Let y i ′ = ε(1 - y i ) + (1 - ε)y i , then we can obtain:
[0020]
[0021] Note: y i is the sample label, either 1 or 0; x i is the sample data; P is the probability function; is the probability that all sample data is y i ;
[0022] Compared with the original cross - entropy formula expression, only y i is replaced by y i ′, and the rest remains unchanged. It is equivalent to: replacing each label y i with yi ', and then perform the regular training process. Therefore, our randomization process does not need to be carried out before training. We only need to replace each label;
[0023]
[0024] That is to say, if the label is 1, it is replaced with a number 1 - ε that is closer. Similarly, when the label is 0, instead of directly putting 0 into training, it is replaced with a relatively small number ε. To see the effect, the expression of the cross - entropy model can be given:
[0025]
[0026] It can be seen from this formula that 1 and 0 have no chance to appear before the output of the cross - entropy model, and ω will continuously expand in the model, causing the cross - entropy model to continuously increase ω, and the output prediction will be as close as possible to 1 or 0. However, this process is contradictory to regularization; in other words, it may cause overfitting. If the labels 0 and 1 are replaced with 1 - ε and ε respectively, after reaching this value, the further optimization output of the model will not be carried out; therefore, the smoothing operation means changing the two extreme values of 0 and 1 into two relatively less extreme values;
[0027] Step 4, experimental settings. The control variable method is used for experiments, and multiple groups of experiments are designed, including:
[0028] (1) Observe whether hierarchical training and hierarchical training with added smoothing are effective;
[0029] (2) Keep the smoothing parameter unchanged and observe the influence of the change in the number of participating layers on the result;
[0030] (3) Keep the number of layers unchanged and change the smoothing parameter; it can be further divided into:
[0031] a. Observe the influence of the same parameters and different parameters,
[0032] b. Observe the influence of all - zero and all - 0.1 label smoothing parameters;
[0033] The experimental conclusions obtained from the three groups are as follows:
[0034] (1) It is effective, but it is not certain whether this effect comes from hierarchical training or label smoothing;
[0035] (2) It can be seen that different layers have inconsistent requirements for label smoothing during hierarchical training;
[0036] (3) It is verified that hierarchical training helps to improve the accuracy of the model; different layers have inconsistent sensitivities to label smoothing.
[0037] Furthermore, to prevent the model from predicting labels too confidently during training and improve the poor generalization ability; using the hierarchical training method combined with label smoothing to weaken the supervision intensity specifically means that during hierarchical training of the network, for the shallow network, label smoothing is used to weaken the supervision intensity, and in the deep layer, the smoothing is tightened to increase the supervision intensity, so that the model has better generalization ability.
[0038] The beneficial effects of the present invention are:
[0039] Compared with the existing graph supervision model ResNet, the use of a feature pyramid, that is, a hierarchical technique, can simultaneously utilize the high-resolution low-level features and the high semantic information of the high-level features. The prediction result is obtained by fusing these features from different layers, and the fusion prediction on each feature layer is carried out separately, which is different from the general traditional feature fusion method; secondly, the present invention uses different stage layers of the number of channels and adds a smooth label technique. By flattening the image, fully connecting, and classifying operations, the supervision intensity is improved, that is, for the shallow network, label smoothing is used to weaken the supervision intensity, and in the deep layer, the smoothing is tightened to increase the supervision intensity, so that the model has better generalization ability, verifying that hierarchical training helps to improve the accuracy of the model and that different layers have inconsistent sensitivities to label smoothing. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is the core network model diagram of the present invention;
[0041] Figure 2 It is the specific flow chart implemented by the present invention;
[0042] Figure 3 It is the experimental effect diagram of smooth layering;
[0043] Figure 4 It is the experimental effect diagram of the influence of the number of smooth layers. DETAILED DESCRIPTION OF THE INVENTION
[0044] The present invention will be further described below in conjunction with the drawings and embodiments.
[0045] As Figure 1 、 2 shown, a method for using a hierarchical training method combined with label smoothing to weaken the supervision intensity includes the following steps:
[0046] Step 1, collect open-source image data from the public cultural cloud;
[0047] Step 2: Assume the image model Resnet is used, and it is determined that the algorithm idea is to perform prediction independently at different feature layers. Using the hierarchical idea of FPN, it is set as the image pyramid algorithm idea, adopting a bottom-up approach, that is, the forward process of the network. In the forward process, the size of the feature map changes after passing through some layers, while it does not change after passing through some other layers. The layers that do not change the size of the feature map are grouped into a stage. Therefore, the features extracted each time are the outputs of the last layer of each stage, and in this way, a feature pyramid can be formed. For the well-layered feature maps, the loss is calculated by pre-calculating the loss of one level during the training of the lower layers for backpropagation, adjusting the hierarchical parameters of the previous lower layers in advance, and adding the losses of each layer to finally obtain a total loss. Using this loss for backpropagation, at the same time, for the layer features of each stage, the same regularization methods as L1, L2, and dropout are added, that is, label smoothing;
[0048] Step 3: Using the determined model and algorithm idea, perform image processing. Flip, blur, and perform other data augmentations on the data, and then send the processed data into the model for data standardization. If the image size is 3*3*3, then the number of channels is 128. During training, the channels of 16, 32, 64, and 128 are each regarded as a stage, which can also be regarded as a layer. Using the network structure of resnet-18, each block is regarded as a layer, and a fully connected layer and classification are added in the middle of each block connecting to the next block. Calculate the hierarchical loss with the label after label smoothing through the cross-entropy loss, that is, when the data regarded as one layer is processed, send the processed layer data into the next layer, and perform the same operation on the layer data. Finally, when passing through the 128 channels, combine the feature information of the previous channel layers, and perform classification after global average pooling and full connection. At this time, the obtained loss is the sum of the losses of the previous layers. Use the total loss for backpropagation. During training, the lower layers first optimize the hierarchical parameters through this additional loss gradient. As the training deepens, gradually optimize the top-layer parameters;
[0049] The loss function of the model uses cross-entropy. For each sample i, its loss function is:
[0050]
[0051] Note: y i is the sample label, either 1 or 0; x i is the sample data; P is the probability function; is the probability that all sample data is y i ;
[0052] After data randomization, the probability of the new label being the same as y i is 1 - ε, and the probability of being different is ε (that is, 1 - y i) Therefore, when using randomly shuffled labels for the training data, with probability \(1 - \varepsilon\), its loss function is the same as the above formula, and with probability \(\varepsilon\), it is:
[0053]
[0054] Note: \(y\) i is the sample label, either 1 or 0; \(x\) i is the sample data; \(P\) is the probability function; is the probability that all sample data is \(y\) i ; \(1 - \varepsilon\) is the probability of the new label being the same as \(y\) i ; \(\varepsilon\) (i.e., \(1 - y\) i ) is the probability of being different;
[0055] By taking a probability-weighted average of the above two formulas, we can obtain:
[0056]
[0057] Note: \(y\) i is the sample label, either 1 or 0; \(x\) i is the sample data; \(P\) is the probability function; is the probability that all sample data is \(y\) i ; \(1 - \varepsilon\) is the probability of the new label being the same as \(y\) i ; \(\varepsilon\) (i.e., \(1 - y\) i ) is the probability of being different;
[0058] Simplifying the above formula, let \(y\) i ' = \(\varepsilon(1 - y\) i )+(1 - \varepsilon)y i , we can obtain:
[0059]
[0060] Note: \(y\) i is the sample label, either 1 or 0; \(x\) i is the sample data; \(P\) is the probability function; is the probability that all sample data is \(y\) i ;
[0061] Compared with the original cross-entropy formula expression, only \(y\) i is replaced by \(y\) i ', and the rest remains unchanged. It is equivalent to: replacing each label \(y\) i with \(y\) i ', and then performing the normal training process. Therefore, we do not need to perform the randomization process before training, but only need to replace each label;
[0062]
[0063] That is to say, if the label is 1, it will be replaced with a closer number 1 - ε. Similarly, when the label is 0, instead of directly putting 0 into training, it will be replaced with a relatively small number ε. To see the effect, the expression of the cross - entropy model can be given:
[0064]
[0065] It can be seen from this formula that 1 and 0 have no chance to appear before the output of the cross - entropy model, while ω will continuously expand in the model, causing the cross - entropy model to continuously increase ω, and the output prediction will be as close as possible to 1 or 0. However, this process is contradictory to regularization; in other words, it may cause overfitting. If the labels 0 and 1 are replaced with 1 - ε and ε respectively, after reaching this value, the further optimization output of the model will not be carried out; therefore, the smoothing operation means changing the two extreme values of 0 and 1 into two relatively less extreme values;
[0066] Step 4, experimental settings. The control variable method is used for experiments, and multiple groups of experiments are designed, including:
[0067] (1) Observe whether there is an effect in hierarchical training and hierarchical training with added smoothing;
[0068] (2) Keep the smoothing parameter unchanged and observe the influence of the change in the number of participating layers on the result;
[0069] (3) Keep the number of layers unchanged and change the smoothing parameter; it can be further subdivided into:
[0070] a. Observe the influence of the same parameters and different parameters,
[0071] b. Observe the influence of the label smoothing parameter being all zero and all 0.1;
[0072] The conclusions of the three groups of experiments are respectively:
[0073] (1) There is an effect, but it is not certain whether this effect comes from hierarchical training or label smoothing;
[0074] (2) It can be seen that different layers have inconsistent requirements for label smoothing during hierarchical training;
[0075] (3) It is verified that hierarchical training helps to improve the accuracy of the model; different layers have inconsistent sensitivities to label smoothing.
[0076] Furthermore, to prevent the model from predicting labels too confidently during training and improve the poor generalization ability, a hierarchical training method combined with label smoothing is used to weaken the supervision intensity. Specifically, during hierarchical training of the network, label smoothing is used to weaken the supervision intensity for the shallow network, and smoothing is tightened in the deep layer to increase the supervision intensity, so that the model has better generalization ability.
[0077] Embodiment
[0078] The visual model ResNet is used, Python is used as the development language, the development tool is VSCode, and the algorithm idea is FPN.
[0079] The experimental environment is as follows:
[0080] ● CPU: AMD Ryzen 9 3900X 12-Core Processor 3.80GHz
[0081] ● Memory: 32BG
[0082] ● Hard Disk: 2TB
[0083] ● Operating System: Windows 10 (64-bit)
[0084] Dataset: Multiple datasets are used to measure the performance respectively. The model learns on the training set, evaluates its performance on the test set, and is evaluated on the validation set.
[0085] Method for measuring the effect: The cross-entropy loss function is used.
[0086] Implementation process: Image data will first undergo image processing, that is, the data will first be augmented by flipping, blurring, etc., and the data will be standardized. The image size is 3*3*3, and the number of channels is 128. The processed data is sent into ResNet. During training, the number of channels 16, 32, 64, and 128 are each regarded as a stage, that is, regarded as a convolutional layer. Then, using the network structure of resnet-18, each block is regarded as a layer, and a fully connected layer and classification are added in the middle of each block connecting to the next block, and the loss of the layer is calculated through the cross-entropy with the label after label smoothing. That is, when the data of one layer is regarded as processed, the processed layer data is passed into the next layer, and the same operation is performed on the next layer of data. Finally, when passing through the number of channels 128, the loss obtained at this time is the sum of the losses of the previous layers. The total loss is used for backpropagation. During training, the lower layers of the layer first optimize the layer parameters through the gradient of this additional loss. As the training deepens, the top layer parameters are gradually optimized. Combining the feature information of the previous number of channels, after global average pooling and fully connected layers, classification is performed.
[0087] As Figure 3 , 4 shown, the control variable method is used for experiments, and three groups of experiments are designed
[0088] (1) Observe the effects of hierarchical training and whether adding smoothed hierarchical training has an effect
[0089] (2) Keep the smoothing parameter unchanged and observe the influence of the change in the number of participating layers on the result
[0090] (3) Keep the number of layers unchanged and change the smoothing parameter
[0091] a. Observe the effects of the same parameters and different parameters
[0092] b. Observe the effects of the label smoothing parameter being all zero and all 0.1
[0093] The conclusions of the grouped experiments are as follows:
[0094] (1) It has an effect, but it is not certain whether this effect comes from hierarchical training or label smoothing
[0095] (2) Different layers have inconsistent requirements for label smoothing during hierarchical training
[0096] (3) It is verified that hierarchical training helps to improve the accuracy of the model, and different layers have inconsistent sensitivities to label smoothing
Claims
1. A method for weakening the supervision intensity by combining a hierarchical training method with label smoothing, characterized in that It includes the following steps: Step 1: Collect open-source image data from the public cultural cloud; Step 2: Use the image model Resnet. The algorithm idea is that the prediction is carried out independently at different feature layers. Utilize the hierarchical idea of FPN. For the image pyramid algorithm idea, adopt a bottom-up idea, that is, the forward process of the network. In the forward process, the size of the feature map will change after passing through some layers, while it will not change after passing through some other layers. Group the layers that do not change the feature map size into one stage. Therefore, the features extracted each time are the outputs of the last layer of each stage, and thus a feature pyramid can be formed; Step 3: Use the determined model and algorithm idea to perform image processing. Perform data augmentation such as flipping and blurring on the data, and send the processed data into the model to standardize the data; Adopt an image size of 3*3*3, and its number of channels is 128. During training, take the number of channels of 16, 32, 64, and 128 as one stage respectively, which can also be regarded as one layer. Utilize the network structure of resnet-18, regard each block as one layer, add a fully connected layer and classification in the middle of connecting each block to the next block, and calculate the loss of the layer through the cross-entropy loss function with the label after label smoothing, that is, consider that the data of one layer is processed. Send the processed layer data into the next layer, and process the layer data with the same operation. Finally, when passing through the 128 channels, combine the feature information of the previous channel layers. After passing through global average pooling and fully connected layers, perform classification, and the loss obtained at this time is the sum of the losses of the previous layers. Use the total loss for backpropagation. During training, the lower layers first optimize the layer parameters through this additional loss gradient. As the training deepens, gradually optimize the top layer parameters; Step 4: Experimental settings. Conduct experiments using the control variable method, and design multiple groups of experiments, including: (1) Observe whether hierarchical training and hierarchical training with added smoothing are effective; (2) Keep the smoothing parameter unchanged and observe the influence of the change in the number of participating layers on the results; (3) Keep the number of layers unchanged and change the smoothing parameter; it can be further divided into: a. Observe the influence of the same parameters and different parameters, b. Observe the influence of the label smoothing parameter being all zero and all 0.1; The conclusions of the three groups of experiments are respectively: (1) It is effective, but it is not certain whether this effect comes from hierarchical training or from label smoothing; (2) It can be seen that the requirements for label smoothing in the process of hierarchical training are inconsistent for different layers; (3) It is verified that hierarchical training helps to improve the accuracy of the model; the sensitivities of different layers to label smoothing are inconsistent.
2. The method for weakening the supervision intensity by combining a hierarchical training method with label smoothing according to claim 1, characterized in that Prevent the model from predicting labels too confidently during training and improve the poor generalization ability; Utilize the hierarchical training method combined with label smoothing to weaken the supervision intensity. Specifically, it means that during hierarchical training of the network, for the shallow network, use label smoothing to weaken the supervision intensity, and tighten the smoothing in the deep layer to increase the supervision intensity, so that the model has better generalization ability.
3. A method for weakening the supervision intensity by combining a hierarchical training method with label smoothing according to claim 1, characterized in that The loss function uses cross-entropy. For each sample , its loss function is: Note: is the sample label, which is 1 or 0; is the sample data; P is the probability function; is the probability that all sample data are ; After data randomization, the probability of having the same new label as is , and the probability of being different is . Therefore, when using label randomization for training data, there is a probability that its loss function is the same as the above formula, and there is a probability of: Note: is the sample label, which is 1 or 0; is the sample data; P is the probability function; is the probability that all sample data are ; is the probability of the new label that is the same as ; is the different probability, ; By probabilistically weighted averaging the above two formulas, we can obtain: Note: is the sample label, being 1 or 0; is the sample data; P is the probability function; is the probability that all sample data are ; is the probability of the new label being the same as ; is the different probability, ; Simplify the above formula and let , then we can get: Note: is the sample label, which is 1 or 0; is the sample data; P is the probability function; is the probability that all sample data are ; Compared with the original cross-entropy formula expression, only is replaced with , and the rest remains unchanged. It is equivalent to: replacing each label with , and then performing the regular training process. Therefore, we do not need to perform randomization before training, and only need to replace each label; That is to say, if the label is 1, replace it with a closer number , similarly, when the label is 0, instead of directly putting 0 into training, replace it with a relatively small number , to see the effect, the expression of the cross-entropy model can be given: From this formula, we can see that 1 and 0 have no chance to appear before the output of the cross entropy model, and It will continue to expand in the model, causing the cross entropy model to continue to increase , the output prediction will be as close to 1 or 0 as possible, but this process is contradictory to regularization; in other words, it may cause overfitting. and Instead, after reaching this value, no further optimization of the model output will be performed; therefore, the smoothing operation refers to changing the two extreme values of 0 and 1 into two relatively less extreme values.
4. A method for weakening the supervision intensity by combining a hierarchical training method with label smoothing according to claim 1, characterized in that, The so-called label smoothing is that for the well-layered feature maps, when training the lower layers, the loss of one level is pre-computed for backpropagation, the layer parameters of the previous lower layers are adjusted in advance, the losses of each layer are added up to finally obtain a total loss, and this loss is used for backpropagation. At the same time, for the layer features in each stage, the same regularization methods as L1, L2, and dropout are added for label smoothing.
Citation Information
Patent Citations
Small-scale equipment part detection method based on weak supervision collaborative learning in open scene of electric power field
CN111444939A
Cultural resource text classification method based on memory network and graph neural network
CN113516198A