Unbalance mark distribution learning method based on asymmetric momentum optimization
By adopting an asymmetric momentum optimization mechanism in the imbalance labeled distribution learning network and independently setting the dominant and auxiliary momentum variables and coefficients, the gradient conflict problem is solved, the stability and prediction accuracy of network training are improved, and predictions that are closer to the true labeled distribution are achieved.
Patent Information
- Application Number
- CN202510883930.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-30
AI Technical Summary
During the training process, the existing imbalance label distribution learning network cannot be effectively optimized due to gradient conflicts between different loss functions, and cannot predict predicted mark distribution close to the real mark distribution.
Asymmetric momentum optimization mechanism is adopted to set independent dominant momentum buffer variables and dominant loss momentum coefficients for dominant loss, and set independent auxiliary momentum buffer variables and auxiliary loss momentum coefficients for auxiliary loss. By updating network parameters according to the momentum variables and coefficients of the current epoch during training, we ensure that dominant loss and auxiliary loss do not interfere with each other in gradient update, and control dominance by the dominant loss momentum coefficient greater than the auxiliary loss momentum coefficient, combining inter-layer alignment loss and labeled distribution prediction supervision loss, we improve the stability and convergence of network training.
It effectively suppresses gradient fluctuations, improves the training stability and robustness of the unbalanced label distribution learning network, and can more accurately predict predicted label distributions close to the true label distribution, especially on the unbalanced dataset, showing higher training stability and prediction accuracy.
Smart Images

Figure CN120387494A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and in particular to an unbalanced label distribution learning method based on asymmetric momentum optimization. Background Art
[0002] Existing networks for alleviating unbalanced label distribution learning usually adopt a distribution alignment strategy. This distribution alignment strategy usually includes multiple losses. When training an unbalanced label distribution learning network, there is usually a problem of gradient conflict between different losses in the total loss function composed of different losses during the convergence process. This will result in an inability to obtain a better parameter unbalanced label distribution learning network model, and further cause the unbalanced label distribution learning network model to be unable to predict a predicted label distribution closer to the true label distribution. For this reason, the applicant has designed an unbalanced label distribution learning method based on asymmetric momentum optimization. Summary of the Invention
[0003] In order to solve the above deficiencies existing in the prior art, the present application proposes an unbalanced label distribution learning method based on momentum optimization.
[0004] An unbalanced label distribution learning method based on asymmetric momentum optimization includes the following steps: S1. Construct an unbalanced label distribution learning network including a feature encoder, a full label encoder, and a decoder. The output ends of the feature encoder and the full label encoder are both connected to the input end of the decoder; wherein, the full label encoder obtains a global label distribution based on the label distribution input to the unbalanced label distribution learning network; S2. Calculate the total optimization loss of the unbalanced label distribution learning network; the total optimization loss includes an inter-layer alignment loss, a feature and label distribution alignment loss, a feature and label similarity matrix alignment loss, and a label distribution prediction supervision loss; S3. Train the unbalanced label distribution learning network based on the training set and the total optimization loss to obtain a network model; In step S3, a dominant momentum buffer variable and a dominant loss momentum coefficient are set for the dominant loss, and an auxiliary momentum buffer variable and an auxiliary loss momentum coefficient are set for the auxiliary loss. Then, during the training process, the network parameters θ in the next epoch are updated according to the dominant momentum buffer variable, the dominant loss momentum coefficient, the auxiliary momentum buffer variable, and the auxiliary loss momentum coefficient in the current epoch; wherein, the dominant loss is the label distribution prediction supervision loss, and the auxiliary losses include the inter-layer alignment loss, the feature and label distribution alignment loss, and the feature and label similarity matrix alignment loss; S4. Obtain the feature vectors and label distributions based on the unbalanced dataset to be predicted, and then input the feature vectors and label distributions into the network model for one forward propagation to obtain the predicted label distribution and the true label distribution.
[0005] Preferably, in step S2, the inter-layer alignment loss refers to the inter-layer alignment loss between the feature encoder and the full label encoder.
[0006] Preferably, in step S2, the inter-layer alignment loss is obtained as follows: Extract the distribution parameters of the i-th encoding layer in the feature encoder and the i-th encoding layer in the full label encoder in the latent representation space; use the output results of each pair of corresponding encoding layers in the feature encoder and the full label encoder as the input for KL divergence calculation to obtain the distribution alignment loss of each corresponding encoding layer, and accumulate the distribution alignment losses of each layer to obtain the inter-layer alignment loss. In this application, the inter-layer alignment loss is used to constrain the structural consistency of the encoding paths of the feature encoder and the full label encoder in the latent representation space, effectively improving the inter-layer alignment effect of the output results of each corresponding encoding layer of the feature encoder and the full label encoder, thereby enhancing the modeling ability of the unbalanced label distribution learning network for the structural relationship between the input feature vectors and label distributions, enabling the unbalanced label distribution learning network to predict a predicted label distribution closer to the true label distribution.
[0007] Preferably, step S3 includes the following steps: Input the feature vectors in the training set into the feature encoder, and input the label distribution in the training set into the full label encoder, perform forward propagation, calculate the total optimization loss of the unbalanced label distribution learning network, and perform backpropagation under the guidance of the total optimization loss to update the weight parameters of the unbalanced label distribution learning network. During the process of updating the weight parameters of the unbalanced label distribution learning network, update the network parameter θ in the next epoch according to the dominant momentum buffer variable, dominant loss momentum coefficient, auxiliary momentum buffer variable, and auxiliary loss momentum coefficient in the current epoch. After iterating 300 epochs, complete the training of one training segment to obtain an unbalanced label distribution learning network model. During the training of the unbalanced label distribution learning network in this application, a total of 10 training segments are trained to obtain 10 unbalanced label distribution learning network models.
[0008] Preferably, in step S3, when the next epoch is the t-th epoch, update the network parameter θ in the training process of the t-th epoch according to the dominant momentum buffer variable, dominant loss momentum coefficient, auxiliary momentum buffer variable, and auxiliary loss momentum coefficient in the training process of the (t - 1)-th epoch, which specifically includes the following steps: a) Update the dominant momentum buffer variable during the t-1th epoch training process based on the dominant loss momentum coefficient and the gradient of the dominant loss and the dominant momentum buffer variable during the t-1th epoch training process; b) Update the auxiliary momentum buffer variable during the t-1th epoch training process based on the auxiliary loss momentum coefficient and the gradient of the auxiliary loss and the leading momentum buffer variable during the t-1th epoch training process; c) Update the network parameters θ during the t-th epoch training process based on the network parameters θ during the t-1th epoch training process and the dominant momentum buffer variables and auxiliary momentum buffer variables during the tth epoch training process.
[0009] Preferably, the leading loss momentum coefficient is greater than the auxiliary loss momentum coefficient.
[0010] Preferably, in step S4, the unbalanced data set to be predicted is a data set containing text data.
[0011] Preferably, in step S4, the predicted dataset is one of a Movie dataset, a SCUT-FBP dataset, an Emotion6 dataset, a FlickrLDL dataset, a RAF-ML dataset, and a Natural Scene dataset.
[0012] Preferably, step S4 includes the following steps: obtaining feature vectors and label distributions based on the imbalanced data set to be predicted, and then inputting the feature vectors and label distributions into the 10 imbalanced label distribution learning network models obtained in step S3 for forward propagation once, and outputting 10 predicted label distributions and 10 corresponding true label distributions.
[0013] Compared with the prior art, the beneficial technical effects of the present invention are: In the method described in this application, an asymmetric momentum optimization mechanism is adopted to set independent leading momentum buffer variables and leading loss momentum coefficients for the leading loss, and independent auxiliary momentum buffer variables and auxiliary loss momentum coefficients for the auxiliary loss. Then, during the training process, the network parameters θ in the next epoch are updated based on the leading momentum buffer variable, leading loss momentum coefficient, auxiliary momentum buffer variable, and auxiliary loss momentum coefficient in the current epoch. The above settings enable this application to effectively ensure that the momentum buffer variables of the leading loss and the auxiliary loss do not interfere with each other during the update process, and control the dominance of the momentum buffer variable of the leading loss in the gradient update by setting the leading loss momentum coefficient to be greater than the auxiliary loss momentum coefficient. Moreover, during the backpropagation process, it can also effectively achieve the smooth adjustment of the momentum buffer variables of the leading loss and the auxiliary loss during the gradient update process, effectively suppressing the gradient fluctuations of the momentum buffer variables of the leading loss and the auxiliary loss during the gradient update process, enhancing the stability of the optimization process of the momentum buffer variables of the leading loss and the auxiliary loss, and contributing to improving the convergence and robustness of the overall training process of the unbalanced label distribution learning network. Especially in the scenario of collaborative optimization of different loss terms, it shows higher training stability and obtains an unbalanced label distribution learning network model with better parameters. Therefore, when using the unbalanced label distribution learning network model obtained by this application to predict an unbalanced data set, it can predict a predicted label distribution that is closer to the true label distribution; In addition, this application also uses the output results of each pair of corresponding coding layers in the feature encoder and the full label encoder as the input for KL divergence calculation to obtain the distribution alignment loss of each corresponding coding layer, and accumulates the distribution alignment losses of each layer to obtain the inter-layer alignment loss L lenc_alig ; The inter-layer alignment loss L lenc_alig is used to constrain the structural consistency of the encoding paths of the feature encoder and the full label encoder in the latent representation space, effectively improving the inter-layer alignment effect of the output results of each corresponding coding layer of the feature encoder and the full label encoder, thereby enhancing the modeling ability of the unbalanced label distribution learning network for the structural relationship between the input feature vector and the label distribution, enabling the unbalanced label distribution learning network to predict a predicted label distribution that is closer to the true label distribution.
[0014] Through testing, it can be seen that on the Movie dataset (with an imbalance coefficient of 11.01), the final error obtained by the method described in this application is reduced by 7.42% compared to the DILDL method; on the SCUT-FBP dataset (with an imbalance coefficient of 19.37), the final error obtained by the method described in this application is reduced by 16.06% compared to the DILDL method; on the Emotion6 dataset (with an imbalance coefficient of 10.46), the final error obtained by the method described in this application is reduced by 9.93% compared to the DILDL method; on the FlickrLDL dataset (with an imbalance coefficient of 57.12), the final error obtained by the method described in this application is reduced by 8.24% compared to the DILDL method; on the RAF-ML dataset (with an imbalance coefficient of 30.33), the final error obtained by the method described in this application is reduced by 9.47% compared to the DILDL method; on the Natural Scene dataset (with an imbalance coefficient of 10.74), the final error obtained by the method described in this application is reduced by 8.75% compared to the DILDL method. This shows that the method described in this application is indeed significantly closer to the true label distribution than the prediction label distribution predicted by the DILDL method. Description of the Drawings
[0015] Figure 1 Network structure diagram of the imbalance label distribution learning network in this application. Detailed Implementation Manner
[0016] Example 1: An imbalance label distribution learning method based on asymmetric momentum optimization. The imbalance dataset prediction method based on momentum optimization predicts the imbalance dataset - the Movie dataset, including the following steps: S1. Construct an imbalance label distribution learning network. The imbalance label distribution learning network includes a feature encoder, a full label encoder, and a decoder. The output ends of the feature encoder and the full label encoder are both connected to the input end of the decoder, as Figure 1 shown; in this application, the structures of the feature encoder and the full label encoder are the same, and both are the same as the encoder structure disclosed in the paper "Imbalanced Label Distribution Learning". In this application, the feature encoder obtains a feature coding representation based on the feature vector input to the imbalance label distribution learning network; the full label encoder obtains a global label distribution based on the label distribution input to the imbalance label distribution learning network; the decoder is used to decode the feature coding representation and the global label distribution respectively to obtain a predicted label distribution and a true label distribution; among them, the predicted label distribution is used as the output result of the imbalance label distribution learning network. Specifically: The feature encoder encodes the input feature vector and maps the encoded feature vector into the latent representation space to obtain a feature encoding representation (μ feature , σ feature ) composed of the mean and variance, where μ feature represents the mean of the feature encoding representation, and σ feature represents the variance of the feature encoding representation; The full-label encoder globally encodes the input label distribution and maps the encoded label distribution into the latent representation space to learn the feature representation of the global label distribution, obtaining a global label distribution (μ label , σ label ) composed of the variance and mean, where μ label represents the mean of the label distribution, and σ label represents the mean of the label distribution; The decoder performs a decoding operation on the feature encoding representation output by the feature encoder to obtain a predicted label distribution, and performs a decoding operation on the global label distribution output by the full-label encoder to obtain a true label distribution.
[0017] S2. Calculate the total optimization loss of the imbalanced label distribution learning network L total ; In this application, the total optimization loss L total of the imbalanced label distribution learning network L lenc_alig includes the inter-layer alignment loss L alig1 , the feature and label distribution alignment loss L alig2 , the feature and label similarity matrix alignment loss L pre , and the label distribution prediction supervision loss to jointly guide the feature and label association, structural consistency, and distribution fitting ability in the learning task of the imbalanced label distribution learning network; L total In this application, the calculation method of the total optimization loss L total is shown in Equation (1) as follows: L alig1 = α L alig2 + β L lenc-alig + λ L pre (1) In formula (1), α, β, and λ are all balance parameters used to adjust the influence degree of the unbalanced label distribution learning network during the training process. In the first embodiment, α is set to 0.1, β is set to 0.2, and λ is set to 0.1.
[0018] In this application, the inter-layer alignment loss L lenc_alig refers to calculating the inter-layer alignment loss between the feature encoder and the full label encoder L lenc_alig , L lenc_alig and its acquisition method is as follows: Extract the distribution parameters of the i-th coding layer in the feature encoder and the i-th coding layer in the full label encoder in the latent representation space, denoted as ([[]] , ) and ([[]] , ), where i represents the i-th coding layer in the feature encoder or the i-th coding layer in the full label encoder; use the output results of each pair of corresponding coding layers in the feature encoder and the full label encoder as the input for KL divergence calculation to obtain the distribution alignment loss of each corresponding coding layer, and accumulate the distribution alignment losses of each layer to obtain the inter-layer alignment loss L lenc_alig ; the inter-layer alignment loss L lenc_alig is used to constrain the structural consistency of the coding paths of the feature encoder and the full label encoder in the latent representation space, effectively improving the inter-layer alignment effect of the output results of each corresponding coding layer of the feature encoder and the full label encoder, thereby enhancing the modeling ability of the unbalanced label distribution learning network for the structural relationship between the input feature vector and the label distribution, enabling the unbalanced label distribution learning network to predict a prediction label distribution closer to the true label distribution.
[0019] In this application, the method of using KL divergence to calculate the feature and label distribution alignment loss feature by using the feature coding representation (μ feature , σ label ) output by the feature encoder and the global label distribution (μ label ) output by the full label encoder L alig1 is only different from the method of using KL divergence to calculate the dominant label distribution information alignment loss by using the feature coding representation and the dominant label distribution disclosed in ZL202510450586.X L alig 1 in that in this application, the global label distribution (μ label , σ label)Replace the dominant marker distribution in ZL202510450586.X; the alignment loss between features and marker distribution in this application L alig1 For constraining the feature encoding representation (μ output by the feature encoder feature , σ feature ) and the global marker distribution (μ output by the full marker encoder label , σ label ) to be closer in the latent representation space.
[0020] In this application, calculate the alignment loss of the feature and marker similarity matrix L alig2 The method is the same as the method for calculating the alignment loss L of the feature and marker representation similarity matrix disclosed in ZL202510450586.X alig alig2 L alig2 The alignment loss of the feature and marker similarity matrix constructed in this application is used to further improve the structural consistency of the feature and marker representation in the latent representation space.
[0021] In this application, calculate the marker distribution prediction supervision loss L pre The method is the same as the method for calculating the alignment loss L between the predicted marker distribution and the true marker distribution disclosed in ZL202510450586.X alig pre L pre The marker distribution prediction supervision loss uses the KL divergence to calculate the difference between the predicted marker distribution output by the decoder and the true marker distribution. The setting of the above loss can effectively constrain the predicted marker distribution output by the decoder to be closer to the true marker distribution, enabling the effective alignment of the predicted marker distribution and the true marker distribution, and effectively improving the distribution fitting ability of the imbalanced marker distribution learning network in the learning task.
[0022] S3. Obtain the training set and the test set, and train the imbalanced marker distribution learning network for 10 training segments based on the training set and the total optimization loss to obtain 10 imbalanced marker distribution learning network models, including the following steps: S3-1. Obtain the training set and the test set, specifically including the following specific steps: S3-1-1. Obtain the original dataset, where the original dataset is the Movie dataset; the Movie dataset is from the paper "Label Distribution Learning", and the Movie dataset is a dataset of user ratings for movies. This dataset contains text annotation data of movies, such as genres, user reviews, rating annotations, etc. The dataset contains 7,755 movies and 54,242,292 rating records from 478,656 different users (the data is sourced from the Netflix platform, and the rating range is from 1 to 5 stars, with a total of 5 rating levels). In the Movie dataset, the rating annotation distribution of each movie is text data, and the rating annotation distribution of each movie is obtained by statistically calculating the voting ratios of all users at each rating level. Feature extraction is based on metadata, covering dimensions such as genres, directors, actors, countries, budgets, etc., where categorical attributes are converted into binary vectors. The final feature vector extracted from each movie is 1,869-dimensional. This original dataset contains 7,755 feature vectors and 5 label distributions.
[0023] S3-1-2. Use the ten-fold cross-validation method to divide the original dataset into 10 subsets, where 9 subsets each include 775 feature vectors and their corresponding label distributions, and the remaining subset includes 780 feature vectors and their corresponding label distributions; during the process of training the imbalanced label distribution learning network in this application, 10 training segments need to be trained, and each training segment includes a training process of 300 epochs. In each training segment, 9 of the subsets are selected as the training set, and the remaining subset is used as the test set. In the 10 training segments of this application, the selected test sets are all different.
[0024] S3-2. Input the feature vectors in the training set into the feature encoder, and input the label distribution in the training set into the full label encoder. Perform forward propagation, calculate the total optimization loss of the imbalanced label distribution learning network, and perform backpropagation under the guidance of the total optimization loss to update the weight parameters of the imbalanced label distribution learning network. During the process of updating the weight parameters of the imbalanced label distribution learning network, update the network parameters θ in the next epoch according to the dominant momentum buffer variable, dominant loss momentum coefficient, auxiliary momentum buffer variable, and auxiliary loss momentum coefficient in the current epoch. After iterating 300 epochs, complete the training of one training segment to obtain an imbalanced label distribution learning network model. In this application, during the training of the imbalanced label distribution learning network, a total of 10 training segments are trained to obtain 10 imbalanced label distribution learning network models. To ensure the stable convergence of the model during training, the learning rate is set to 0.001, and the batch size is set to 50. This not only ensures high training efficiency but also helps the network fully learn the data features. In this application, the asymmetric momentum optimization mechanism combined with the settings of the learning rate and batch size ensures the convergence speed while avoiding training instability caused by too large or too small gradient updates.
[0025] In this application, the label distribution prediction supervision loss L pre is regarded as the dominant loss, and other losses (other losses include the inter-layer alignment loss L lenc_alig and the feature-label distribution alignment loss L alig1 as well as the feature-label similarity matrix alignment loss L alig2 ) are regarded as auxiliary losses. During the training process of the imbalanced label distribution learning network, this application adopts an asymmetric momentum optimization mechanism for the dominant loss and the auxiliary losses, that is, set independent dominant momentum buffer variables and dominant loss momentum coefficients for the dominant loss, and call the dominant loss momentum coefficient the dominant loss momentum coefficient, and set independent auxiliary momentum buffer variables and auxiliary loss momentum coefficients for the auxiliary losses, and call the auxiliary loss momentum coefficient the auxiliary loss momentum coefficient. Then, during the training process, update the network parameters θ in the next epoch according to the dominant momentum buffer variable, dominant loss momentum coefficient, auxiliary momentum buffer variable, and auxiliary loss momentum coefficient in the current epoch. Taking the next epoch as the t-th epoch as an example, the update method of the network parameters θ during the training process of the t-th epoch includes the following steps: a). Based on the dominant loss momentum coefficient and the gradient of the dominant loss during the training process of the (t - 1)-th epoch And update the dominant momentum buffer variable in the first momentum buffer during the training process of the (t - 1)-th epoch using the dominant momentum buffer variable updated during the training process of the (t - 1)-th epoch Update the dominant momentum buffer variable in the first momentum buffer during the training process of the (t - 1)-th epoch The method is as shown in Equation (2), and this first momentum buffer is regarded as the dominant momentum buffer; (2) In Equation (2), represents the dominant loss momentum coefficient, represents the dominant momentum buffer variable during the training process of the (t - 1)-th epoch, represents the dominant momentum buffer variable during the training process of the t-th epoch, represents the gradient of the dominant loss during the training process of the (t - 1)-th epoch; b), Based on the auxiliary loss momentum coefficient , the gradient of the auxiliary loss during the training process of the (t - 1)-th epoch And update the auxiliary momentum buffer variable in the second momentum buffer during the training process of the (t - 1)-th epoch using the dominant momentum buffer variable updated during the training process of the (t - 1)-th epoch Update the dominant momentum buffer variable in the first momentum buffer during the training process of the (t - 1)-th epoch The method is as shown in Equation (3), and this second momentum buffer is regarded as the auxiliary momentum buffer; (3) In Equation (3), represents the auxiliary loss momentum coefficient, represents the auxiliary momentum buffer variable during the training process of the (t - 1)-th epoch, represents the auxiliary momentum buffer variable during the training process of the t-th epoch, represents the gradient of the auxiliary loss during the training process of the (t - 1)-th epoch. In this application, during the training process of the same epoch, all the auxiliary momentum buffer variables of the auxiliary losses are also the same; and during the training process of all epochs, the dominant loss momentum coefficient is fixed, the auxiliary loss momentum coefficient is also fixed, and all the auxiliary loss momentum coefficients of the auxiliary losses have the same value, and the value of is greater than to achieve the control of the optimization direction of the dominant loss. In this embodiment, the dominant loss momentum coefficient The value is 0.95, the auxiliary loss momentum coefficient is 0.85; c), Based on the network parameters θ in the training process of the (t - 1)-th epoch, the dominant momentum buffer variable in the training process of the t-th epoch and the auxiliary momentum buffer variable in the training process of the t-th epoch Update the network parameters θ in the training process of the t-th epoch, define the network parameters θ in the training process of the t-th epoch as θ', and the calculation method for updating θ' is as described in Equation (4): (4) In Equation (4), γ is the learning rate, θ is the network parameter in the training process of the (t - 1)-th epoch, and are the dominant momentum buffer variable and the auxiliary momentum buffer variable respectively. The calculation method for updating the network parameters θ in the training process of the t-th epoch reflects the joint influence of the dominant momentum buffer variable and the auxiliary momentum buffer variable on the update of the current network parameters. In this application, α, β, and λ are essentially hyperparameters set manually, and the network parameter θ is the parameter learned and updated during the training of the unbalanced label distribution learning network.
[0026] In this application, the above-mentioned asymmetric momentum optimization mechanism uses the dominant loss to drive parameter changes with a higher update weight during the training process, while the gradient of the auxiliary loss is assisted and adjusted under a smaller momentum, thereby effectively avoiding gradient conflicts, improving the stability of the training of the unbalanced label distribution learning network, enhancing the learning effect of the unbalanced label distribution learning network on the unbalanced label distribution dataset, and enabling the unbalanced label distribution learning network model obtained using this application to predict a predicted label distribution that is closer to the true label distribution when predicting an unbalanced dataset. In this application, the asymmetric momentum optimization mechanism is lightly deployed by using param_group or hook in mainstream deep learning frameworks (such as PyTorch) to ensure the efficiency and stability of the optimization process.
[0027] The asymmetric momentum optimization mechanism adopted in this application sets independent leading momentum buffer variables and leading loss momentum coefficients for the leading loss, and independent auxiliary momentum buffer variables and auxiliary loss momentum coefficients for the auxiliary loss. Then, during the training process, the network parameters θ in the next epoch are updated based on the leading momentum buffer variables, leading loss momentum coefficients, auxiliary momentum buffer variables, and auxiliary loss momentum coefficients in the current epoch. The above settings enable this application to effectively ensure that the momentum buffer variables of the leading loss and the auxiliary loss do not interfere with each other during the update process, and control the dominance of the momentum buffer variable of the leading loss in the gradient update by setting the leading loss momentum coefficient to be greater than the auxiliary loss momentum coefficient. Moreover, during the backpropagation process, it can also effectively achieve the smooth adjustment of the momentum buffer variables of the leading loss and the auxiliary loss during the gradient update process, effectively suppressing the gradient fluctuations of the momentum buffer variables of the leading loss and the auxiliary loss during the gradient update process, enhancing the stability of the optimization process of the momentum buffer variables of the leading loss and the auxiliary loss, and contributing to improving the convergence and robustness of the overall training process of the imbalanced label distribution learning network. Especially in the scenario of collaborative optimization of different loss terms, it shows higher training stability, obtaining an imbalanced label distribution learning network model with better parameters. Therefore, when using the imbalanced label distribution learning network model obtained by this application to predict an imbalanced dataset, it can predict a predicted label distribution that is closer to the true label distribution.
[0028] S4. Respectively obtain feature vectors and label distributions based on an imbalanced dataset with an imbalanced movie rating label distribution, and then input the feature vectors and label distributions into the 10 imbalanced label distribution learning network models obtained in step S3 for one forward propagation, outputting 10 predicted label distributions and their corresponding 10 true label distributions. Among them, the imbalanced dataset refers to a dataset with an imbalanced movie rating label distribution; the method for respectively obtaining feature vectors and label distributions based on the imbalanced dataset with an imbalanced movie rating label distribution in this application is the same as the method for obtaining feature vectors and label distributions described in the section "A.1.1 Description of Datasets • Movie (movie rating)" of the paper "Imbalanced Label Distribution Learning". Since the predicted label distribution output by the decoder in this application contains the rating of the movie, existing movie recommendation systems (such as the movie recommendation systems publicly available on MovieLens and Netflix Prize) can predict the rating of the movie based on the above 10 predicted label distributions, so as to recommend movies with higher predicted ratings to users.
[0029] Testing: Compare the method described in this application (abbreviated as Ours method in Table 1) with the existing SA-BFGS method (from Xin Geng's paper "Label Distribution Learning"), EDLRL method (from Xiuyi Jia et al.'s paper "Facial Emotion Distribution Learning by Exploiting Low-Rank Label Correlations Locally"), LDLSF method (from Tingting Ren et al.'s paper "Label Distribution Learning with Label-Specific Features"), LDL-LCLR method (from Tingting Ren et al.'s paper "Label Distribution Learning with Label Correlations via Low-Rank Approximation"), Adam-LDL-SCL method (from Xiuyi Jia et al.'s paper "Label Distribution Learning with Label Correlations on Local Samples"), LDL-LDM method (from Jing Wang et al.'s paper "Label Distribution Learning by Exploiting Label Distribution Manifold"), OFR-FL method (from Tsung-Yi Lin et al.'s paper "Focal Loss for Dense Object Detection"), OFR-CB method (from Tsung-Yi Lin et al.'s paper "Focal Loss for Dense Object Detection"), OFR-DB method (from Tsung-Yi Lin et al.'s paper "Focal Loss for Dense Object Detection"), RDA method (from Xingyu Zhao et al.'s paper "Imbalanced Label Distribution Learning"), and DILDL method (from ZL202510450586.The unbalanced label distribution learning methods such as the "X public decoupling operation-based unbalanced label distribution learning method" are tested using the same test strategy, and the test results are shown in Table 1. The test strategy adopted in this application is as follows: Test is performed based on the test set in 10 training segments and the corresponding unbalanced label distribution learning network model respectively. Calculate the Chebyshev distance value between the predicted label distribution and the true label distribution output by each unbalanced label distribution learning network model based on the Chebyshev distance method. Then, average these Chebyshev distance values to obtain the average Chebyshev distance. Then, calculate the absolute value of the difference between each Chebyshev distance value and the average Chebyshev distance. Take the maximum absolute value as the fluctuation error, and add this fluctuation error to the average Chebyshev distance to obtain the final error. This final error can show the difference between the predicted label distribution and the true label distribution.
[0030] Embodiment 2: Compared with Embodiment 1, the difference between Embodiment 2 and Embodiment 1 is as follows: 1). The unbalanced dataset used in Embodiment 2 is the SCUT-FBP dataset. The SCUT-FBP dataset in this application is the same as the SCUT-FBP dataset disclosed in the paper "SCUT-FBP: A Benchmark Dataset for Facial Beauty Perception". The SCUT-FBP dataset focuses on the research of facial beauty perception. The SCUT-FBP dataset contains text annotation data of facial images. The text annotation data includes information such as gender, age, and scoring marks of facial beauty, as well as description degree. In this dataset, each image is scored by multiple annotators for facial beauty according to subjective feelings. The value range of the scoring level is from 1 point to 5 points, and the description degree is finally calculated according to the proportion of each scoring level in the annotation.
[0031] 2). In step S3-1-2 of Embodiment 2, the way to obtain subsets is as follows: Use the ten-fold cross-validation method to divide the original dataset into 10 subsets, and each subset includes 3600 feature vectors and their corresponding label distributions.
[0032] In Embodiment 2, the method described in this application is tested using the same test strategy as existing unbalanced label distribution learning methods such as the SA-BFGS method, EDLRL method, LDLSF method, LDL-LCLR method, Adam-LDL-SCL method, LDL-LDM method, OFR-FL method, OFR-CB method, OFR-DB method, RDA method, and DILDL method. The test results are shown in Table 1.
[0033] Example 3: Compared with Example 1, the difference between Example 3 and Example 1 lies in: 2) The imbalanced dataset used in Example 3 is the Emotion6 dataset; the Emotion6 dataset is the same as the Emotion6 dataset disclosed in the paper "A mixed bag of emotions: Model, predict, and transfer emotion distributions". Emotion6 contains text annotation data formed by the emotional responses subjectively generated by human observers when viewing images. The text annotation data covers seven emotion categories (i.e., anger, disgust, joy, fear, sadness, surprise, and neutral) and their corresponding description degrees. Each image is selected by multiple annotators according to subjective emotion judgment for emotion categories, and finally the description degree is calculated according to the voting ratio of each emotion category.
[0034] 2) In step S3-1-2 of Example 3, the way to obtain subsets is as follows: The original dataset is divided into 10 subsets using the ten-fold cross-validation method, and each subset includes 1782 feature vectors and their corresponding label distributions.
[0035] In Example 3, the method described in this application is tested with the same test strategy as existing imbalanced distribution learning methods such as the SA-BFGS method, EDLRL method, LDLSF method, LDL-LCLR method, Adam-LDL-SCL method, LDL-LDM method, OFR-FL method, OFR-CB method, OFR-DB method, RDA method, and DILDL method. The test results are shown in Table 1.
[0036] Example 4: Compared with Example 1, the difference between Example 4 and Example 1 lies in: 1). The imbalanced dataset used in Example 4 is the Flickr LDL dataset; the Flickr LDL dataset is the same as the Flickr LDL dataset disclosed in the paper "Learning Visual Sentiment Distributions via Augmented Conditional Probability Neural Network"; the Flickr LDL dataset contains text annotation data formed by the subjective emotional responses of human observers when viewing images, and the text annotation data includes annotation information of eight emotion categories (i.e., awe, disgust, fear, amusement, sadness, contentment, excitement, and anger) and their corresponding degrees of description. Each image is selected for emotion categories by multiple annotators, and finally the degree of description is calculated according to the voting proportion of each emotion category.
[0037] 2). In step S3-1-2 of Example 4, the subset acquisition method is as follows: The original dataset is divided into 10 subsets using the ten-fold cross-validation method, and each subset includes 10,035 feature vectors and their corresponding label distributions.
[0038] In Example 4, the method described in this application is tested using the same test strategy as existing imbalanced label distribution learning methods such as the SA-BFGS method, EDLRL method, LDLSF method, LDL-LCLR method, Adam-LDL-SCL method, LDL-LDM method, OFR-FL method, OFR-CB method, OFR-DB method, RDA method, and DILDL method. The test results are shown in Table 1.
[0039] Example 5: Compared with Example 1, the difference between Example 5 and Example 1 is as follows: 1). The imbalanced dataset used in Example 5 is the RAF-ML dataset; the RAF-ML dataset is the same as the RAF-ML dataset disclosed in the paper "Blended Emotion in-the-Wild: Multi-label Facial Expression Recognition Using Crowdsourced Annotations and Deep Locality Feature Learning"; the RAF-ML dataset contains text label data of facial images, and the text label data includes six emotion categories represented by the facial images (i.e., happy, sad, surprised, fearful, angry, and neutral) and their corresponding degrees of description. In this dataset, the degree of description of the emotion label is calculated based on the percentage of each emotion category in the annotator votes. 2), in step S3-1-2 of the fifth embodiment, the method for obtaining subsets is as follows: The original dataset is divided into 10 subsets using the ten-fold cross-validation method. Among them, 9 subsets each include 4417 feature vectors and their corresponding label distributions, and the remaining subset includes 4419 feature vectors and their corresponding label distributions.
[0040] In the fourth embodiment, the method of the present application and existing unbalanced label distribution learning methods such as the SA-BFGS method, EDL-LRL method, LDLSF method, LDL-LCLR method, Adam-LDL-SCL method, LDL-LDM method, OFR-FL method, OFR-CB method, OFR-DB method, RDA method, and DILDL method are tested using the same test strategy. The test results are shown in Table 1.
[0041] In the fifth embodiment, the method of the present application and existing unbalanced distribution learning methods such as the SA-BFGS method, EDLRL method, LDLSF method, LDL-LCLR method, Adam-LDL-SCL method, LDL-LDM method, OFR-FL method, OFR-CB method, OFR-DB method, RDA method, and DILDL method are tested using the same test strategy. The test results are shown in Table 1.
[0042] Embodiment Six: Compared with the first embodiment, the difference between the sixth embodiment and the first embodiment is as follows: 1), The unbalanced dataset used in the sixth embodiment is the Natural Scene dataset; the Natural Scene dataset is the same as the Natural Scene dataset disclosed in the paper "Label Distribution Learning"; the Natural Scene dataset contains text annotation data of natural scene images, and the text annotation data covers nine ground object categories (i.e., sun, cloud, sky, building, water, mountain, snow, desert, and plant) and their corresponding description degrees. In this dataset, the description degrees of the ground object categories are calculated by a nonlinear programming method based on the inconsistent label sorting results given by multiple annotators.
[0043] 2), In step S3-1-2 of the sixth embodiment, the method for obtaining subsets is as follows: The original dataset is divided into 10 subsets using the ten-fold cross-validation method, and each subset includes 1800 feature vectors and their corresponding label distributions.
[0044] In Embodiment 6, the method of the present application is tested using the same test strategy as existing imbalance label distribution learning methods such as the SA-BFGS method, the EDLRL method, the LDLSF method, the LDL-LCLR method, the Adam-LDL-SCL method, the LDL-LDM method, the OFR-FL method, the OFR-CB method, the OFR-DB method, the RDA method, and the DILDL method. The final error results of the test are shown in Table 1.
[0045] Table 1 Final Errors of Different Methods on Different Datasets
[0046] As can be seen from Table 1, the method of the present invention has improved performance on the above six datasets. On the Movie dataset (imbalance coefficient of 11.01), the final error obtained by the method of the present application is reduced by 7.42% compared to the DILDL method (the best-performing imbalance distribution learning method among the above existing imbalance distribution learning methods). On the SCUT-FBP dataset (imbalance coefficient of 19.37), the final error obtained by the method of the present application is reduced by 16.06% compared to the DILDL method. On the Emotion6 dataset (imbalance coefficient of 10.46), the final error obtained by the method of the present application is reduced by 9.93% compared to the DILDL method. On the FlickrLDL dataset (imbalance coefficient of 57.12), the final error obtained by the method of the present application is reduced by 8.24% compared to the DILDL method. On the RAF-ML dataset (imbalance coefficient of 30.33), the final error obtained by the method of the present application is reduced by 9.47% compared to the DILDL method. On the Natural Scene dataset (imbalance coefficient of 10.74), the final error obtained by the method of the present application is reduced by 8.75% compared to the DILDL method.
[0047] The above effects that can be achieved by the method described in this application are mainly due to the asymmetric momentum optimization mechanism introduced in this application and the inter-layer alignment loss calculated for the feature encoder and the full-label encoder. In the above settings, the inter-layer alignment loss of the feature encoder and the full-label encoder can effectively constrain the structural consistency of the encoding paths of the feature encoder and the full-label encoder in the latent representation space, effectively improving the inter-layer alignment effect of the output results of each corresponding encoding layer of the feature encoder and the full-label encoder. The setting of the asymmetric momentum optimization mechanism enables this application to effectively ensure that the momentum buffer variables of the main loss and the auxiliary loss do not interfere with each other during the update process, and controls the dominance of the momentum buffer variable of the main loss in the gradient update by setting the main loss momentum coefficient to be greater than the auxiliary loss momentum coefficient. Moreover, during the backpropagation process, it can also effectively achieve the smooth adjustment of the momentum buffer variables of the main loss and the auxiliary loss during the gradient update process, effectively suppressing the gradient fluctuations of the momentum buffer variables of the main loss and the auxiliary loss during the gradient update process, enhancing the stability of the optimization process of the momentum buffer variables of the main loss and the auxiliary loss, contributing to improving the convergence and robustness of the overall training process of the imbalanced label distribution learning network. Especially in the scenario of collaborative optimization of different loss terms, it shows higher training stability and obtains an imbalanced label distribution learning network model with better parameters. This also means that when using the imbalanced label distribution learning network model obtained by this application to predict an imbalanced dataset, it can predict a predicted label distribution that is closer to the true label distribution.
[0048] To verify the contributions of the inter-layer alignment loss between the feature encoder and the label encoder and the setting of the asymmetric momentum optimization mechanism in this application to the method described in this application, this application specifically conducted targeted ablation experiments based on the Movie dataset and used six evaluation metrics, namely Chebyshev distance, Clark distance, Canberra measure, KL divergence, cosine coefficient, and Intersection similarity, which are used to evaluate the distance or similarity between the true label distribution and the predicted label distribution, as the performance of the imbalanced label distribution learning network model. The test results are shown in Table 2; Table 2 Test Results of Ablation Experiments
[0049] In Table 2, the smaller the test results of Chebyshev distance, Clark distance, Canberra measure, and KL divergence, the better the performance; the larger the cosine coefficient and Intersection similarity, the better the performance. In this application, KL divergence is used to measure the difference between the predicted label distribution and the true label distribution. The smaller the KL divergence value, the closer the predicted label distribution is to the true label distribution.
[0050] As can be seen from Table 2, when the unbalanced label distribution learning network calculates the inter-layer alignment loss and adopts the asymmetric momentum optimization mechanism, that is, the method described in this application, evaluation metrics such as Chebyshev distance, Clark distance, Canberra measure, KL divergence, cosine coefficient, and Intersection similarity have all achieved significant performance improvements. For example, under the condition that the unbalanced label distribution learning network calculates the inter-layer alignment loss and adopts the asymmetric momentum optimization mechanism (i.e., the method described in this application), the test result of the KL divergence metric is 0.182. While in the case where the unbalanced label distribution learning network adopts the asymmetric momentum optimization mechanism and omits the calculation of the inter-layer alignment loss, the test result of the KL divergence metric is 0.187. Obviously, when the unbalanced label distribution learning network adopts the asymmetric momentum optimization mechanism and omits the calculation of the inter-layer alignment loss, the KL divergence metric increases by approximately 2.75% compared to the case where the unbalanced label distribution learning network calculates the inter-layer alignment loss and adopts the asymmetric momentum optimization mechanism. Moreover, as can also be seen from Table 2, when the unbalanced label distribution learning network omits the calculation of the inter-layer alignment loss and only adopts the asymmetric momentum optimization mechanism, compared to the method described in this application where the unbalanced label distribution learning network calculates the inter-layer alignment loss and adopts the asymmetric momentum optimization mechanism, the Chebyshev distance metric increases by 5.56% (from 0.162 to 0.171), the Clark distance metric increases by 2.32% (from 0.690 to 0.706), and the Canberra measure metric increases by 0.85% (from 1.295 to 1.306). In terms of the similarity metrics, the cosine coefficient decreases by 0.81% (from 0.864 to 0.857), and the Intersection similarity decreases by 0.40% (from 0.758 to 0.755). This shows that the predicted label distribution obtained when the unbalanced label distribution learning network described in this application calculates the inter-layer alignment loss and adopts the asymmetric momentum optimization mechanism is closer to the true label distribution than the predicted label distribution obtained when the unbalanced label distribution learning network omits the calculation of the inter-layer alignment loss and only adopts the asymmetric momentum optimization mechanism.
[0051] It can also be seen from Table 2 that when the unbalanced label distribution learning network only has the encoder alignment module and removes the asymmetric momentum optimization mechanism, compared with the method described in this application where the unbalanced label distribution learning network calculates the inter-layer alignment loss and adopts the asymmetric momentum optimization mechanism, the KL divergence index increases from 0.182 to 0.186, with a 2.20% increase in KL divergence, a 5.56% increase in the Chebyshev distance index, a 2.32% increase in the Clark distance index, and a 0.85% increase in the Canberra measure index; in terms of the similarity index, the cosine coefficient decreases by 0.81% and the Intersection similarity decreases by 0.40%. This shows that the predicted label distribution obtained by the method described in this application when the unbalanced label distribution learning network calculates the inter-layer alignment loss and adopts the asymmetric momentum optimization mechanism is closer to the true label distribution than the predicted label distribution obtained by the unbalanced label distribution learning network when omitting the calculation of the inter-layer alignment loss and only adopting the asymmetric momentum optimization mechanism.
Claims
1. An unbalanced label distribution learning method based on asymmetric momentum optimization, characterized in that: It includes the following steps: S1. Construct an imbalance label distribution learning network including a feature encoder, a full label encoder, and a decoder; S2. Calculate the total optimization loss of the imbalance label distribution learning network; the total optimization loss includes an inter-layer alignment loss, a feature-label distribution alignment loss, a feature-label similarity matrix alignment loss, and a label distribution prediction supervision loss; S3. Train the imbalance label distribution learning network based on the training set and the total optimization loss to obtain a network model; In step S3, a dominant momentum buffer variable and a dominant loss momentum coefficient are set for the dominant loss, and an auxiliary momentum buffer variable and an auxiliary loss momentum coefficient are set for the auxiliary loss. Then, during the training process, the network parameters θ in the next epoch are updated according to the dominant momentum buffer variable, the dominant loss momentum coefficient, the auxiliary momentum buffer variable, and the auxiliary loss momentum coefficient in the current epoch; the dominant loss is the label distribution prediction supervision loss, and the auxiliary loss includes the inter-layer alignment loss, the feature-label distribution alignment loss, and the feature-label similarity matrix alignment loss; S4. Obtain a feature vector and a label distribution based on the unbalanced dataset to be predicted, and then input them into the network model for one forward propagation to obtain a predicted label distribution and a true label distribution.
2. The method for learning the unbalance label distribution based on asymmetric momentum optimization according to claim 1, wherein: In step S2, the inter-layer alignment loss refers to the inter-layer alignment loss between the feature encoder and the full label encoder.
3. The method for learning the unbalance label distribution based on asymmetric momentum optimization according to claim 2, wherein: In step S2, the inter-layer alignment loss is obtained as follows: Extract the distribution parameters of the i-th encoding layer in the feature encoder and the i-th encoding layer in the full label encoder in the latent representation space; use the output results of each pair of corresponding encoding layers in the feature encoder and the full label encoder as the input for KL divergence calculation to obtain the distribution alignment loss of each corresponding encoding layer, and accumulate the distribution alignment losses of each layer to obtain the inter-layer alignment loss.
4. The method for learning the unbalance marker distribution based on asymmetric momentum optimization according to claim 1, characterized in that: Step S3 includes the following steps: Input the feature vectors in the training set into the feature encoder, and input the label distribution in the training set into the full label encoder, perform forward propagation, calculate the total optimization loss of the imbalance label distribution learning network, and perform backpropagation under the guidance of the total optimization loss to update the weight parameters of the imbalance label distribution learning network. During the process of updating the weight parameters of the imbalance label distribution learning network, the network parameters θ in the next epoch are updated according to the dominant momentum buffer variable, the dominant loss momentum coefficient, the auxiliary momentum buffer variable, and the auxiliary loss momentum coefficient in the current epoch. After 300 epochs of iteration, the training of one training segment is completed to obtain an imbalance label distribution learning network model. Among them, during the training of the imbalance label distribution learning network, a total of 10 training segments are trained to obtain 10 imbalance label distribution learning network models.
5. The method for learning the unbalance label distribution based on asymmetric momentum optimization according to claim 1 or 4, characterized in that: In step S3, when the next epoch is the t-th epoch, the network parameters θ in the training process of the t-th epoch are updated according to the dominant momentum buffer variable, the dominant loss momentum coefficient, the auxiliary momentum buffer variable, and the auxiliary loss momentum coefficient in the training process of the (t - 1)-th epoch. Specifically, it includes the following steps: a), Update the dominant momentum buffer variable in the training process of the (t - 1)-th epoch based on the dominant loss momentum coefficient, the gradient of the dominant loss, and the dominant momentum buffer variable in the training process of the (t - 1)-th epoch; b), Update the auxiliary momentum buffer variable in the training process of the (t - 1)-th epoch based on the auxiliary loss momentum coefficient, the gradient of the auxiliary loss, and the dominant momentum buffer variable in the training process of the (t - 1)-th epoch; c), Update the network parameters θ in the training process of the t-th epoch based on the network parameters θ in the training process of the (t - 1)-th epoch, the dominant momentum buffer variable, and the auxiliary momentum buffer variable in the training process of the t-th epoch.
6. The method for learning the unbalance label distribution based on asymmetric momentum optimization according to claim 5, characterized in that: In step S4, the dominant loss momentum coefficient is greater than the auxiliary loss momentum coefficient.
7. The method for learning unbalanced label distribution based on asymmetric momentum optimization according to claim 1, characterized in that: In step S4, the unbalanced dataset to be predicted is a dataset containing text data.
8. The method for learning the unbalance label distribution based on asymmetric momentum optimization according to claim 7, wherein: In step S4, the predicted dataset is one of the Movie dataset, the SCUT-FBP dataset, the Emotion6 dataset, the FlickrLDL dataset, the RAF-ML dataset, and the Natural Scene dataset.
9. The method for learning the unbalance label distribution based on asymmetric momentum optimization according to claim 1, characterized in that: Step S4 includes the following steps: Obtain the feature vector and the label distribution respectively based on the unbalanced dataset to be predicted, and then input the feature vector and the label distribution into the 10 unbalanced label distribution learning network models obtained in step S3 for one forward propagation, and output 10 predicted label distributions and their corresponding 10 true label distributions.
Citation Information
Patent Citations
Association rule determination method and device for unbalanced sample
CN113282686A
Disease prediction method based on unbalanced fundus image data
CN117576012A
Multi-mark image classification method and device based on saliency features and readable medium
CN119169386A
Unbalance mark distribution learning method based on decoupling operation
CN119962705A
Systems and methods for partially supervised learning with momentum prototypes
US20220067506A1
Cited By
Unbalanced mark distribution learning method suitable for movie recommendation
CN120952112A
Imbalanced Label Distribution Learning Method for Movie Recommendation
CN120952112B
Unbalance mark distribution learning method suitable for face attraction evaluation
CN121789016A