Unbalanced Label Distribution Learning Method Based on Decoupled Operations
By adopting the decoupling operation and loss balance method in the unbalanced label distribution learning, the problem of excessive focus on the distribution of dominant labels in the prior art is solved, and a more accurate prediction of label distribution and better learning effect of non-dominant label distribution is achieved.
Patent Information
- Application Number
- CN202510450586.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-11
AI Technical Summary
In imbalanced label distribution learning, the prior art is prone to over-focusing on the dominant label distribution and ignoring the non-dominant label distribution, resulting in the neural network being unable to effectively learn and utilize label information.
A method of unbalanced mark distribution learning based on decoupling operation is proposed. By constructing a network including encoder I, encoder II and encoder III, the initial decoupling and secondary decoupling operations are used to obtain advantageous and non-advantage mark distributions, calculate various losses and balance parameters, and ensure the gradient transmission and learning effect of non-advantage mark distributions.
It effectively avoids mutual interference between the dominant marker distribution and the non-dominant marker distribution, ensures that the predicted marker distribution is closer to the real marker distribution, and improves the learning effect of the non-dominant marker distribution and the overall performance of the model.
Smart Images

Figure CN119962705B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and particularly relates to an unbalanced label distribution learning method based on decoupling operations. Background Art
[0002] In the wide application of machine learning, label distribution learning (LDL) is a technical means aimed at constructing an effective mapping relationship between instances (i.e., feature vectors) and label distributions. At present, it has shown great potential in many fields such as movie rating analysis, facial expression analysis, video parsing, and age estimation. However, the label distributions in actual application scenarios often exhibit significant imbalance. Taking the field of movie rating analysis as an example, in an ideal situation, it is expected to obtain a balanced distribution of sentiment label distributions. However, due to many factors such as limited number of annotators, background differences, subjective opinion differences, and annotation noise, the actual obtained label distribution often presents an unbalanced distribution state. In this state where the labels are unbalancedly distributed, the dominant label distribution occupies most of the description space, making the description degree of the non-dominant label distribution extremely low, seriously affecting the comprehensive learning and utilization of label information by the neural network.
[0003] Therefore, the applicant has designed an unbalanced label distribution learning method that can effectively avoid over-focusing on the dominant label distribution and ignoring the non-dominant label distribution in unbalanced label distribution learning. Summary of the Invention
[0004] In order to solve the above deficiencies existing in the prior art, the present application proposes an unbalanced label distribution learning method based on decoupling operations.
[0005] An unbalanced label distribution learning method based on decoupling operations includes the following steps:
[0006] S1. Construct an unbalanced label distribution learning network including Encoder I, Encoder II, and Encoder III. Encoder I and Encoder II respectively obtain feature encoding representations and global label distributions based on the input feature vectors and input label distributions. Encoder III is used to encode the data obtained by initially decoupling the input label distribution to obtain the dominant label distribution and the non-dominant label distribution. The decoder is used to obtain the predicted label distribution from the feature encoding representation, and is also used to decode the global label distribution to obtain the true label distribution. The predicted label distribution is secondarily decoupled to obtain the predicted dominant class label distribution and non-dominant class label distribution. The input feature vector is the feature vector input into the unbalanced label distribution learning network, and the input label distribution is the label distribution input into the unbalanced label distribution learning network.
[0007] S2. Calculate the alignment loss of the dominant marker distribution information and the alignment loss of the non-dominant marker distribution information respectively based on the feature encoding representation, the dominant marker distribution, and the non-dominant marker distribution; construct the feature similarity matrix A based on the feature encoding representation, the global marker distribution, and the cosine similarity mn and the marker similarity matrix Z mn , and calculate the alignment loss of the feature and marker representation similarity matrix based on matrices A mn and Z mn ; calculate the alignment loss based on the predicted marker distribution and the true marker distribution; calculate the dominant marker distribution loss and the non-dominant marker distribution loss respectively based on the predicted dominant class marker distribution, non-dominant class marker distribution, and true marker distribution; use the parameters α, β, γ, and λ to balance the above losses to obtain the total optimization loss;
[0008] S3. Train the imbalanced marker distribution learning network based on the training set and the total optimization loss to obtain a network model;
[0009] S4. Obtain the feature vector and marker distribution based on the imbalanced dataset to be predicted, and then input them into the network model for forward propagation once respectively to obtain the predicted marker distribution and the true marker distribution.
[0010] Preferably, in step S1, the steps of the first decoupling and the second decoupling are the same. The first decoupling decouples the input marker distribution, and the second decoupling decouples the predicted marker distribution. The second decoupling specifically includes the following steps:
[0011] Set the description degree of the dominant marker (the dominant marker is the marker with the highest description degree) in the predicted marker distribution to 1, and the description degrees of other markers to 0 to obtain the dominant class marker distribution; set the description degree of the dominant marker in the predicted marker distribution to 0, and normalize the non-dominant markers (non-dominant markers refer to markers other than the dominant marker) using softmax so that the total description degree of the non-dominant markers is 1 to obtain the non-dominant marker distribution.
[0012] Preferably, step S2 specifically includes the following steps:
[0013] S2-1. Calculate the alignment loss L alig 1 of the dominant marker distribution information using the KL divergence between the feature encoding representation and the dominant marker distribution;
[0014] S2-2. Calculate the alignment loss L alig 2 of the non-dominant marker distribution information using the KL divergence between the feature encoding representation and the non-dominant marker distribution;
[0015] S2-3. Add noise δ feature to the variance σ feature of the feature encoding representation, and then add it to the mean μ of the feature encoding representationfeature Through reparameterization, perform the calculation shown in (1) to generate the feature representation r feature ;
[0016] r feature = μ feature + σ feature δ feature , δ feature ~N(0,1) (1)
[0017] In equation (1), N(0,1) represents the standard normal distribution;
[0018] Then, use cosine similarity to calculate the similarity between the feature representation corresponding to the m-th instance and the feature representation corresponding to the n-th instance After that, construct a two-dimensional feature similarity matrix A based on the above similarity mn ;
[0019] Add noise δ to the variance σ of the global label distribution whole , and then add it to the mean μ of the global label distribution whole , and then through reparameterization, perform the calculation shown in (2) to generate the global label distribution representation r whole ; labelDis
[0020]
[0020] r labelDis = μ whole + σ whole δ whole , δ whole ~N(0,1) (2)
[0021] In equation (2), N(0,1) represents the standard normal distribution;
[0022] Then, use cosine similarity to calculate the similarity between the feature representation corresponding to the m-th instance and the feature representation corresponding to the n-th instance After that, construct a two-dimensional label similarity matrix Z based on the above similarity mn .
[0023] Use the KL divergence to calculate the alignment loss L of the feature and label representation similarity matrix for matrices A mn and Z mn 3; alig 3;
[0024] S2-4. Align the predicted label distribution and the true label distribution using the KL divergence, and calculate the alignment loss L between the predicted label distribution and the true label distributionalig 4;
[0025] S2-5. Align the predicted dominant class label distribution with the true label distribution using KL divergence, and calculate the dominant label distribution loss L DDL ; Align the predicted non-dominant class label distribution with the true label distribution using KL divergence, and calculate the non-dominant label distribution loss L NDDL ; The above settings can effectively avoid the mutual interference during the alignment process of the predicted dominant class label distribution with the true label and the alignment process of the non-dominant class label distribution with the true label distribution.
[0026] Preferably, in step S3, the imbalance label distribution learning network is trained for 10 training segments based on the training set and the total optimization loss to obtain 10 network models.
[0027] Preferably, step S3 includes the following steps:
[0028] S3-1. Obtain the training set and the test set;
[0029] S3-2. Input the feature vectors in the training set into Encoder I, input the label distribution in the training set into Encoder II and Encoder III, perform forward propagation, calculate the total optimization loss of the imbalance label distribution learning network, and perform backpropagation under the guidance of the total optimization loss to update the weight parameters of the imbalance label distribution learning network, completing the training process of one epoch. After iterating 300 epochs of the training process, the training of one training segment is completed to obtain an imbalance label distribution learning network model; during the process of training the imbalance label distribution learning network in this application, 10 training segments are trained.
[0030] Preferably, in step S3-1, the original data set is divided into 10 subsets using the ten-fold cross-validation method; in each training segment, 9 of these subsets are selected as the training set, and the remaining subset is used as the test set. In the 10 training segments of this application, the selected test sets are all different.
[0031] Preferably, in step S4, the feature vectors and label distribution are obtained based on the imbalance data set to be predicted, and then are respectively input into the 10 network models for one-time forward propagation to obtain 10 predicted label distributions and 10 corresponding true label distributions.
[0032] Compared with the prior art, the beneficial technical effects of the present invention are:
[0033] In the method of the present invention, the settings of the primary decoupling and secondary decoupling of the present invention enable the dominant marker distribution and the non-dominant marker distribution to effectively avoid interference with each other during the alignment-1 process, the alignment-2 process, and the alignment process between the predicted marker distribution output by the decoder and the true marker distribution, so that the predicted marker distribution output by the decoder of the present application is closer to the true marker distribution;
[0034] Moreover, in the present application, when calculating the total optimization loss, the balance parameter of the dominant marker distribution loss L DDL and the non-dominant marker distribution loss L NDDL is set to α, and the balance parameter of the dominant marker distribution information alignment loss L alig 1 (the loss generated during the alignment-1 process) and the non-dominant marker distribution information alignment loss L alig 2 (the loss generated during the alignment-2 process) is also set to α, where α = 0.1 to 0.9. The above settings enable the gradient of the non-dominant marker distribution to still be effectively transmitted during the learning process of the unbalanced marker distribution learning network, thus ensuring the learning effect of the non-dominant marker distribution; moreover, through the settings of the balance parameters α, β, γ, and λ in the present application, the dominant marker distribution loss L DDL , the non-dominant marker distribution loss L NDDL 、 the dominant marker distribution information alignment loss L alig 1, the non-dominant marker distribution information alignment loss L alig 2, the feature and marker representation similarity matrix alignment loss L alig 3, and the predicted marker distribution and true marker distribution alignment loss L alig 4 together constitute the total optimization loss ; in the present application, the construction of the total optimization loss can make the description degree of the non-dominant marker distribution independent of the description degree of the dominant marker distribution during the learning process of the unbalanced marker distribution learning network, and further enable the unbalanced marker distribution learning network to better learn the description degree of the dominant marker distribution and the description degree of the non-dominant marker distribution, so that the predicted marker distribution is closer to the true marker distribution. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a network structure diagram of the unbalanced marker distribution learning network in the present application. DETAILED DESCRIPTION OF THE INVENTION
[0036] Example 1:
[0037] A method for learning the distribution of unbalanced labels based on decoupling operations. The method for predicting an unbalanced dataset based on decoupling operations predicts an unbalanced dataset, namely the Movie dataset, and includes the following steps:
[0038] S1. Construct an unbalanced label distribution learning network including Encoder I, Encoder II, and Encoder III, as Figure 1 shown; in this application, Encoder I and Encoder II respectively obtain feature encoding representations and global label distributions based on the input feature vectors and input label distributions; Encoder III is used to encode the data obtained by initially decoupling the input label distribution to obtain a dominant label distribution and a non-dominant label distribution; the decoder is used to decode the feature encoding representation and the global label distribution to obtain a predicted label distribution and a true label distribution; the predicted label distribution is secondarily decoupled to obtain a predicted dominant class label distribution and a non-dominant class label distribution; the input feature vector refers to the feature vector input to the unbalanced label distribution learning network; the input label distribution refers to the label distribution input to the unbalanced label distribution learning network; specifically:
[0039] Encoder I encodes the input feature vector and maps the encoded feature vector into the latent space to obtain a feature encoding representation (μ feature , σ feature ) composed of the mean and variance, where μ feature represents the mean of the feature encoding representation, and σ feature represents the variance of the feature encoding representation;
[0040] The input label distribution is initially decoupled to respectively obtain a distribution with dominant labels and a distribution with non-dominant labels. The distribution with dominant labels and the distribution with non-dominant labels are respectively encoded by Encoder III to obtain a dominant label distribution (μ DDL , σ DDL ) and a non-dominant label distribution (μ NDDL , σ NDDL ) composed of variance and mean; where μ DDL represents the mean of the dominant label distribution, and σ DDL represents the variance of the dominant label distribution; μ NDDL represents the mean of the non-dominant label distribution, and σ NDDL represents the variance of the non-dominant label distribution;
[0041] Encoder II globally encodes the input label distribution and maps the encoded label distribution into the latent space to learn the feature representation of the global label distribution, obtaining a global label distribution (μ whole , σ whole ) composed of variance and mean, μ wholerepresents the mean of the global label distribution, σ whole represents the mean of the global label distribution;
[0042] The decoder decodes the feature encoding representation output by Encoder I to obtain the predicted label distribution, and decodes the global label distribution (μ whole , σ whole ) output by Encoder II to obtain the true label distribution;
[0043] Then, the predicted label distribution is decoupled twice to obtain the predicted dominant class label distribution and the predicted non-dominant class label distribution.
[0044] In this application, in step S1, the steps of the first decoupling and the second decoupling are the same. The first decoupling decouples the input label distribution, and the second decoupling decouples the predicted label distribution. Taking the second decoupling as an example, the decoupling steps of this application are introduced. The second decoupling specifically includes the following steps: Denote the description degree of the dominant label (the dominant label is the label with the highest description degree) in the predicted label distribution as 1, and the description degrees of other labels as 0 to obtain the dominant class label distribution; Denote the description degree of the dominant label in the predicted label distribution as 0, and normalize the non-dominant labels (the non-dominant labels refer to the labels other than the dominant label) using softmax so that the sum of the description degrees of the non-dominant labels is 1 to obtain the non-dominant label distribution.
[0045] In this application, the output end of Encoder I is connected to the input end of the decoder, the output end of Encoder III is connected to the input end of the decoder, and the output end of Encoder II is connected to the input end of the decoder; In this application, Encoder I is essentially a feature encoder, Encoder II is essentially a full label encoder, and Encoder III is essentially a dominant / non-dominant label distribution encoder; Among them, the structures of Encoder I, Encoder II, and Encoder III are the same and are all consistent with the encoder disclosed in the paper "Imbalanced Label Distribution Learning"; In this application, the label refers to the tag, the dominant label is the dominant tag, and the non-dominant label is the non-dominant tag;
[0046] S2. Calculate the total optimization loss of the imbalanced label distribution learning network; specifically as follows:
[0047] S2-1. Use the KL divergence to calculate the alignment loss L feature , σ feature ) of the dominant label distribution between the feature encoding representation (μ DDL , σ DDL ); In this application, during the training process of the imbalanced label distribution learning network, the existence of the alignment loss Lalig1 of the dominant label distribution can constrain the feature encoding representation (μ alig 1;feature ,σ feature ) and the dominant marker distribution (μ DDL ,σ DDL ) distribution deviation, so that the eigenvector (μ feature ,σ feature ) and the dominant marker distribution (μ DDL ,σ DDL ) are closer in the latent space to achieve the feature vector (μ feature ,σ feature ) and the dominant marker distribution (μ DDL ,σ DDL ) in the latent space, called alignment-1;
[0048] S2-2, the feature encoding representation (μ feature ,σ feature ) and non-dominant marker distribution (μ NDDL ,σ NDDL ) Use KL divergence to calculate the non-dominant label distribution information alignment loss L alig 2; In this application, during the training of the imbalanced label distribution learning network, the non-dominant label distribution information alignment loss L alig The existence of 2 can constrain the eigenvector (μ feature ,σ feature ) and the distribution of non-dominant markers (μ NDDL ,σ NDDL ) distribution deviation, so that the eigenvector (μ feature ,σ feature ) and the distribution of non-dominant markers (μ NDDL ,σ NDDL ) are closer in the latent space to achieve the feature vector (μ feature ,σ feature ) and the distribution of non-dominant markers (μ NDDL ,σ NDDL ) in the latent space, referred to as alignment-2;
[0049] S2-3. Variance σ of feature encoding representation feature Add noise δ feature , and then with the mean μ represented by the feature code feature By reparameterization, the calculation shown in (1) is performed to generate the feature representation r feature ;
[0050] r feature = μ feature + σ feature δ feature ,δ feature ~N(0,1) (1)
[0051] In formula (1), N(0,1) represents the standard normal distribution;
[0052] Then, the cosine similarity is used to calculate the feature representation corresponding to the m-th instance and the feature representation corresponding to the n-th instance The similarity between them is calculated. Then, a two-dimensional feature similarity matrix is constructed based on the above similarity This matrix is denoted as matrix A mn ;
[0053] The variance σ of the global label distribution whole Add noise δ whole Then, it is reparameterized and calculated as shown in (2) with the mean μ of the global label distribution whole to generate the global label distribution representation r labelDis ;
[0054] r labelDis = μ whole + σ whole δ whole , δ whole ~N(0,1) (2)
[0055] In formula (2), N(0,1) represents the standard normal distribution;
[0056] Then, the cosine similarity is used to calculate the feature representation corresponding to the m-th instance and the feature representation corresponding to the n-th instance The similarity between them is calculated. Then, a two-dimensional label similarity matrix is constructed based on the above similarity This matrix is denoted as matrix Z mn .
[0057] The matrices A mn and Z mn are used to calculate the alignment loss L of the feature and label representation similarity matrix using the KL divergence alig 3;
[0058] S2-4. Align the predicted label distribution and the true label distribution using the KL divergence, and calculate the alignment loss L of the predicted label distribution and the true label distribution alig 4;
[0059] S2-5. Align the predicted dominant class label distribution and the true label distribution using the KL divergence, and calculate the dominant label distribution loss L DDLAlign the predicted non-dominant class label distribution with the true label distribution using KL divergence to calculate the non-dominant label distribution loss L NDDL ; The above settings can effectively avoid the mutual interference during the alignment process of the predicted dominant class label distribution with the true label and the alignment process of the non-dominant class label distribution with the true label distribution;
[0060] By setting the balance parameters α, β, γ, and λ, the dominant label distribution loss L DDL and the non-dominant label distribution loss L NDDL 、 The dominant label distribution information alignment loss L alig 1, the non-dominant label distribution information alignment loss L alig 2, the feature and label representation similarity matrix alignment loss L alig 3 and the predicted label distribution and true label distribution alignment loss L alig 4 together constitute the total optimization loss ; In this application, the construction of the above loss functions can make the description degree of the non-dominant label distribution independent of the description degree of the dominant label distribution during the learning process of the imbalanced label distribution learning network, so that the imbalanced label distribution learning network can better learn the description degree of the dominant label distribution and the non-dominant label distribution, thereby making the predicted label distribution closer to the true label distribution.
[0061] In this application, the calculation method of the total optimization loss is shown in Equation (3):
[0062] + (3)
[0063] In Equation (3), α, β, γ, and λ are all balance parameters, which are used to adjust the weights of each part of the loss during the training process of the imbalanced label distribution learning network to ensure that the imbalanced label distribution learning network model can achieve the best performance balance in different datasets and tasks; in the first embodiment, the values of α, β, γ, and λ are 0.4, 0.1, 0.1, and 0.1 respectively.
[0064] In Equation (3), the dominant label distribution loss is used to calculate the loss generated during the alignment process of the dominant class label distribution with the true label distribution using KL divergence. The setting of the dominant label distribution loss can enable the imbalanced label distribution learning network to effectively avoid over-focusing on the dominant label distribution and ignoring the non-dominant label distribution under the imbalanced label distribution; the dominant label distribution loss The calculation method is consistent with Formula (1) disclosed in "Imbalanced Label Distribution Learning".
[0065] Non-dominant label distribution loss It is used to calculate the loss generated during the alignment process of the non-dominant class label distribution and the true label distribution using KL divergence; non-dominant label distribution loss The calculation method is consistent with Formula (1) disclosed in "Imbalanced Label Distribution Learning".
[0066] Dominant label distribution information alignment loss Calculate the loss generated during the alignment process of the feature vector (μ feature , σ feature ) and the dominant label distribution (μ DDL , σ DDL ) in the latent space, so that the imbalanced label distribution learning network can better capture the relationship between the dominant label distribution and the feature vector during the learning process; among them, the dominant label distribution information alignment loss The calculation method is as shown in Equation (4):
[0067] (4)
[0068] In Equation (4), represents the variance ratio calculated from the i-th dominant label distribution, represents the sum of the squares of the mean differences calculated from the i-th dominant label distribution;
[0069] In Equation (4), The calculation method is as shown in Equation (5):
[0070] = (5)
[0071] In Equation (5), represents the variance calculated from the i-th feature coding representation output by Encoder Ⅰ, represents the variance calculated from the i-th dominant label distribution output by the dominant / non-dominant label encoder.
[0072] In Equation (4), The calculation method is as shown in Equation (6):
[0073] = (6)
[0074] In Equation (6), represents the mean calculated from the i-th feature encoding representation output by Encoder I, represents the mean calculated from the i-th dominant label distribution output by the dominant / non-dominant label encoder, represents the variance calculated from the k-th dominant label distribution output by the dominant / non-dominant label encoder; in this application, the dominant label is essentially the leading label, and the non-dominant label is essentially the non-leading label;
[0075] Non-dominant label distribution information alignment loss Calculate the loss generated during the alignment process of the feature vector (μ feature , σ feature ) and the non-dominant label distribution (μ NDDL , σ NDDL ) in the latent space through KL divergence; setting the above loss helps the unbalanced label distribution learning network learn more accurate association information between the non-dominant label distribution and the feature vector; Non-dominant label distribution information alignment loss is calculated as shown in Equation (7);
[0076] (7)
[0077] In Equation (7), represents the variance ratio calculated from the i-th non-dominant label distribution, represents the sum of the squares of the mean differences calculated from the i-th non-dominant label distribution;
[0078] In Equation (7), is calculated as shown in Equation (8):
[0079]
[0080] In Equation (8), represents the variance calculated from the i-th feature encoding representation output by Encoder I, represents the variance calculated from the i-th non-dominant label distribution output by the dominant / non-dominant label encoder;
[0081] In Equation (7), is calculated as shown in Equation (9):
[0082] (9)
[0083] In Equation (9), represents the mean calculated from the i-th feature encoding representation output by Encoder I, represents the mean calculated from the i-th non-dominant label distribution output by the dominant / non-dominant label encoder, Denotes the variance calculated from the k-th non-dominant token distribution of the dominant / non-dominant token encoder output;
[0084] Alignment loss of the feature and token representation similarity matrix By calculating the two-dimensional token similarity matrix Z mn Aligning to the two-dimensional feature similarity matrix A mn The loss generated during the alignment process; in this application, obtaining the two-dimensional feature similarity matrix A mn And the two-dimensional token similarity matrix Z mn The purpose is to calculate the alignment loss Lalig3 of the feature and token representation similarity matrix; among them, the calculation method of the similarity matrix alignment loss Lalig3 is as shown in Equation (10),
[0085] (10)
[0086] In Equation (10), M represents the number of feature vectors input to the unbalanced token distribution learning network.
[0087] Alignment loss between the predicted token distribution and the true token distribution The KL divergence is used to calculate the difference between the predicted token distribution output by the decoder and the true token distribution. The setting of the above loss can effectively align the predicted token distribution and the true token distribution; Alignment loss between the predicted token distribution and the true token distribution The calculation method is as shown in Equation (11).
[0088] (11)
[0089] In Equation (11), r labelDis Denotes adding noise δ whole To the variance σ of the global token distribution whole And then calculating through reparameterization with the mean μ of the global token distribution whole To generate the global token distribution; D represents the true token distribution; () represents the decoder; Denotes the KL divergence.
[0090] S3. Obtain the training set and the test set, and train the unbalanced token distribution learning network for 10 training segments based on the training set and the total optimization loss to obtain 10 unbalanced token distribution learning network models, including the following steps:
[0091] S3-1. Obtain the training set and the test set, specifically including the following specific steps:
[0092] S3-1-1. Obtain the original dataset, where the original dataset is the Movie dataset [Geng, 2016]; the Movie dataset [Geng, 2016] is from the paper "Label Distribution Learning", and this dataset contains basic movie-related annotation data. The Movie dataset [Geng, 2016] is a dataset of user ratings for movies. This dataset contains 7,755 movies and 54,242,292 rating records from 478,656 different users (the data is sourced from the Netflix platform, and the rating range is from 1 to 5 stars, with a total of 5 levels). The rating label distribution for each movie is calculated by the percentage of each rating level. Feature extraction is based on metadata, covering dimensions such as genre, director, actor, country, budget, etc., where categorical attributes are converted into binary vectors. The final feature vector extracted from each movie is 1,869-dimensional. This original dataset contains 7,755 feature vectors and 5 label distributions.
[0093] S3-1-2. Use the ten-fold cross-validation method to divide the original dataset into 10 subsets, where 9 subsets each include 775 feature vectors and their corresponding label distributions, and the remaining subset includes 780 feature vectors and their corresponding label distributions; during the process of training the unbalanced label distribution learning network in this application, 10 training segments need to be trained, and each training segment includes a training process of 300 epochs. In each training segment, 9 of the subsets are selected as the training set, and the remaining subset is used as the test set. In the 10 training segments of this application, the selected test sets are all different.
[0094] S3-2. Input the feature vectors in the training set into Encoder I, and input the label distributions in the training set into Encoder II and Encoder III. Perform forward propagation to calculate the total optimization loss of the unbalanced label distribution learning network, and perform backpropagation under the guidance of the total optimization loss to update the weight parameters of the unbalanced label distribution learning network, completing one epoch of the training process. After iterating 300 epochs of the training process, one training segment of the training is completed, and an unbalanced label distribution learning network model is obtained; during the process of training the unbalanced label distribution learning network in this application, 10 training segments need to be trained. Therefore, during the training process of the unbalanced label distribution learning network in this application, 10 unbalanced label distribution learning network models can be obtained. During the process of training the unbalanced label distribution learning network in this application, the learning rate is set to 0.001 to ensure that the model can converge stably during the training process; the batch size is set to 50, which takes into account both the training efficiency and the full learning of the data by the unbalanced label distribution learning network model.
[0095] S4. Respectively obtain the feature vectors and label distributions based on the imbalanced dataset with an imbalanced movie rating label distribution. Then, input the feature vectors and label distributions into the 10 imbalanced label distribution learning network models obtained in step S3 for one forward propagation, and output 10 predicted label distributions and their corresponding 10 true label distributions. Among them, the imbalanced dataset refers to the dataset with an imbalanced movie rating label distribution; in this application, the method of respectively obtaining the feature vectors and label distributions based on the imbalanced dataset with an imbalanced movie rating label distribution to be predicted is the same as the method of obtaining the feature vectors and label distributions described in the part of A.1.1 Description of Datasets •Movie (movie rating) in the paper "Imbalanced Label Distribution Learning". In this application, since the predicted label distribution output by the decoder contains the rating of the movie, existing movie recommendation systems (such as the movie recommendation systems publicly available on MovieLens and Netflix Prize) can predict the rating of the movie based on the above 10 predicted label distributions, so as to recommend movies with higher predicted ratings to users.
[0096] Test:
[0097] The method described in this application (abbreviated as Ours method in Table 1) is tested using the same test strategy as existing imbalanced distribution learning methods such as the SA-BFGS method (from Xin Geng's paper "Label Distribution Learning"), the EDLRL method (from Xiuyi Jia et al.'s paper "Facial Emotion Distribution Learning by Exploiting Low-Rank Label Correlations Locally"), the LDLSF method (from Tingting Ren et al.'s paper "Label Distribution Learning with Label-Specific Features"), the LDL-LCLR method (from Tingting Ren et al.'s paper "Label Distribution Learning with Label Correlations via Low-Rank Approximation"), the Adam-LDL-SCL method (from Xiuyi Jia et al.'s paper "Label Distribution Learning with Label Correlations on Local Samples"), the LDL-LDM method (from Jing Wang et al.'s paper "Label Distribution Learning by Exploiting Label Distribution Manifold"), the OFR-FL method (from Tsung-Yi Lin et al.'s paper "Focal Loss for Dense Object Detection"), the OFR-CB method (from Tsung-Yi Lin et al.'s paper "Focal Loss for Dense Object Detection"), the OFR-DB method (from Tsung-Yi Lin et al.'s paper "Focal Loss for Dense Object Detection"), the RDA method (from Xingyu Zhao et al.'s paper "Imbalanced Label Distribution Learning"), and the DILDL method (from Xingyu Zhao et al.'s paper "Imbalanced Label Distribution Learning"). The test results are shown in Table 1;The test strategy adopted in this application is as follows: Based on the test sets in 10 training segments and the corresponding unbalanced label distributions, the learning network models are tested respectively, and the Chebyshev distance values between the predicted label distribution and the true label distribution output by each unbalanced label distribution learning network model are calculated based on the Chebyshev distance method. Then, the average value of these Chebyshev distances is obtained, which is the average Chebyshev distance. Then, the absolute value of the difference between each Chebyshev distance value and the average Chebyshev distance is calculated, and the maximum absolute value is taken as the fluctuation error. Adding this fluctuation error to the average Chebyshev distance gives the final error, which can show the difference between the predicted label distribution and the true label distribution.
[0098] Embodiment 2:
[0099] Compared with Embodiment 1, the difference between Embodiment 2 and Embodiment 1 is as follows:
[0100] 1). The unbalanced dataset adopted in Embodiment 2 is the SCUT - FBP dataset; the SCUT - FBP dataset in this application is the same as the SCUT - FBP dataset disclosed in the paper "SCUT-FBP: A Benchmark Dataset for Facial Beauty Perception". This dataset focuses on the research of facial beauty perception and contains the beauty score annotations of facial images, with the score range from 1 to 5.
[0101] 2). In step S3-1-2 of Embodiment 2, the subset acquisition method is as follows: The original dataset is divided into 10 subsets using the ten-fold cross-validation method, and each subset includes 3600 feature vectors and their corresponding label distributions.
[0102] In Embodiment 2, the method of this application and existing unbalanced label distribution learning methods such as SA-BFGS method, EDLRL method, LDLSF method, LDL-LCLR method, Adam-LDL-SCL method, LDL-LDM method, OFR-FL method, OFR-CB method, OFR-DB method, RDA method, and DILDL method are tested using the same test strategy, and the test results are shown in Table 1.
[0103] Embodiment 3:
[0104] Compared with Embodiment 1, the difference between Embodiment 3 and Embodiment 1 is as follows:
[0105] 1), The imbalanced dataset used in Example 2 is the Emotion6 dataset; the Emotion6 dataset is the same as the Emotion6 dataset disclosed in the paper "A mixed bag of emotions: Model, predict, and transfer emotion distributions", and this dataset contains labeled data of seven emotion categories (anger, disgust, joy, fear, sadness, surprise, neutral).
[0106] 2), In step S3-1-2 of Example 2, the way to obtain subsets is: using the ten-fold cross-validation method to divide the original dataset into 10 subsets, and each subset includes 1782 feature vectors and their corresponding label distributions.
[0107] In Example 3, the method of the present application and existing imbalanced distribution learning methods such as the SA-BFGS method, EDLRL method, LDLSF method, LDL-LCLR method, Adam-LDL-SCL method, LDL-LDM method, OFR-FL method, OFR-CB method, OFR-DB method, RDA method, and DILDL method are tested using the same test strategy, and the test results are shown in Table 1.
[0108] Example 4:
[0109] Compared with Example 1, the difference between Example 4 and Example 1 is that:
[0110] 1), The imbalanced dataset used in Example 2 is the Flickr LDL dataset; the Flickr LDL dataset is the same as the Flickr LDL dataset disclosed in the paper "Learning Visual Sentiment Distributions via Augmented Conditional Probability Neural Network"; this dataset contains labeled data of eight emotion categories (namely awe, disgust, fear, amusement, sadness, contentment, excitement, and anger).
[0111] 2), In step S3-1-2 of Example 2, the way to obtain subsets is: using the ten-fold cross-validation method to divide the original dataset into 10 subsets, and each subset includes 10035 feature vectors and their corresponding label distributions.
[0112] In the fourth embodiment, the method of the present application is tested using the same test strategy as existing unbalanced label distribution learning methods such as the SA-BFGS method, the EDLRL method, the LDLSF method, the LDL-LCLR method, the Adam-LDL-SCL method, the LDL-LDM method, the OFR-FL method, the OFR-CB method, the OFR-DB method, the RDA method, and the DILDL method. The test results are shown in Table 1.
[0113] Embodiment Five:
[0114] Compared with Embodiment One, the difference between Embodiment Five and Embodiment One is as follows:
[0115] 1). The unbalanced dataset used in Embodiment Two is the RAF-ML dataset; the RAF-ML dataset is the same as the RAF-ML dataset disclosed in the paper "Blended Emotion in-the-Wild: Multi-label Facial Expression Recognition Using Crowdsourced Annotations and Deep Locality Feature Learning"; this dataset contains annotation data of six expression distributions (i.e., happy, sad, surprised, fearful, angry, and neutral).
[0116] 2). In step S3-1-2 of Embodiment Two, the subset acquisition method is as follows: the original dataset is divided into 10 subsets using the ten-fold cross-validation method. Among them, 9 subsets each include 4417 feature vectors and their corresponding label distributions, and the remaining subset includes 4419 feature vectors and their corresponding label distributions.
[0117] In Embodiment Five, the method of the present application is tested using the same test strategy as existing unbalanced distribution learning methods such as the SA-BFGS method, the EDLRL method, the LDLSF method, the LDL-LCLR method, the Adam-LDL-SCL method, the LDL-LDM method, the OFR-FL method, the OFR-CB method, the OFR-DB method, the RDA method, and the DILDL method. The test results are shown in Table 1.
[0118] Embodiment Six:
[0119] Compared with Embodiment One, the difference between Embodiment Six and Embodiment One is as follows:
[0120] 1), the imbalanced dataset used in Embodiment 2 is the Natural Scene dataset; the Natural Scene dataset is the same as the Natural Scene dataset disclosed in the paper "Label Distribution Learning"; this dataset contains nine related labels, including sun, cloud, sky, building, water, mountain, snow, desert, and plant;
[0121] 2), in step S3-1-2 of Embodiment 2, the way to obtain subsets is: using the ten-fold cross-validation method to divide the original dataset into 10 subsets, and each subset includes 1800 feature vectors and their corresponding label distributions.
[0122] In Embodiment 6, the method of the present application is tested with the same test strategy as existing imbalanced label distribution learning methods such as the SA-BFGS method, EDLRL method, LDLSF method, LDL-LCLR method, Adam-LDL-SCL method, LDL-LDM method, OFR-FL method, OFR-CB method, OFR-DB method, RDA method, and DILDL method. The test results are shown in Table 1.
[0123] Table 1 Test results of different methods on different datasets
[0124]
[0125] As can be seen from Table 1, the method of the present invention has improved performance on the above six datasets. On the Movie dataset, the percentage increase in the final error obtained by the method of the present application compared to the RDA method (the best-performing imbalanced distribution learning method among the above existing imbalanced distribution learning methods) is 10.7%. On the SCUT-FBP dataset, the final error obtained by the method of the present application is 5.8% higher than that of the RDA method. On the Emotion6 dataset, the final error obtained by the method of the present application is 3.1% higher than that of the RDA method. On the FlickrLDL dataset, the final error obtained by the method of the present application is 3.5% higher than that of the RDA method. On the RAF-ML dataset, the final error obtained by the method of the present application is 7.2% higher than that of the RDA method. On the Natural Scene dataset, the final error obtained by the method of the present application is 3.8% higher than that of the RDA method. In particular, the present invention has a better effect of obtaining the predicted label distribution on the Movie dataset, and the obtained predicted label distribution is closer to the true label distribution.
[0126] To verify the contributions of the primary decoupling and secondary decoupling in this application to the method described in this application, this application specifically conducted targeted ablation experiments based on the Movie dataset and used six evaluation metrics, namely Chebyshev distance, Clark distance, Canberra measure, KL divergence, cosine coefficient, and Intersection similarity, to evaluate the distance or similarity between the true label distribution and the predicted label distribution as the performance of the imbalanced label distribution learning network model. The test results are shown in Table 2;
[0127] Table 2 Test Results of Ablation Experiments
[0128]
[0129] In Table 2, the smaller the test results of Chebyshev distance, Clark distance, Canberra measure, and KL divergence, the better the performance; the larger the cosine coefficient and Intersection similarity, the better the performance.
[0130] As can be seen from Table 2, when both primary decoupling and secondary decoupling are performed simultaneously, significant performance improvements are achieved in evaluation metrics such as Chebyshev distance, Clark distance, Canberra measure, KL divergence, cosine coefficient, and Intersection similarity; for example: as can be seen from Table 2, when primary decoupling and secondary decoupling are performed, the test result of the KL divergence metric is 0.221, while when primary decoupling is removed, the test result of the KL divergence metric is 0.246. Obviously, compared with the situation of performing both primary decoupling and secondary decoupling, the KL divergence metric has increased by 11.31% when primary decoupling is removed. In this application, KL divergence is used to measure the difference between the predicted label distribution and the true label distribution. The smaller the KL divergence value, the closer the predicted label distribution is to the true label distribution. The primary decoupling in this application decomposes the label distribution input into the imbalanced label distribution learning network into the dominant label distribution (μ DDL , σ DDL ) and the non-dominant label distribution (μ NDDL , σ NDDL ), which helps the imbalanced label distribution learning network to more accurately capture the features of different labels during the learning process, reduce the information loss when approximating the true distribution with a single distribution, and thus obtain a smaller KL divergence value;
[0131] In addition, it can also be seen from Table 1 that compared with the case of performing primary decoupling and secondary decoupling simultaneously, the Chebyshev distance index has increased by 6.8% in the case of removing primary decoupling; compared with the case of performing primary decoupling and secondary decoupling simultaneously, the Clark distance index has increased by 3.1% in the case of removing primary decoupling; compared with the case of performing primary decoupling and secondary decoupling simultaneously, the Canberra measure index has increased by 1.5% in the case of removing primary decoupling; compared with the case of performing primary decoupling and secondary decoupling simultaneously, the cosine coefficient index has decreased by 2.2% in the case of removing primary decoupling; compared with the case of performing primary decoupling and secondary decoupling simultaneously, the Intersection similarity index has decreased by 2.7% in the case of removing primary decoupling.
[0132] Similarly, it can be known that compared with the case of performing primary decoupling and secondary decoupling simultaneously, the KL divergence index has increased by 9.05% in the case of removing secondary decoupling, which fully highlights the key role of secondary decoupling in improving the prediction label distribution closer to the true label distribution in this application.
Claims
1. A method for learning imbalanced label distribution based on decoupled operation, characterized by: The following steps are involved: S1. Construct an unbalanced label distribution learning network including encoders I, II and III. Encoders I and II obtain feature encoding representation and global label distribution based on the input feature vector and label distribution respectively; Encoder III is used to encode the data after the initial decoupling of the input label distribution to obtain the dominant and non-dominant label distribution; The decoder is used to decode the feature encoding representation and the global label distribution to obtain the predicted and true label distribution; The predicted label distribution is decoupled twice to obtain the predicted dominant label distribution and non-dominant label distribution; S2, based on the feature encoding representation, dominant marker distribution and non-dominant marker distribution, calculate the dominant and non-dominant marker distribution information alignment loss respectively; Based on the feature encoding representation, global label distribution and cosine similarity, the feature similarity matrix and label similarity matrix are constructed respectively. Based on these two matrices, the alignment loss of the feature and label representation similarity matrix is calculated; the alignment loss is calculated based on the predicted label distribution and the true label distribution; the dominant label distribution loss and the non-dominant label distribution loss are calculated based on the predicted dominant label distribution, the non-dominant label distribution and the true label distribution; the parameters α, β, γ and λ are used to balance the above losses and obtain the total optimization loss; S3, training the imbalanced labeled distribution learning network based on the training set and the total optimization loss to obtain a network model; Wherein, step S3 includes S3-1 obtaining a training set and a test set, and step S3-1 specifically includes the following specific steps: S3-1-1. Obtain an original data set, wherein the original data set is a Movie data set; the Movie data set is a data set about users' ratings of movies, and the distribution of rating marks for each movie is calculated by calculating the percentage of each rating level; S3-1-2, using the ten-fold cross validation method to divide the original data set into 10 subsets, 9 of which include 775 feature vectors and their corresponding label distributions, and the remaining subset includes 780 feature vectors and their corresponding label distributions; S4. Obtain feature vectors and label distribution based on the unbalanced data set to be predicted, and then input them into the network model for forward propagation once to obtain predicted label distribution and true label distribution.
2. The unbalanced label distribution learning method based on decoupled operation according to claim 1, characterized in that: In step S1, the steps of the primary decoupling and the secondary decoupling are the same. The primary decoupling decouples the input label distribution, and the secondary decoupling decouples the predicted label distribution. The secondary decoupling specifically includes the following steps: The dominant marker description in the predicted marker distribution is recorded as 1, and the other marker descriptions are recorded as 0 to obtain the dominant marker distribution; the dominant marker description in the predicted marker distribution is recorded as 0, and the non-dominant markers are normalized using softmax to make the sum of the non-dominant marker descriptions 1 to obtain the non-dominant marker distribution.
3. The unbalanced label distribution learning method based on decoupled operation according to claim 1, characterized in that: Step S2 specifically includes the following steps: S2-1. Use KL divergence to calculate the dominant marker distribution information alignment loss L for feature encoding representation and dominant marker distribution alig 1; S2-2, the feature encoding representation and the non-dominant label distribution are calculated using KL divergence to align the non-dominant label distribution information loss L alig 2; S2-3. Variance σ of feature encoding representation feature Add noise δ feature , and then with the mean μ represented by the feature code feature By reparameterization, the calculation shown in (1) is performed to generate the feature representation r feature ; r feature = m feature + s feature d feature ,d feature ~N(0,1) (1) In formula (1), N(0,1) represents the standard normal distribution; Then, cosine similarity is used to calculate the feature representation corresponding to the mth instance The feature representation corresponding to the nth instance Then, based on the above similarity, a two-dimensional feature similarity matrix A is constructed. mn ; The variance σ of the global label distribution whole Add noise δ whole , and then compared with the mean μ of the global label distribution whole By reparameterization, the calculation shown in (2) is performed to generate the global label distribution representation r labelDis ; r labelDis = m whole + s whole d whole ,d whole ~N(0,1) (2) In formula (2), N(0,1) represents the standard normal distribution; Then, cosine similarity is used to calculate the feature representation corresponding to the mth instance The feature representation corresponding to the nth instance Then, based on the above similarity, a two-dimensional label similarity matrix Z is constructed. mn ; The matrix A mn and Z mn Use KL divergence to calculate the feature and label representation similarity matrix alignment loss L alig 3; S2-4. Align the predicted label distribution with the true label distribution using KL divergence, and calculate the alignment loss L between the predicted label distribution and the true label distribution alig 4; S2-5. Align the predicted dominant class label distribution with the true label distribution using KL divergence and calculate the dominant label distribution loss L DDL ; Align the predicted non-dominant class label distribution with the true label distribution using KL divergence and calculate the non-dominant label distribution loss L NDDL .
4. The unbalanced label distribution learning method based on decoupled operation according to claim 1, characterized in that: In step S3, the imbalanced labeled distribution learning network is trained for 10 training segments based on the training set and the total optimization loss to obtain 10 network models.
5. The unbalanced label distribution learning method based on decoupled operation according to claim 4, characterized in that: Step S3 also includes the following steps: S3-2. Input the feature vector in the training set into encoder I, input the label distribution into encoder II and encoder III, perform forward propagation, calculate the total optimization loss of the unbalanced label distribution learning network, perform back propagation under the guidance of the total optimization loss, update the weight parameters of the unbalanced label distribution learning network, complete the training process of one epoch, and complete the training of one training segment after iterating the training process for 300 epochs to obtain an unbalanced label distribution learning network model; in the process of training the unbalanced label distribution learning network, train 10 training segments.
6. The unbalanced label distribution learning method based on decoupled operation according to claim 5, characterized in that: Based on the unbalanced data set to be predicted, the feature vector and label distribution are obtained, and then they are input into 10 network models for forward propagation once to obtain 10 predicted label distributions and 10 corresponding true label distributions.
Citation Information
Patent Citations
Disease prediction method based on unbalanced fundus image data
CN117576012A
Noise partial mark learning model for open set scene and image processing device
CN118864950A