Unbalance mark distribution learning method based on decoupling operation

By adopting decoupling operations and multi-loss function optimization methods in imbalance label distribution learning, the excessive focus on dominant label distribution in the prior art is solved, and the learning effect and alignment of label distribution are improved.

CN119962705AActive Publication Date: 2025-05-09QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)

Patent Information

Application Number
CN202510450586.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-09
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

In imbalanced label distribution learning, the prior art is prone to over-focusing on the dominant label distribution and ignoring the non-dominant label distribution, resulting in the neural network being unable to fully learn and utilize label information.

Method used

A method of unbalanced label distribution learning based on decoupling operation is proposed. By constructing a network including an encoder and a decoder, feature encoding, label distribution decoupling and reconstruction, multiple loss functions are calculated and parameters are balanced to optimize the network.

Benefits of technology

The mutual interference between the dominant mark distribution and the non-dominant mark distribution is effectively avoided, and the alignment between the predicted mark distribution and the real mark distribution is improved, so that the learning effect of the non-dominant mark distribution is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962705A_ABST
    Figure CN119962705A_ABST
Patent Text Reader

Abstract

The invention discloses an unbalance mark distribution learning method based on decoupling operation, and relates to the technical field of machine learning. The method comprises the following steps of: performing primary decoupling on input mark distribution, and encoding by using an encoder III to obtain dominant mark distribution and non-dominant mark distribution; decoding the feature coding representation and the global mark distribution by using a decoder to obtain predicted mark distribution and real mark distribution; and carrying out secondary decoupling on the predicted mark distribution to obtain predicted dominant class mark distribution and non-dominant class mark distribution. Through the arrangement of primary decoupling and secondary decoupling, the dominant mark distribution and the non-dominant mark distribution can be effectively prevented from interfering with each other in the aligning-1 process, the aligning-2 process and the alignment process of the predicted mark distribution output by the decoder and the real mark distribution; therefore, the predicted mark distribution output by the decoder is closer to the real mark distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular to an unbalanced label distribution learning method based on decoupling operation. Background Art

[0002] In the wide application of machine learning, label distribution learning (LDL) is a technical means to construct an effective mapping relationship between instances (i.e., feature vectors) and label distributions. At present, it has shown great potential in many fields such as movie score analysis, expression analysis, video parsing, and age estimation. However, the label distribution in actual application scenarios often shows significant imbalance. Taking the field of movie score analysis as an example, in an ideal situation, it is expected to obtain a balanced distribution of emotional labels. However, due to many factors such as the limited number of labelers, background differences, subjective opinion differences, and label noise, the actual label distribution often shows an unbalanced distribution state. In this state of unbalanced label distribution, the dominant label distribution occupies most of the description space, making the non-dominant label distribution extremely low in description, which seriously affects the comprehensive learning and utilization of label information by the neural network.

[0003] Therefore, the applicant has designed an unbalanced label distribution learning method that can effectively avoid over-focusing on the dominant label distribution and ignoring the non-dominant label distribution in unbalanced label distribution learning. Summary of the invention

[0004] In order to solve the above-mentioned deficiencies in the prior art, the present application proposes an unbalanced label distribution learning method based on decoupling operation.

[0005] A method for learning imbalanced labeled distribution based on decoupled operation, comprising the following steps: S1. Construct an unbalanced label distribution learning network including encoder I, encoder II and encoder III. Encoder I and encoder II respectively obtain feature coding representation and global label distribution based on input feature vector and input label distribution; encoder III is used to encode the data after the initial decoupling of input label distribution to obtain dominant label distribution and non-dominant label distribution; the decoder is used to obtain predicted label distribution from feature coding representation, and is also used to decode the global label distribution to obtain the real label distribution; the predicted label distribution is decoupled twice to obtain the predicted dominant label distribution and non-dominant label distribution; the input feature vector is the feature vector of the input unbalanced label distribution learning network; the input label distribution is the label distribution of the input unbalanced label distribution learning network; S2. Based on the feature encoding representation, dominant marker distribution and non-dominant marker distribution, the dominant marker distribution information alignment loss and non-dominant marker distribution information alignment loss are calculated respectively; based on the feature encoding representation, global marker distribution and cosine similarity, the feature similarity matrix A is constructed respectively mn And the label similarity matrix Z mn , based on the matrix A mn and Z mn Calculate the alignment loss of the feature and marker representation similarity matrix; calculate the alignment loss based on the predicted marker distribution and the true marker distribution; calculate the dominant marker distribution loss and the non-dominant marker distribution loss based on the predicted dominant marker distribution, the non-dominant marker distribution and the true marker distribution respectively; use the parameters α, β, γ and λ to balance the above losses and obtain the total optimization loss; S3, training the imbalanced labeled distribution learning network based on the training set and the total optimization loss to obtain a network model; S4. Obtain feature vectors and label distribution based on the unbalanced data set to be predicted, and then input them into the network model for forward propagation once to obtain predicted label distribution and true label distribution.

[0006] Preferably, in step S1, the steps of the primary decoupling and the secondary decoupling are the same, the primary decoupling decouples the input label distribution, and the secondary decoupling decouples the predicted label distribution, and the secondary decoupling specifically includes the following steps: The descriptive degree of the dominant marker in the predicted marker distribution (the dominant marker is the marker with the highest descriptive degree) is recorded as 1, and the descriptive degree of other markers is recorded as 0, so as to obtain the dominant marker distribution; the descriptive degree of the dominant marker in the predicted marker distribution is recorded as 0, and the non-dominant markers (non-dominant markers are markers other than the dominant markers) are normalized using softmax so that the sum of the descriptive degrees of the non-dominant markers is 1, so as to obtain the non-dominant marker distribution.

[0007] Preferably, step S2 specifically includes the following steps: S2-1. Use KL divergence to calculate the dominant marker distribution information alignment loss L for feature encoding representation and dominant marker distribution alig 1; S2-2, the feature encoding representation and the non-dominant label distribution are calculated using KL divergence to align the non-dominant label distribution information loss L alig 2; S2-3. Variance σ of feature encoding representation feature Add noise δ feature , and then with the mean μ represented by the feature code feature By reparameterization, the calculation shown in (1) is performed to generate the feature representation r feature ; r feature= μ feature + σ feature δ feature , δ feature ~N(0,1) (1) In formula (1), N(0,1) represents the standard normal distribution; Then, cosine similarity is used to calculate the feature representation corresponding to the mth instance The feature representation corresponding to the nth instance Then, based on the above similarity, a two-dimensional feature similarity matrix A is constructed. mn ; The variance σ of the global label distribution whole Add noise δ whole , and then compared with the mean μ of the global label distribution whole By reparameterization, the calculation shown in (2) is performed to generate the global label distribution representation r labelDis ; r labelDis = μ whole + σ whole δ whole , δ whole ~N(0,1) (2) In formula (2), N(0,1) represents the standard normal distribution; Then, cosine similarity is used to calculate the feature representation corresponding to the mth instance The feature representation corresponding to the nth instance Then, based on the above similarity, a two-dimensional label similarity matrix Z is constructed. mn .

[0008] The matrix A mn and Z mn Use KL divergence to calculate the feature and label representation similarity matrix alignment loss L alig 3; S2-4. Align the predicted label distribution with the true label distribution using KL divergence, and calculate the alignment loss L between the predicted label distribution and the true label distribution alig 4; S2-5. Align the predicted dominant class label distribution with the true label distribution using KL divergence and calculate the dominant label distribution loss L DDL ; Align the predicted non-dominant class label distribution with the true label distribution using KL divergence and calculate the non-dominant label distribution loss L NDDLThe above setting can effectively avoid mutual interference in the process of aligning the predicted dominant class label distribution with the true label distribution and in the process of aligning the non-dominant class label distribution with the true label distribution.

[0009] Preferably, in step S3, the imbalanced labeled distribution learning network is trained for 10 training segments based on the training set and the total optimization loss to obtain 10 network models.

[0010] Preferably, step S3 comprises the following steps: S3-1, obtain training set and test set; S3-2. Input the feature vector in the training set into encoder I, input the label distribution in the training set into encoder II and encoder III, perform forward propagation, calculate the total optimization loss of the unbalanced label distribution learning network, and perform backpropagation under the guidance of the total optimization loss to update the weight parameters of the unbalanced label distribution learning network, complete the training process of one epoch, and complete the training of one training segment after iterating the training process for 300 epochs to obtain an unbalanced label distribution learning network model; in the process of training the unbalanced label distribution learning network in this application, 10 training segments are trained.

[0011] Preferably, in step S3-1, the original data set is divided into 10 subsets using a ten-fold cross validation method; in each training segment, 9 subsets are selected as training sets and the remaining subsets are selected as test sets. In the 10 training segments of the present application, the selected test sets are different.

[0012] Preferably, in step S4, feature vectors and label distributions are obtained based on the unbalanced data set to be predicted, and then respectively input into 10 network models for forward propagation once to obtain 10 predicted label distributions and 10 corresponding true label distributions.

[0013] Compared with the prior art, the beneficial technical effects of the present invention are: In the method of the present invention, the settings of the primary decoupling and the secondary decoupling of the present invention can effectively avoid mutual interference between the dominant label distribution and the non-dominant label distribution during the alignment-1 process, the alignment-2 process, and the alignment process between the predicted label distribution output by the decoder and the true label distribution, thereby making the predicted label distribution output by the decoder of the present application closer to the true label distribution; Moreover, in this application, when calculating the total optimization loss, the advantage marker distribution loss L DDL and loss of distribution of non-dominant markers L NDDL The balance parameter is set to α, and the advantage mark distribution information alignment loss L alig1 (the loss generated during the alignment -1 process) and the non-dominant marker distribution information alignment loss L alig The balance parameter of 2 (the loss generated in the alignment-2 process) is also set to α, α = 0.1 ~ 0.9. The above setting enables the gradient of the non-dominant marker distribution to be effectively transmitted in the learning process of the unbalanced marker distribution learning network, thereby ensuring the learning effect of the non-dominant marker distribution; Moreover, the present application sets the balance parameters α, β, γ and λ to make the dominant marker distribution loss L DDL , non-dominant marker distribution loss L NDDL 、 Advantageous label distribution information alignment loss L alig 1. Non-dominant label distribution information alignment loss L alig 2. Feature and tag representation similarity matrix alignment loss L alig 3 and the alignment loss L between the predicted label distribution and the true label distribution alig 4 together constitute the total optimization loss ; In this application, the total optimization loss The construction can make the descriptive degree of non-dominant marker distribution independent of the descriptive degree of dominant marker distribution in the learning process of imbalanced marker distribution learning network, so that the imbalanced marker distribution learning network can better learn the descriptive degree of dominant marker distribution and the descriptive degree of non-dominant marker distribution, so that the predicted marker distribution is closer to the real marker distribution. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 Network structure diagram of the imbalanced label distribution learning network in this application. DETAILED DESCRIPTION

[0015] Embodiment 1: A method for learning imbalanced label distribution based on decoupling operation, wherein the imbalanced data set prediction method based on decoupling operation is to predict the imbalanced data set, i.e., a Movie data set, and comprises the following steps: S1. Construct an unbalanced label distribution learning network including encoder I, encoder II and encoder III, such as Figure 1As shown; in the present application, encoder I and encoder II respectively obtain feature coding representation and global label distribution based on input feature vector and input label distribution; encoder III is used to encode the data after the initial decoupling of input label distribution to obtain dominant label distribution and non-dominant label distribution; the decoder is used to decode the feature coding representation and the global label distribution to obtain the predicted label distribution and the true label distribution; the predicted label distribution is decoupled twice to obtain the predicted dominant label distribution and non-dominant label distribution; the input feature vector refers to the feature vector of the input imbalance label distribution learning network; the input label distribution refers to the label distribution of the input imbalance label distribution learning network; specifically: Encoder I encodes the input feature vector and maps the encoded feature vector to the latent space to obtain a feature encoding representation consisting of the mean and variance (μ feature ,σ feature ), where μ feature represents the mean of the feature encoding, σ feature represents the variance of the feature encoding representation; The input marker distribution is initially decoupled to obtain the distribution with dominant markers and the distribution with non-dominant markers. The distribution with dominant markers and the distribution with non-dominant markers are encoded by encoder III to obtain the dominant marker distribution (μ DDL ,σ DDL ) and the distribution of non-dominant markers (μ NDDL ,σ NDDL ), where μ DDL represents the mean of the dominant marker distribution, σ DDL represents the variance of the dominant marker distribution; μ NDDL represents the mean of the distribution of non-dominant markers, σ NDDL represents the variance of the distribution of nondominant markers; Encoder II encodes the input tag distribution as a whole and maps the encoded tag distribution into the latent space to learn the feature representation of the global tag distribution and obtain the global tag distribution (μ whole ,σ whole ), μ whole represents the mean of the global label distribution, σ whole represents the mean of the global label distribution; The decoder decodes the feature encoding representation output by encoder I to obtain the predicted tag distribution, and performs a global tag distribution (μ whole ,σ whole ) performs decoding operation to obtain the true label distribution; Then, the predicted marker distribution is decoupled twice to obtain the predicted dominant marker distribution and the predicted non-dominant marker distribution.

[0016] In the present application, in step S1, the steps of the primary decoupling and the secondary decoupling are the same. The primary decoupling decouples the input label distribution, and the secondary decoupling decouples the predicted label distribution. Taking the secondary decoupling as an example, the decoupling steps of the present application are introduced. The secondary decoupling specifically includes the following steps: the description degree of the dominant label (the dominant label is the label with the highest descriptive degree) in the predicted label distribution is recorded as 1, and the description degrees of other labels are recorded as 0 to obtain the dominant label distribution; the description degree of the dominant label in the predicted label distribution is recorded as 0, and the non-dominant labels (non-dominant labels refer to labels other than the dominant labels) are normalized using softmax so that the sum of the description degrees of the non-dominant labels is 1 to obtain the non-dominant label distribution.

[0017] In the present application, the output end of encoder I is connected to the input end of the decoder, the output end of encoder III is connected to the input end of the decoder, and the output end of encoder II is connected to the input end of the decoder; in the present application, encoder I is essentially a feature encoder, encoder II is essentially a full-label encoder, and encoder III is essentially a dominant / non-dominant label distribution encoder; wherein, encoder I, encoder II, and encoder III have the same structure, which is consistent with the encoder disclosed in the paper "Imbalanced Label Distribution Learning"; in the present application, label refers to label, dominant label is dominant label, and non-dominant label is non-dominant label; S2. Calculate the total optimization loss of the imbalanced labeled distribution learning network; the details are as follows: S2-1, the feature encoding representation (μ feature ,σ feature ) and dominant marker distribution (μ DDL ,σ DDL ) Use KL divergence to calculate the dominant label distribution information alignment loss L alig 1; In this application, in the training process of the imbalanced label distribution learning network, the existence of the dominant label distribution information alignment loss Lalig1 can constrain the feature encoding representation (μ feature ,σ feature ) and the dominant marker distribution (μ DDL ,σ DDL ) distribution deviation, so that the eigenvector (μ feature ,σ feature ) and the dominant marker distribution (μ DDL ,σ DDL ) are closer in the latent space to achieve the feature vector (μ feature ,σ feature ) and the dominant marker distribution (μ DDL,σ DDL ) in the latent space, called alignment-1; S2-2, the feature encoding representation (μ feature ,σ feature ) and non-dominant marker distribution (μ NDDL ,σ NDDL ) Use KL divergence to calculate the non-dominant label distribution information alignment loss L alig 2; In this application, during the training of the imbalanced label distribution learning network, the non-dominant label distribution information alignment loss L alig The existence of 2 can constrain the eigenvector (μ feature ,σ feature ) and the distribution of non-dominant markers (μ NDDL ,σ NDDL ) distribution deviation, so that the eigenvector (μ feature ,σ feature ) and the distribution of non-dominant markers (μ NDDL ,σ NDDL ) are closer in the latent space to achieve the feature vector (μ feature ,σ feature ) and the distribution of non-dominant markers (μ NDDL ,σ NDDL ) in the latent space, referred to as alignment-2; S2-3. Variance σ of feature encoding representation feature Add noise δ feature , and then with the mean μ represented by the feature code feature By reparameterization, the calculation shown in (1) is performed to generate the feature representation r feature ; r feature = μ feature + σ feature δ feature , δ feature ~N(0,1) (1) In formula (1), N(0,1) represents the standard normal distribution; Then, cosine similarity is used to calculate the feature representation corresponding to the mth instance The feature representation corresponding to the nth instance Then, a two-dimensional feature similarity matrix is ​​constructed based on the above similarity. , and record this matrix as matrix A mn ; The variance σ of the global label distribution whole Add noise δ whole , and then compared with the mean μ of the global label distribution wholeBy reparameterization, the calculation shown in (2) is performed to generate the global label distribution representation r labelDis ; r labelDis = μ whole + σ whole δ whole , δ whole ~N(0,1) (2) In formula (2), N(0,1) represents the standard normal distribution; Then, cosine similarity is used to calculate the feature representation corresponding to the mth instance The feature representation corresponding to the nth instance Then, a two-dimensional label similarity matrix is ​​constructed based on the above similarity. , and denote this matrix as matrix Z mn .

[0018] The matrix A mn and Z mn Use KL divergence to calculate the feature and label representation similarity matrix alignment loss L alig 3; S2-4. Align the predicted label distribution with the true label distribution using KL divergence, and calculate the alignment loss L between the predicted label distribution and the true label distribution alig 4; S2-5. Align the predicted dominant class label distribution with the true label distribution using KL divergence and calculate the dominant label distribution loss L DDL ; Align the predicted non-dominant class label distribution with the true label distribution using KL divergence and calculate the non-dominant label distribution loss L NDDL ; The above setting can effectively avoid mutual interference in the process of aligning the predicted dominant class label distribution with the true label distribution and in the process of aligning the non-dominant class label distribution with the true label distribution; By setting the balance parameters α, β, γ and λ, the dominant marker distribution loss L DDL , non-dominant marker distribution loss L NDDL 、 Advantageous label distribution information alignment loss L alig 1. Non-dominant label distribution information alignment loss L alig 2. Feature and tag representation similarity matrix alignment loss L alig 3 and the alignment loss L between the predicted label distribution and the true label distribution alig 4 together constitute the total optimization loss ; In the present application, the construction of the above-mentioned loss functions can make the descriptive degree of non-dominant marker distribution independent of the descriptive degree of dominant marker distribution in the learning process of the imbalanced marker distribution learning network, thereby enabling the imbalanced marker distribution learning network to better learn the descriptive degree of dominant marker distribution and the descriptive degree of non-dominant marker distribution, thereby making the predicted marker distribution closer to the true marker distribution.

[0019] In this application, the total optimization loss The calculation method is as shown in formula (3): + (3) In formula (3), α, β, γ and λ are all balance parameters, which are used to adjust the weights of the losses of each part during the training process of the imbalanced label distribution learning network to ensure that the imbalanced label distribution learning network model can achieve the best performance balance in different data sets and tasks; in this embodiment 1, the values ​​of α, β, γ and λ are 0.4, 0.1, 0.1 and 0.1 respectively.

[0020] In formula (3), the dominant marker distribution loss Used to calculate the loss generated during the alignment process of the dominant class label distribution and the true label distribution using KL divergence, dominant label distribution loss The setting of can make the imbalanced label distribution learning network effectively avoid over-focusing on the dominant label distribution and ignoring the non-dominant label distribution under the imbalanced label distribution; the dominant label distribution loss The calculation method of is consistent with the formula (1) published in "Imbalanced Label Distribution Learning".

[0021] Non-dominant marker distribution loss Used to calculate the loss generated during the alignment process of the non-dominant class label distribution and the true label distribution using KL divergence; non-dominant label distribution loss The calculation method of is consistent with the formula (1) published in "Imbalanced Label Distribution Learning".

[0022] Advantageous Marker Distribution Information Alignment Loss The feature vector (μ) is calculated by KL divergence feature ,σ feature ) and the dominant marker distribution (μ DDL ,σ DDL ) is the loss generated in the alignment process in the latent space, so that the imbalanced label distribution learning network can better capture the relationship between the dominant label distribution and the feature vector during the learning process; among them, the dominant label distribution information alignment loss The calculation method is as shown in formula (4): (4) In formula (4), represents the variance ratio calculated from the distribution of the ith dominant marker, represents the sum of squares of mean differences calculated for the distribution of the ith dominant marker; In formula (4), The calculation method of is shown in formula (5): = (5) In formula (5), The i-th feature encoding output by encoder I represents the calculated variance, represents the variance calculated from the distribution of the ith dominant marker output by the dominant / non-dominant marker encoder.

[0023] In formula (4), The calculation method of is shown in formula (6): = (6) In formula (6), The i-th feature encoding output by encoder I represents the calculated mean, represents the mean value calculated from the distribution of the ith dominant marker output by the dominant / non-dominant marker encoder, represents the variance calculated from the distribution of the kth dominant marker output by the dominant / non-dominant marker encoder; in this application, the dominant marker is essentially the dominant label, and the non-dominant marker is essentially the non-dominant label; Non-dominant marker distribution information alignment loss The feature vector (μ) is calculated by KL divergence feature ,σ feature ) and the distribution of non-dominant markers (μ NDDL ,σ NDDL ) is the loss generated during the alignment process in the latent space; the above loss setting helps the imbalanced label distribution learning network to learn more accurate association information between the non-dominant label distribution and the feature vector; non-dominant label distribution information alignment loss The calculation method of is shown in formula (7); (7) In formula (7), represents the variance ratio calculated for the distribution of the i-th non-dominant marker, represents the sum of squares of mean differences calculated for the distribution of the i-th non-dominant marker; In formula (7), The calculation method of is shown in formula (8):

[0024] In formula (8), The i-th feature encoding output by encoder I represents the calculated variance, represents the variance calculated from the distribution of the ith non-dominant marker output by the dominant / non-dominant marker encoder; In formula (7), The calculation method of is shown in formula (9): (9) In formula (9), The i-th feature encoding output by encoder I represents the calculated mean, represents the mean value calculated from the distribution of the ith non-dominant marker output by the dominant / non-dominant marker encoder, represents the variance calculated from the distribution of the kth non-dominant marker output by the dominant / non-dominant marker encoder; Feature and token representation similarity matrix alignment loss By calculating the two-dimensional label similarity matrix Z mn To the two-dimensional feature similarity matrix A mn The loss generated during the alignment process; in this application, the two-dimensional feature similarity matrix A is obtained mn and the two-dimensional label similarity matrix Z mn The purpose is to calculate the alignment loss Lalig3 of the similarity matrix between the feature and the mark; the calculation method of the similarity matrix alignment loss Lalig3 is shown in formula (10): (10) In formula (10), M represents the number of feature vectors of the input imbalanced label distribution learning network.

[0025] Alignment loss between predicted label distribution and true label distribution The difference between the predicted label distribution and the true label distribution output by the decoder is calculated using KL divergence. The above loss setting can effectively align the predicted label distribution with the true label distribution. The loss of aligning the predicted label distribution with the true label distribution is The calculation method of is shown in formula (11).

[0026] (11) In formula (11), r labelDis represents the variance σ of the global label distribution whole Add noise δ whole , and then compared with the mean μ of the global label distributionwhole The global label distribution is calculated by reparameterization; D represents the true label distribution; ( ) indicates a decoder; represents the KL divergence.

[0027] S3, obtaining a training set and a test set, and training the imbalanced labeled distribution learning network for 10 training segments based on the training set and the total optimization loss to obtain 10 imbalanced labeled distribution learning network models, including the following steps: S3-1. Obtaining training sets and test sets includes the following specific steps: S3-1-1. Get the original data set, where the original data set is the Movie data set [Geng, 2016]; the Movie data set [Geng, 2016] comes from the paper "Label Distribution Learning", which contains basic movie-related annotation data. The Movie data set [Geng, 2016] is a data set about users' ratings of movies. The data set contains 7,755 movies and 54,242,292 rating records from 478,656 different users (the data comes from the Netflix platform, with a rating range of 1 to 5 stars, a total of 5 levels). The rating label distribution of each movie is calculated by the percentage of each rating level. Feature extraction is based on metadata, covering dimensions such as genre, director, actor, country, budget, etc., where the classification attributes are converted into binary vectors. The feature vector finally extracted from each movie is 1,869-dimensional. The original data set contains 7755 feature vectors and 5 label distributions.

[0028] S3-1-2. Use the ten-fold cross-validation method to divide the original data set into 10 subsets, 9 of which include 775 feature vectors and their corresponding label distributions, and the remaining subset includes 780 feature vectors and their corresponding label distributions; in the process of training the unbalanced label distribution learning network in this application, 10 training segments are required, each training segment includes a training process of 300 epochs, and in each training segment, 9 subsets are selected as training sets, and the remaining subset is used as a test set. In the 10 training segments of this application, the selected test sets are different.

[0029] S3-2, input the feature vector in the training set into encoder I, and input the label distribution in the training set into encoder II and encoder III, forward propagate, calculate the total optimization loss of the unbalanced label distribution learning network, and perform back propagation under the guidance of the total optimization loss, update the weight parameters of the unbalanced label distribution learning network, complete the training process of one epoch, and complete the training of one training segment after iterating the training process of 300 epochs, and obtain an unbalanced label distribution learning network model; in the process of training the unbalanced label distribution learning network in this application, 10 training segments are required to be trained, so in the process of training the unbalanced label distribution learning network in this application, 10 unbalanced label distribution learning network models can be obtained. In the process of training the unbalanced label distribution learning network in this application, the learning rate is set to 0.001 to ensure that the model can converge stably during the training process; the batch size is set to 50, while ensuring the training efficiency, taking into account the full learning of the data by the unbalanced label distribution learning network model.

[0030] S4. Based on the unbalanced dataset with an unbalanced distribution of the movie rating labels to be predicted, the feature vector and label distribution are obtained respectively, and then the feature vector and label distribution are respectively input into the 10 unbalanced label distribution learning network models obtained in step S3 for forward propagation once, and 10 predicted label distributions and 10 corresponding true label distributions are output. Among them, the unbalanced dataset refers to a dataset with an unbalanced distribution of movie rating labels; the method of obtaining the feature vector and label distribution based on the unbalanced dataset with an unbalanced distribution of the movie rating labels to be predicted in this application is the same as the method of obtaining the feature vector and label distribution described in the A.1.1 Description of Datasets •Movie (movie rating) section of the paper "Imbalanced Label Distribution Learning". In this application, since the predicted label distribution output by the decoder includes the rating of the movie, the existing movie recommendation system (such as the movie recommendation system disclosed by MovieLens and Netflix Prize) can predict the rating of the movie based on the above 10 predicted label distributions, thereby recommending movies with higher predicted ratings to users.

[0031] test: The method described in this application (referred to as Ours method in Table 1) is compared with the existing SA-BFGS method (from Xin Geng's paper "Label Distribution Learning"), the EDLRL method (from Xiuyi Jia et al.'s paper "Facial Emotion Distribution Learning by Exploiting Low-Rank Label Correlations Locally"), the LDLSF method (from Tingting Ren et al.'s paper "Label Distribution Learning with Label-Specific Features"), the LDL-LCLR method (from Tingting Ren et al.'s paper "Label Distribution Learning with Label Correlations via Low-Rank Approximation"), the Adam-LDL-SCL method (from Xiuyi Jia et al.'s paper "Label Distribution Learning with Label Correlations on Local Samples"), the LDL-LDM method (from Jing Wang et al.'s paper "Label Distribution Learning by Exploiting Label Distribution Manifold"), the OFR-FL method (from Tsung-Yi Lin et al.'s paper "Focal Loss for Dense Object The imbalanced distribution learning methods such as OFR-CB method (from Tsung-Yi Lin et al.'s paper "Focal Loss for Dense Object Detection"), OFR-DB method (from Tsung-Yi Lin et al.'s paper "Focal Loss for Dense Object Detection"), RDA method (from Xingyu Zhao et al.'s paper "Imbalanced Label Distribution Learning") and DILDL method (from Xingyu Zhao et al.'s paper "Imbalanced Label Distribution Learning") were tested using the same test strategy. The test results are shown in Table 1.The test strategy adopted in this application is: based on the test sets in 10 training segments and the corresponding unbalanced label distribution learning network models, the Chebyshev distance value between the predicted label distribution output by each unbalanced label distribution learning network model and the true label distribution is calculated based on the Chebyshev distance method, and then the Chebyshev distances are averaged to obtain the Chebyshev distance average, and then the absolute value of the difference between each Chebyshev distance value and the Chebyshev distance average is calculated, the maximum absolute value is the fluctuation error, and the fluctuation error is added to the Chebyshev distance average to obtain the final error, which can show the difference between the predicted label distribution and the true label distribution. ;

[0032] Embodiment 2: Compared with the first embodiment, the second embodiment differs from the first embodiment in that: 1) The unbalanced dataset used in Example 2 is the SCUT-FBP dataset; the SCUT-FBP dataset in this application is consistent with the SCUT-FBP dataset disclosed in the paper "SCUT-FBP: A Benchmark Dataset for Facial Beauty Perception", which focuses on facial beauty perception research and includes beauty score annotations for facial images, with a score range of 1 to 5 points; 2) In step S3-1-2 of the second embodiment, the subset is obtained by using a ten-fold cross validation method to divide the original data set into 10 subsets, each of which includes 3600 feature vectors and their corresponding label distributions.

[0033] In Example 2, the method described in the present application is tested with the same testing strategy as the existing unbalanced labeled distribution learning methods such as the SA-BFGS method, EDLRL method, LDLSF method, LDL-LCLR method, Adam-LDL-SCL method, LDL-LDM method, OFR-FL method, OFR-CB method, OFR-DB method, RDA method and DILDL method, and the test results are shown in Table 1.

[0034] Embodiment three: Compared with the first embodiment, the third embodiment differs from the first embodiment in that: 1) The unbalanced dataset used in Example 2 is the Emotion6 dataset; the Emotion6 dataset is consistent with the Emotion6 dataset published in the paper Amixed bag of emotions: Model, predict, and transfer emotion distributions, which contains annotated data of seven emotion categories (anger, disgust, joy, fear, sadness, surprise, and neutrality); 2) In step S3-1-2 of the second embodiment, the subset is obtained by dividing the original data set into 10 subsets using a ten-fold cross validation method, each subset including 1782 feature vectors and their corresponding label distributions.

[0035] In Example 3, the method described in the present application is tested with the same testing strategy as the existing unbalanced distribution learning methods such as the SA-BFGS method, EDLRL method, LDLSF method, LDL-LCLR method, Adam-LDL-SCL method, LDL-LDM method, OFR-FL method, OFR-CB method, OFR-DB method, RDA method and DILDL method, and the test results are shown in Table 1.

[0036] Embodiment 4: Compared with the first embodiment, the fourth embodiment differs from the first embodiment in that: 1) The unbalanced dataset used in Example 2 is the Flickr LDL dataset; the Flickr LDL dataset is consistent with the Flickr LDL dataset disclosed in the paper "Learning Visual Sentiment Distributions via Augmented Conditional Probability Neural Network"; the dataset contains labeled data of eight emotion categories (i.e., awe, disgust, fear, entertainment, sadness, satisfaction, excitement, and anger); 2) In step S3-1-2 of the second embodiment, the subset is obtained by dividing the original data set into 10 subsets using a ten-fold cross validation method, each subset including 10035 feature vectors and their corresponding label distributions.

[0037] In Example 4, the method described in the present application is tested with the same testing strategy as the existing unbalanced labeled distribution learning methods such as the SA-BFGS method, EDLRL method, LDLSF method, LDL-LCLR method, Adam-LDL-SCL method, LDL-LDM method, OFR-FL method, OFR-CB method, OFR-DB method, RDA method and DILDL method, and the test results are shown in Table 1.

[0038] Embodiment five: Compared with the first embodiment, the fifth embodiment differs from the first embodiment in that: 1) The imbalanced dataset used in Example 2 is the RAF-ML dataset; the RAF-ML dataset is consistent with the RAF-ML dataset published in the paper "Blended Emotion in-the-Wild: Multi-label Facial Expression Recognition Using Crowdsourced Annotations and Deep Locality Feature Learning"; the dataset contains annotated data of six expression distributions (i.e., happiness, sadness, surprise, fear, anger, and neutrality); 2) In step S3-1-2 of Example 2, the subset is obtained by using a ten-fold cross validation method to divide the original data set into 10 subsets, 9 of which include 4417 feature vectors and their corresponding label distributions, and the remaining subset includes 4419 feature vectors and their corresponding label distributions.

[0039] In Example 5, the method described in the present application is tested with the same testing strategy as the existing unbalanced distribution learning methods such as the SA-BFGS method, EDLRL method, LDLSF method, LDL-LCLR method, Adam-LDL-SCL method, LDL-LDM method, OFR-FL method, OFR-CB method, OFR-DB method, RDA method and DILDL method, and the test results are shown in Table 1.

[0040] Embodiment six: Compared with the first embodiment, the sixth embodiment differs from the first embodiment in that: 1) The unbalanced dataset used in Example 2 is the Natural Scene dataset. The Natural Scene dataset is consistent with the Natural Scene dataset published in the paper Label Distribution Learning. The dataset contains nine associated tags, including sun, cloud, sky, building, water, mountain, snow, desert, and plant. 2) In step S3-1-2 of the second embodiment, the subset is obtained by dividing the original data set into 10 subsets using a ten-fold cross validation method, each subset including 1800 feature vectors and their corresponding label distributions.

[0041] In Example 6, the method described in the present application is tested with the same testing strategy as the existing unbalanced labeled distribution learning methods such as the SA-BFGS method, EDLRL method, LDLSF method, LDL-LCLR method, Adam-LDL-SCL method, LDL-LDM method, OFR-FL method, OFR-CB method, OFR-DB method, RDA method and DILDL method, and the test results are shown in Table 1.

[0042] Table 1 Test results of different methods on different datasets

[0043] As can be seen from Table 1, the performance of the method described in the present invention on the above six data sets is improved. On the Movie data set, the final error obtained by the method described in the present application is improved by 10.7% compared with the RDA method (the best imbalanced distribution learning method among the above existing imbalanced distribution learning methods), and the final error obtained by the method described in the present application on the SCUT-FBP data set is improved by 5.8% compared with the RDA method. On the Emotion6 data set, the final error obtained by the method described in the present application is improved by 3.1% compared with the RDA method. On the FlickrLDL data set, the final error obtained by the method described in the present application is improved by 3.5% compared with the RDA method, on the RAF-ML data set, the final error obtained by the method described in the present application is improved by 7.2% compared with the RDA method, and on the Natural Scene data set, the final error obtained by the method described in the present application is improved by 3.8% compared with the RDA method. In particular, the present invention has a better effect in obtaining the predicted label distribution on the Movie data set, and the obtained predicted label distribution is closer to the true label distribution.

[0044] In order to verify the contribution of the primary decoupling and secondary decoupling in this application to the method described in this application, this application specifically conducts targeted ablation experiments based on the Movie dataset, and uses six evaluation indicators, namely Chebyshev distance, Clark distance, Canberra measure, KL divergence, cosine coefficient and Intersection similarity, to evaluate the distance or similarity between the true label distribution and the predicted label distribution as the performance of the unbalanced label distribution learning network model. The test results are shown in Table 2. Table 2 Ablation experiment test results

[0045] In Table 2, the smaller the test results of Chebyshev distance, Clark distance, Canberra measure and KL divergence, the better the performance; the larger the cosine coefficient and Intersection similarity, the better the performance.

[0046] It can be seen from Table 2 that when the initial decoupling and the secondary decoupling are performed simultaneously, the evaluation indicators such as Chebyshev distance, Clark distance, Canberra measure, KL divergence, cosine coefficient and Intersection similarity have achieved significant performance improvements; for example: it can be seen from Table 2 that the test result of the KL divergence indicator is 0.221 during the initial decoupling and the secondary decoupling, and when the initial decoupling is removed, the test result of the KL divergence indicator is 0.246. Obviously, the KL divergence indicator is improved by 11.31% when the initial decoupling is removed compared with the case where the initial decoupling and the secondary decoupling are performed simultaneously. In this application, KL divergence is used to measure the difference between the predicted label distribution and the true label distribution. The smaller the KL divergence value, the closer the predicted label distribution is to the true label distribution. The initial decoupling of this application decomposes the label distribution input into the imbalanced label distribution learning network into a dominant label distribution (μ DDL ,σ DDL ) and the distribution of non-dominant markers (μ NDDL ,σ NDDL ), which helps the unbalanced labeled distribution learning network to more accurately capture the characteristics of different labels during the learning process, and reduces the information loss when a single distribution is used to approximate the true distribution, thereby obtaining a smaller KL divergence value; In addition, it can be seen from Table 1 that the Chebyshev distance index is improved by 6.8% when the initial decoupling is removed compared with the case where the initial decoupling and the secondary decoupling are performed simultaneously; the Clark distance index is improved by 3.1% when the initial decoupling is removed compared with the case where the initial decoupling and the secondary decoupling are performed simultaneously; the Canberra measurement index is improved by 1.5% when the initial decoupling is removed compared with the case where the initial decoupling and the secondary decoupling are performed simultaneously; the cosine coefficient index is reduced by 2.2% when the initial decoupling is removed compared with the case where the initial decoupling and the secondary decoupling are performed simultaneously; the Intersection similarity index is reduced by 2.7% when the initial decoupling is removed compared with the case where the initial decoupling and the secondary decoupling are performed simultaneously.

[0047] Similarly, it can be seen that the KL divergence index is improved by 9.05% when the secondary decoupling is removed compared to the case where both the primary decoupling and the secondary decoupling are performed simultaneously. This fully highlights the key role of secondary decoupling in improving the application to obtain a predicted label distribution that is closer to the true label distribution.

Claims

1. A method for learning imbalanced label distribution based on decoupled operation, characterized by: The following steps are involved: S1. Construct an unbalanced label distribution learning network including encoders I, II and III. Encoders I and II obtain feature encoding representation and global label distribution based on the input feature vector and label distribution respectively; Encoder III is used to encode the data after the initial decoupling of the input label distribution to obtain the dominant and non-dominant label distribution; The decoder is used to decode the feature encoding representation and the global label distribution to obtain the predicted and true label distribution; The predicted label distribution is decoupled twice to obtain the predicted dominant label distribution and non-dominant label distribution; S2, based on the feature encoding representation, dominant marker distribution and non-dominant marker distribution, calculate the dominant and non-dominant marker distribution information alignment loss respectively; Based on the feature encoding representation, global label distribution and cosine similarity, the feature similarity matrix and label similarity matrix are constructed respectively. Based on these two matrices, the alignment loss of the feature and label representation similarity matrix is ​​calculated; the alignment loss is calculated based on the predicted label distribution and the true label distribution; the dominant label distribution loss and the non-dominant label distribution loss are calculated based on the predicted dominant label distribution, the non-dominant label distribution and the true label distribution; the parameters α, β, γ and λ are used to balance the above losses and obtain the total optimization loss; S3, training the imbalanced labeled distribution learning network based on the training set and the total optimization loss to obtain a network model; S4. Obtain feature vectors and label distribution based on the unbalanced data set to be predicted, and then input them into the network model for forward propagation once to obtain predicted label distribution and true label distribution.

2. The unbalanced label distribution learning method based on decoupled operation according to claim 1, characterized in that: In step S1, the steps of the primary decoupling and the secondary decoupling are the same. The primary decoupling decouples the input label distribution, and the secondary decoupling decouples the predicted label distribution. The secondary decoupling specifically includes the following steps: The dominant marker description in the predicted marker distribution is recorded as 1, and the other marker descriptions are recorded as 0 to obtain the dominant marker distribution; the dominant marker description in the predicted marker distribution is recorded as 0, and the non-dominant markers are normalized using softmax to make the sum of the non-dominant marker descriptions 1 to obtain the non-dominant marker distribution.

3. The unbalanced label distribution learning method based on decoupled operation according to claim 1, characterized in that: Step S2 specifically includes the following steps: S2-1. Use KL divergence to calculate the dominant marker distribution information alignment loss L for feature encoding representation and dominant marker distribution alig 1; S2-2, the feature encoding representation and the non-dominant label distribution are calculated using KL divergence to align the non-dominant label distribution information loss L alig 2; S2-3. Variance σ of feature encoding representation feature Add noise δ feature , and then with the mean μ represented by the feature code feature By reparameterization, the calculation shown in (1) is performed to generate the feature representation r feature ; r feature = m feature + s feature d feature , d feature ~N(0,1) (1) In formula (1), N(0,1) represents the standard normal distribution; Then, cosine similarity is used to calculate the feature representation corresponding to the mth instance The feature representation corresponding to the nth instance Then, based on the above similarity, a two-dimensional feature similarity matrix A is constructed. mn ; The variance σ of the global label distribution whole Add noise δ whole , and then compared with the mean μ of the global label distribution whole By reparameterization, the calculation shown in (2) is performed to generate the global label distribution representation r labelDis ; r labelDis = m whole + s whole d whole , d whole ~N(0,1) (2) In formula (2), N(0,1) represents the standard normal distribution; Then, cosine similarity is used to calculate the feature representation corresponding to the mth instance The feature representation corresponding to the nth instance Then, based on the above similarity, a two-dimensional label similarity matrix Z is constructed. mn ; The matrix A mn and Z mn Use KL divergence to calculate the feature and label representation similarity matrix alignment loss L alig 3; S2-4. Align the predicted label distribution with the true label distribution using KL divergence, and calculate the alignment loss L between the predicted label distribution and the true label distribution alig 4; S2-5. Align the predicted dominant class label distribution with the true label distribution using KL divergence and calculate the dominant label distribution loss L DDL ; Align the predicted non-dominant class label distribution with the true label distribution using KL divergence and calculate the non-dominant label distribution loss L NDDL .

4. The unbalanced label distribution learning method based on decoupled operation according to claim 1, characterized in that: In step S3, the imbalanced labeled distribution learning network is trained for 10 training segments based on the training set and the total optimization loss to obtain 10 network models.

5. The unbalanced label distribution learning method based on decoupled operation according to claim 4, characterized in that: Step S3 includes the following steps: S3-1, obtain training set and test set; S3-2. Input the feature vector in the training set into encoder I, input the label distribution into encoder II and encoder III, perform forward propagation, calculate the total optimization loss of the unbalanced label distribution learning network, perform back propagation under the guidance of the total optimization loss, update the weight parameters of the unbalanced label distribution learning network, complete the training process of one epoch, and complete the training of one training segment after iterating the training process for 300 epochs to obtain an unbalanced label distribution learning network model; in the process of training the unbalanced label distribution learning network, train 10 training segments.

6. The unbalanced label distribution learning method based on decoupled operation according to claim 5, characterized in that: In step S3-1, the original data set is divided into 10 subsets using the ten-fold cross validation method; in each training segment, 9 subsets are selected as training sets and the remaining subsets are selected as test sets; in the 10 training segments, the selected test sets are different.

7. The unbalanced label distribution learning method based on decoupled operation according to claim 5, characterized in that: Based on the unbalanced data set to be predicted, the feature vector and label distribution are obtained, and then they are input into 10 network models for forward propagation once to obtain 10 predicted label distributions and 10 corresponding true label distributions.

Citation Information

Patent Citations

  • Disease prediction method based on unbalanced fundus image data

    CN117576012A

  • Noise partial mark learning model for open set scene and image processing device

    CN118864950A

  • Autoencoder-based graph construction for semi-supervised learning

    WO2021157863A1

Cited By

  • Unbalance mark distribution learning method based on asymmetric momentum optimization

    CN120387494A

  • Imbalanced Label Distribution Learning Method Based on Asymmetric Momentum Optimization

    CN120387494B

  • Unbalanced mark distribution learning method suitable for movie recommendation

    CN120952112A

  • Imbalanced Label Distribution Learning Method for Movie Recommendation

    CN120952112B

  • Unbalance mark distribution learning method suitable for face attraction evaluation

    CN121789016A