Long-tail multi-label feature space optimization method and device based on comparative learning

By applying the contrast learning method in the long-tail multi-label data scenario, the feature space is optimized to deal with the problem of overlapping labels of multi-label samples, and the traditional method has solved the problem of performance degradation and insufficient generalization ability in the processing of long-tail multi-label data, and a more robust and generalization ability feature representation is achieved.

CN120180079APending Publication Date: 2025-06-20NORTH CHINA ELECTRICAL POWER RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510235487.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When traditional classification methods and supervised learning methods deal with long-tail multi-label data, there are problems such as degradation in model performance and difficulty in generalizing effectively.

Method used

The long-tail multi-label feature space optimization method based on contrast learning is adopted. Feature extraction is performed through the comparative learning framework, the similarity of sample pairs and the overlap coefficient of multi-label sample labels are calculated, and the class averaging and class supplementary operations are performed, and the comparison learning loss function is optimized to obtain the balanced feature space.

Benefits of technology

Effectively distinguishing different labels and capturing the feature space of label co-occurrence, improving the robustness and generalization capabilities of the model, especially in long-tail and multi-label scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180079A_ABST
    Figure CN120180079A_ABST
Patent Text Reader

Abstract

The invention discloses a comparative learning-based long-tail multi-label feature space optimization method and device. The method comprises the following steps of: inputting long-tail multi-label data subjected to data enhancement processing into a first branch and a second branch of a comparative learning framework, and respectively carrying out feature extraction to obtain a first sample pair feature and a second sample pair feature; calculating the similarity between the first sample pair feature and the second sample pair feature; based on the similarity, calculating to obtain a multi-label sample label overlapping coefficient of the first sample pair feature and the second sample pair feature; and based on the multi-label sample label overlapping coefficient, carrying out class averaging and class supplementing operation to obtain an optimized feature space. According to the method, the corresponding multi-label sample label overlapping coefficient is designed for the multi-label sample label overlapping problem and is added into the loss item to optimize the learning process of the comparative learning model, and the comparative learning-based long-tail multi-label feature space is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of feature space optimization, and in particular, to a long-tail multi-label feature space optimization method and device based on contrast learning. Background Art

[0002] With the development of deep learning and big data, more and more application scenarios present highly complex data features. Among them, the long-tail distribution and multi-label problem have become important challenges in the field. The long-tail distribution means that in a dataset, the vast majority of labels (categories) only have a small number of samples, while a small number of labels have a large number of samples. This phenomenon of data imbalance widely exists in many practical tasks. The multi-label problem refers to the fact that for each example, it may be associated with multiple labels, rather than belonging to only a single label. For these complex real-world scenarios, traditional classification methods and supervised learning methods face great challenges, especially when dealing with multi-label data that follows a long-tail class distribution, there are problems such as a decline in model performance and difficulty in effective generalization. Summary of the Invention

[0003] In view of the problems in the prior art, embodiments of the present invention provide a long-tail multi-label feature space optimization method and device based on contrast learning. Starting from the feature level, the present invention attempts to extend the traditional contrast learning strategy to the long-tail multi-label scenario, aiming to learn a feature space that can effectively distinguish different labels and capture label co-occurrence. At the same time, the present invention can also be combined with other optimization processing technologies at different stages, having good flexibility and adaptability.

[0004] Embodiments of the present invention provide a long-tail multi-label feature space optimization method based on contrast learning, including:

[0005] Input the long-tail multi-label data after data augmentation processing into the first branch and the second branch of the contrast learning framework respectively to extract the first sample pair feature and the second sample pair feature;

[0006] Calculate the similarity between the first sample pair feature and the second sample pair feature;

[0007] Based on the similarity, calculate the multi-label sample label overlap coefficient between the first sample pair feature and the second sample pair feature;

[0008] Based on the multi-label sample label overlap coefficient, perform class averaging and class supplementation operations to obtain a balanced multi-label contrast learning loss function;

[0009] Use the balanced multi-label contrast learning loss function for supervised training to obtain an optimized feature space.

[0010] In one embodiment, the class supplementation operation includes performing a class supplementation operation on an initial balanced multi-label contrast learning loss function using a class prototype to obtain a final balanced multi-label contrast learning loss function.

[0011] In one embodiment, the class prototype is obtained through the following steps:

[0012] Collect a first sample pair feature matrix, a first sample pair label matrix, a second sample pair feature matrix, and a second sample pair label matrix;

[0013] Perform class prototype estimation on the first sample pair feature matrix, the first sample pair label matrix, the second sample pair feature matrix, and the second sample pair label matrix to obtain an initial class prototype;

[0014] Update the initial class prototype using the exponential moving average method to obtain a final class prototype.

[0015] In one embodiment, the multi-label sample label overlap coefficient is calculated using the Jaccard coefficient.

[0016] In one embodiment, the data augmentation processing includes any one of random cropping, resizing the target image to a preset resolution, color distortion, and Gaussian blur.

[0017] In one embodiment, inputting the long-tail multi-label data after data augmentation processing into the first branch and the second branch of the contrast learning framework respectively for feature extraction to obtain the first sample pair feature and the second sample pair feature includes: inputting the long-tail multi-label data after the first data augmentation processing into the first branch of the contrast learning framework for feature extraction to obtain the first sample pair feature, and inputting the long-tail multi-label data after the second data augmentation processing into the second branch of the contrast learning framework for feature extraction to obtain the second sample pair feature.

[0018] The present invention also provides a long-tail multi-label feature space optimization device based on contrast learning, including:

[0019] A feature extraction module for inputting the long-tail multi-label data after data augmentation processing into the first branch and the second branch of the contrast learning framework respectively for feature extraction to obtain the first sample pair feature and the second sample pair feature;

[0020] A first calculation module for calculating the similarity between the first sample pair feature and the second sample pair feature;

[0021] A second calculation module for calculating the multi-label sample label overlap coefficient between the first sample pair feature and the second sample pair feature based on the similarity;

[0022] A class averaging and class complementing operation module for performing class averaging and class complementing operations based on the multi-label sample label overlap coefficient to obtain a balanced multi-label contrast learning loss function;

[0023] A supervised training module for performing supervised training using the balanced multi-label contrast learning loss function to obtain an optimized feature space.

[0024] An embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned method is implemented.

[0025] An embodiment of the present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned method is implemented.

[0026] An embodiment of the present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above-mentioned method is implemented.

[0027] An embodiment of the present invention provides a method and device for optimizing a long-tail multi-label feature space based on contrast learning. Based on the supervised contrast learning method, the present invention designs a corresponding multi-label sample label overlap coefficient for the problem of multi-label sample label overlap, and adds it to the loss term to optimize the learning process of the contrast learning model, and optimizes the long-tail multi-label feature space based on contrast learning. At the same time, the present invention balances the expression of features of each category through class averaging operations, minimizes the dominant role of the head category in model learning as much as possible, and prevents over-biasing towards the mainstream category. Finally, the present invention is trained through the optimized loss function to obtain a more robust feature representation, effectively supporting the application of the model in downstream tasks such as classification and retrieval, improving the fitting ability of the contrast learning model to complex real-world data, and extending the long-tail technology to the long-tail multi-label task scenario. Description of the Drawings

[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:

[0029] Figure 1 It is a flowchart of a method for optimizing a long-tail multi-label feature space based on contrast learning in an embodiment of the present invention;

[0030] Figure 2 It is a flowchart of a method for obtaining class prototypes in an embodiment of the present invention;

[0031] Figure 3 It is a flowchart of another long-tail multi-label feature space optimization method based on contrastive learning in an embodiment of the present invention;

[0032] Figure 4 It is a schematic structural diagram of a long-tail multi-label feature space optimization device based on contrastive learning in an embodiment of the present invention;

[0033] Figure 5 It is a schematic structural diagram of a computer device entity in an embodiment of the present invention. Detailed implementation manners

[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but not to limit the present invention.

[0035] To facilitate the understanding of the technical solutions provided by the present invention, the research background of the technical solutions of the present invention will be briefly described below.

[0036] In the scenario of long-tail distribution, a small number of labels have a large amount of data, forming the head classes; while most labels have only a small amount of data, forming the tail classes. This distribution pattern is very common in fields such as natural language processing, computer vision, and recommendation systems. For example, in an image classification task, certain classes (such as "cat", "dog") may have thousands of samples, while other classes (such as rare animals or objects) may have only a few sample images. This problem of the imbalance in the number of samples between classes poses challenges to the training of classification models. If not addressed, the trained model is likely to be biased towards the head classes with a large amount of training data, resulting in poor performance on the tail classes with scarce data.

[0037] In a multi-label learning scenario, a sample may be associated with multiple labels at the same time. This situation is very common in many practical applications. For example, in text classification tasks, a document may involve multiple topics or categories; in image annotation tasks, a picture may contain multiple objects or scenes at the same time. Traditional single-label learning methods are difficult to handle such complex multi-label scenarios. Multi-label learning not only requires the model to correctly predict multiple labels of samples, but also requires balancing the learning of head and tail labels in the context of long-tail distribution. In particular, when some labels appear in the head class and others appear in the tail class, the model is likely to favor frequent labels and ignore uncommon labels. Contrastive learning effectively improves the representation ability of the model by optimizing the similarities and differences between samples. In view of the excellent performance of contrastive learning in tasks such as image classification, object detection, natural language processing, and recommendation systems, the present invention is based on the application of this technology in the field of long-tail multi-label, aiming to maximize the classification effect of the model.

[0038] Traditional classification algorithms usually rely on a balanced distribution of labels. When the data presents a long-tail distribution, the model is more likely to learn the head class and ignore the tail class. This phenomenon causes the model to perform poorly on the tail class and lack generalization ability. To address this problem, existing methods include data resampling and cost-sensitive learning, but these methods usually have their limitations. For example, resampling solves the problem from the perspective of data balance, which may lead to overfitting of the head class data, while cost-sensitive learning requires complex adjustments to the class weights from the perspective of objective function design.

[0039] In addition, traditional long-tail learning research is usually based on the setting of single-label scenarios, which simplifies the properties and structure of the data to some extent. In the real world, data is often multi-labeled, which reflects the situation where different labels may co-appear in the same example. This means that the label set of an example may overlap or be correlated, and different labels are not independent of each other. As a result, the direct application of traditional long-tail strategies based on single-label settings will not be able to effectively handle long-tail multi-label tasks, resulting in poor performance of image classification models under this task.

[0040] Starting from the feature level, the present invention attempts to extend the traditional contrastive learning strategy to the long-tail multi-label scenario, aiming to learn a feature space that can effectively distinguish different labels and capture label co-occurrence. At the same time, the present invention can also be combined with other optimization processing techniques at different stages (such as data resampling, cost-sensitive learning, model and test end integration), which is flexible and adaptable.

[0041] like Figure 1The following is a flowchart of a method for optimizing the long-tail multi-label feature space based on contrast learning in an embodiment of the present invention. The embodiment of the present invention provides a method for optimizing the long-tail multi-label feature space based on contrast learning, including:

[0042] S1. Input the long-tail multi-label data after data augmentation processing into the first branch and the second branch of the contrast learning framework respectively to extract the first sample pair feature and the second sample pair feature.

[0043] S2. Calculate the similarity between the first sample pair feature and the second sample pair feature.

[0044] S3. Based on the similarity, calculate the multi-label sample label overlap coefficient between the first sample pair feature and the second sample pair feature.

[0045] S4. Based on the multi-label sample label overlap coefficient, perform class averaging and class supplementation operations to obtain a balanced multi-label contrast learning loss function.

[0046] S5. Use the balanced multi-label contrast learning loss function for supervised training to obtain an optimized feature space.

[0047] Specifically, as can be seen from the Figure 1 flow shown, the present invention first inputs the long-tail multi-label data after data augmentation processing into the first branch and the second branch of the contrast learning framework respectively to extract the first sample pair feature and the second sample pair feature; then calculates the similarity between the first sample pair feature and the second sample pair feature; based on the similarity, further calculates the multi-label sample label overlap coefficient between the first sample pair feature and the second sample pair feature; finally, based on the multi-label sample label overlap coefficient, performs class averaging and class supplementation operations to obtain a balanced multi-label contrast learning loss function, and uses the balanced multi-label contrast learning loss function for supervised training to obtain an optimized feature space. Based on the supervised contrast learning method, the present invention designs a corresponding multi-label sample label overlap coefficient for the problem of multi-label sample label overlap, and adds it to the loss term to optimize the learning process of the contrast learning model, and optimizes the long-tail multi-label feature space based on contrast learning. At the same time, the present invention balances the expression of features of each category through class averaging operations, minimizes the dominant role of the head category in the model learning as much as possible, and prevents over-biasing towards the mainstream category. Finally, the present invention is trained through the optimized loss function to obtain a more robust feature representation, effectively supports the application of the model in downstream tasks such as classification and retrieval, and improves the fitting ability of the contrast learning model to complex real-world data.

[0048] In one embodiment, the multi-label sample label overlap coefficient is calculated by the Jaccard coefficient.

[0049] Specifically, the supervised contrastive loss function of the present invention can be designed as follows:

[0050]

[0051] where i represents the index of the anchor sample. For a sample example x in the batch sample B i , its corresponding label is y i , and the feature representation after passing through the encoder is z i ; B q is a subset of the batch sample B, which contains all samples of class q; |·| represents the number of samples in the set; τ is a scalar representing the temperature hyperparameter, which is used to control the sensitivity between similar samples.

[0052] Based on the above supervised contrastive loss function, the label overlap coefficient s i,p (the coefficient value between the i-th sample and the p-th sample) is introduced. s i,p ∈[0,1] represents the overlap degree between the two groups of labels of the i-th sample and the p-th sample. If the label sets of the two samples are the same, then s i,p =1, otherwise s i,p <1. When s i,p =0, it means that there are no common labels between the samples.

[0053] The present invention calculates s i,p based on the Jaccard coefficient. The specific formula is as follows:

[0054]

[0055] where represents the label of the n-th class of the i-th sample, which is 1 if the class n is included, and 0 otherwise.

[0056] At this time, the optimized supervised contrastive loss function can be defined as:

[0057]

[0058] where the set Ω(i) = {p ∈ A(i): s i,p >0}, and A(i) = B\{i}. The set Ω(i) contains samples with s i,p greater than the threshold 0, and this threshold reflects the judgment of the similarity degree of the label sets in the contrastive loss.

[0059] In one embodiment, the data augmentation processing includes any one of random cropping, adjusting the target image to a preset resolution, color distortion, and Gaussian blur.

[0060] In one embodiment, inputting the long-tail multi-label data after data augmentation processing into the first branch and the second branch of the contrastive learning framework respectively for feature extraction to obtain the first sample pair feature and the second sample pair feature includes: inputting the long-tail multi-label data after the first data augmentation processing into the first branch of the contrastive learning framework for feature extraction to obtain the first sample pair feature, and inputting the long-tail multi-label data after the second data augmentation processing into the second branch of the contrastive learning framework for feature extraction to obtain the second sample pair feature.

[0061] Specifically, the first branch of the contrastive learning framework of the present invention can adopt the first data augmentation processing and perform feature extraction to obtain the first sample pair feature. The second branch of the contrastive learning framework can adopt the second data augmentation processing and perform feature extraction to obtain the second sample pair feature. Alternatively, the first branch of the contrastive learning framework can adopt the second data augmentation processing and perform feature extraction to obtain the first sample pair feature. The second branch of the contrastive learning framework can adopt the first data augmentation processing and perform feature extraction to obtain the second sample pair feature. The two branches of the contrastive learning framework adopt different data augmentation strategies to increase the robustness and generalization ability of the contrastive learning model to the data, so as to learn more representative features. The first data augmentation processing can be standard data augmentation such as random cropping (including any one of cropping, resizing, and flipping) and adjusting the image to a preset resolution (for example, a resolution of 224×224 pixels). The second data augmentation processing can be SimAugment, which adopts three means: random cropping (including any one of cropping, resizing, and flipping), color distortion, and Gaussian blur.

[0062] Specifically, the two branches represent different perspectives of the same input sample. By applying different data augmentations to the same image, two views with different semantics but from the same image can be generated. The purpose of doing this is to make the model pull closer the feature vectors of different perspectives to learn a more robust representation that is independent of a specific perspective. In addition, the diversity of features is increased, enhancing the robustness of the contrastive learning model and preventing the contrastive learning model from overfitting to specific features. Finally, the present invention can also adopt an adaptive augmentation strategy (i.e., dynamically adjusting the data augmentation intensity according to the magnitude of the contrastive loss. For example, if the similarity between the two views under the contrastive loss is too low, the augmentation intensity can be weakened, and vice versa) and a multi-scale augmentation strategy (i.e., performing data augmentation at different scales to guide the model to pay attention to local information and global information simultaneously). The present invention can be applied to the long-tail multi-label scenario. In the long-tail multi-label scenario, since the same sample contains one or more categories and there is label co-occurrence, the relationship between labels needs to be considered additionally. Therefore, traditional single-label methods cannot be directly applied to the multi-label scenario. For example, under single-label, pictures of the same dog category will be pulled closer in the feature space, while in a multi-label image, a picture can have both a dog and a cat. In this case, if we still pull closer according to the same category, only the pictures that contain both a dog and a cat will be pulled closer as samples, while samples with all other category combinations will be pulled farther away. Therefore, the present invention designs a multi-label sample label overlap coefficient where label relationships need to be processed, enabling effective application in long-tail multi-label tasks and filling the gap in the field of methods that do not consider from the feature level.

[0063] As Figure 2 shown in the flowchart of a method for class averaging and class supplementation operations in an embodiment of the present invention. In one embodiment, the class supplementation operation includes performing a class supplementation operation on the initial balanced multi-label contrastive learning loss function using the class prototype to obtain the final balanced multi-label contrastive learning loss function.

[0064] Specifically, the present invention performs class supplementation and class averaging operations on the optimized supervised contrastive loss function (that is, the initial balanced multi-label contrastive learning loss function). The class supplementation operation is to ensure that all classes appear in each batch to avoid unstable optimization. The present invention introduces the class prototype feature representation to perform the class supplementation operation on the loss function. The key idea of the class averaging operation is to average the instances of each class in the batch, so that each class has an approximate contribution to the optimization. Intuitively, it reduces the proportion of the head classes in the denominator and emphasizes the importance of the tail.

[0065] At this time, the final balanced multi-label contrastive learning loss function is obtained

[0066]

[0067] Among them, is the set of labels in the batch sample B, and B j is all the samples in the batch sample B that contain the category j, and t j is the index of the category prototype, and T ′ ={p∈T: s i,p >0} is a subset of the category prototype index set T={t1, t2,..., t c}.

[0068] As Figure 3 shown in the flowchart of a method for obtaining a category prototype in an embodiment of the present invention, in one embodiment, the category prototype is obtained through the following steps:

[0069] S401. Collect the first sample pair feature matrix, the first sample pair label matrix, the second sample pair feature matrix, and the second sample pair label matrix.

[0070] S402. Perform category prototype estimation on the first sample pair feature matrix, the first sample pair label matrix, the second sample pair feature matrix, and the second sample pair label matrix to obtain an initial category prototype.

[0071] S403. Update the initial category prototype using the exponential moving average method to obtain the final category prototype.

[0072] Specifically, for the learning of the category prototype, the present invention regards the multi-label sample as a combination of category prototypes selected based on relevant labels. Assume that Z t ∈R N×d is a d-dimensional feature matrix with N samples, L∈{0, 1} N×C is its corresponding label matrix, and CP t ∈R C×d is the category prototype matrix in a training iteration t. Then Z t =L·CP t +ε, where ε is the residual noise term. Assume that the label noise is unbiased and independent of the label matrix L. Then the category prototype can be approximately expressed as At this time, the category prototype is updated through learning iteration as:

[0073]

[0074] Among them, ρ is the decay hyperparameter in the exponential moving average. The exponential moving average (EMA) avoids the collapse of the category prototype to a certain extent during the training iteration process.

[0075] In addition, all the representations of contrastive learning are l2-normalized to ensure that the feature space is within a unit hypersphere.

[0076] In one embodiment, the long-tail multi-label feature space optimization method based on contrast learning further includes: using a balanced multi-label contrast learning loss function to supervise and train the encoder to obtain a relatively balanced feature space.

[0077] Specifically, in supervised contrast learning, the construction of the feature space mainly depends on the attraction between similar samples and the repulsion between dissimilar samples. Similar samples will be pulled closer to each other in the space, and this process is independent of the class frequency (i.e., the number of samples in each class). Therefore, regardless of whether the dataset distribution is balanced, samples of the same class will gather together as much as possible, forming good intra-class compactness. However, the process of pulling different class samples apart is affected by the class frequency. The more numerous head classes dominate the training process because, as negative samples, the head classes with a large number have a greater cumulative gradient contribution. This phenomenon causes the model to mainly focus on optimizing head class samples and relatively insufficient learning of tail class samples. This imbalance directly affects the separability of the feature space, resulting in a significant decline in the representation ability and classification performance of tail classes under long-tail data distributions. To address this problem, the present invention optimizes the supervised contrast loss function, reducing the gradient contribution of head class samples, thereby giving higher attention to tail classes in feature learning. This improvement is also further adapted to the multi-label scenario, better handling the label overlap problem, enabling supervised contrast learning to be better applied to real long-tail multi-label data distributions. Therefore, by supervising and training the encoder, a more balanced and highly separable feature space can be obtained, improving performance in long-tail and multi-label scenarios.

[0078] It should be noted that during the training process of the present invention, the contrast learning model is trained in batches, and each round of training is represented by t. A batch contains N samples, and after passing through the feature extractor, a feature matrix Z can be obtained. t . Combine Z t with the corresponding label matrix L, and calculate an approximation of the class prototype through an approximate formula It is an approximate expression of CP t . Starting from the second round, the class prototype CP t of each round is updated by combining the of the current batch with the class prototype CP t-1 of the previous round, that is, in the first round, it is calculated). After each subsequent round, CP (class prototype) is the Multiply by the coefficient (1 - ρ) and add the CP calculated in the previous round multiplied by ρ. This update method utilizes the idea of Exponential Moving Average (EMA), balancing the weights of the current batch information and historical information through the two coefficients ρ and 1 - ρ. The CP calculated in each round will be used to balance the class compensation operation in the multi-label contrastive learning loss function and then for the calculation of the loss. This iterative update method ensures that the class prototypes can be gradually optimized during training, taking into account the dynamic adjustment of the current batch features and historical features.

[0079] In one embodiment, the method for optimizing the long-tail multi-label feature space based on contrastive learning further includes: performing backpropagation parameter update on the optimized feature space and then outputting the finally optimized feature space.

[0080] In one embodiment, the SGD optimizer is used to perform backpropagation parameter update on the optimized feature space and then output the finally optimized feature space.

[0081] As Figure 4 shown is the flowchart of another method for optimizing the long-tail multi-label feature space based on contrastive learning in the embodiments of the present invention. Specifically, the present invention inputs long-tail multi-label data into the contrastive learning framework. The first branch of the contrastive learning framework adopts the first data augmentation process (standard data augmentation), and the second branch of the contrastive learning framework adopts the second data augmentation process (SimAugment). Then, through the encoder and the projection layer, the first sample pair feature vector and the second sample pair feature vector are obtained for feature extraction. Loss calculation is performed on the first sample pair feature vector and the second sample pair feature vector. Among them, first, the similarity between the first sample pair feature vector and the second sample pair feature vector is calculated, and then the multi-label sample label overlap coefficient is calculated based on the similarity. Based on the multi-label sample label overlap coefficient, class averaging and class supplementation operations are performed to obtain the optimized feature space. The SGD optimizer is used to perform backpropagation parameter update on the optimized feature space and then output the finally optimized feature space. The finally optimized feature space is used for downstream tasks after contrastive learning.

[0082] When the present invention applies the relatively balanced feature space for each category (here refers to the feature space after long-tail multi-label contrastive loss learning, so it is no longer a long-tail distribution) to downstream tasks, it can effectively improve the generalization ability of the model, especially in classification tasks. Based on the learned feature space, one or more classifiers are introduced to map the feature vectors to specific class labels. The classifiers can effectively learn the discriminant boundaries suitable for each category under the supervision of the classification loss during the training process, thereby improving the classification accuracy on different categories and the generalization ability of the overall model.

[0083] It should be noted that in downstream tasks, a fully connected classifier based on Sigmoid can be adopted, and the classification loss can adopt BCE loss (the most basic multi-label classification loss); Focal loss (focusing on difficult-to-separate samples); ASL (focusing on the imbalance between positive and negative samples, applying larger weights to samples of each category to make the model focus on the learning of samples); DB Focalloss assigns larger weights to tail classes by balancing the distribution weights and tolerating regularization of head samples, enabling the model to focus on their learning. An adaptive strategy for adjusting the weights of positive and negative samples is added based on DB Focal loss. Among them, the adaptive strategy dynamically adjusts based on the probability difference between positive and negative samples of each category, dynamically applying larger weights to samples, and reducing the probability difference between positive and negative sample predictions, which reflects that the model balance pays more attention to sample learning. DB Focalloss takes into account the weights for balancing the long-tail multi-label distribution, head sample tolerance regularization, and positive and negative sample balance, thus maximizing the utilization of the learned feature space.

[0084] For the long-tail multi-label classification method based on contrast learning, the present invention extends the contrast learning strategy to the long-tail multi-label scenario. By means of the class supplementation strategy, it ensures that each training batch contains samples of all categories, thus avoiding unstable optimization; the class averaging operation balances the gradients of category samples, correcting the bias that the model tends to mispredict tail class samples as head class samples, in order to learn a feature representation space that can effectively distinguish different labels. This method enhances the generalization ability of the model under the long-tail distribution, especially showing a significant improvement in the performance of tail classes. At the same time, this method takes into account the label co-occurrence problem in multi-label learning. By directly dealing with label overlap, it optimizes the distribution of multi-label data in the feature space and improves the robustness of the model to multi-label tasks.

[0085] As Figure 4 shown is a schematic structural diagram of an apparatus for optimizing the long-tail multi-label feature space according to an embodiment of the present invention. The embodiment of the present invention also provides an apparatus for optimizing the long-tail multi-label feature space, including:

[0086] A feature extraction module, configured to input the long-tail multi-label data after data augmentation processing into the first branch and the second branch of the contrast learning framework respectively for feature extraction to obtain the first sample pair feature and the second sample pair feature;

[0087] A first calculation module, configured to calculate the similarity between the first sample pair feature and the second sample pair feature;

[0088] A second calculation module, configured to calculate the multi-label sample label overlap coefficient between the first sample pair feature and the second sample pair feature based on the similarity;

[0089] The class-average and class-supplementary operation module is used to perform class-average and class-supplementary operations based on the multi-label sample label overlap coefficient to obtain a balanced multi-label contrastive learning loss function;

[0090] The supervised training module is used to perform supervised training using the balanced multi-label contrastive learning loss function to obtain an optimized feature space.

[0091] It should be noted that the method embodiments of the present invention can be implemented one-to-one corresponding to the foregoing virtual device embodiments, and will not be elaborated herein.

[0092] The present invention can also be designed as a flexible and pluggable feature space learning module. By learning the feature space, it increases the robustness to handle complex data characteristics in the real world and has wide application scenario adaptability. Due to the generality of its feature representation, this method can be used as a plug-in module and can be easily integrated into various deep learning applications such as existing classification models, object detection, and recommendation systems. Its modular design allows it to be combined with other optimization techniques (such as data resampling, cost-sensitive learning, etc.) to further enhance the performance of the model in different tasks. Through this plug-in, the system can effectively handle the challenges in long-tail distribution and multi-label learning, showing extremely high flexibility and adaptability, enabling the model to have better generalization ability in diverse data and application environments.

[0093] An embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned method is implemented.

[0094] As Figure 5 shown is a schematic diagram of the physical structure of the computer device provided by an embodiment of the present invention. The computer device includes a processor, a memory, and a bus. Among them, the processor and the memory complete communication with each other through the bus.

[0095] The processor is used to call program instructions in the memory to execute the methods provided by the above-mentioned method embodiments.

[0096] An embodiment of the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned method is implemented.

[0097] An embodiment of the present invention also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the above-mentioned method is implemented.

[0098] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0099] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0100] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realize the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0101] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0102] The specific embodiments described above further elaborate on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A long-tail multi-label feature space optimization method based on contrastive learning, characterized in that: include: The long-tail multi-label data after data enhancement is input into the first branch and the second branch of the contrastive learning framework for feature extraction to obtain the first sample pair feature and the second sample pair feature respectively; Calculate the similarity between the first sample pair feature and the second sample pair feature; Based on the similarity, a multi-label sample label overlap coefficient between the first sample pair feature and the second sample pair feature is calculated; Based on the multi-label sample label overlap coefficient, class averaging and class supplementation operations are performed to obtain a balanced multi-label contrastive learning loss function; The balanced multi-label contrastive learning loss function is used to perform supervised training to obtain an optimized feature space.

2. The method according to claim 1, characterized in that: The class supplementation operation includes using category prototypes to perform a class supplementation operation on an initial balanced multi-label contrastive learning loss function to obtain a final balanced multi-label contrastive learning loss function.

3. The method according to claim 2, characterized in that: The category prototype is obtained by the following steps: Collecting a first sample pair feature matrix, a first sample pair label matrix, a second sample pair feature matrix, and a second sample pair label matrix; Performing category prototype estimation on the first sample pair feature matrix, the first sample pair label matrix, the second sample pair feature matrix, and the second sample pair label matrix to obtain an initial category prototype; The exponential moving average method is used to update the initial category prototype to obtain the final category prototype.

4. The method according to claim 1, characterized in that: The multi-label sample label overlap coefficient is calculated by the Jaccard coefficient.

5. The method according to claim 1, characterized in that: The data enhancement processing includes any one of random cropping, adjusting the target image to a preset resolution, color distortion, and Gaussian blur.

6. The method according to claim 5, characterized in that The method of inputting the long-tail multi-label data after data enhancement processing into the first branch and the second branch of the contrastive learning framework to respectively extract features to obtain first sample pair features and second sample pair features includes: inputting the long-tail multi-label data after the first data enhancement processing into the first branch of the contrastive learning framework to extract features to obtain first sample pair features, and inputting the long-tail multi-label data after the second data enhancement processing into the second branch of the contrastive learning framework to extract features to obtain second sample pair features.

7. A long-tail multi-label feature space optimization device based on contrastive learning, characterized in that: include: A feature extraction module is used to input the long-tail multi-label data after data enhancement processing into the first branch and the second branch of the contrastive learning framework to respectively extract features to obtain first sample pair features and second sample pair features; A first calculation module, used for calculating the similarity between the first sample pair feature and the second sample pair feature; A second calculation module is used to calculate the multi-label sample label overlap coefficient of the first sample pair feature and the second sample pair feature based on the similarity; A class averaging and class supplementation operation module, used to perform class averaging and class supplementation operations based on the multi-label sample label overlap coefficient to obtain a balanced multi-label contrastive learning loss function; The supervised training module is used to perform supervised training using the balanced multi-label contrastive learning loss function to obtain an optimized feature space.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.