Pseudo label generation and contrast learning joint optimization method for target detection

Through the joint optimization method of pseudo-label generation and comparison learning, the problems of insufficient quality screening and insufficient feature optimization in the existing technology are solved, which significantly improves the cross-domain performance and robustness of the object detection model and reduces the dependence on labeled data.

CN120032161APending Publication Date: 2025-05-23HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411956560.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-29
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing object detection methods lack effective pseudo-label quality screening mechanism when using pseudo-labels to train target domain models, which may lead to negative migration; at the same time, intra-class consistency and inter-class differences are not fully utilized, and the comprehensive optimization strategy is insufficient, resulting in limited model performance improvement.

Method used

A joint optimization method for pseudo-label generation and comparison learning is proposed. By generating pseudo-labels with confidence scores and performing high-quality screening, combining intra-class and inter-class comparison learning, the consistency and distinction of features are optimized, and the pseudo-label classification loss, bounding box regression loss and comparison learning loss are integrated through comprehensive loss optimization.

Benefits of technology

It significantly improves the cross-domain performance and robustness of the target detection model, reduces dependence on labeled data, effectively solves the negative impact of inter-domain differences, and significantly improves the detection performance of the model under the condition of unlabeled data in the target domain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032161A_ABST
    Figure CN120032161A_ABST
Patent Text Reader

Abstract

A false label generation and comparative learning combined optimization method for target detection comprises the following steps: firstly, reasoning target domain data by using a target detection model, generating a false label with a confidence score, and selecting a high-quality false label for training through a filtering mechanism; secondly, designing a comparative learning module, performing positive and negative sample comparison on target-level features in a source domain and a target domain, and improving the feature characterization capability of the model by maximizing the feature similarity between similar samples and minimizing the feature similarity between heterogeneous samples; and finally, false label classification training and bounding box regression training of real labels are combined to jointly optimize the target detection model. According to the invention, the problem of model performance reduction caused by domain offset in a target detection task can be solved; the influence of pseudo label noise is effectively reduced, and the generalization ability of a target detection model in a cross-domain task is improved. The method is suitable for target detection tasks in the fields of unmanned driving, security monitoring, intelligent manufacturing and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing and deep learning, and in particular to a joint optimization method of pseudo label generation and contrastive learning for target detection. Background Art

[0002] As one of the important research directions in the field of computer vision, object detection plays a key role in many fields such as autonomous driving, security monitoring, and smart medical care. In recent years, with the rapid development of deep learning technology, object detection algorithms represented by the YOLO series have been widely used due to their advantages such as end-to-end training, strong real-time performance, and high detection accuracy. However, the training of object detection models usually relies on large-scale annotated datasets, which brings significant challenges in practical applications.

[0003] The performance of the target detection model is closely related to the quality of the training data. The quantity and quality of the labeled data directly determine the generalization ability of the model. However, for different target detection tasks (such as different scenarios or fields), building large-scale, high-quality labeled datasets usually requires a lot of manpower and time, and the labeling cost is high. In addition, some specific fields or scenarios, such as severe weather conditions and special industrial environments, may face difficulties in data collection, which further limits the improvement of model performance.

[0004] In practical applications, object detection models usually need to be migrated from a well-annotated source domain to a sparsely annotated target domain. However, since the source and target domains may have significant differences in image distribution, background complexity, lighting conditions, etc., i.e., inter-domain differences, directly migrating the trained model to the target domain will result in a significant decrease in detection performance. This problem is particularly prominent in cross-scenario or cross-device object detection tasks.

[0005] In response to the situation where there is no labeled data in the target domain, recent studies have proposed methods for unsupervised learning using pseudo labels. Pseudo labels are generated by a trained model through reasoning on the target domain data, representing the preliminary annotation of the target domain samples. However, the quality of pseudo labels is often unstable, especially when the target domain data distribution is complex and the initial model performance is insufficient. Low-quality pseudo labels may have a negative impact on subsequent training, which is the negative transfer problem. Therefore, how to effectively screen high-quality pseudo labels and make full use of pseudo label information in model training has become a key research direction.

[0006] In recent years, contrastive learning, as an unsupervised feature learning method, has demonstrated excellent performance in tasks such as object detection and image classification. By optimizing the feature similarity between samples based on the characteristics of similar intra-class sample features and different inter-class sample features, contrastive learning can enhance the feature expression ability of the model. However, how to effectively combine contrastive learning with pseudo-label training to further optimize the performance of the object detection model remains a hot and difficult issue in current research. Summary of the Invention

[0007] Aiming at the technical problems existing in the above-mentioned existing object detection methods in training the target domain model using pseudo-labels, such as the lack of an effective screening mechanism for pseudo-label quality, which may lead to negative transfer; in terms of feature optimization, the intra-class consistency and inter-class differences are not fully utilized; and the comprehensive optimization strategy is insufficient, and the pseudo-label loss, bounding box regression loss, and feature contrast loss cannot be effectively combined, resulting in limited improvement in model performance. This technical solution provides a joint optimization method for pseudo-label generation and contrastive learning for object detection, which combines pseudo-label screening and contrastive learning, effectively utilizes unlabeled data in the target domain, and improves the cross-domain performance and robustness of the object detection model; and through pseudo-label generation and screening, feature contrast learning, and comprehensive loss optimization, the performance of the detection model can be significantly improved under the condition of unlabeled data in the target domain; this method can not only reduce the dependence on labeled data, but also effectively solve the negative impact brought by domain differences; it can effectively solve the above problems.

[0008] The present invention is achieved through the following technical solutions:

[0009] A joint optimization method for pseudo-label generation and contrastive learning for object detection. First, use the object detection model to infer the target domain data, generate pseudo-labels with confidence scores, and select high-quality pseudo-labels for training through a filtering mechanism; secondly, design a contrastive learning module to perform positive and negative sample contrast on the object-level features in the source domain data and the target domain data, and improve the feature representation ability of the model by maximizing the feature similarity between the same-class samples and minimizing the feature similarity between different-class samples; finally, combine the pseudo-label classification training with the bounding box regression training of the real labels to jointly optimize the object detection model; including the following steps:

[0010] Initial training: Train the YOLOv5 model on the source domain labeled data to obtain a preliminary detection model M init ;

[0011] Pseudo-label generation: Use the preliminary detection model M init to infer the unlabeled data in the target domain, generate pseudo-labels, and screen high-quality pseudo-labels through a confidence threshold;

[0012] Feature extraction: During the model training process, the YOLOv5 backbone network CSPDarknet53 is used to extract features from the input image. CSPDarknet53 generates category predictions, confidence predictions, and bounding box predictions for each grid cell through a multi-level feature extraction mechanism.

[0013] Sample screening: The pseudo labels generated in step S2 are screened according to the confidence level to ensure the quality of the training data. The specific operation method of screening the pseudo labels is as follows:

[0014] After feature extraction, the confidence prediction of each target is evaluated, and the target is divided into positive samples and negative samples according to the confidence prediction value; targets with high confidence and matching categories are judged as positive samples, and targets with low confidence or mismatched categories are judged as negative samples;

[0015] Contrastive learning: Through intra-class and inter-class contrastive learning, the contrastive learning loss is calculated to optimize the consistency and discrimination of features;

[0016] Comprehensive loss optimization: Pseudo-labels and 1% of the labeled data in the target domain are used as the true value to calculate the loss with the model prediction result. The loss function includes pseudo-label classification loss, bounding box regression loss and contrastive learning loss, and performs joint optimization to improve model performance.

[0017] Furthermore, the specific operation mode of the initial training is:

[0018] Let the source domain labeled data be D src ={(x i ,y i )}, where x i is the input image, y i The annotation information corresponding to the input image, including target category, confidence and bounding box information;

[0019] The initial model is obtained through training, which optimizes the source domain data through the standard YOLO loss function, including classification loss, confidence loss and bounding box regression loss;

[0020] Through this process, a model with preliminary target detection capability is obtained, namely the preliminary detection model M init ; Provides a basis for subsequent pseudo-label generation and reasoning about unlabeled data in the target domain.

[0021] Furthermore, the specific operation method of generating the pseudo-label is as follows: let the target domain unlabeled data be D src ={X j}, where X j For source domain samples, using the preliminary detection model M init There is no labeled data D for the target domain src= {X j} to perform inference and generate preliminary pseudo-labels Preliminary pseudo-labels include target category, confidence, and bounding box information; when generating pseudo-labels, a confidence threshold c is set to filter out targets with confidence higher than the threshold c as valid pseudo-labels and eliminate low-quality pseudo-labels with confidence lower than the threshold c; the filtered pseudo-label set is: After passing through the conf(Y j ) activation function to map it to the interval [0, 1], Y j is the model's predicted label, and the calculation formula for conf(Y j ) is as follows:

[0022]

[0023] In the above formula, e is the natural exponential constant.

[0024] Furthermore, in the intra-class contrast learning in the contrastive learning, the specific operation method for calculating the intra-class contrast learning loss is:

[0025] Extract positive sample features: Extract the corresponding features in the feature map according to the positive sample index and perform normalization processing on them;

[0026] Calculate the intra-class similarity matrix: Calculate the cosine similarity between positive samples using matrix multiplication; Exclude the similarity of a sample to itself through a diagonal mask;

[0027] Scale the similarity using the temperature parameter to control the sensitivity of the model to the similarity difference. A smaller temperature parameter will amplify the gap between high and low similarities;

[0028] The formula for calculating the intra-class contrast learning loss for each positive sample is as follows:

[0029]

[0030] In the above formula, x i represents the feature vector of the current target sample, the sample concerned in the intra-class contrast learning, and the goal is to maximize its similarity with positive samples among other samples of the same class while minimizing its similarity with negative samples among other samples of different classes; x j represents the feature vector of other samples of the same class as x i ; the constraint j≠i means that x j and x i are different samples of the same class; x k represents the feature vector of a sample of a different class from x i , and the constraint k≠i means that x k and x iare samples of different categories; N is the total number of samples, N p is the number of positive samples in the current category, that is, samples x i The total number of samples of the same category; N n is the number of positive samples in the current category, belonging to sample x i The total number of samples of different categories; sim(x i ,x j ) is the sample x i ,x j The similarity measure is usually cosine similarity, and its calculation formula is: τ is the temperature parameter, which is used to adjust the scaling of similarity and control the range of gradient variation;

[0031] The intra-class contrastive learning loss is obtained, and the average of the losses of all positive samples is taken as the intra-class loss of the current category; in this process, the features of the positive samples are normalized, and the similarity between samples of the same category is calculated and optimized; by maximizing the similarity of samples within the class, the feature distribution differences between samples of the same category are narrowed.

[0032] Furthermore, in the inter-class contrastive learning in the contrastive learning, the specific operation method of calculating the contrastive inter-class learning loss is as follows:

[0033] Extract the positive sample features within each category and normalize each feature to the same unit vector;

[0034] For each category of positive samples, calculate the center vector / mean vector of the positive samples in each category as the feature representative of the category;

[0035] The loss objective is to minimize the similarity of the same category and maximize the similarity of different categories;

[0036] The formula for calculating the inter-class contrastive learning loss for each positive sample is as follows:

[0037]

[0038] In the above formula, N is the total number of samples, usually the total number of samples in the current batch; C is the number of categories, that is, there are C categories in the classification task; is the feature vector of the nth sample in the ath class, indicating a sample belonging to class a; The feature vector of the sample belonging to the sth class represents a class sample different from the current class a; is the number of samples in class a; is the number of samples in the sth class; for The similarity between them is usually measured by cosine similarity, and its formula is

[0039] After the above steps, we get the inter-class contrastive learning loss. Inter-class contrastive learning aims to enhance the feature discrimination between different categories. It optimizes the separation of inter-class features by calculating the similarity between positive samples from different categories. This process makes the features of different categories more dispersed in the feature space, thereby improving the ability to distinguish between categories, reducing confusion between categories and improving classification accuracy.

[0040] Furthermore, the loss function in the comprehensive loss optimization includes:

[0041] Pseudo-label classification loss: used to measure the classification accuracy of the model on pseudo-labels, prompting the model to make accurate category predictions on target domain data;

[0042] Bounding box regression loss: used to optimize the model's prediction of target position and size and improve the bounding box regression accuracy;

[0043] Contrastive learning loss: It includes intra-class and inter-class contrast losses, which aims to optimize the cohesion and discrimination of sample features, thereby improving the representation ability of the model in the target domain.

[0044] Furthermore, the source domain data and target domain data adopt the Cityscape dataset and Sim10k dataset respectively.

[0045] Furthermore, the Cityscapes dataset is a source domain dataset; the Cityscapes dataset mainly contains high-quality urban street scene images, which are suitable for target detection in traffic scenes and have rich annotation information. The dataset contains more than 5,000 high-resolution images, covering 20 categories, such as vehicles, pedestrians, buildings, etc. In order to align with the target domain, only vehicles and pedestrians are selected.

[0046] Furthermore, the Sim10k dataset is a target domain dataset; the Sim10k dataset contains synthetic images from a virtual driving environment with traffic scenes similar to Cityscapes. The images generated by this dataset are different in appearance from images in the real world, and there are some inter-domain deviations, such as weather, lighting, and perspective. The labels of the Sim10k dataset are similar to those of Cityscapes, mainly including target categories such as vehicles and pedestrians.

[0047] A device comprises a memory and a processor, wherein:

[0048] A memory for storing computer programs that can be run on the processor;

[0049] The processor, when running the computer program, can execute the steps of the above-mentioned method for joint optimization of pseudo label generation and contrastive learning for target detection.

[0050] A storage medium having a computer program stored thereon, wherein when the computer program is executed by at least one processor, the steps of the above-mentioned method for joint optimization of pseudo label generation and contrastive learning for target detection can be implemented.

[0051] Beneficial Effects

[0052] The present invention proposes a joint optimization method for pseudo label generation and contrastive learning for target detection, which has the following beneficial effects compared with the prior art:

[0053] (1) This technical solution adopts a high-quality pseudo-label generation and screening mechanism; by introducing an initial detection model, the target domain unlabeled data is inferred to generate pseudo-labels, and combined with a confidence threshold screening mechanism to ensure the high quality of pseudo-labels. By eliminating low-confidence and low-quality pseudo-labels, the negative transfer problem caused by low-quality pseudo-labels on model training is effectively avoided, thereby improving the training effect of the model.

[0054] (2) This technical solution adopts a feature optimization strategy that combines intra-class and inter-class comparative learning. In the feature learning process, both intra-class consistency and inter-class discrimination are optimized. By optimizing the similarity of intra-class positive sample features, the difference in intra-class distribution is reduced, and the robustness of the model is improved; by optimizing the difference in positive sample features between different categories, the model's ability to distinguish different target categories is enhanced. This strategy of combining intra-class and inter-class comparative learning significantly improves the feature expression ability and detection performance of the target detection model.

[0055] (3) This technical solution adopts a comprehensive loss optimization method to integrate pseudo-label classification loss, bounding box regression loss and contrastive learning loss. By reasonably allocating the weights of each loss, the detection accuracy and feature learning efficiency are balanced, thereby achieving multi-task collaborative optimization and further improving the training effect of the model.

[0056] (4) This technical solution has strong cross-domain adaptability and effectively solves the problem of inter-domain differences. Through the joint application of pseudo-label screening and contrastive learning, the model can better adapt to the image distribution characteristics of the target domain. When there is no annotation or limited annotation data in the target domain, high-precision target detection can still be achieved, which significantly improves the performance of cross-domain tasks.

[0057] (5) This technical solution significantly reduces the labeling cost. It adopts pseudo-label generation and screening technology and combines it with unsupervised comparative learning methods to make full use of the unlabeled data in the target domain, thereby reducing the dependence on large-scale labeled data and significantly reducing the cost of manual labeling, providing technical support for the application and promotion of target detection tasks in actual scenarios.

[0058] (6) This technical solution adopts modular design, which is easy to integrate and expand. It is based on the YOLOv5 model framework. Its pseudo-label generation and screening, contrastive learning module and comprehensive loss optimization strategy are all modularly designed, with good integration and scalability. This method can be easily applied to other target detection models or deep learning tasks, and has strong versatility and applicability.

[0059] (7) This technical solution can significantly improve model performance and stability when the target domain data distribution is complex or there are differences between domains. By optimizing feature distribution through the screening of high-quality pseudo-labels and comparative learning, it can also enhance the stability and robustness of the model in different tasks and scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 It is a schematic diagram of the overall process of the present invention. DETAILED DESCRIPTION

[0061] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. The described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Under the premise of not departing from the design concept of the present invention, various modifications and improvements made by ordinary persons in the art to the technical solutions of the present invention should all fall within the protection scope of the present invention.

[0062] Embodiment 1:

[0063] A joint optimization method of pseudo-label generation and contrastive learning for target detection. The implementation process of contrastive learning based on YOLOv5 grid confidence can be divided into data preparation, model design and training optimization. First, YOLOv5 is used to grid the image, and the confidence of each grid is extracted, that is, the product of the target confidence and the category probability. The confidence threshold is set to divide the grid into positive samples and negative samples, and the corresponding feature embedding is extracted. The model adds a contrastive learning module based on YOLOv5. The positive and negative samples optimize the embedding space by designing contrastive learning loss, so that the features of positive samples are closer and the features of negative samples are farther away. During training, the weighted sum of the detection loss and the contrastive learning loss is calculated by forward propagation, and the model parameters are updated by back propagation. At the same time, the division strategy of positive and negative samples is dynamically adjusted to adapt to the training process. The total loss function combines the weights of the detection loss and the contrastive learning loss of YOLOv5 to achieve the coordinated optimization of feature representation learning and target detection tasks. The specific operation method is as follows Figure 1 As shown, the steps include:

[0064] S1. Prepare source domain dataset and target domain dataset. The source domain dataset includes images and corresponding annotation files, bounding boxes and category labels; the target domain dataset only includes image data.

[0065] The source domain dataset is the Cityscapes dataset; the Cityscapes dataset mainly contains high-quality urban street scene images, which are suitable for target detection in traffic scenes and have rich annotation information. The dataset contains more than 5,000 high-resolution images, covering 20 categories such as vehicles, pedestrians, buildings, etc. In order to align with the target domain, only vehicles and pedestrians are selected.

[0066] The target domain dataset is the Sim10k dataset; the Sim10k dataset contains synthetic images from a virtual driving environment with traffic scenes similar to Cityscapes. The images generated by this dataset are different in appearance from images in the real world, and there are some inter-domain deviations, such as weather, lighting, and perspective. The labels of the Sim10k dataset are similar to those of Cityscapes, mainly including target categories such as vehicles and pedestrians.

[0067] S2. Based on the source domain labeled data, use the YOLOv5 model for preliminary training. Let the source domain labeled data be D src ={(x i ,y i )}, where x i is the input image, y iis the annotation information corresponding to the input image, including target category, confidence and bounding box information. The initial model is obtained through training. The model optimizes the source domain data through the standard YOLO loss function, including classification loss, confidence loss and bounding box regression loss. Through this process, a model with preliminary target detection capabilities is obtained, namely the preliminary detection model M init ; Provides a basis for subsequent pseudo-label generation and reasoning about unlabeled data in the target domain.

[0068] S3. Let the target domain have unlabeled data as D src ={X j}, where X j For source domain samples, use the preliminary detection model M init There is no labeled data D for the target domain tgt ={X j} Perform inference and get the result pred_d. Iterate through the prediction list of each image one by one. Each prediction consists of the detection results of multiple targets and generates preliminary pseudo labels. Preliminary pseudo-labeling Including target category cls, confidence conf and bounding box coordinates xyxy; the size of the image is normalized and calculated through the tensor gn to ensure the universality of results for images of different resolutions.

[0069] When generating pseudo labels, a confidence threshold c is set, and targets with confidence higher than the threshold c are selected as valid pseudo labels, and low-quality pseudo labels with confidence lower than the threshold c are removed; the pseudo label set after selection is: After conf(Y j ) activation function to map it to the interval [0,1], Y j Predict labels for the model, conf(Y j ) is calculated as follows:

[0070]

[0071] In the above formula, e is the natural exponential constant.

[0072] S4. During the training process of the model, the YOLOv5 backbone network CSPDarknet53 is used to extract features of the input image. CSPDarknet53 generates category predictions, confidence predictions, and bounding box predictions for each grid cell through a multi-level feature extraction mechanism.

[0073] When extracting target information, by judging whether the confidence is higher than the set threshold conf_threshold, the detection results with low confidence are filtered out, thereby improving the reliability of the target detection results.

[0074] S5. After the target information that meets the conditions is converted into tensor format through torch.tensor, it is added to the prediction list predictions_list one by one to ensure the scalability of the results and the ease of subsequent calculation.

[0075] S6, YOLOv5 prediction output pred_d contains multi-scale target detection box prediction values, and separates the prediction boxes of three scales of 80×80, 40×40 and 20×20 by index. The specific layer of the network model extracts the corresponding feature map, such as the 80×80 feature map. The feature map is flattened to convert the spatial dimensions (H, W) into a single dimension for easy association with the prediction box.

[0076] S7. After feature extraction, the confidence prediction of each target is evaluated. According to the value of the confidence prediction, the class probability class_probs and confidence conf_scores are extracted from the prediction value, and the screening conditions are set to distinguish positive and negative samples. The positive and negative samples are selected by Boolean mask. The target with high confidence and matching category is judged as positive sample, and the target with low confidence or mismatching category is judged as negative sample.

[0077] The positive sample screening conditions are: the category of the predicted box is equal to the current target category cls, and the confidence of the predicted box is higher than the set threshold conf_upper_threshold;

[0078] The screening conditions for negative samples are: the category of the predicted box is not equal to the current target category cls. The confidence is between conf_lower_threshold and conf_min_threshold.

[0079] Use torch.nonzero to extract the index of positive and negative samples and ensure the uniqueness of the index.

[0080] S8. Extract positive sample features and perform intra-class comparative learning. The purpose of intra-class comparative learning is to enhance the similarity between samples of the same category and optimize the consistency of intra-class features. In this process, the features of positive samples are normalized, and the similarity between samples of the same category is calculated and optimized. By maximizing the similarity of samples within a class, the feature distribution differences between samples of the same category are reduced. The specific operation method is:

[0081] Extract the feature vector corresponding to the positive sample from the flattened feature map, and use the L2 norm to normalize the positive sample features so that the length of the feature vector is standardized to 1 to avoid the influence of the eigenvalue scale on the similarity calculation. Calculate the cosine similarity between all positive samples through matrix multiplication to obtain the similarity matrix. Exclude the similarity between itself through the diagonal mask; scale the similarity using the temperature parameter to control the sensitivity of the model to the similarity difference. A smaller temperature parameter will amplify the gap between high and low similarities. The goal of intra-class contrastive learning is to maximize the similarity between positive samples. The formula for calculating the intra-class contrastive learning loss for each positive sample is as follows:

[0082]

[0083] In the above formula, x i Represents the feature vector of the current target sample, the sample of interest in intra-class contrastive learning, the goal is to maximize its similarity with other positive samples of the same type, while minimizing its similarity with negative samples of other classes; x j Represents x i The feature vectors of other samples of the same type; the constraint of j≠i means x j and x i are different samples of the same type; x k Represents x i The feature vectors of samples of different categories, the constraint k≠i represents x k and x i are samples of different categories; N is the total number of samples, N p is the number of positive samples in the current category, that is, samples x i The total number of samples of the same category; N n is the number of positive samples in the current category, belonging to sample x i The total number of samples of different categories; sim(x i ,x j ) is the sample x i ,x j The similarity measure is usually cosine similarity, and its calculation formula is: τ is a temperature parameter, which is used to adjust the scaling of similarity and control the range of gradient variation.

[0084] After the above steps, we get the intra-class contrastive learning loss, and take the average of the loss of all positive samples as the intra-class loss of the current category; the purpose of intra-class contrastive learning is to enhance the similarity between samples of the same category and optimize the consistency of intra-class features. In this process, the features of positive samples are normalized, and the similarity between samples of the same category is calculated and optimized; by maximizing the similarity of samples within a class, the feature distribution differences between samples of the same category are reduced.

[0085] S9, inter-class contrastive learning: Inter-class contrastive learning aims to enhance the feature discrimination between different categories. By calculating the similarity between positive samples from different categories, the separation of inter-class features is optimized. This process improves the ability to distinguish between categories by making the features of different categories more dispersed in the feature space, thereby reducing confusion between categories and improving classification accuracy. The specific operation method is:

[0086] Extract the features of positive samples of the current category and other categories (each category), and normalize each feature to the same unit vector to facilitate the subsequent calculation of cosine similarity; for the positive samples of each category, calculate the center vector / mean vector of the positive samples in each category as the feature representative of the category; the loss goal is to minimize the similarity of the same category and maximize the similarity of different categories;

[0087] The formula for calculating the inter-class contrastive learning loss for each positive sample is as follows:

[0088]

[0089] In the above formula, N is the total number of samples, usually the total number of samples in the current batch; C is the number of categories, that is, there are C categories in the classification task; is the feature vector of the nth sample in the ath class, indicating a sample belonging to class a; The feature vector of the sample belonging to the sth class represents a class sample different from the current class a; is the number of samples in class a; is the number of samples in the sth class; for The similarity between them is usually measured by cosine similarity, and its formula is

[0090] After the above steps, we get the inter-class contrastive learning loss. Inter-class contrastive learning aims to enhance the feature discrimination between different categories. It optimizes the separation of inter-class features by calculating the similarity between positive samples from different categories. This process makes the features of different categories more dispersed in the feature space, thereby improving the ability to distinguish between categories, reducing confusion between categories and improving classification accuracy.

[0091] S10, using pseudo-labels and 1% of the labeled data in the target domain as the true value and the model prediction result to calculate the loss, comprehensively optimize the pseudo-label classification loss, bounding box regression loss, intra-class and inter-class comparison loss, and uniformly optimize the model weights. The total loss function is:

[0092] L total =L cls +L box +L conf+λ 1 L intra +λ 2 L inter

[0093] Among them, λ 1 and λ 2 is a hyperparameter that controls the weight of intra-class and inter-class contrast losses.

[0094] L cls For classification loss, BCE (Binary Cross-Entropy) loss is usually used, which is suitable for multi-label classification tasks. The formula is as follows:

[0095]

[0096] Among them, y c is the true label of the category; The class probabilities predicted by the model.

[0097] L conf For the target confidence loss, BCE loss is used, and its formula is as follows:

[0098]

[0099] in, is the confidence probability predicted by the model, y i is the true label.

[0100] L box For the bounding box loss, GIOU loss is used, and its formula is as follows:

[0101] L box =1-GIOU(B p ,B t )

[0102] Among them, B p ,B t are the values ​​of the prediction box and the label box respectively; GIOU is defined as follows:

[0103]

[0104] Among them, A is the predicted box, B is the real box, and C is the minimum closed rectangular box containing the predicted box A and the real box B.

[0105] Through this total loss function, joint optimization is performed to improve the performance and generalization ability of the target detection model in the target domain.

[0106] Experimental verification:

[0107] In order to verify the feasibility and superiority of this embodiment, the inventor conducted experiments to verify and evaluate the effect of the method proposed in this embodiment.

[0108] The following common target detection performance indicators are used in the experiment:

[0109] mAP (mean Average Precision): Calculates the average precision of the detection results and measures the overall performance of the model in each category;

[0110] IoU (Intersection over Union): Calculates the overlap between the predicted box and the true box to evaluate the accuracy of bounding box regression;

[0111] Recall: measures the proportion of true targets detected by the model;

[0112] Precision: Measures the accuracy of positive samples in the model detection results.

[0113] The specific values ​​are shown in the following table:

[0114]

[0115]

[0116] This experiment uses Cityscapes as the source domain and Sim10k as the target domain. By comparing different optimization methods of the YOLOv5 model, it verifies the role of pseudo-label generation, intra-class contrastive learning, inter-class contrastive learning and comprehensive loss optimization in improving model performance. The experimental results show that with the gradual introduction of optimization methods, the accuracy and recall rate of the model have been significantly improved, and the mAP has increased from 65.4% of the original model to 76.8%. Precision and Recall also show obvious improvements.

[0117] Embodiment 2:

[0118] A device comprises a memory and a processor, wherein:

[0119] A memory for storing computer programs that can be run on the processor;

[0120] The processor, when running the computer program, can execute the steps of the method for joint optimization of pseudo label generation and contrastive learning for target detection described in Example 1.

[0121] Embodiment 3:

[0122] A storage medium, on which a computer program is stored, and when the computer program is executed by at least one processor, it can implement the steps of a method for jointly optimizing pseudo-label generation and contrastive learning for object detection described in Embodiment 1.

Claims

1. A joint optimization method for pseudo-label generation and contrastive learning for target detection, characterized in that: First, the target detection model is used to infer the target domain data, generate pseudo labels with confidence scores, and select high-quality pseudo labels for training through a filtering mechanism. Second, a contrastive learning module is designed to compare the positive and negative samples of the target-level features in the source domain data and the target domain data. By maximizing the feature similarity between samples of the same type and minimizing the feature similarity between samples of different types, the feature representation ability of the model is improved. Finally, the pseudo-label classification training is combined with the bounding box regression training of the real label to jointly optimize the object detection model; the following steps are included: Initial training: Train the YOLOv5 model on the source domain labeled data to obtain the preliminary detection model M init ; Pseudo label generation: Using the preliminary detection model M init Reasoning on unlabeled data in the target domain, generating pseudo labels, and filtering high-quality pseudo labels through confidence thresholds; Feature extraction: During the model training process, the YOLOv5 backbone network CSPDarknet53 is used to extract features from the input image. CSPDarknet53 generates category predictions, confidence predictions, and bounding box predictions for each grid cell through a multi-level feature extraction mechanism. Sample screening: The pseudo labels generated in step S2 are screened according to the confidence level to ensure the quality of the training data. The specific operation method of screening the pseudo labels is as follows: After feature extraction, the confidence prediction of each target is evaluated, and the target is divided into positive samples and negative samples according to the confidence prediction value; targets with high confidence and matching categories are judged as positive samples, and targets with low confidence or mismatched categories are judged as negative samples; Contrastive learning: Through intra-class and inter-class contrastive learning, the contrastive learning loss is calculated to optimize the consistency and discrimination of features; Comprehensive loss optimization: Pseudo-labels and 1% of the labeled data in the target domain are used as the true value to calculate the loss with the model prediction result. The loss function includes pseudo-label classification loss, bounding box regression loss and contrastive learning loss, and performs joint optimization to improve model performance.

2. The method for joint optimization of pseudo label generation and contrastive learning for target detection according to claim 1, characterized in that: The specific operation method of the initial training is as follows: Let the source domain labeled data be D src ={X i , Y i )}, where X i is the input image, Y i The annotation information corresponding to the input image, including target category, confidence and bounding box information; The initial model is obtained through training, which optimizes the source domain data through the standard YOLO loss function, including classification loss, confidence loss and bounding box regression loss; Through this process, a model with preliminary target detection capability is obtained, namely the preliminary detection model M init ; Provides a basis for subsequent pseudo-label generation and reasoning about unlabeled data in the target domain.

3. The method for joint optimization of pseudo label generation and contrastive learning for target detection according to claim 1, characterized in that: The specific operation method of generating pseudo labels is as follows: Let the target domain unlabeled data be D src ={X j }, where X j For source domain samples, use the preliminary detection model M init There is no labeled data D for the target domain src ={X j } Perform reasoning and generate preliminary pseudo labels Preliminary pseudo-labeling Including target category, confidence and bounding box information; when generating pseudo labels, set a confidence threshold c, filter out targets with confidence higher than the threshold c as valid pseudo labels, and remove low-quality pseudo labels lower than the threshold c; the filtered pseudo label set is: After conf(Y j ) activation function to map it to the interval [0, 1], Y j Predict labels for the model, conf(Y j ) is calculated as follows: In the above formula, e is the natural exponential constant.

4. The method for joint optimization of pseudo label generation and contrastive learning for target detection according to claim 1, characterized in that: The specific operation method of calculating the contrastive intra-class learning loss in the contrastive learning is as follows: Extract positive sample features: extract the corresponding features in the feature map according to the positive sample index and normalize them; Calculate the intra-class similarity matrix: use matrix multiplication to calculate the cosine similarity between positive samples; exclude the similarity between itself and itself through diagonal mask; Scaling similarity uses a temperature parameter to scale similarity, controlling the model's sensitivity to similarity differences. A smaller temperature parameter will amplify the gap between high and low similarities. The formula for calculating the intra-class contrastive learning loss for each positive sample is as follows: In the above formula, x i Represents the feature vector of the current target sample, the sample of interest in intra-class contrastive learning, the goal is to maximize its similarity with other positive samples of the same type, while minimizing its similarity with negative samples of other classes; x j Represents x i The feature vectors of other samples of the same type; the constraint of j≠i means x j and x i are different samples of the same type; x k Represents x i The feature vectors of samples of different categories, the constraint k≠i represents x k and x i are samples of different categories; N is the total number of samples, N p is the number of positive samples in the current category, that is, samples x i The total number of samples of the same category; N n is the number of positive samples in the current category, belonging to sample x i The total number of samples of different categories; sim(x i ,x j ) is the sample x i ,x j The similarity measure is usually cosine similarity, and its calculation formula is: τ is the temperature parameter, which is used to adjust the scaling of similarity and control the range of gradient variation; The intra-class contrastive learning loss is obtained, and the average of the losses of all positive samples is taken as the intra-class loss of the current category; in this process, the features of the positive samples are normalized, and the similarity between samples of the same category is calculated and optimized; by maximizing the similarity of samples within the class, the feature distribution differences between samples of the same category are narrowed.

5. The method for joint optimization of pseudo label generation and contrastive learning for target detection according to claim 1, characterized in that: The specific operation method of calculating the contrastive inter-class learning loss in the contrastive learning is as follows: Extract the positive sample features within each category and normalize each feature to the same unit vector; For each category of positive samples, calculate the center vector / mean vector of the positive samples in each category as the feature representative of the category; The loss objective is to minimize the similarity of the same category and maximize the similarity of different categories; The formula for calculating the inter-class contrastive learning loss for each positive sample is as follows: In the above formula, N is the total number of samples, usually the total number of samples in the current batch; C is the number of categories, that is, there are C categories in the classification task; is the feature vector of the nth sample in the ath class, indicating a sample belonging to class a; The feature vector of the sample belonging to the sth class represents a class sample different from the current class a; is the number of samples in class a; is the number of samples in the sth class; for The similarity between them is usually measured by cosine similarity, and its formula is 6. The method for joint optimization of pseudo label generation and contrastive learning for target detection according to claim 1, characterized in that: The loss function in the comprehensive loss optimization includes: Pseudo-label classification loss: used to measure the classification accuracy of the model on pseudo-labels, prompting the model to make accurate category predictions on target domain data; Bounding box regression loss: used to optimize the model's prediction of target position and size and improve the bounding box regression accuracy; Contrastive learning loss: It includes intra-class and inter-class contrast losses, which aims to optimize the cohesion and discrimination of sample features, thereby improving the representation ability of the model in the target domain.

7. The method for joint optimization of pseudo label generation and contrastive learning for target detection according to claim 1, characterized in that: The source domain data and target domain data are respectively Cityscape dataset and Sim10k dataset.

8. The method for joint optimization of pseudo label generation and contrastive learning for target detection according to claim 7, characterized in that: The Cityscapes dataset contains high-quality urban street scene images, with more than 5,000 high-resolution images, covering 20 categories such as vehicles, pedestrians, buildings, etc.; in order to align with the target domain, only vehicles and pedestrians are selected; the Sim10k dataset contains synthetic images from a virtual driving environment, with traffic scenes similar to Cityscapes; the images generated by this dataset differ in appearance from images in the real world, and there are some inter-domain deviations, such as weather, lighting, and viewing angle; the labels of the Sim10k dataset are similar to those of Cityscapes, including target categories of vehicles and pedestrians.

9. A device comprising a memory and a processor, characterized in that: The memory is used to store a computer program that can be run on the processor; The processor, when running a computer program, can execute the steps of a method for joint optimization of pseudo label generation and contrastive learning for target detection as described in any one of claims 1 to 8.

10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by at least one processor, it can implement the steps of a method for joint optimization of pseudo label generation and contrastive learning for target detection as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Prototype and confidence driven passive self-adaptive medical image classification system and method

    CN121190878A