Active learning method of cross-domain target detection model and sample screening method thereof

By employing an active learning approach in cross-domain object detection models, this method utilizes visual and textual feature similarity assessment of regions of interest to select samples and combines Mean-Teacher self-training. This addresses the issue of semantic domain differences in cross-domain object detection, thereby improving the model's performance and reliability in real-world scenarios.

CN120976677APending Publication Date: 2025-11-18SOUTHWESTERN UNIV OF FINANCE & ECONOMICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511116877.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing methods cannot effectively capture semantic domain differences in cross-domain object detection, leading to a decline in model performance in real-world scenarios. In particular, the reliability and applicability of object detection systems are affected by domain shifts caused by factors such as changes in lighting conditions.

Method used

We employ an active learning approach for cross-domain object detection models. By evaluating the similarity between visual and textual features of regions of interest, we select samples with the most certain scores for labeling. Combined with a Mean-Teacher self-training mechanism, we overcome the problem of false label quality and improve model performance.

Benefits of technology

It effectively captures domain differences at the object level, breaks through the limitations of single visual features, improves the model's semantic recognition ability in cross-domain scenarios, reduces manual annotation costs, and enhances the model's performance in the target domain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976677A_ABST
    Figure CN120976677A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computer vision, and provides an active learning method of a cross-domain target detection model and a sample screening method thereof. The sample screening method comprises the following steps: firstly, based on a model trained by a source domain, identifying a region of interest of a target domain image sample; then, on the basis of the visual features of the regions of interest of the target domain image sample and the text features of the classes, the similarity between the two classes of features is utilized, and the certainty score of the class classification of the regions of interest of the image sample is evaluated; and then according to the deterministic score, screening to obtain an active learning sample, so that the domain difference of an object level can be effectively captured, the limitation of a single visual feature is broken through, and the problem that an existing method cannot capture the domain difference of a semantic level is effectively solved. According to the active learning method, the sample screening method is integrated into a complete active learning framework, and the false label quality problem is effectively solved through a Mean-Teamer self-training mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computer vision, in particular to an active learning method of a cross-domain target detection model and a sample screening method thereof. BACKGROUND

[0002] As a core task in the field of computer vision, target detection plays a crucial role in many frontier fields such as autonomous driving, intelligent security, medical image analysis, and edge artificial intelligence. However, with the development, the demand for high-precision and high-reliability target detection systems is increasingly urgent, and at the same time, a common and thorny challenge has been exposed: domain shift problem.

[0003] Domain shift problem refers to the significant difference between the training environment and the deployment environment. This difference may be caused by changes in factors such as lighting conditions, weather changes, camera angles, or device performance. For example: the source domain data used for training comes from a laboratory with uniform lighting, while the target domain for actual deployment is a factory site with variable lighting. Domain shift problem often leads to a sharp decline in model performance, seriously affecting the reliability and applicability of the target detection system in actual scenarios.

[0004] To solve this problem, the most direct method is to obtain a large amount of high-quality labeled data in the target domain. However, data labeling requires a large amount of manpower and material resources, resulting in high costs, especially in fields that require professional knowledge, such as medical diagnosis or industrial detection. Therefore, how to minimize the cost of manual labeling while effectively improving the performance of the model in the target domain has become a key problem in the field of target detection that needs to be solved.

[0005] Active learning, English full name: Active Learning, abbreviated as AL, is a machine learning paradigm that selects the most valuable samples for labeling to maximize model performance with the least labeling cost. Its core idea is: let the model itself choose the most uncertain and most performance-improving samples. Due to the above characteristics of active learning, in solving the domain shift problem, it can maximize the performance improvement of the model under limited labeling budget, and has received widespread attention in recent years. For example, the uncertainty-based method selects the most uncertain samples predicted by the model for labeling.

[0006] The above uncertainty-based method is typically "clustering + uncertainty". Specifically, first, based on visual features, the target domain samples are divided into several clusters through clustering; then, in each cluster, the most "difficult" samples are selected through uncertainty measures such as entropy, confidence, etc. However, it should be noted that the above clustering operation processes the visual features of the entire image and cannot effectively capture the domain differences of local objects, so it can achieve significant results in image-level tasks such as image classification. However, the target detection task has the complex characteristics of multiple instances and multiple tasks, and needs to predict the position of the bounding box and the class label at the same time. When using the above method, there are certain limitations because it cannot capture the semantic domain differences, for example, the same object may have different visual appearances in different domains but maintain semantic consistency. SUMMARY

[0007] The technical problem to be solved by the present application is to provide an active learning method for a cross-domain target detection model and a sample screening method, aiming to solve the problem that existing methods cannot capture semantic domain differences.

[0008] The technical solution adopted by the present application to solve the above technical problem is:

[0009] A sample screening method for active learning of a cross-domain target detection model, comprising the following steps:

[0010] A1, based on the target detection model trained in the source domain, a source domain model is constructed;

[0011] Each image sample contained in the target domain dataset is input into the source domain model to identify each region of interest it contains; the identification of the region of interest includes labeling the bounding box of the region of interest; the image region framed by each bounding box constitutes the image block corresponding to each region of interest;

[0012] A2, using a visual feature extraction module, the visual features of the image blocks corresponding to each region of interest identified in step A1 are obtained; the visual feature extraction module includes a pre-trained visual encoder, and is obtained by fine-tuning using the regions of interest contained in the target domain as training samples;

[0013] Using a text feature extraction module, based on the class prompt text of each class contained in the target domain dataset, the text features of each class contained in the target domain dataset are obtained; the text feature extraction module includes a pre-trained text encoder; the class prompt text of each class is constructed based on its class name;

[0014] A3, for each image sample included in the target domain dataset, respectively, using the visual features of the image block corresponding to the interest region included in the image sample and the text features of each category included in the target domain dataset, based on the similarity between the two types of features, evaluating the certainty score of the category classification of the interest region included in the image sample;

[0015] A4, according to the certainty score of the category classification of the interest region included in each image sample included in the target domain dataset, filtering according to the preset condition to obtain the sample for active learning of the cross-domain target detection model.

[0016] Further, in step A3, the first fusion and then evaluation or the first evaluation and then fusion are adopted, and the certainty score of the category classification of the interest region included in the image sample is evaluated based on the similarity between the two types of features;

[0017] Wherein, the first fusion and then evaluation mode includes:

[0018] A31, fuse the visual features of the image block corresponding to the interest region included in the image sample to obtain the visual features of the image sample;

[0019] A32, using the visual features of the image sample, respectively, calculating the similarity between the visual features and the text features of each category included in the target domain dataset,

[0020] A33, based on the maximum similarity between the visual features of the image sample and the text features of each category, as the certainty score of the category classification of the interest region included in the image sample;

[0021] Wherein, the first evaluation and then fusion mode includes:

[0022] A31, respectively calculating the similarity between the visual features of the image block corresponding to each interest region included in the image sample and the text features of each category included in the target domain dataset;

[0023] A32, for each interest region included in the image sample, respectively, taking the maximum similarity between the visual features and the text features of each category as the certainty score of the category classification of the interest region;

[0024] A33, fuse the certainty scores of the category classification of each interest region included in the image sample as the certainty score of the category classification of the interest region included in the image sample.

[0025] Further, in step A3, the first fusion and then evaluation mode is adopted;

[0026] Wherein, in step A31, the visual features of the image block corresponding to the interest region included in the image sample are fused according to the following formula to obtain the visual features of the image sample;

[0027]

[0028] wherein, denotes the number of regions of interest contained in the i-th image sample of the target domain dataset, denotes the i-th region of interest of the i-th image sample of the target domain dataset, denotes the visual feature of the j-th region of interest of the i-th image sample of the target domain dataset, denotes the visual feature of the i-th image sample of the target domain dataset; Step A33, based on the maximum similarity between the visual feature of the i-th image sample and the text feature of each category, as the certainty score of the category classification of the region of interest contained in the i-th image sample

[0029] Step A33, based on the maximum similarity between the visual feature of the i-th image sample and the text feature of each category, as the certainty score of the category classification of the region of interest contained in the i-th image sample

[0030]

[0031] wherein, denotes the similarity between the visual feature of the i-th image sample of the target domain dataset and the enhanced text feature of the j-th category.

[0032] Further, the visual feature extraction module comprises a pre-trained visual encoder and a contrastive learning head.

[0033] The visual feature extraction module is used to extract visual features, including:

[0034] First, the pre-trained visual encoder is used to obtain the initial visual feature of the input image block.

[0035] Then, the initial visual feature of the input image block is projected to the contrastive space through the contrastive learning head to obtain the final visual feature of the input image block.

[0036] Further, the contrastive learning head comprises a first linear layer, a ReLU activation layer and a second linear layer in sequence.

[0037] Further, the text feature extraction module comprises a pre-trained text encoder, a learnable feature matrix and a fusion module, and is trained with the regions of interest contained in the target domain as training samples; the learnable feature matrix contains learnable embedding features of each category contained in the target domain dataset.

[0038] ​​​​​​​The text feature extraction module is used to obtain text features of each category based on category prompt texts of each category contained in the target domain dataset, including:

[0039] First, a pre-trained text encoder is used to extract text embedding features of each category contained in the target domain dataset based on the category prompt texts of each category contained in the target domain dataset.

[0040] Then, a fusion module is used to fuse the text embedding features and the learnable embedding features of each category to obtain the text features of each category.

[0041] Further, the category prompt text is constructed based on the category name of the corresponding category according to the template of the pre-trained text encoder.

[0042] Further, the pre-trained visual encoder and the pre-trained text encoder are the visual encoder and the text encoder of the pre-trained CLIP model respectively; the category prompt text is constructed based on the category name of the corresponding category according to the template "a photo of category name" of the text encoder of the pre-trained CLIP model.

[0043] Further, the training of the text feature extraction module is synchronized with the fine-tuning of the visual feature extraction module, including:

[0044] B1, construct a fine-tuning set and initialize learnable features; the training samples contained in the fine-tuning set are interest regions, including image blocks corresponding to the interest regions and pseudo labels of the interest regions; the interest regions of the training samples constituting the fine-tuning set are obtained from image samples contained in the target domain dataset by the source domain model recognition;

[0045] B2, use the visual feature extraction module to obtain the visual features of the image blocks corresponding to each interest region of the fine-tuning set;

[0046] B3, respectively, calculate the similarity between the visual features of the image blocks corresponding to each interest region of the fine-tuning set and the text features of each category of the target domain dataset;

[0047] B4, use the pseudo labels of each interest region of the fine-tuning set to calculate the contrast loss based on the similarity obtained in step B3 based on the label information; update the parameters of the visual feature extraction module and the parameters of the text feature extraction module except the pre-trained text encoder according to the obtained contrast loss;

[0048] B5, determine whether the training is completed, if yes, complete the training of the text feature extraction module and the fine-tuning of the visual feature extraction module; otherwise, return to step B2.

[0049] The similarity between the visual features and the text features is calculated as follows:

[0050] First, normalize the visual features and text features to be calculated according to the following formulas respectively:

[0051]

[0052]

[0053] wherein, denotes the visual feature to be calculated, denotes the text feature to be calculated, denotes the two-norm;

[0054] Then, calculate the similarity between the visual features and the text features to be calculated according to the following formula:

[0055]

[0056] wherein, is a temperature hyperparameter;

[0057] In step B1, the regions of interest of the training samples constituting the fine-tuning set are obtained by identifying the image samples contained in the target domain dataset by the source domain model, and the confidence is higher than the set confidence threshold;

[0058] In step B2, the text features of each category are obtained by fusing the text embedding features and the learnable embedding features of each category using the fusion module according to the following formula:

[0059]

[0060] wherein, denotes the text feature of the i-th category, denotes the learnable embedding feature of the i-th category, denotes the text embedding feature of the i-th category, is a learnable fusion coefficient;

[0061] In step B4, the contrastive loss based on label information is calculated based on the similarity obtained in step B3 using the pseudo labels of each region of interest in the fine-tuning set according to the following formula:

[0062]

[0063] wherein, is the number of samples contained in the fine-tuning set, denotes the visual feature of the i-th region of interest in the fine-tuning set, denotes the visual feature of the j-th region of interest in the fine-tuning set,​​​​ a similarity between the text features of the classes, a similarity between the text features of the class corresponding to the visual features of the image block in the fine-tuning set corresponding to the visual features of the image block in the fine-tuning set corresponding to the visual features of the image block in the fine-tuning set

[0064] In another aspect, the present application also provides an active learning method of a cross-domain object detection model, comprising the following steps:

[0065] S1, using a source domain dataset, performing supervised training to obtain an initial object detection model;

[0066] S2, taking the initial object detection model as a student model; copying the parameters of the student model to construct a teacher model;

[0067] S3, taking the teacher model as a source domain model, using the sample screening method of the cross-domain object detection model active learning as described above to obtain samples for cross-domain object detection model active learning;

[0068] S4, labeling the samples for cross-domain object detection model active learning obtained in step S3 to obtain a labeled sample set and an unlabeled sample set of a target domain; the labeled sample set , wherein, is a target domain dataset;

[0069] S5, based on the average teacher self-training framework, training according to the following steps:

[0070] S51, calculating the loss:

[0071] For each unlabeled sample contained in the unlabeled sample set , generate a first random augmented sample and a second random augmented sample; use the teacher model to identify the pseudo label and the bounding box of the region of interest of each first random augmented sample, and obtain the visual features of the region of interest of the first random augmented sample, and sample the visual features to a preset size as the first object features;

[0072] Based on the bounding box of the region of interest of the first random augmented sample identified by the teacher model, according to the correspondence between the first random augmented sample and the second random augmented sample, obtain the image block framed by each bounding box in the second random augmented sample; use the student model to obtain the visual features of each image block, and sample the visual features to a preset size as the second object features;

[0073] Use the labeled sample set , perform supervised training on the student model to obtain a supervised loss Based on the regions of interest identified by the teacher model in the first randomly augmented sample, regions of interest with confidence levels meeting a preset threshold are selected as samples for unsupervised training. The student model is then trained unsupervised using the image patches and pseudo-labels corresponding to the selected unsupervised training samples to obtain the unsupervised loss. The aforementioned supervised loss and unsupervised losses All of them use the same loss function, including bounding box regression loss and classification loss;

[0074] Using the features of the first and second objects, the consistency loss is calculated according to the following formula. :

[0075]

[0076] in, The number of regions of interest identified by the teacher model. In order to be with the first A set of interest regions containing objects of the same category, wherein the category of objects contained in the interest region is determined by pseudo-labels generated by the teacher model; Represents a set The number of regions of interest included. Represents a set The included first Areas of interest; Represents a set The included first The first object feature of a region of interest This indicates the first teacher model that has been identified. The first object feature of a region of interest This indicates the first student model that has been identified. Second object features of each region of interest; Temperature coefficient;

[0077] Calculate the total loss using the following formula. :

[0078]

[0079] in, These are weight hyperparameters;

[0080] S52. Update the parameters of the student model using the total loss calculated in step S51.

[0081] S53. Based on the student model updated in step S52, update the parameters smoothly using the EMA formula as follows:

[0082]

[0083] wherein, is a parameter of the teacher model before updating, is a parameter of the teacher model after updating, is a parameter of the student model after updating in step S52, is a smoothing coefficient;

[0084] S54, determining whether the training is completed, if yes, ending the training, otherwise, returning to step S51.

[0085] The beneficial effects of the present application are:

[0086] The sample screening method of the present application can effectively capture domain differences at the object level based on the interest region. Secondly, the introduction of text features breaks through the limitations of single visual features and can overcome the problem that only relying on visual features cannot accurately identify domain bias at the semantic level in cross-domain scenarios. For example, the same object may exhibit different visual appearances but maintain semantic consistency in different domains. Therefore, the present application effectively solves the problem that existing methods cannot capture domain differences at the semantic level.

[0087] Based on the sample screening method, the present application further provides an active learning method for a cross-domain target detection model, which integrates the above-mentioned sample screening method into a complete active learning framework and effectively overcomes the problem of pseudo-label quality through a Mean-Teacher self-training mechanism. BRIEF DESCRIPTION OF DRAWINGS

[0088] Figure 1 FIG. 1 is a flowchart of an active learning method for a cross-domain target detection model according to the present application. DETAILED DESCRIPTION

[0089] The present application aims to provide a sample screening method for active learning of a cross-domain target detection model, which solves the problem that existing methods cannot capture domain differences at the semantic level. Specifically, the method comprises the following steps:

[0090] A1, constructing a source domain model based on a target detection model that has completed source domain training;

[0091] Each image sample contained in the target domain dataset is input into the source domain model to identify each interest region contained therein. The identification of the interest region includes labeling the boundary box of the interest region. The image region framed by each boundary box constitutes an image block corresponding to each interest region.

[0092] A2, obtaining the visual features of the image blocks corresponding to each interest region identified in step A1 by using a visual feature extraction module. The visual feature extraction module comprises a pre-trained visual encoder, which is fine-tuned by taking the interest regions contained in the target domain as training samples.

[0093] The text feature extraction module is used to obtain text features of each category contained in the target domain dataset based on category prompt texts of each category contained in the target domain dataset. The text feature extraction module comprises a pre-trained text encoder. The category prompt texts of each category are respectively constructed based on category names thereof.

[0094] A3. For each image sample contained in the target domain dataset, the certainty score of the category classification of the interest region contained in the image sample is respectively evaluated based on the similarity between the visual features of the image block corresponding to the interest region contained in the image sample and the text features of each category contained in the target domain dataset.

[0095] A4. The samples for active learning of the cross-domain target detection model are obtained by screening according to a preset condition based on the certainty scores of the category classifications of the interest regions contained in each image sample contained in the target domain dataset.

[0096] In step A1, the source domain model is constructed using the target detection model that has completed source domain training, and then the interest region is recognized using the source domain model, so that in subsequent sample screening, feature extraction can be performed based on the interest region, i.e., at the object level, to adapt to the characteristics of multi-instance and multi-task target detection and effectively capture domain differences at the object level.

[0097] In step A2, the visual features of each interest region, i.e., the visual features at the object level, are obtained using the visual feature extraction module. The category prompt texts are constructed based on the category names, and the text features of each category are obtained using the text feature extraction module. By introducing text features based on visual features, a multi-modal framework is constructed, breaking through the limitations of single visual features.

[0098] After obtaining the features, deterministic evaluation can be performed based on the matching of the two types of features, and then sample screening can be performed. When screening image samples, the uncertainty of an image sample depends on the uncertainty of the interest region contained therein, i.e., the higher the classification uncertainty of the interest region, i.e., the object, contained in the image sample, the more uncertain the judgment of the model, and the greater the amount of potential information contained in the image sample, and the higher the value of the model performance improvement.

[0099] Therefore, in step A3, the certainty score of the category classification of the interest region contained in the image sample is respectively evaluated based on the similarity between the visual features of the image block corresponding to the interest region contained in the image sample and the text features of each category contained in the target domain dataset. In step A4, the samples for active learning of the cross-domain target detection model are obtained by screening according to a preset condition based on the certainty scores of the category classifications of the interest regions contained in each image sample contained in the target domain dataset.

[0100] In step A3, the calculation and evaluation can be performed in any manner as long as the evaluation of the image sample can be achieved based on the certainty of the category classification of the region of interest. As an optional solution, the cross-domain from the object level to the image level can be achieved through a fusion operation, and the fusion operation can be performed in a manner of fusion first and evaluation second or evaluation first and fusion second.

[0101] The manner of fusion first and evaluation second includes:

[0102] A31, fusing the visual features of the image blocks corresponding to the regions of interest included in the image sample to obtain the visual feature of the image sample;

[0103] A32, respectively calculating the similarity between the visual feature of the image sample and the text features of each category included in the target domain dataset,

[0104] A33, taking the maximum similarity between the visual feature of the image sample and the text features of each category as the certainty score of the category classification of the region of interest included in the image sample.

[0105] The manner of evaluation first and fusion second includes:

[0106] A31, respectively calculating the similarity between the visual features of the image blocks corresponding to each region of interest included in the image sample and the text features of each category included in the target domain dataset;

[0107] A32, respectively taking the maximum similarity between the visual feature of each region of interest included in the image sample and the text features of each category as the certainty score of the category classification of the region of interest;

[0108] A33, fusing the certainty scores of the category classifications of each region of interest included in the image sample as the certainty score of the category classification of the region of interest included in the image sample.

[0109] The fusion in the above manners can be extreme value, mean value, weight sum, or MLP, and the specific implementation can be selected according to the application scenario.

[0110] In subsequent embodiments, since the target domain is made by adding artificially synthesized fog on the basis of the source domain, the manner of fusion first and evaluation second is preferably adopted considering the following aspects:

[0111] First, the key difference of the cross-domain scene is reflected in the global feature rather than the local target attribute, in addition to the fog, such as light.

[0112] Secondly, if the pseudo label of a certain interest region is wrong, such as a false positive of the source domain model, the similarity score may be abnormally high or low, thus dominating the score of the entire image. When the sampling is evaluated and then fused, it will lead to misjudgment.

[0113] Thirdly, target detection contains multiple instances. If an image sample contains multiple same instances, such as multiple buses, because the features are relatively concentrated in the embedding space, the semantic information can still be preserved after fusion. In the calculation of the similarity score, the similarity of the bus class will be significantly higher than that of other classes, that is, the certainty score is higher. If an image sample contains multiple different instances, such as buses and people, the visual features of different interest regions form a mixed representation after fusion, which deviates from the text features of any single class, resulting in a decrease in the certainty score. In the target detection task, the image samples containing multiple different instances are most likely to be misjudged. Therefore, improving the probability of selection of this type of image sample is beneficial to reflecting the complexity of the scene and the uncertain area of the model.

[0114] Fourthly, on a large-scale data set, each interest region of each image sample will have cross-overlapping conditions. Redundant calculation will reduce the efficiency of active learning and waste computing resources.

[0115] In the present application, a visual feature extraction module including a pre-trained visual encoder is used. Although the pre-trained visual encoder is trained on a large-scale data, its general nature cannot fully capture the specific features in the domain-adaptive target detection task, especially when there is a large difference between the source domain and the target domain. At the same time, in the target detection task, the features that need to be focused on are the object-level, i.e. the interest region, rather than the entire image. Therefore, before feature extraction, the interest regions contained in the target domain are first used as training samples to fine-tune the visual feature extraction module including the pre-trained visual encoder, so that it can better adapt to and capture the specific features of the target domain and better adapt to the object-level feature extraction for the interest region.

[0116] In the present application, the class prompt text is constructed based on the class name. The class name can be directly used, or a fixed text combined with the class name can be used. However, considering that isolated class names are prone to semantic ambiguity and lack of scene description, leading to loose distribution of feature space, therefore, preferably, a fixed text combined with the class name is used. It is recommended that the class prompt text is constructed based on the class name of the corresponding class according to the template of the pre-trained text encoder. For example: the pre-trained visual encoder and the pre-trained text encoder are the visual encoder and the text encoder of the pre-trained CLIP model respectively; the class prompt text is constructed based on the class name of the corresponding class according to the template "a photo of class name" of the text encoder of the pre-trained CLIP model.

[0117] Meanwhile, considering that the fixed category prompt text cannot express domain-specific environmental conditions and context features, the same object may have different semantic representations in different domains in a cross-domain scenario, for example: "car" in an autonomous driving image set should have different semantic descriptions in sunny and foggy driving scenarios; meanwhile, when the domain offset is large, the pre-defined category prompt text cannot be effectively aligned with the visual features obtained by the fine-tuned visual encoder, resulting in inaccurate uncertainty evaluation. Therefore, further, the text feature extraction module includes a pre-trained text encoder, a learnable feature matrix and a fusion module, and the interest region contained in the target domain is taken as a training sample, and the learnable feature matrix is obtained after training. The learnable feature matrix contains the learnable embedding features of each category contained in the target domain dataset. Using the text feature extraction module, the text features of each category are obtained based on the category prompt text of each category contained in the target domain dataset, including:

[0118] Firstly, the pre-trained text encoder is used to extract the text embedding features of each category contained in the target domain dataset based on the category prompt text of each category contained in the target domain dataset;

[0119] Then, the fusion module is used to fuse the text embedding features and the learnable embedding features of each category to obtain the text features of each category.

[0120] On the other hand, in the cross-domain target detection task, the traditional Mean-Teacher relies on the teacher model to generate pseudo-labels, but when the domain difference is large, the prediction of the teacher model will produce high-noise pseudo-labels due to domain offset, resulting in incorrect alignment of the student model; similarly, the student model trained only with the source domain data has very weak representation ability for the target domain in the initial stage, and if it is directly trained relying on the pseudo-labels, it will fall into a vicious cycle of "error accumulation". Therefore, the present application also proposes an active learning method for a cross-domain target detection model, comprising the following steps:

[0121] S1, using the source domain dataset for supervised training to obtain an initial target detection model;

[0122] S2, taking the initial target detection model as a student model; copying the parameters of the student model to construct a teacher model;

[0123] S3, taking the teacher model as a source domain model, using the sample screening method of the present application for active learning of a cross-domain target detection model to obtain samples for active learning of the cross-domain target detection model;

[0124] S4, labeling the samples for active learning of the cross-domain target detection model obtained in step S3 to obtain a labeled sample set and an unlabeled sample set of the target domain; the ;

[0125] S5, training based on the average teacher self-training framework.

[0126] The method provides high-quality initial annotations for the model by preferentially annotating samples with the largest amount of information in the target domain, alleviates the misleading of training by pseudo-label noise, and thus breaks the cold start deadlock.

[0127] Further description is made below in combination with embodiments.

[0128] Embodiments

[0129] The present embodiment provides an active learning method for a cross-domain target detection model, as shown in Figure 1 , which comprises the following steps.

[0130] S1, source domain training

[0131] In this step of the present embodiment, the source domain dataset is used for supervised training to obtain an initial target detection model, and the process is the same as that of the existing method. The loss function is .

[0132]

[0133]

[0134] wherein, is a multi-class cross-entropy loss, is a bounding box regression loss, is the true value of the coordinates of the bounding box, is the mean value of the coordinates of the bounding box predicted by the model, is the variance of the coordinates of the bounding box predicted by the model.

[0135] S2, initialize student model and teacher model

[0136] In this step, the initial target detection model is used as the student model, and the parameters of the student model are copied to construct the teacher model.

[0137] S3, screen active learning samples

[0138] This step in the present embodiment comprises the following steps.

[0139] S31, identify the region of interest of the target domain

[0140] In this step, the teacher model is used as the source domain model.

[0141] The image samples included in the target domain dataset are respectively input into the source domain model to identify each interest region included in the image samples. The identification of the interest region includes labeling a bounding box of the interest region and a pseudo label. An image region framed by each bounding box constitutes an image block corresponding to the interest region.

[0142] S32, feature extraction

[0143] In this step, the visual feature extraction module is used to obtain the visual features of the image blocks corresponding to the interest regions identified in step S31. The text feature extraction module is used to obtain the text features of each category included in the target domain dataset based on the category prompt text of each category included in the target domain dataset.

[0144] In this embodiment, the visual feature extraction module includes a pre-trained visual encoder and a contrastive learning head, and is obtained by fine-tuning using the interest regions included in the target domain as training samples. The contrastive learning head is a projection head in contrastive learning, which is used to map the high-dimensional features extracted by the encoder to a lower-dimensional space that is more suitable for calculating the contrastive loss, which can significantly improve the accuracy of subsequent evaluation. In this embodiment, the contrastive learning head uses a nonlinear projection head, which includes a first linear layer, a ReLU activation layer, and a second linear layer in sequence.

[0145] Since the contrastive learning head is added, the visual features are extracted by the visual feature extraction module, including:

[0146] First, the pre-trained visual encoder is used to obtain the initial visual features of the input image block;

[0147] Then, the initial visual features of the input image block are projected to the contrastive space by the contrastive learning head to obtain the final visual features of the input image block.

[0148] In this embodiment, the text feature extraction module includes a pre-trained text encoder, a learnable feature matrix, and a fusion module, and is obtained by training using the interest regions included in the target domain as training samples. The learnable feature matrix contains learnable embedding features of each category included in the target domain dataset.

[0149] Since the learnable feature matrix and the fusion module are added, the text features of each category are obtained by the text feature extraction module based on the category prompt text of each category included in the target domain dataset, including:

[0150] First, the pre-trained text encoder is used to extract the text embedding features of each category included in the target domain dataset based on the category prompt text of each category included in the target domain dataset;

[0151] Then, using the fusion module, text embedding features and learnable embedding features of various categories are fused together to obtain text features of various categories.

[0152] The aforementioned fusion module can sample using weighted summation or MLP methods. In this embodiment, a weighted summation method is used. The fusion module fuses text embedding features and learnable embedding features of each category according to the following formula to obtain text features for each category:

[0153]

[0154] in, Indicates the first Text features of each category, Indicates the first Learnable embedding features for each category Indicates the first Text embedding features for each category, The fusion coefficient is a learnable coefficient.

[0155] The training of the visual feature extraction module and the text feature extraction module can be performed according to existing technologies and scenarios, either separately or simultaneously. To facilitate feature alignment, in this embodiment, the training of the text feature extraction module is performed synchronously with the fine-tuning of the visual feature extraction module, including:

[0156] B1. Data Preparation

[0157] In this step, a fine-tuning set is constructed, and learnable features are initialized. The training samples contained in the fine-tuning set are regions of interest, including image patches corresponding to the regions of interest and pseudo-labels for the regions of interest; the regions of interest constituting the training samples of the fine-tuning set are obtained by the source domain model from the image samples contained in the target domain dataset.

[0158] To ensure sample quality, a further requirement is that the regions of interest constituting the training samples of the fine-tuned set are obtained from image samples contained in the target domain dataset through source domain model identification, and their confidence level is higher than a set confidence threshold. In this embodiment, the confidence threshold is 0.85.

[0159] B2. Feature Extraction

[0160] In this step, the visual feature extraction module is used to obtain the visual features of the image patches corresponding to each region of interest in the fine-tuning set; the text feature extraction module is used to obtain the text features of each category based on the category prompt text contained in the target domain dataset.

[0161] B3. Calculating Similarity

[0162] In this step, the similarity between the visual features of the image patches corresponding to each region of interest in the fine-tuning set and the text features of each category in the target domain dataset is calculated separately according to the following formula. :

[0163]

[0164]

[0165]

[0166] in, Indicates fine-tuning set The Visual features of image patches corresponding to each region of interest. Indicates the first Text features of each category, Represents the L2 norm; In this embodiment, the temperature hyperparameter is... The value is 0.07.

[0167] B4. Parameter Update

[0168] In this step, firstly, using the pseudo-labels of each region of interest in the fine-tuning set, and based on the similarity obtained in step B3, the contrast loss based on label information is calculated according to the following formula. :

[0169]

[0170] in, The number of samples contained in the fine-tuning set. Indicates the first in the fine-tuning set The visual features of the image patch corresponding to the first region of interest are similar to those of the second region of interest. The similarity between text features of each category, Indicates the first in the fine-tuning set The visual features of each region of interest corresponding to an image patch and the category corresponding to its pseudo-label. The similarity between text features.

[0171] Then, based on the obtained contrast loss The parameters of the visual feature extraction module and the parameters of the text feature extraction module, excluding the pre-trained text encoder, are updated. The parameters of the pre-trained text encoder are frozen here, primarily due to the small number of class samples.

[0172] B5. Iterative Decision

[0173] It is determined whether the training is completed. If yes, the training of the text feature extraction module and the fine-tuning of the visual feature extraction module are completed. Otherwise, the step B2 is returned.

[0174] In this step, the category prompt text of each category is constructed based on the category name of the category. Specifically, in this embodiment, the category prompt text is constructed based on the category name of the corresponding category according to the template of the adopted pre-trained text encoder. Further, in this embodiment, the pre-trained visual encoder and the pre-trained text encoder are respectively the visual encoder and the text encoder of the pre-trained CLIP model; and the category prompt text is constructed based on the category name of the corresponding category according to the template "a photo of category name" of the text encoder of the pre-trained CLIP model.

[0175] S33, calculating the certainty score

[0176] In this step, for each image sample included in the target domain dataset, the certainty score of the category classification of the interest region included in the image sample is evaluated based on the similarity between the visual features of the image blocks corresponding to the interest region included in the image sample and the text features of each category included in the target domain dataset.

[0177] In this step of this embodiment, the following steps are included:

[0178] S331, the visual features of the image blocks corresponding to the interest region included in the image sample are fused by using the mean fusion method to obtain the visual features of the image sample according to the following formula:

[0179]

[0180] wherein, represents the number of interest regions included in the i-th image sample of the target domain dataset, represents the visual features of the j-th interest region corresponding to the i-th image sample of the target domain dataset, represents the visual features of the i-th image sample of the target domain dataset.

[0181] S332, the similarity between the visual features of the image sample and the text features of each category included in the target domain dataset is calculated respectively according to the following formula:

[0182]

[0183] ​​​​​​​

[0184]

[0185] wherein, denotes the fine set The visual features of the image block corresponding to the interest region of the denotes the text features of the is a temperature hyperparameter, in the embodiment, 0.07.

[0186] S333, according to the following formula, based on the maximum similarity between the visual features of the image sample and the text features of each category, as the certainty score of the category classification of the interest region contained in the

[0187]

[0188] wherein, denotes the similarity between the visual features of the image sample of the target domain dataset and the enhanced text features of the

[0189] S34, screening samples

[0190] In this step, according to the certainty score of the category classification of the interest region contained in each image sample contained in the target domain dataset, the samples for active learning of the cross-domain target detection model are obtained by screening according to the preset condition. In the embodiment, according to the certainty score of the category classification of each image sample contained in the target domain dataset, the image samples are sorted from small to large, and the top 5% of the image samples are selected as the samples for active learning of the cross-domain target detection model.

[0191] S4, labeling

[0192] In this step, the samples for active learning of the cross-domain target detection model obtained in step S3 are labeled to obtain a labeled sample set and an unlabeled sample set of the target domain; the wherein, is the target domain dataset.

[0193] S5, based on the average teacher self-training framework, the following steps are performed for training:

[0194] S51, calculate the loss:​​​​​

[0195] For unlabeled sample sets Each unlabeled sample is used to generate a first random augmented sample and a second random augmented sample. Using the teacher model, the pseudo-labels and bounding boxes of the regions of interest of each first random augmented sample are identified, and the visual features of the regions of interest of the first random augmented sample are obtained. The visual features are sampled to a preset size and used as the first object features.

[0196] Based on the bounding boxes of the regions of interest identified by the teacher model in the first random augmented sample, and according to the correspondence between the first and second random augmented samples, the image blocks selected by each bounding box in the second random augmented sample are obtained; using the student model, the visual features of each image block are obtained, and the visual features are sampled to a preset size as the second object features.

[0197] In this embodiment, the first randomly enhanced sample uses weak enhancement, such as slight cropping or flipping; the second randomly enhanced sample uses strong enhancement, such as color perturbation or cutout.

[0198] Using labeled sample sets Supervised training is performed on the student model to obtain supervised loss. .

[0199] Based on the regions of interest (ROIs) identified by the teacher model in the first randomly augmented sample, ROIs with confidence levels meeting a preset threshold are selected as unsupervised training samples. The student model is then trained unsupervised using the image patches and pseudo-labels corresponding to these selected unsupervised training samples to obtain the unsupervised loss. In this embodiment, the preset threshold is 0.85.

[0200] The aforementioned supervised loss and unsupervised loss All of them use the same loss function as in step S1, including bounding box regression loss and classification loss.

[0201] Using the features of the first and second objects, the consistency loss is calculated according to the following formula. :

[0202]

[0203] in, The number of regions of interest identified by the teacher model. In order to be with the first A set of interest regions containing objects of the same category, wherein the category of objects contained in the interest region is determined by pseudo-labels generated by the teacher model; Represents a set the number of included regions of interest, representing a set of first object features of the first region of interest included; representing a set of first object features of the second region of interest included, representing a set of first object features of the first region of interest identified by the teacher model, representing a set of second object features of the second region of interest identified by the student model; is a temperature coefficient, in the present embodiment,

[0204] takes a value of 0.07. The total loss is calculated according to the following formula

[0205]

[0206] wherein, is a weight hyperparameter, in the present embodiment, takes a value of 0.05.

[0207] S52, update the parameters of the student model using the total loss calculated in step S51;

[0208] S53, based on the updated student model in step S52, update the parameters by EMA smoothing according to the following formula:

[0209]

[0210] wherein, is the parameter of the teacher model before updating, is the parameter of the teacher model after updating, is the parameter of the student model after updating in step S52, is a smoothing coefficient, in the present embodiment, takes a value of 0.9996.

[0211] S54, determine whether the training is completed, if yes, end the training, otherwise, return to step S51.

[0212] To verify the effect of the present application, the inventors compared the scheme of the above-mentioned embodiment with existing methods. Among them, the scheme of the embodiment adopts Faster R-CNN in the Detectron2 model library as the target detection model, and its backbone network is VGG-16; as a comparison, the existing scheme includes PT and CMT, wherein the PT model, also known as Prob-T, is derived from Learning domain adaptive object detection with probabilistic teacher, which is an unsupervised cross-domain object detection framework (UDA-OD), and its core contribution is to explicitly model the uncertainty of pseudo labels into the teacher-student self-training process; the CMT model is derived from Contrastive Mean Teacher for DomainAdaptive Object Detectors, which seamlessly integrates Mean-Teacher self-training and target-level contrastive learning to solve the two major pain points of UDA-OD, i.e., large pseudo label noise and domain feature drift.

[0213] Experiment 1

[0214] The Cityscapes dataset is used as the source domain dataset, and the Foggy Cityscapes dataset is used as the target domain dataset.

[0215] Among them, the Cityscapes dataset includes 2975 training pictures with their corresponding bounding boxes and class labels. The Foggy Cityscapes dataset is made by adding artificially synthesized fog to the Cityscapes dataset, and there is a domain shift between the source domain data.

[0216] According to different visibility ranges, the Foggy Cityscapes dataset simulates three fog densities of 0.02, 0.01 and 0.005, which are the attenuation coefficients β in the atmospheric scattering model. In the experiment, the most challenging 0.02 bin and all bins are used. All bins, i.e. ALL in Table 1, include a complete dataset of three fog concentration levels.

[0217] The results of the present experiment are shown in Table 1, which use AP and mAP to represent the performance of the model. The specific implementation of the indicators is as follows:

[0218] First, the threshold of the intersection over union (IoU) is set to 0.5 according to the PASCAL VOC standard, and if the IoU of the predicted box and the real box is ≥0.5, it is considered as correct detection, and then the following is obtained:

[0219] TP: True Positive, i.e. correctly predicted positive samples, whose IoU is greater than or equal to the threshold;

[0220] FP: False Positive, i.e. incorrectly predicted positive samples, whose IoU is less than the threshold;

[0221] FN: False Negative, i.e. undetected real targets.

[0222] Further, the following indicators are obtained:

[0223] 1. Precision: TP / (TP + FP), which represents the proportion of true positive samples among all predicted positive samples.

[0224] 2. Recall: TP / (TP + FN) = TP / total number of real positive samples, which represents the proportion of correctly detected real positive samples among all real positive samples.

[0225] Then, by integrating the Precision-Recall curve, AP is calculated to measure the detection accuracy of a single class. The average mAP of all classes is calculated to reflect the overall detection performance.

[0226] In Table 1, "Type" is the implementation type of the method: MT represents the Mean Teacher paradigm, and AL represents the Active Learning paradigm. In Table 1, "Source" refers to the model trained on the source domain without target domain adaptation; "Oracle" refers to the model trained on the target domain using real labels, representing the theoretical upper limit performance.

[0227] Experiment Two

[0228] In this experiment, the same performance indicators as in Experiment One are used, with the difference being the data set and the comparison scheme.

[0229] In this experiment, the KITTI dataset is used as the source domain dataset, and the Cityscapes dataset is used as the target domain dataset.

[0230] The KITTI dataset is another street view dataset different from the Cityscapes dataset, and its data is collected using cameras in cities. In the experiment, only the car category common to KITTI and Cityscapes is considered.

[0231] In contrast, in addition to PT and CMT, MeGA-CDA, TIA and SIGMA are also included. Among them, MeGA-CDA, derived from Memory guided attention for category-aware unsupervised domain adaptive object detection, proposes a "category-aware memory-guided attention" framework to solve the core pain point of "class-agnostic alignment → negative transfer" in traditional UDA-OD; TIA, derived from Task-specific inconsistency alignment for domain adaptive object detection, first proposes a UDA-OD framework that aligns classification and localization tasks separately in task space rather than feature space; SIGMA, derived from Semantic-complete graph matching for domain adaptive object detection, proposes to align the class conditional distribution between different domains through graph matching technology, fill in the gap of semantic mismatch, and achieve fine-grained domain adaptation.

[0232] The test results are shown in Table 2.

[0233] Table 1, experimental results of Experiment 1

[0234]

[0235] Table 2, experimental results of Experiment 2

[0236]

[0237] Finally, it should be noted that the above embodiments are only preferred embodiments and do not limit the present application. It should be noted that for those skilled in the art, without departing from the purpose of the present application and the scope of the claims, several modifications, equivalent replacements, improvements, etc. can be made, which should be included in the protection scope of the present application.

Claims

1. A sample selection method for active learning of a cross-domain target detection model, characterized in that, Includes the following steps: A1. Construct a source domain model based on the target detection model that has been trained in the source domain; Each image sample contained in the target domain dataset is input into the source domain model to identify each region of interest contained therein; the identification of the region of interest includes annotating the bounding boxes of the region of interest; the image regions selected by each bounding box constitute the image blocks corresponding to each region of interest. A2. Using the visual feature extraction module, obtain the visual features of the image blocks corresponding to each region of interest identified in step A1; the visual feature extraction module includes a pre-trained visual encoder, which is fine-tuned using the regions of interest contained in the target domain as training samples. Using a text feature extraction module, text features for each category of the target domain dataset are obtained based on the category prompt texts for each category. The text feature extraction module includes a pre-trained text encoder. The category prompt texts for each category are constructed based on their respective category names. A3. For each image sample contained in the target domain dataset, respectively, using the visual features of the image patch corresponding to the region of interest contained in the image sample and the text features of each category contained in the target domain dataset, evaluate the deterministic score of the category classification of the region of interest contained in the image sample based on the similarity between the two types of features. A4. Based on the deterministic scores of the category classification of the regions of interest contained in each image sample in the target domain dataset, samples for active learning by the cross-domain object detection model are obtained by filtering according to preset conditions.

2. The sample selection method for active learning of a cross-domain target detection model as described in claim 1, characterized in that: In step A3, either the method of fusing first and then evaluating or evaluating first and then fusing is adopted. Based on the similarity between the two types of features, the deterministic score of the category classification of the interest region contained in the image sample is evaluated. The approach of integrating first and then evaluating includes: A31. By fusing the visual features of the image blocks corresponding to the regions of interest contained in the image sample, the visual features of the image sample are obtained. A32. Using the visual features of the image sample, calculate the similarity between it and the text features of each category contained in the target domain dataset, respectively. A33. The maximum similarity between the visual features of the image sample and the text features of each category is used as the deterministic score for the category classification of the region of interest contained in the image sample. The approach of evaluating first and then integrating includes: A31. Calculate the visual features of the image patches corresponding to each region of interest contained in the image sample and the similarity between them and the text features of each category contained in the target domain dataset; A32. For each region of interest contained in the image sample, the maximum similarity between its visual features and the text features of each category is used as the deterministic score for the category classification of the region of interest. A33. The deterministic scores of the category classifications of each region of interest contained in the image sample are fused together and used as the deterministic scores of the category classifications of the regions of interest contained in the image sample.

3. The sample selection method for active learning of a cross-domain target detection model as described in claim 2, characterized in that: Step A3 adopts a method of fusion followed by evaluation; In step A31, the visual features of the image sample are obtained by fusing the visual features of the image blocks corresponding to the region of interest contained in the image sample according to the following formula. in, Represents the first of the target domain datasets The number of regions of interest contained in a single image sample Represents the first of the target domain datasets The first image sample Each region of interest corresponds to a visual feature of an image patch. Represents the first of the target domain datasets Visual features of an image sample; In step A33, based on the following formula, the first... The highest similarity between the visual features of an image sample and the text features of each category is used as the first... Deterministic score of the category classification of the region of interest contained in each image sample ; in, Represents the first of the target domain datasets The visual features of the first image sample and the second Similarity between enhanced text features of each category.

4. The sample selection method for active learning of a cross-domain target detection model as described in claim 1, characterized in that: The visual feature extraction module includes a pre-trained visual encoder and a contrastive learning head; The visual feature extraction module is used to extract visual features, including: First, the initial visual features of the input image patch are obtained using a pre-trained visual encoder; Then, the initial visual features of the input image patch are projected into the contrast space using a contrast learning head to obtain the final visual features of the input image patch.

5. The sample selection method for active learning of a cross-domain target detection model as described in claim 4, characterized in that: The contrastive learning head comprises, in sequence, a first linear layer, a ReLU activation layer, and a second linear layer.

6. A sample selection method for active learning of a cross-domain target detection model as described in any one of claims 1 to 5, characterized in that: The text feature extraction module includes a pre-trained text encoder, a learnable feature matrix, and a fusion module, and is trained using regions of interest contained in the target domain as training samples; the learnable feature matrix contains learnable embedding features of various categories contained in the target domain dataset. Using the text feature extraction module, based on the category prompt texts of each category contained in the target domain dataset, text features for each category are obtained, including: First, using a pre-trained text encoder, text embedding features for each category of the target domain dataset are extracted based on the category prompt texts for each category contained in the target domain dataset. Then, using the fusion module, text embedding features and learnable embedding features of various categories are fused together to obtain text features of various categories.

7. The sample selection method for active learning of a cross-domain target detection model as described in claim 6, characterized in that: The category prompt text is constructed based on the category name of the corresponding category and the template of the pre-trained text encoder used.

8. The sample selection method for active learning of a cross-domain target detection model as described in claim 7, characterized in that: The pre-trained visual encoder and pre-trained text encoder are the visual encoder and text encoder of the pre-trained CLIP model, respectively; the category prompt text is constructed based on the category name of the corresponding category and according to the template "a photo of category name" of the text encoder of the pre-trained CLIP model.

9. The sample selection method for active learning of a cross-domain target detection model as described in claim 6, characterized in that: The training of the text feature extraction module is performed synchronously with the fine-tuning of the visual feature extraction module, including: B1. Construct a fine-tuning set and initialize learnable features; the training samples contained in the fine-tuning set are regions of interest, including image patches corresponding to the regions of interest and pseudo-labels of the regions of interest; the regions of interest constituting the training samples of the fine-tuning set are obtained by the source domain model from the image samples contained in the target domain dataset. B2. Using the visual feature extraction module, obtain the visual features of the image patches corresponding to each region of interest in the fine-tuning set; B3. Specifically, calculate the similarity between the visual features of the image patches corresponding to each region of interest in the fine-tuning set and the text features of each category in the target domain dataset; B4. Using the pseudo-labels of each region of interest in the fine-tuning set, calculate the contrast loss based on the label information based on the similarity obtained in step B3; update the parameters of the visual feature extraction module and the parameters of the text feature extraction module except for the pre-trained text encoder according to the obtained contrast loss. B5. Determine whether training is complete. If yes, complete the training of the text feature extraction module and the fine-tuning of the visual feature extraction module; otherwise, return to step B2.

10. The sample selection method for active learning of a cross-domain target detection model as described in claim 9, characterized in that: Calculate the similarity between visual features and text features using the following steps: First, normalize the visual features and text features to be computed using the following formulas: in, Represents the visual features to be calculated. Represents the text features to be calculated. Represents the L2 norm; Then, the similarity between the visual features and text features to be calculated is obtained using the following formula: in, This refers to temperature hyperparameters. In step B1, the regions of interest constituting the training samples of the fine-tuning set are obtained by the source domain model from the image samples contained in the target domain dataset, and the confidence level is higher than the set confidence threshold. In step B2, the fusion module is used to fuse text embedding features and learnable embedding features of each category according to the following formula to obtain text features of each category: in, Indicates the first Text features of each category, Indicates the first Learnable embedding features for each category, Indicates the first Text embedding features for each category, The learnable fusion coefficient; In step B4, using the pseudo-labels of each region of interest in the fine-tuning set, and based on the similarity obtained in step B3, the contrast loss based on label information is calculated according to the following formula. : in, The number of samples contained in the fine-tuning set. Indicates the first in the fine-tuning set The visual features of the image patch corresponding to the first region of interest are similar to those of the second region of interest. The similarity between text features of each category, Indicates the first in the fine-tuning set The visual features of each region of interest corresponding to an image patch and the category corresponding to its pseudo-label. The similarity between text features.

11. An active learning method for a cross-domain target detection model, characterized in that, Includes the following steps: S1. Using the source domain dataset, perform supervised training to obtain the initial object detection model; S2. Use the initial object detection model as the student model; copy the parameters of the student model to construct the teacher model; S3. Using the teacher model as the source domain model, and employing a sample selection method for active learning of cross-domain target detection model as described in any one of claims 1 to 10, obtain samples for active learning of cross-domain target detection model. S4. Label the samples actively learned by the cross-domain object detection model obtained in step S3 to obtain a set of labeled samples in the target domain. and unlabeled sample set The ,in, For the target domain dataset; S5. Based on the average teacher self-training framework, train according to the following steps: S51. Calculate the loss: For unlabeled sample sets Each unlabeled sample is used to generate a first random augmented sample and a second random augmented sample. Using the teacher model, the pseudo-labels and bounding boxes of the regions of interest of each first random augmented sample are identified, and the visual features of the regions of interest of the first random augmented sample are obtained. The visual features are sampled to a preset size and used as the first object features. Based on the bounding boxes of the regions of interest identified by the teacher model in the first random augmented sample, and according to the correspondence between the first and second random augmented samples, the image blocks selected by each bounding box in the second random augmented sample are obtained; using the student model, the visual features of each image block are obtained, and the visual features are sampled to a preset size as the second object features; Using labeled sample sets Supervised training is performed on the student model to obtain supervised loss. Based on the regions of interest identified by the teacher model in the first randomly augmented sample, regions of interest with confidence levels meeting a preset threshold are selected as samples for unsupervised training. The student model is then trained unsupervised using the image patches and their pseudo-labels corresponding to the selected unsupervised training samples to obtain the unsupervised loss. The aforementioned supervised loss and unsupervised loss All of them use the same loss function, including bounding box regression loss and classification loss; Using the features of the first and second objects, the consistency loss is calculated according to the following formula. : in, The number of regions of interest identified by the teacher model. In order to be with the first A set of interest regions containing objects of the same category, wherein the category of objects contained in the interest region is determined by pseudo-labels generated by the teacher model; Represents a set The number of regions of interest included. Represents a set The included first Areas of interest; Represents a set The included first The first object feature of a region of interest This indicates the first teacher model that has been identified. The first object feature of a region of interest This indicates the first student model that has been identified. Second object features of each region of interest; Temperature coefficient; Calculate the total loss using the following formula. : in, These are weight hyperparameters; S52. Update the parameters of the student model using the total loss calculated in step S51. S53. Based on the student model updated in step S52, update the parameters smoothly using the EMA formula as follows: in, The parameters of the teacher model before the update. For the parameters of the updated teacher model, The parameters of the student model after step S52 are updated. For smoothing coefficients; S54. Determine whether training is complete. If yes, end training; otherwise, return to step S51.