Cross-Domain Object Detection Method and System Based on Semi-Supervised Learning in Extreme Weather

By distilling the bounding box regression knowledge of the teacher model to the student model in extreme weather, and using the diffusion model for feature adversarial training, and fine-tuning the student model with dynamic category pseudo-label screening strategy, the problems of insufficient data sets and low pseudo-label quality in object detection technology in extreme weather are solved, significantly improving detection performance and maintaining the stability of inference speed.

CN119863615BActive Publication Date: 2025-06-17SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510345865.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-17
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

The lack of sufficient labeled data sets in the prior art under extreme weather conditions has led to the inability to achieve targeted training, and the single filtering strategy based on classification confidence in semi-supervised learning cannot effectively filter out high-quality pseudo-labels.

Method used

By distilling the bounding box regression knowledge of the trained teacher model into the student model, a pre-trained detection model with regression ability in extreme weather is obtained. Then, intermediate images are generated using the diffusion model for instance-level feature adversarial training, and the feature difference between the source domain and the target domain is narrowed. At the same time, the dynamic category pseudo-label filtering strategy is used to filter the pseudo-label, and the selected high-quality pseudo-labels are used to fine-tune the student model.

Benefits of technology

It effectively solves the problems of insufficient data sets and low pseudo-label quality in object detection technology in extreme weather, significantly improves the performance of the detection model in extreme weather, reduces missed and missed detection phenomena, and maintains the stability of the inference speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119863615B_ABST
    Figure CN119863615B_ABST
Patent Text Reader

Abstract

The cross-domain object detection method and system based on semi-supervised learning under extreme weather provided by the present invention, the method comprising: enabling the student model to learn the bounding box regression knowledge of the trained teacher model through distillation to have the regression ability under extreme weather, using the intermediate state images generated by the diffusion model that are closer to the target domain and have less noise, performing instance-level feature adversarial training on the pre-trained detection model with the regression ability under extreme weather obtained in the distillation stage, determining a new teacher model and a new student model based on the detection model after adversarial training, in the fine-tuning stage, using a dynamic class pseudo-label screening strategy to screen the pseudo-labels generated by the new teacher model, and training the new student model with the screened target domain pseudo-labels to obtain an object detection model for detecting images to be detected under extreme weather, which can reduce the missed detection and misdetection phenomena in the extreme weather scenario of the prior art and maintain the original inference speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and particularly to a cross-domain target detection method and system based on semi-supervised learning under extreme weather conditions. Background Art

[0002] At present, the research on target detection under extreme weather conditions is of great significance, which is mainly manifested in the following aspects: (1) It can improve safety: The highly generalizable target detection technology can help vehicles and other means of transportation better identify obstacles and pedestrians on the road under low visibility conditions, thus significantly improving traffic safety; (2) It contributes to the development of autonomous driving technology: Autonomous vehicles need to accurately identify and locate surrounding objects under various weather conditions. Target detection in bad weather is an important challenge in the development of autonomous driving technology, and its research is crucial for the realization of an all-weather autonomous driving system; (3) It will assist in disaster emergency management: Extreme weather may cause traffic delays and accidents, especially at key traffic nodes such as airports and seaports.

[0003] However, compared with conventional target detection, there is relatively little work on target detection under extreme weather conditions, and there is a lack of a sufficient labeled extreme weather dataset. And due to the lack of a corresponding dataset, images are usually de-rained or de-fogged (off line), and then the preprocessed images are input into a detection model, such as the IA-YOLO detection model. In this detection model, a differentiable image processing module and a learnable small neural network are used to adaptively repair the images, and then the images are input into the YOLOv3 detection network. Although this detection method can improve the overall quality of the images, it may also lose semantic information important for detection, resulting in a large number of missed detections. On this basis, a symbiotic model of repair and detection is also built, and the feature information extracted by a unified backbone network is used for repair and detection work. For example, Together-Net uses YOLOX as the detection model, and a symbiotic repair network is built on its backbone network to perform repair and detection tasks simultaneously. However, limited by the network complexity, it is difficult to achieve a balance between the two tasks of repair and detection. Another example is that the mean teacher model is a semi-supervised learning method widely used in cross-domain target detection tasks. It can be fully trained specifically on the target domain. Although it can improve the quality of pseudo-labels to a certain extent, in scenes with heavy rain or thick fog, the quality of pseudo-labels is still relatively low, and a single filtering strategy based on classification confidence cannot screen out high-quality pseudo-labels.

[0004] In view of this, providing a solution to solve the above technical problems has already been an urgent concern for those skilled in the art. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a cross-domain object detection method and system based on semi-supervised learning under extreme weather conditions in view of the above-mentioned defects of the prior art, aiming to solve the problems that the existing object detection technology under extreme weather cannot achieve targeted training due to the lack of a sufficient number of labeled extreme weather datasets, and that the single filtering strategy based on classification confidence in semi-supervised learning cannot screen out high-quality pseudo-labels.

[0006] The technical solution adopted by the present invention to solve the technical problem is as follows:

[0007] A cross-domain object detection method based on semi-supervised learning under extreme weather conditions, wherein the method includes:

[0008] Distill and transfer the bounding box regression knowledge of the trained teacher model to the student model to obtain a pre-trained detection model with regression ability under extreme weather conditions; wherein, the student model is a detection model with detection ability under normal weather conditions, and the trained teacher model is a model trained by using a preset number of extreme weather images with labels to a detection model superior to the student model in terms of bounding box regression ability;

[0009] Input the obtained original source domain images under normal weather conditions and original target domain images under extreme weather conditions into the diffusion model respectively, and input the intermediate images between the source domain and the target domain output by the diffusion model into the pre-trained detection model for instance-level feature adversarial training to obtain an adversarial trained detection model;

[0010] Determine a new teacher model and a new student model based on the adversarial trained detection model, and use a preset dynamic class pseudo-label screening method to screen the pseudo-labels generated by the new teacher model based on the original target domain images, and use the screened target domain pseudo-labels to fine-tune the new student model to obtain a trained target detection model;

[0011] Use the trained target detection model to perform object detection on the images to be detected under extreme weather conditions to obtain corresponding detection results.

[0012] In one implementation, the distilling and transferring the bounding box regression knowledge of the trained teacher model to the student model to obtain a pre-trained detection model with regression ability under extreme weather conditions includes:

[0013] Use the progressive knowledge distillation strategy to distill and transfer the bounding box regression knowledge of the trained teacher model to the student model to obtain a pre-trained detection model with regression ability under extreme weather conditions;

[0014] Among them, the progressive knowledge distillation strategy is a distillation strategy that gradually increases the weight of the regression distillation loss according to a preset loss weight adjustment strategy during the knowledge distillation process, so that the student model gradually has the regression ability in extreme weather, and the initial weight of the regression distillation loss is not greater than the preset weight value.

[0015] In one implementation, the calculation formula of the regression distillation loss is:

[0016] ;

[0017] Among them, represents the regression distillation loss, β represents the weight of the regression distillation loss, N represents the total number of samples, represents the regression output of the trained teacher model for the i-th sample, represents the regression output of the student model for the i-th sample, represents the mean square error between the regression output of the trained teacher model for the i-th sample and the regression output of the student model for the i-th sample.

[0018] In one implementation, the teacher model is the RT-DETR model, and the student model is the YOLOv11 model.

[0019] In one implementation, the diffusion model is a denoising probability diffusion model, and the diffusion model generates the intermediate state image by adding noise and gradually denoising, and adds content loss and style loss during the denoising process of the diffusion model;

[0020] Among them, the content loss is calculated by comparing the features of the intermediate state image and the original target domain image or the original source domain image on the convolutional layer, and the style loss is calculated by comparing the Gram matrix of the intermediate state image and the original source domain image or the original target domain image.

[0021] In one implementation, the cross-domain object detection method based on semi-supervised learning in extreme weather further includes:

[0022] During the feature adversarial training, extract the target box area from the intermediate state image, based on the mask of the target box area, and perform instance-level feature alignment through the confrontation between the generator and the discriminator;

[0023] Among them, the discriminator is used to judge whether the features of the target box areas corresponding to the source domain and the target domain are distinguishable, and the generator is used to minimize the judgment ability of the discriminator.

[0024] In one implementation, in the process of fine-tuning and training the new student model with the selected target domain pseudo-labels to obtain a trained object detection model, it further includes:

[0025] Obtaining the parameters of the new student model in real time for each training round;

[0026] Based on a preset adaptive double exponential moving average strategy, and dynamically updating the parameters of the new teacher model by using the parameters of the new student model in each training round;

[0027] Wherein, the decay rate in the preset adaptive double exponential moving average strategy is dynamically adjusted based on the training round to dynamically control the parameter update speed of the new teacher model;

[0028] And, the calculation formula of the decay rate is:

[0029] ;

[0030] Wherein, represents the current training round number, represents the decay rate at the current training round number, represents the total number of training rounds, represents the minimum value of the decay rate, represents the maximum value of the decay rate.

[0031] In one implementation, the use of a preset dynamic class pseudo-label screening method to screen the pseudo-labels generated by the new teacher model based on the original target domain images includes:

[0032] Performing class statistics and distribution statistics on the original target domain images to obtain corresponding statistical results, and performing targeted data augmentation processing on the original target domain images based on the statistical results to obtain new target domain images;

[0033] Using a preset dynamic class pseudo-label screening method to screen the pseudo-labels generated by the new teacher model based on the original target domain images and the new target domain images.

[0034] In one implementation, the preset dynamic class pseudo-label screening method is:

[0035] Respectively comparing the classification confidence scores corresponding to when the pseudo-labels indicate that the objects to be detected belong to the target class with a first threshold and a second threshold corresponding to the target class to obtain corresponding comparison results; wherein, the first threshold is not less than the second threshold, and during the fine-tuning training process, the first threshold and the second threshold are dynamically adjusted according to a preset threshold dynamic adjustment strategy;

[0036] When the comparison result indicates that the classification confidence score is higher than the first threshold, the pseudo-label is determined as a positive sample to obtain the target domain pseudo-label;

[0037] When the comparison result indicates that the classification confidence score is lower than the second threshold, the pseudo-label is determined as a negative sample, and the negative sample is regularized to obtain a regularized negative sample for fine-tuning training of the new student model;

[0038] When the comparison result indicates that the classification confidence score is between the first threshold and the second threshold, it is determined whether the pseudo-label is consistent with the pseudo-label predicted by a preset large model;

[0039] When the pseudo-label is consistent with the pseudo-label predicted by the preset large model, the pseudo-label is determined as a positive sample to obtain the target domain pseudo-label;

[0040] When the pseudo-label is inconsistent with the pseudo-label predicted by the preset large model, the pseudo-label is determined as a negative sample, and the negative sample is regularized to obtain a regularized negative sample for fine-tuning training of the new student model.

[0041] The present invention also discloses a cross-domain object detection system based on semi-supervised learning under extreme weather, wherein the system includes:

[0042] A model distillation module for distilling and transmitting the bounding box regression knowledge possessed by a trained teacher model to a student model to obtain a pre-trained detection model with regression ability under extreme weather; wherein, the student model is a detection model with detection ability under normal weather, and the trained teacher model is a model obtained by training a detection model with better bounding box regression ability than the student model using a preset number of extreme weather images with labels;

[0043] An adversarial training module for respectively inputting the obtained original source domain images under normal weather and original target domain images under extreme weather into a diffusion model, and inputting the intermediate state images between the source domain and the target domain output by the diffusion model into the pre-trained detection model for instance-level feature adversarial training to obtain an adversarially trained detection model;

[0044] A model fine-tuning module for determining a new teacher model and a new student model based on the adversarially trained detection model, and using a preset dynamic class pseudo-label screening method to screen the pseudo-labels generated by the new teacher model based on the original target domain images, and fine-tuning the new student model using the screened target domain pseudo-labels to obtain a trained target detection model;

[0045] A target detection module, configured to perform target detection on an image to be detected under extreme weather by using the trained target detection model to obtain corresponding detection results.

[0046] The present invention also discloses a terminal, which includes: a memory, a processor, and a cross-domain target detection program based on semi-supervised learning under extreme weather stored in the memory and executable on the processor. When the cross-domain target detection program based on semi-supervised learning under extreme weather is executed by the processor, the steps of the cross-domain target detection method based on semi-supervised learning under extreme weather as described above are implemented.

[0047] The present invention also discloses a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the cross-domain target detection method based on semi-supervised learning under extreme weather as described above.

[0048] The cross - domain object detection method and system based on semi - supervised learning under extreme weather provided by the present invention. The cross - domain object detection method based on semi - supervised learning under extreme weather includes: distilling and transferring the bounding box regression knowledge of the trained teacher model to the student model to obtain a pre - trained detection model with regression ability under extreme weather. Among them, the student model is a detection model with detection ability under normal weather, and the trained teacher model is a model obtained by training a detection model with better bounding box regression ability than the student model using a preset number of extreme weather images with labels. Input the obtained original source - domain images under normal weather and original target - domain images under extreme weather into the diffusion model respectively, and input the intermediate - state images between the source domain and the target domain output by the diffusion model into the pre - trained detection model for instance - level feature adversarial training to obtain an adversarially trained detection model. Determine a new teacher model and a new student model based on the adversarially trained detection model, and use a preset dynamic class pseudo - label screening method to screen the pseudo - labels generated by the new teacher model based on the original target - domain images, and use the screened target - domain pseudo - labels to fine - tune and train the new student model to obtain a trained target detection model. Use the trained target detection model to perform object detection on the images to be detected under extreme weather to obtain corresponding detection results. Thus, it can be seen that the present invention enables the student model to have the regression ability of the trained teacher model under extreme weather through distillation. Then, the diffusion model is used to generate intermediate - state images that are closer to the target domain and have less noise, and the intermediate - state images are used to perform instance - level feature adversarial training on the pre - trained detection model that has learned the bounding box regression knowledge of the trained teacher model and has the regression ability under extreme weather to reduce the feature difference between the source domain and the target domain to obtain an adversarially trained detection model. Then, based on the adversarially trained detection model, a new teacher model and a new student model are determined. When the new teacher model generates pseudo - labels in the follow - up, the quality of the pseudo - labels generated by the new teacher model can be improved. Then, in the fine - tuning stage, the dynamic class pseudo - label screening strategy is used to screen the pseudo - labels generated by the new teacher model, and then the screened high - quality target - domain pseudo - labels are used to train the new student model to obtain a trained target detection model for object detection of the images to be detected under extreme weather, thereby being able to significantly reduce the missed detection and misdetection phenomena in the current technology in extreme weather scenarios, and the inference speed remains basically unchanged. Brief Description of the Drawings

[0049] Figure 1 It is a flowchart of a preferred embodiment of the cross - domain object detection method based on semi - supervised learning under extreme weather in the present invention;

[0050] Figure 2 It is a schematic diagram of a specific knowledge distillation method disclosed by the present invention;

[0051] Figure 3 It is a schematic diagram of a specific model pre-training method disclosed by the present invention;

[0052] Figure 4 It is a schematic diagram of a specific pseudo-label dynamic screening method disclosed by the present invention;

[0053] Figure 5 It is a flowchart of a specific model fine-tuning method disclosed by the present invention;

[0054] Figure 6 It is a functional principle block diagram of a preferred embodiment of a cross-domain object detection system based on semi-supervised learning under extreme weather in the present invention;

[0055] Figure 7 It is a functional principle block diagram of a preferred embodiment of a terminal in the present invention. Detailed implementation manners

[0056] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.

[0057] Please refer to Figure 1 , Figure 1 It is a flowchart of a cross-domain object detection method based on semi-supervised learning under extreme weather in the present invention. As Figure 1 shown, the cross-domain object detection method based on semi-supervised learning under extreme weather described in the embodiments of the present invention includes:

[0058] Step S11: Distill and transfer the bounding box regression knowledge possessed by the trained teacher model to the student model to obtain a pre-trained detection model with regression ability under extreme weather; wherein, the student model is a detection model with detection ability under normal weather, and the trained teacher model is a model obtained by training a detection model with better bounding box regression ability than the student model using a preset number of extreme weather images with labels.

[0059] In this embodiment, first, a student model and a teacher model are determined. The teacher model is a detection model that is superior to the student model in terms of bounding box regression ability and is more complex than the student model. However, its complex structure can provide more comprehensive knowledge for the student model. Then, a preset number of extreme weather images with labels are used to train the teacher model to obtain a trained teacher model. Next, the bounding box regression knowledge possessed by the trained teacher model is distilled and transferred to the student model to obtain a pre-trained detection model with regression ability in extreme weather, that is, through knowledge distillation, the student model is made to possess the regression ability of the trained teacher model in extreme weather. Among them, the preset number of extreme weather images with labels can be 1000, and extreme weather refers to abnormal meteorological phenomena with low occurrence probability but strong destructiveness, usually caused by climate change, natural factors or human activities, such as heavy rain, heavy snow, hail, typhoon, strong cold wave, sandstorm, fog and other bad weather. In such bad weather, when performing object detection, there are likely to be more misdetections and missed detections, resulting in low detection accuracy.

[0060] It can be understood that by training the teacher model with a small number of labeled extreme weather images and then using the more accurately regressed bounding box regression knowledge output by the trained teacher model to guide the output regression box of the student model, the student model is made to possess the regression ability of the teacher model. For example, the positioning performance of the student model is very poor. To improve the positioning performance of the student model, the method of knowledge distillation is used to improve the regression accuracy of the model.

[0061] Specifically, the bounding box regression knowledge possessed by the trained teacher model is distilled and transferred to the student model by using a progressive knowledge distillation strategy to obtain a pre-trained detection model with regression ability in extreme weather; among them, the progressive knowledge distillation strategy is a distillation strategy in which during the process of knowledge distillation, the weight of the regression distillation loss is gradually increased according to a preset loss weight adjustment strategy, so that the student model gradually possesses the regression ability in extreme weather, and the initial weight of the regression distillation loss is not greater than a preset weight value. It can be understood that since the difference between the student model and the teacher model is relatively large in the initial stage of training, in order to allow the student model to have an adaptation process, the progressive knowledge distillation method can be used to make the student model gradually possess the regression ability of the trained teacher model in extreme weather, and the preset weight value can be 0.1, that is to say, the initial weight of the regression distillation loss is not greater than 0.1.

[0062] For example, at the beginning stage of distillation, a smaller regression distillation loss weight, such as 0.05, is used to encourage the student model to only learn some basic features of the teacher model. As the student model gradually approaches the regression ability of the teacher model, more complex knowledge begins to be transferred to the student model, and the weight of the regression distillation loss gradually increases according to the preset loss weight adjustment strategy. The content learned by the student model gradually becomes more complex. The student model not only learns the bounding box regression of simple samples but also begins to learn complex and difficult-to-separate bounding box regressions. Then, in the later stage of the distillation process, more high-order knowledge of the teacher model is introduced, such as fine-grained object detection, occlusion handling, and small object regression. At this time, the student model already has strong capabilities and can better handle the object detection task, and learns the inference process, fine-grained features, and decision-making ability of the teacher model through distillation.

[0063] Moreover, the calculation formula for the regression distillation loss is as follows:

[0064] ;

[0065] Where, represents the regression distillation loss, β represents the weight of the regression distillation loss, N represents the total number of samples, represents the regression output of the trained teacher model for the i-th sample, represents the regression output of the student model for the i-th sample, represents the square of the L2 norm, that is, the mean squared error (MSE, Mean Squared Error), represents the mean squared error between the regression output of the trained teacher model for the i-th sample and the regression output of the student model for the i-th sample.

[0066] That is to say, during the training process, the weight β of the regression distillation loss can be dynamically adjusted according to the number of training rounds, so that the student model can gradually learn complex knowledge from the teacher model while avoiding relying too much on the complex knowledge of the teacher model in the early stage.

[0067] It should be noted that the teacher model can be an RT-DETR (Real-Time Detection Transformers, that is, real-time object detection based on the Transformer architecture) model, and the student model can be a YOLOv11 model. Given the advantages of the RT-DETR model in the regression (bounding box regression) part, that is, it can handle more complex object distributions, occlusions, and small object detections, but its training complexity is high and it is difficult to deploy. See Figure 2As shown, the RT-DETR model (teacher model) trained under a small number of labeled extreme weather images outputs relatively accurate bounding boxes to guide the output bounding boxes of the YOLOv11 model (student model) trained under a large number of normal weather images. That is to say, through the knowledge distillation method, the regression part (bounding box regression knowledge) of the DETR model is distilled into the YOLOv11 model, so that the YOLOv11 model can still be efficient in inference speed while learning the regression ability of the RT-DETR model under extreme weather. That is, by gradually simulating the output bounding boxes of the RT-DETR model by the student model to gradually fine-tune its output bounding boxes, making its regression boxes more accurate. That is, the localization ability of YOLOv11 in the target domain is still very weak. Therefore, through progressive bounding box distillation, using the powerful bounding box regression ability of RT-DETR, after training with a small number of labeled extreme weather images, the regression knowledge of the trained RT-DETR is passed to YOLOv11.

[0068] Moreover, when the teacher model can be the RT-DETR model and the student model is the YOLOv11 model, the calculation formula of the above regression distillation loss can be:

[0069] ;

[0070] Wherein, represents the regression output of the trained RT-DETR model for the i-th sample, represents the regression output of the YOLOv11 model for the i-th sample, represents the mean square error between the regression output of the trained RT-DETR model for the i-th sample and the regression output of the YOLOv11 model for the i-th sample.

[0071] Step S12: Input the obtained original source domain images under normal weather and the original target domain images under extreme weather into the diffusion model respectively, and input the intermediate state images between the source domain and the target domain output by the diffusion model into the pre-trained detection model for instance-level feature adversarial training to obtain the detection model after adversarial training.

[0072] In this embodiment, through the knowledge distillation method, the pre-trained detection model already has the detection ability in extreme weather scenarios. Then, in the adversarial training stage, the obtained original source domain images under normal weather and the original target domain images under extreme weather can be input into the diffusion model respectively, and the intermediate state images between the source domain and the target domain output by the diffusion model are input into the pre-trained detection model for instance-level feature adversarial training to obtain the detection model after adversarial training, thereby further narrowing the difference between the source domain and the target domain.

[0073] For example, seeFigure 3 As shown, the original source domain image (image under normal weather conditions) and the original target domain image (image under extreme weather conditions) are respectively input into the diffusion model for training. The diffusion model generates the target domain image in the source domain style and the source domain image in the target domain style, that is, the intermediate state images between the two domains. Then, the target domain image in the source domain style and the source domain image in the target domain style are input into the pre-trained detection model (YOLOv11 model with regression ability under extreme weather) for instance-level adversarial training. Through adversarial learning, the domain difference between normal weather and extreme weather is further reduced, so that when using the new teacher model to generate pseudo-labels subsequently, the quality of the pseudo-labels generated by the new teacher model under extreme weather can be further improved. Among them, the method based on the diffusion model adopted in this application is more stable during training. The generated intermediate state images have less noise and are closer to the target domain, thus reducing the model difference, improving the classification accuracy and further improving the localization ability. And instance-level feature adversarial training is introduced. Compared with the full-image feature alignment and pixel-level feature alignment in the prior art, it effectively reduces the interference of background information, greatly reduces the difference between the source domain and the target domain, and improves the detection performance of the model in the target domain.

[0074] It should be noted that the diffusion model is a type of probability-based generative model. It generates images by simulating the gradual "noising" and "denoising" process of images. That is, the diffusion model generates intermediate state images by adding noise and gradually denoising. That is to say, in the image generation task, the diffusion model generates new images by learning the process of recovering real images from noise. Compared with the method of using GAN (Generative Adversarial Network) in the prior art, the training of the diffusion model is relatively stable. The training process of the GAN method is easily affected by mode collapse and training imbalance, and the generator may focus on generating a few styles of images and ignore other types of image styles. However, since the optimization goal of the diffusion model is to recover the image by gradually denoising, and the objective function of the diffusion model is usually based on maximizing the likelihood estimation. Therefore, both the loss function and the training process are relatively smooth and are not prone to training instability. So the diffusion model obtains the generated target domain image in the source domain style through multiple steps of denoising and optimization. It retains the content of the target domain image and at the same time integrates the style of the source domain image. The same is true for generating the source domain image in the target domain style.

[0075] Specifically, the diffusion model can adopt the denoising probability diffusion model. To obtain the target domain image in the source domain style, the generated image is initialized from the target domain image, that is, by adding a certain amount of noise to the target domain image for initialization, and then the image is recovered through a gradual denoising process. That is to say, the diffusion model learns the distribution of images by adding noise and gradually denoising.

[0076] And to achieve style transfer, style features of the source domain or the target domain are introduced in the denoising process at each step, that is, content loss and style loss are added in the denoising process of the diffusion model; among them, the content loss is calculated by comparing the features of the intermediate image (generated image) and the original target domain image or the original source domain image on the convolutional layer, and the style loss is calculated by comparing the Gram matrix of the intermediate image and the original source domain image or the original target domain image.

[0077] For example, it is desired that the image generated by the diffusion model can maintain the content features of the target domain image. Therefore, when generating a target domain image with the style of the source domain, the content loss between the image generated by the diffusion model and the target domain image is calculated, which can be completed by comparing the features of the images on certain convolutional layers, that is:

[0078] ;

[0079] Among them, represents the content loss, represents the generated image, represents the target domain image, represents the features of the generated image extracted by the convolutional neural network, represents the features of the target domain image extracted by the convolutional neural network, and in order to transform the style of the target domain image into the style of the source domain, the style loss between the source domain image and the generated image needs to be calculated.

[0080] Among them, the style loss can be calculated by comparing the Gram matrix of the images, that is:

[0081] ;

[0082] Among them, represents the style loss, which measures the style difference between the intermediate image and the source domain image, represents the feature map of the generated image in the th layer of the convolutional neural network, represents the source domain image, represents the feature map of the source domain image in the th layer of the convolutional neural network, represents the Gram matrix of the feature map of the generated image in the th layer, represents the Gram matrix of the feature map of the source domain image in the th layer, and the Gram matrix describes the correlation of features in the image. The Gram matrix is used to capture the style features of the image, such as texture and style statistical characteristics.

[0083] It should be noted that considering the large background differences between the target domain and the source domain, to prevent incorrect alignment, this application focuses on the target itself and performs instance-level feature alignment through the mask of the target region, that is, instance-level feature adversarial training. The core idea of instance-level feature adversarial training is to introduce perturbations in the feature space to reduce the feature distribution differences between the source domain and the target domain. Different from traditional full-image feature alignment, the instance-level alignment method only focuses on the target region in the image, that is, only performs feature alignment on the object region in the image.

[0084] Specifically, during the feature adversarial training process, the target box region is extracted from the intermediate-state image. Based on the mask of the target box region, instance-level feature alignment is performed through the adversarial manner between the generator and the discriminator to reduce the feature distribution differences between the source domain and the target domain. Among them, the discriminator is used to judge whether the features of the target box regions corresponding to the source domain and the target domain are distinguishable, and the generator is used to make the features of the target box regions corresponding to the source domain and the target domain indistinguishable by minimizing the discrimination ability of the discriminator.

[0085] For example, first, the target boxes need to be extracted from the intermediate-state image. These target boxes specify the regions of interest, and the mask is used to represent these regions. The mask is usually a binary image, where the part corresponding to the target region is 1 and the rest is 0. Through the mask, only the features of the target region are concerned. Using the mask of the target box, the features of the non-target regions in the feature map can be masked, and the features of the target region are retained, that is, the mask M s and M t act on the feature maps of the source image and the target image to process only the target region. Through mask application, the attention is focused on the features of the target box region, enabling the model to focus on the key features of the target and ignore the background and other irrelevant parts when performing feature alignment between the source domain and the target domain. This way can significantly improve the detection performance of the model in the target domain, especially in the case of large domain differences.

[0086] Another example is that the maximum mean discrepancy (MMD) can be used to measure the distribution difference between the source domain and the target domain. By maximizing the similarity of the feature distributions of the target regions in the source domain and the target domain and minimizing their MMD values, that is:

[0087] ;

[0088] ;

[0089] ;

[0090] Among them, The distribution alignment loss representing the maximum mean difference representing the source domain data feature distribution expectation of representing the target domain data feature distribution expectation of representing the kernel function, such as the Gaussian kernel, calculating the distribution difference through the masked features representing the source domain masked feature map representing the target domain masked feature map representing the source domain masking matrix representing the target domain masking matrix representing the source domain feature extraction function representing the source domain feature map representing the target domain feature extraction function representing the target domain feature map representing element-wise multiplication, used to multiply the feature map and the masking matrix element-wise to retain the features of the key regions and suppress the irrelevant regions

[0091] It should be noted that the L2 norm adopted in this application is a simplified form of the MMD formula, rather than the conventional Hilbert space representation form

[0092] Then, by minimizing , the model can learn how to align the target box region features of the source domain and the target domain

[0093] Moreover, adversarial training reduces the feature difference between the source domain and the target domain through the adversarial process between the generator and the discriminator. In the adversarial training of the target box region, the discriminator judges whether the features of the target box regions in the source domain and the target domain are distinguishable, while the generator minimizes the discriminator's judgment ability to make the target region features of the source domain and the target domain indistinguishable. The loss of adversarial training i.e.:

[0094] ;

[0095] where D represents the discriminator, which is responsible for judging whether the features of the target box regions in the source domain and the target domain are distinguishable, and realizes the feature alignment of the target box region through masking, which can effectively reduce the domain difference between the source domain and the target domain, especially in the object detection task

[0096] Step S13: Determine a new teacher model and a new student model based on the adversarially trained detection model, and use a preset dynamic class pseudo-label screening method to screen the pseudo-labels generated by the new teacher model based on the original target domain images, and use the screened target domain pseudo-labels to fine-tune and train the new student model to obtain a trained object detection model.

[0097] In this embodiment, the performance of the model is improved through the distillation and adversarial training phases. Then, a new teacher model and a new student model can be determined based on the adversarially trained detection model. For example, an adversarially trained detection model can be used as the new teacher model, and an adversarially trained detection model can also be used as the new student model. Then, in the semi-supervised training phase, use a preset dynamic class pseudo-label screening method to screen the pseudo-labels generated by the new teacher model based on the original target domain images, and then use the screened target domain pseudo-labels to fine-tune and train the new student model to obtain a trained object detection model. Among them, the adopted dynamic class pseudo-label screening method can not only screen out high-quality target domain pseudo-labels of all categories, but also effectively utilize low-quality pseudo-label samples for internal mining, improve the robustness of the model, and enhance the detection ability of the model in extreme weather scenarios.

[0098] Specifically, perform class statistics and distribution statistics on the original target domain images to obtain corresponding statistical results, and perform targeted data augmentation processing on the original target domain images based on the statistical results to obtain new target domain images; use a preset dynamic class pseudo-label screening strategy to screen the pseudo-labels generated by the new teacher model based on the original target domain images and the new target domain images. It can be understood that for targeted data augmentation, that is, the number of classes in the head category is not increased, and the tail category is subjected to targeted data augmentation through methods such as mixup (mixing augmentation), copy-paste (copy-pasting), and cutout (cropping). Compared with the general data augmentation of the prior art, it effectively reduces the impact of class imbalance on the model. It is difficult for the single-threshold strategy of the existing method to screen out high-quality pseudo-labels of the tail category.

[0099] Among them, the preset dynamic category pseudo-label screening method is as follows: compare the classification confidence scores corresponding to the pseudo-labels indicating that the target to be detected belongs to the target category with the first threshold and the second threshold corresponding to the target category respectively to obtain the corresponding comparison results; among them, the first threshold is not less than the second threshold, and during the fine-tuning training process, the first threshold and the second threshold are dynamically adjusted according to the preset threshold dynamic adjustment strategy; when the comparison result shows that the classification confidence score is higher than the first threshold, the pseudo-label is determined as a positive sample to obtain the target domain pseudo-label; when the comparison result shows that the classification confidence score is lower than the second threshold, the pseudo-label is determined as a negative sample, and the negative sample is regularized to obtain the regularized negative sample for the fine-tuning training of the new student model; when the comparison result shows that the classification confidence score is between the first threshold and the second threshold, it is judged whether the pseudo-label is consistent with the pseudo-label predicted by the preset large model; when the pseudo-label is consistent with the pseudo-label predicted by the preset large model, the pseudo-label is determined as a positive sample to obtain the target domain pseudo-label; when the pseudo-label is inconsistent with the pseudo-label predicted by the preset large model, the pseudo-label is determined as a negative sample, and the negative sample is regularized to obtain the regularized negative sample for the fine-tuning training of the new student model.

[0100] It can be understood that the dynamic category pseudo-label screening strategy is an important strategy used in knowledge distillation or semi-supervised learning, and its purpose is to screen high-quality target domain pseudo-labels according to the classification confidence scores of each category. During the training process, the new student model usually relies on the pseudo-labels provided by the new teacher model, and the quality of these pseudo-labels is crucial for the training of the new student model. In the dynamic category pseudo-label screening, a dynamic threshold is set for each category, which can be adjusted based on the classification confidence scores of each category, and the large model is combined for auxiliary judgment to screen out high-quality target domain pseudo-labels to guide the training of the new student model, so as to ensure that only high-quality target domain pseudo-labels are used for the training of the new student model, further improving the performance of the new student model. And after the low-quality negative samples are regularized, they can also be used for the fine-tuning training of the new student model, enabling the new student model to adapt to the low-quality target domain samples.

[0101] For example, see Figure 4As shown in the figure, for the output of the new teacher model, first, the part with a classification confidence score lower than the classification threshold is directly filtered as the background. For example, the classification threshold can be 0.1. Then, for each category C, two thresholds are set, namely the high threshold (the first threshold) and the low threshold (the second threshold). Among them, as the training progresses, the high threshold will gradually increase, and samples higher than this high threshold will be considered positive samples and directly used as target domain pseudo-labels for the training of the new student model; pseudo-labels lower than the low threshold will be considered negative samples and need to be regularized for negative samples and internally mined. And for those intermediate-state samples with classification confidence scores between the low threshold and the high threshold, a large model is introduced to assist in judgment, such as grounding Dino (a large object detection model). Through the assistance of the large model for judgment, that is, if the pseudo-label predicted by the large model is consistent with the prediction of the current model, then this pseudo-label can be used as a positive sample; if the large model and the current model have inconsistent predictions, then this pseudo-label is regarded as a negative sample and is regularized for negative samples and internally mined. Among them, through the assistance of the large model for judgment, unreliable pseudo-labels can be avoided from being added to the training, ensuring more accurate label selection for intermediate-state samples.

[0102] It should be noted that the adjustment goal of the high threshold is to gradually improve the screening criteria as the training progresses, ensuring that the model only relies on high-quality pseudo-labels in the later stage of training. For example, in this application, an adjustment strategy based on model convergence is adopted. When the model has tended to converge on certain categories, that is, when the classification confidence scores for certain categories reach a relatively high level, the increase of the high threshold can be accelerated, so that the model relies more on reliable pseudo-labels. When the model converges quickly, that is, when the model can already stably identify certain categories, the increase of the high threshold can be accelerated to reduce the influence of uncertain samples; when the convergence is slow, a lower high threshold can be maintained to allow more samples to participate in the training. The high threshold for category C That is:

[0103] ;

[0104] Among them, represents the initial high threshold for category C, represents the dynamic adjustment factor, which maps the confidence through the Sigmoid function to the threshold increase ratio to ensure a smooth transition, represents the exponential decay term, is the average confidence of the current category c, reflecting the prediction stability of the model for this category, is the confidence benchmark value, used to control the trigger point of threshold adjustment, and k is the convergence rate factor, used to control the sensitivity of the high threshold to the increase of confidence, is the maximum increase coefficient, used to limit the threshold increase amplitude.

[0105] Moreover, the main function of the low threshold is to avoid misjudging samples with low confidence. Therefore, it also needs to be adjusted gradually according to the progress of training. For example, the dynamic adjustment of the low threshold can be achieved by adopting a strategy of adjusting based on the difficulty of class samples, that is, the low threshold can be dynamically adjusted according to the difficulty of each class. By analyzing the learning progress of the model for each class, the low threshold can be dynamically reduced for more difficult-to-identify classes. This strategy can ensure that those more difficult-to-identify classes can participate in the training more, avoiding the model relying too much on easily identifiable classes. Therefore, for more difficult-to-identify classes, that is, low-confidence classes, the training can be enhanced by reducing the low threshold. The low threshold for class C That is:

[0106] ;

[0107] Among them, represents the initial low threshold for class C, represents the decay coefficient, represents the maximum confidence for class C, represents the average confidence for class C.

[0108] It should also be noted that in the regularization of negative samples, the goal is to regularize those negative samples whose classification confidence scores are lower than the low threshold to prevent the model from relying too much on a small number of high-quality positive samples, that is, high-confidence positive samples. This regularization method helps the model avoid overfitting to easily classified high-quality positive samples by introducing some constraints and techniques, thereby improving its learning ability for low-confidence (i.e., uncertain) samples and enhancing the generalization ability of the model. By enhancing the model's learning of low-confidence negative samples, it helps the model avoid over-reliance on a small number of high-quality positive samples, thereby improving its learning ability and generalization ability for difficult samples. For example, higher weights can be assigned to low-confidence negative samples, making the contributions of these low-confidence negative samples to model training larger, so that the model cannot rely solely on those high-confidence positive samples but must learn low-confidence negative samples to improve its performance in difficult situations. The weight That is:

[0109] ;

[0110] Among them, represents the weight of the i-th negative sample below the low threshold, M represents the total number of negative samples below the low threshold, represents the pseudo-label of the i-th negative sample below the low threshold, represents the predicted output of the new student model for the i-th sample, represents the loss function of the i-th negative sample below the low threshold. For low-confidence samples ( Larger) assigns a higher loss weight to force the model to learn difficult samples, and reduces the weight for samples with high confidence ( Smaller) to avoid the model over-relying on simple samples.

[0111] It can be seen that the dynamic class pseudo-label screening strategy can more accurately screen pseudo-labels by setting dynamically adjusted high and low thresholds for each class, combining large model-assisted judgment and negative sample regularization, avoiding the model over-relying on high-quality positive samples, and preventing the model from falling into the noise of pseudo-labels. That is, by introducing large model judgment and the mining of intermediate state samples, the stability of training and the quality of pseudo-labels are improved, so that the model can be better trained in unsupervised learning tasks, and significantly reduce the missed detection of the model in extreme weather.

[0112] It should be noted that in the mean teacher model method, the parameters of the new teacher model are calculated from the parameters of the new student model and updated smoothly using exponential moving average. Compared with the traditional method, this can reduce the parameter fluctuations during training, so that the new teacher model can provide learning signals more stably. However, due to the deficiencies of exponential moving average such as excessive smoothing and lag, adaptive double exponential moving average can be adopted, which adaptively adjusts the decay rate according to the learning state of the new student model, enabling the new teacher model to better adapt to the changes of the new student model during the training process. Specifically, the parameters of the new student model in each training round are obtained in real time; based on the preset adaptive double exponential moving average strategy, and using the parameters of the new student model in each training round to dynamically update the parameters of the new teacher model; among them, the decay rate in the preset adaptive double exponential moving average strategy is dynamically adjusted based on the training round to dynamically control the parameter update speed of the new teacher model;

[0113] And, the calculation formula for the decay rate is:

[0114] ;

[0115] Where, represents the current training round number, represents the decay rate at the current training round number, represents the total number of training rounds, represents the minimum value of the decay rate, represents the maximum value of the decay rate.

[0116] That is, in order to make the decay rate of EMA (Exponential Moving Average) gradually adapt to the training progress, the number of training epochs can be used to dynamically adjust the decay rate, so as to ensure that in the initial stage of training, the new teacher model can quickly respond to the changes of the new student model, while in the later stage of training, the new teacher model tends to be stable.

[0117] It should also be noted that the preset adaptive double exponential moving average strategy is an improved exponential method. Its core idea is to introduce two exponential moving averages and enable them to be dynamically adjusted according to the training progress. This improvement solves the possible "over-smoothing" or "slow response" problems in the traditional exponential moving average method and enhances the adaptability of the new teacher model in knowledge distillation. In the standard EMA, the update of the teacher model is based on the weighted average of historical parameters. The update speed of the parameters is controlled by a fixed decay rate α. Generally, a larger decay rate, such as a decay rate close to 1, means a slower update speed, which helps to smooth the training process of the model, but may also lead to too slow a response of the new teacher model to keep up with the changes of the new student model in a timely manner. On the contrary, a smaller decay rate means that the new teacher model adapts to the new student model faster, but may lead to excessive fluctuations of the new teacher model. Therefore, the introduction of an adaptive decay rate aims to dynamically adjust this decay rate so as to quickly follow the changes of the new student model in the initial stage of training, while smoothing the update of the teacher model in the later stage of training and maintaining the stability of training. By introducing an adaptive decay rate and two exponential moving update processes, the adaptability and stability of the new teacher model are enhanced. Compared with the traditional exponential moving average, the adaptive double exponential moving average can better balance the flexibility and stability of the new teacher model, adapt to the changes of the new student model during the training process, and optimize the training process by dynamically adjusting the decay rate.

[0118] For example, the process of adaptive double exponential moving average is a process that uses two different exponential moving averages, namely the fast exponential moving average and the slow exponential moving average. The fast exponential moving average responds quickly to the changes of the new student model and can quickly update the parameters of the new teacher model. The slow exponential moving average updates more slowly and maintains the stability of the new teacher model to avoid excessive changes.

[0119] Fast update exponential As shown in the following formula:

[0120] ;

[0121] Slow exponential update As shown in the following formula:

[0122] ;

[0123] Among them, and represent the decay rate for adaptive adjustment of the current training round, which varies dynamically with the training process; represents the parameters of the new student model.

[0124] That is, in this embodiment, by introducing an adaptive decay rate and two exponential moving update processes, the adaptability and stability of the new teacher model are enhanced. Compared with the traditional exponential moving average, the adaptive double exponential moving average can better balance the flexibility and stability of the new teacher model, adapt to the changes of the new student model during the training process, and optimize the training process by dynamically adjusting the decay rate.

[0125] For example, as shown in Figure 5 , perform category statistics and distribution statistics on the original target domain images to obtain corresponding statistical results, and perform targeted data augmentation processing on the original target domain images based on the statistical results to obtain new target domain images. For example, use methods such as mixup, copy-paste, and cutout for targeted data augmentation. For source domain images, use ordinary data augmentation methods to process them to obtain new source domain images, such as using conventional data augmentation methods such as rotation, translation, and mosaic. Input the original target domain images and the new target domain images into the new teacher model and the new student model, and input the original source domain images and the new source domain images into the new student model for training. Among them, during the training process, the new student model depends on the pseudo-labels provided by the new teacher model, and the quality of these pseudo-labels is crucial for the training of the new student model. Therefore, a preset dynamic category pseudo-label screening strategy is used to screen out high-quality target domain pseudo-labels output by the new teacher model for the training of the new student model. And in order to make the new teacher model better adapt to the progress of training, an adaptive double exponential moving average strategy is adopted, which adaptively adjusts the decay rate according to the learning state of the new student model. Moreover, during the process of screening pseudo-labels, auxiliary judgment can also be carried out by combining a large model to screen out high-quality pseudo-labels to guide the training of the new student model, further improving the model performance.

[0126] It should be noted that during the training stage, for target domain images, all targeted data augmentation methods are adopted. For source domain images, conventional data augmentation methods such as rotation, translation, and mosaic are still used. As the training progresses, the model gradually converges, and some data augmentation methods can be gradually discarded to make the model tend to be stable. Finally, only mixup and translation are retained as the two augmentation methods for target domain images and source domain images to make the model converge smoothly.

[0127] Step S14: Use the trained object detection model to perform object detection on the to-be-detected image under extreme weather conditions to obtain corresponding detection results.

[0128] It can be understood that by training the student model using techniques such as additional bounding box distillation, generating intermediate images with a diffusion model, and dynamic class pseudo-label screening strategies adopted in the distillation stage, adversarial training stage, and fine-tuning training stage for object detection under extreme weather conditions, the localization and classification capabilities in extreme weather scenarios can be significantly improved with only a small increase in the number of parameters and computational complexity, and the phenomena of missed detection and misdetection can be significantly reduced.

[0129] It can be seen that in the embodiments of the present invention, distillation enables the student model to possess the regression ability of the trained teacher model under extreme weather. Then, a diffusion model is used to generate intermediate images that are closer to the target domain and have less noise, and the intermediate images are used to perform instance-level feature adversarial training on the pre-trained detection model that has learned the bounding box regression knowledge of the trained teacher model and thus has the regression ability under extreme weather to reduce the feature difference between the source domain and the target domain to obtain the detection model after adversarial training. Then, based on the detection model after adversarial training, a new teacher model and a new student model are determined. When generating pseudo-labels using the new teacher model in the subsequent process, the quality of the pseudo-labels generated by the new teacher model can be improved. Then, in the fine-tuning stage, the dynamic class pseudo-label screening strategy is used to screen the pseudo-labels generated by the new teacher model, and the screened high-quality target domain pseudo-labels are used to train the new student model to obtain a trained target detection model for object detection of the image to be detected under extreme weather, thereby being able to significantly reduce the phenomena of missed detection and misdetection in the extreme weather scenario of the current technology, and the inference speed remains basically unchanged.

[0130] That is, compared with the prior art, the present invention introduces an additional knowledge distillation stage on the basis of the average teacher model, namely a three-stage training paradigm. First, the positioning ability of the pre-trained detection model is significantly improved through progressive regression box distillation. Second, the prior art all adopts adversarial learning methods in the image migration stage, which are very unstable. The present invention uses a generation method based on a diffusion model for more stable training, and the generated intermediate images have less noise and are closer to the target domain, thereby reducing the model difference, improving the classification accuracy and further improving the positioning ability. An instance-level feature adversarial training is introduced, which effectively reduces the interference of background information and greatly reduces the difference between the source domain and the target domain compared with the global feature alignment and pixel-level feature alignment in the prior art, improving the detection performance of the detection model in the target domain. Finally, in the fine-tuning training stage, the present invention adopts targeted decay data augmentation, which effectively reduces the impact of class imbalance on the model and ensures that the model can converge smoothly compared with the data augmentation in the prior art. The flexibility and stability of the new teacher model are better balanced through the adaptive double exponential moving average method to adapt to the changes of the new student model in the training process, and the training process is optimized by dynamically adjusting the decay rate. Moreover, it is difficult to screen out high-quality pseudo-labels of tail classes by the single-threshold strategy in the prior art. The present invention adopts a dynamic pseudo-label screening strategy of large model assisted judgment and negative sample regularization, which can not only screen out high-quality positive samples of all classes, but also effectively utilize low-quality negative samples for internal mining, improve the model robustness, and enhance the detection ability of the model in extreme weather scenarios.

[0131] In one embodiment, as Figure 6 shown, based on the above cross-domain object detection method based on semi-supervised learning in extreme weather, the present invention also correspondingly provides a cross-domain object detection system based on semi-supervised learning in extreme weather, including:

[0132] A model distillation module 11, configured to distill and transfer the bounding box regression knowledge possessed by the trained teacher model to the student model, so as to obtain a pre-trained detection model with regression ability in extreme weather; wherein, the student model is a detection model with detection ability in normal weather, and the trained teacher model is a model obtained by training a detection model with better bounding box regression ability than the student model using a preset number of extreme weather images with labels;

[0133] An adversarial training module 12, configured to respectively input the obtained original source domain images in normal weather and original target domain images in extreme weather into a diffusion model, and input the intermediate images between the source domain and the target domain output by the diffusion model into the pre-trained detection model for instance-level feature adversarial training to obtain an adversarially trained detection model;

[0134] The model fine-tuning module 13 is configured to determine a new teacher model and a new student model based on the detection model after adversarial training, and use a preset dynamic class pseudo-label screening method to screen the pseudo-labels generated by the new teacher model based on the original target domain images, and use the screened target domain pseudo-labels to fine-tune the new student model to obtain a trained object detection model;

[0135] The object detection module 14 is configured to perform object detection on the image to be detected in extreme weather using the trained object detection model to obtain corresponding detection results.

[0136] Figure 7 The figure is a schematic structural diagram of the terminal provided by the embodiment of the present application. The terminal may include:

[0137] A memory 501, a processor 502, and a computer program stored on the memory 501 and executable on the processor 502.

[0138] When the processor 502 executes the program, it implements the cross-domain object detection method based on semi-supervised learning in extreme weather provided in the above embodiment.

[0139] Further, the terminal further includes:

[0140] A communication interface 503 for communication between the memory 501 and the processor 502.

[0141] The memory 501 is used to store a computer program executable on the processor 502.

[0142] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0143] If the memory 501, the processor 502, and the communication interface 503 are implemented independently, the communication interface 503, the memory 501, and the processor 502 may be connected to each other through a bus and complete communication with each other. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0144] Optionally, in a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are integrated on a single chip, the memory 501, the processor 502, and the communication interface 503 can communicate with each other through an internal interface.

[0145] The processor 502 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0146] This embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned cross-domain object detection method based on semi-supervised learning under extreme weather is implemented.

[0147] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are to be regarded as exemplary only, and the true scope and spirit of the present invention are pointed out by the claims.

[0148] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples. The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can read and execute instructions from the instruction execution system, apparatus, or device), or in connection with these instruction execution systems, apparatus, or devices.

[0149] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0150] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A cross-domain target detection method based on semi-supervised learning in extreme weather, characterized in that: The method comprises: Distilling the bounding box regression knowledge of the trained teacher model to the student model, and obtaining a pre-trained detection model with regression capability under extreme weather conditions; wherein the student model is a detection model with detection capability under normal weather conditions, and the trained teacher model is a model obtained by training a detection model with a bounding box regression capability superior to the student model using a preset number of extreme weather images with labels; The original source domain image under normal weather and the original target domain image under extreme weather are respectively input into the diffusion model, and the intermediate state image between the source domain and the target domain output by the diffusion model is input into the pre-trained detection model for instance-level feature adversarial training to obtain the detection model after adversarial training; Determine a new teacher model and a new student model based on the detection model after the adversarial training, and use a preset dynamic category pseudo-label screening method to screen the pseudo-labels generated by the new teacher model based on the original target domain image, and use the screened target domain pseudo-labels to fine-tune the new student model to obtain a trained target detection model; The trained target detection model is used to perform target detection on the image to be detected under extreme weather conditions to obtain corresponding detection results.

2. The cross-domain target detection method based on semi-supervised learning under extreme weather conditions according to claim 1 is characterized in that: The bounding box regression knowledge distillation of the trained teacher model is transferred to the student model to obtain a pre-trained detection model with regression capability under extreme weather conditions, including: The bounding box regression knowledge of the trained teacher model is distilled and transferred to the student model using a progressive knowledge distillation strategy, thus obtaining a pre-trained detection model with regression capabilities under extreme weather conditions. Among them, the progressive knowledge distillation strategy is a distillation strategy that gradually increases the weight of the regression distillation loss according to the preset loss weight adjustment strategy during the knowledge distillation process, so that the student model gradually has the regression ability under extreme weather conditions, and the initial weight of the regression distillation loss is not greater than the preset weight value.

3. The cross-domain target detection method based on semi-supervised learning under extreme weather conditions according to claim 2 is characterized in that: The calculation formula of the regression distillation loss is: ; in, represents the regression distillation loss, β represents the weight of the regression distillation loss, N represents the total number of samples, represents the regression output of the trained teacher model for the i-th sample, represents the regression output of the student model for the i-th sample, It represents the mean square error between the regression output of the trained teacher model for the i-th sample and the regression output of the student model for the i-th sample.

4. The cross-domain target detection method based on semi-supervised learning under extreme weather conditions according to claim 1 is characterized in that: The teacher model is the RT-DETR model, and the student model is the YOLOv11 model.

5. The cross-domain target detection method based on semi-supervised learning under extreme weather conditions according to claim 1 is characterized in that: The diffusion model is a denoising probabilistic diffusion model, and the diffusion model generates the intermediate state image by adding noise and gradually denoising, and adds content loss and style loss in the denoising process of the diffusion model; The content loss is calculated by comparing the features of the intermediate image with the original target domain image or the original source domain image on the convolutional layer, and the style loss is calculated by comparing the Gram matrix of the intermediate image with the original source domain image or the original target domain image.

6. The cross-domain target detection method based on semi-supervised learning under extreme weather conditions according to claim 1 is characterized in that: Also includes: During the feature adversarial training, a target frame region is extracted from the intermediate state image, and instance-level feature alignment is performed based on a mask of the target frame region and in an adversarial manner between a generator and a discriminator; The discriminator is used to determine whether the features of the target frame area corresponding to the source domain and the target domain are distinguishable, and the generator is used to minimize the determination ability of the discriminator.

7. The cross-domain target detection method based on semi-supervised learning under extreme weather conditions according to claim 1 is characterized in that: In the process of fine-tuning the new student model using the filtered target domain pseudo labels to obtain a trained target detection model, the process also includes: Obtaining the parameters of the new student model in each training round in real time; Based on a preset adaptive double exponential moving average strategy, the parameters of the new teacher model are dynamically updated using the parameters of the new student model in each training round; wherein the decay rate in the preset adaptive double exponential moving average strategy is dynamically adjusted based on the training rounds to dynamically control the parameter update speed of the new teacher model; And, the calculation formula of the attenuation rate is: ; in, Indicates the current training round number, Indicates the decay rate under the current number of training rounds, represents the total number of training rounds, represents the minimum value of the attenuation rate, Indicates the maximum value of the decay rate.

8. The cross-domain target detection method based on semi-supervised learning in extreme weather according to any one of claims 1 to 7, characterized in that: The method of using a preset dynamic category pseudo-label screening method to screen the pseudo-labels generated by the new teacher model based on the original target domain image includes: Performing category statistics and distribution statistics on the original target domain image to obtain corresponding statistical results, and performing targeted data enhancement processing on the original target domain image based on the statistical results to obtain a new target domain image; The pseudo labels generated by the new teacher model based on the original target domain image and the new target domain image are screened using a preset dynamic category pseudo label screening method.

9. The cross-domain target detection method based on semi-supervised learning in extreme weather according to claim 8, characterized in that: The preset dynamic category pseudo label screening method is: Compare the classification confidence score corresponding to when the pseudo label indicates that the target to be detected belongs to the target category with the first threshold and the second threshold corresponding to the target category to obtain corresponding comparison results; wherein the first threshold is not less than the second threshold, and during the fine-tuning training, the first threshold and the second threshold are dynamically adjusted according to a preset threshold dynamic adjustment strategy; When the comparison result indicates that the classification confidence score is higher than the first threshold, the pseudo label is determined as a positive sample to obtain the target domain pseudo label; When the comparison result indicates that the classification confidence score is lower than the second threshold, the pseudo label is determined as a negative sample, and the negative sample is regularized to obtain a regularized negative sample for fine-tuning training of the new student model; When the comparison result shows that the classification confidence score is between the first threshold and the second threshold, determining whether the pseudo label is consistent with the pseudo label predicted by the preset large model; When the pseudo label is consistent with the pseudo label predicted by the preset large model, the pseudo label is determined as a positive sample to obtain the target domain pseudo label; When the pseudo label is inconsistent with the pseudo label predicted by the preset large model, the pseudo label is determined as a negative sample, and the negative sample is regularized to obtain a regularized negative sample for fine-tuning training of the new student model.

10. A cross-domain target detection system based on semi-supervised learning in extreme weather, characterized in that: The system comprises: A model distillation module, for distilling and transferring bounding box regression knowledge possessed by a trained teacher model to a student model, to obtain a pre-trained detection model with regression capability under extreme weather conditions; wherein the student model is a detection model with detection capability under normal weather conditions, and the trained teacher model is a model obtained by training a detection model with a bounding box regression capability superior to that of the student model using a preset number of extreme weather images with labels; An adversarial training module is used to input the original source domain images under normal weather conditions and the original target domain images under extreme weather conditions into the diffusion model respectively, and input the intermediate state images between the source domain and the target domain output by the diffusion model into the pre-trained detection model to perform instance-level feature adversarial training, so as to obtain a detection model after adversarial training; A model fine-tuning module is used to determine a new teacher model and a new student model based on the detection model after the adversarial training, and use a preset dynamic category pseudo-label screening method to screen the pseudo-labels generated by the new teacher model based on the original target domain image, and use the screened target domain pseudo-labels to fine-tune the new student model to obtain a trained target detection model; The target detection module is used to use the trained target detection model to perform target detection on the image to be detected under extreme weather conditions to obtain corresponding detection results.

Citation Information

Patent Citations

  • Automatic driving automobile cross-weather target detection method based on feature adversarial learning

    CN119169580A

  • Target detection domain adaptation method based on random context consistency reasoning

    CN119313866A