Cross-domain target detection method based on hierarchical and multi-scale adversarial training

By adopting hierarchical and multi-scale adversarial training strategies in the YOLOv5 detection model, the problem of performance degradation of the object detection model between data in different fields is solved, and high-precision real-time cross-domain object detection is achieved.

CN120047947APending Publication Date: 2025-05-2710TH RES INST OF CETC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510108644.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing target detection model has the problem of performance degradation or failure between data in different fields, mainly due to the distribution differences in data in different fields and the dependence on labeled data.

Method used

A cross-domain object detection method based on hierarchical and multi-scale adversarial training is adopted. By designing a hierarchical adversarial training strategy and a multi-scale adversarial training strategy in the YOLOv5 detection model, cross-domain feature distribution alignment is performed, domain distribution differences are reduced, and negative migration is suppressed.

Benefits of technology

The performance and generalization capabilities of the model in cross-domain object detection tasks are improved, the dependence on labeled data is reduced, and high-precision real-time object detection is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047947A_ABST
    Figure CN120047947A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-domain target detection method based on hierarchical and multi-scale adversarial training, and the method comprises the steps: inputting a sample into a to-be-trained target detection model for training, and obtaining a trained target detection model; the sample is composed of a source domain image and a target domain image; inputting the target domain image into a trained target detection model for detection to obtain a target detection result; the target detection result comprises a target bounding box and a target category. According to the method, data labels of the target field are not needed, and the detection accuracy of the model for the unlabeled field of interest is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly to a cross-domain object detection method based on hierarchical and multi-scale adversarial training. Background Art

[0002] Object detection is one of the core technologies in the field of computer vision, aiming to identify and locate objects of interest from images or videos, which can help people extract key information from complex visual environments. In recent years, deep neural networks can hierarchically learn robust and generalized deep feature representations, and meet the end-to-end requirements in practical applications. Therefore, more and more object detection methods begin to be based on deep neural networks, especially convolutional neural networks that are good at processing raster data such as images and videos. Among them, the emergence of mainstream detectors such as Faster RCNN (Faster Region-based Convolutional Neural Networks), YOLO (You Only Look Once), and Single Shot Multi-box Detector (SSD) has greatly promoted the development of object detection technology, met the detection requirements of different scenarios, and has been widely applied in fields such as intelligent video surveillance, intelligent transportation, medical diagnosis, and autonomous driving.

[0003] Although the application of deep learning in the field of object detection has achieved great success, the excellent performance of deep detection models depends on large-scale high-quality labeled data. In practical applications, especially when dealing with large-scale data sets or real-time applications, the cost of obtaining labeled data will further increase; in addition, if the existing detection models are directly applied to data in other fields after being trained on specific data, the performance often degrades or even fails. The fundamental reason is that there are distribution differences between data in different fields, and detection models trained based on supervised learning usually have difficulty adapting to the data distribution and characteristics of new fields, resulting in a significant reduction in the generalization ability on target data.

[0004] The above problems seriously affect and limit the use and deployment of object detection models in different scenarios. To solve the domain shift problem, domain adaptation methods have emerged. By mining the knowledge of existing data and the correlation between data in different domains, this technology can reduce the distribution differences between different domains, enabling the model to learn domain-invariant features and enhancing the discriminative and generalization abilities of the model for target data and tasks. Domain adaptation methods in the deep learning era can be divided into two categories: (1) Deep domain adaptation methods based on discrepancy measurement use predefined distribution metrics to weigh the differences in feature distributions between different domains. By adding an adaptation layer in the network, the feature distributions of different domains are explicitly pulled closer. Typical distribution measurement methods are the maximum mean discrepancy and its various variants. There are also studies that propose using the second-order statistics of features for distribution measurement; (2) Deep domain adaptation methods based on adversarial training usually consist of a domain discriminator and a feature extractor. The former is trained to distinguish samples from different domains, while the latter confuses the judgment of the domain discriminator during training. The two oppose each other to implicitly learn cross-domain invariant features. There are also studies that use a dual-classifier architecture to align feature distributions by maximizing the difference between classifiers.

[0005] Compared with image classification and semantic segmentation tasks, object detection is a more complex task because it requires predicting the location of the bounding box and the category of the corresponding object simultaneously, which poses a more severe challenge to cross-domain object detection tasks and makes cross-domain object detection one of the hot topics in the current field of computer vision research. Specifically, cross-domain object detection means that given a source domain with labeled information and a target domain without labeled information, a detector is trained to learn cross-domain invariant features through transfer learning methods (such as domain adaptation), so that the detector has good detection performance in the target domain with domain shift. The domain adaptation Faster RCNN method marks the beginning of cross-domain object detection. It proposes image-level and instance-level adaptation strategies and constrains model training by designing a consistency regularization term to adapt to the data distribution differences in different domains. A large number of cross-domain object detection methods are studied with the two-stage detector architecture of Faster RCNN as the core. However, with the requirements of practical applications and the limitations of two-stage detectors in terms of computational complexity and inference speed, in order to improve the efficiency of cross-domain detection of the model, cross-domain object detection methods based on single-stage detectors have also been proposed one after another. Single-stage detectors can simultaneously complete candidate region generation and object classification and localization. Compared with two-stage detectors, they usually have faster inference speed and lower computational cost, making them more suitable for practical applications and deployments. Existing methods basically complete distribution alignment based on adversarial training. Summary of the Invention

[0006] Aiming at the problems that target detection methods rely on large-scale high-quality labeled data, the trained detector cannot adapt to the decline in detection performance caused by domain differences, and negative transfer caused by blind knowledge transfer, etc., this application provides a cross-domain target detection method based on hierarchical and multi-scale adversarial training, which is applicable to various cross-domain target detection tasks, especially scenarios such as different weather conditions, different sensors, heterogeneous data, and simulation data to real data, solves the problem of a sharp drop in model performance when facing new domain data and tasks, and reduces the dependence of model training on labeled data.

[0007] This application discloses a cross-domain target detection method based on hierarchical and multi-scale adversarial training, which includes:

[0008] Input the samples into the target detection model to be trained for training to obtain a trained target detection model; the samples are composed of source domain images and target domain images;

[0009] Input the target domain images into the trained target detection model for detection to obtain target detection results; the target detection results include target bounding boxes and the categories of the targets.

[0010] Furthermore, the source domain images are composed of labeled data; the labeled data includes labeled target bounding boxes and category labels of the targets; the target domain images are composed of raw data; the raw data is unlabeled data; the target bounding boxes are used to indicate the positions where the targets are located.

[0011] Furthermore, the target detection model to be trained includes a backbone network, a neck network, a head network, and an adversarial training adaptation strategy module; the adversarial training adaptation strategy module includes two groups of domain discriminators and a gradient reversal layer; the first group of domain discriminators is connected to the backbone network; the second group of domain discriminators is connected to the head network; the backbone network is connected to the head network through the neck network; the backbone network is used to extract shallow features, middle features, and deep features of the input image; the head network is used to perform multi-scale detection on the input image to identify targets of different scale sizes in the image; both the first group of domain discriminators and the second group of domain discriminators are used to distinguish whether the features come from the source domain or the target domain; the gradient reversal layer is used to unify the adversarial learning process.

[0012] Furthermore, during the process of training the target detection model to be trained, input the source domain images and target domain images into the backbone network, and perform hierarchical adversarial training on the feature maps extracted by each C3 module; in the training stage of the head network, use the features of the three layers of the head network as three different-scale feature maps for cross-domain adaptation, and calculate the multi-scale adaptation loss of the head network; based on the category labels of the source domain labeled data, use the training loss function of the YOLOv5 network to train the discriminative ability of the target detection model.

[0013] Further, the hierarchical adversarial training of the feature maps extracted by each C3 module includes:

[0014] Taking the features of three layers in the backbone network as shallow features, middle-layer features, and deep features respectively, and calculating the overall adaptation loss of the backbone network.

[0015] Further, the calculation of the overall adaptation loss of the backbone network includes:

[0016] Performing pixel-level cross-domain adaptation on the shallow network features based on the least square loss to obtain a shallow adaptation loss function;

[0017] Performing image-level cross-domain adaptation on the middle-layer network features based on the binary cross-entropy loss to obtain a middle-layer adaptation loss function;

[0018] Performing image-level cross-domain adaptation on the deep network features based on the focal loss to obtain a deep adaptation loss function;

[0019] Based on the shallow adaptation loss function, the middle-layer adaptation loss function, and the deep adaptation loss function, obtaining the overall adaptation loss of the backbone network.

[0020] Further, the performing pixel-level cross-domain adaptation on the shallow network features based on the least square loss to obtain a shallow adaptation loss function includes:

[0021] Performing pixel-level cross-domain adaptation on the shallow network features using the least square loss. Assuming that the shallow feature maps of the i-th source domain image and the j-th target domain image are f i s and where u and v represent the pixel coordinates in the shallow feature map, the shallow adaptation loss function is:

[0022]

[0023] where D sa is the domain discriminator, is the shallow adaptation loss function;

[0024] The performing image-level cross-domain adaptation on the middle-layer network features based on the binary cross-entropy loss to obtain a middle-layer adaptation loss function includes:

[0025] The middle-layer features contain local shapes and simple objects. Performing image-level cross-domain adaptation on the middle-layer network features using the binary cross-entropy loss. Assuming that the middle-layer feature map of the i-th sample is h i and its corresponding domain label is d i then the middle-layer adaptation loss function is:

[0026]

[0027] Among them, D ma is the domain discriminator, is the middle-layer adaptation loss function;

[0028] Performing image-level cross-domain adaptation on the deep network features based on focal loss to obtain the deep adaptation loss function, including:

[0029] Assume that the deep feature map of the i-th sample is z i , and its corresponding domain label is d i , and γ is the weight control coefficient, then the deep adaptation loss function is:

[0030]

[0031] Among them, D da is the domain discriminator, is the deep adaptation loss function;

[0032] The adaptation loss function of the backbone network part of the YOLOv5 network is:

[0033]

[0034] Among them, is the adaptation loss function of the backbone network part.

[0035] Furthermore, calculating the multi-scale adaptation loss function of the head network, including:

[0036] Assume that the k-th scale feature map of the i-th sample is c i , and its corresponding domain label is d i , and u, v represent the pixel points of the feature map, then the adaptation loss function of the k-th scale is:

[0037]

[0038] Among them, is the adaptation loss function of the k-th scale, is the domain discriminator of the k-th scale;

[0039] The multi-scale adaptation loss function of the head network is obtained by summing all scales:

[0040]

[0041] Among them, is the multi-scale adaptation loss function of the head network.

[0042] Furthermore, the training loss function of the YOLOv5 network is:

[0043]

[0044] Among them, is the training loss function of the YOLOv5 network, X s , B s , Y s respectively represent sets; is the localization loss, is the classification loss, is the confidence loss.

[0045] Furthermore, the overall optimization loss function of the YOLOv5 network is:

[0046]

[0047] Among them, is the training loss function of the YOLOv5 network, is the adaptation loss function of the backbone network part, is the multi-scale adaptation loss function of the head network, and λ is a trade-off parameter.

[0048] Due to the adoption of the above technical solutions, the present application has the following advantages:

[0049] (1) Abandon the outdated Faster RCNN detector. The main body is a single-stage YOLOv5 detection model, which can perform high-precision real-time target detection;

[0050] (2) Through the hierarchical adversarial training strategy of the YOLOv5 backbone network, fully align the feature distributions in different domains and suppress the negative transfer phenomenon caused by blind adaptation. Among them, the shallow feature adaptation helps local feature alignment, the middle feature adaptation can complement the shallow and deep adaptations, and the deep feature adaptation effectively alleviates the negative transfer phenomenon;

[0051] (3) Through the multi-scale adversarial training strategy of the YOLOv5 head network, achieve multi-scale matching of feature distributions, and ensure the discrimination and generalization ability of the model for multi-scale targets;

[0052] (4) After the model training is completed, the adversarial training-related modules can be discarded. When performing the target detection task, only the YOLOv5 main body weights need to be loaded, which can greatly improve the performance of the model in the cross-domain target detection task without increasing the model calculation amount;

[0053] (5) This application provides a method that does not require data labels in the target domain and reduces the domain distribution difference through a cross-domain adversarial training strategy to improve the detection performance of the model for target data. Taking the YOLOv5 network as the main body, by designing a hierarchical adversarial training strategy for the backbone network and a multi-scale adversarial training strategy for the head network, according to the difference in the characterization information of feature maps at different levels, cross-domain feature distribution alignment is fully carried out to alleviate the negative transfer problem caused by blind adaptation, and the cross-domain discrimination and generalization ability of the model are enhanced. At the same time, a multi-scale adversarial training strategy is designed at the head network stage, and the rich discriminative information contained in the feature map to be detected is used to perform pixel-by-pixel adaptation to reduce local instance differences, ensuring the multi-scale detection ability of the detection model and improving the detection accuracy of the model for the unannotated area of interest. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of this application, and those of ordinary skill in the art can also obtain other drawings based on these drawings.

[0055] Figure 1 Schematic diagram of a target detection model to be trained in an embodiment of this application;

[0056] Figure 2 Schematic diagram of the pixel-level domain discriminator structure in an embodiment of this application;

[0057] Figure 3 Schematic diagram of the image-level domain discriminator structure in an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] The present application will be further described in conjunction with the drawings and embodiments. The described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art shall fall within the scope of protection of the embodiments of the present application.

[0059] The main limitations and challenges of current cross-domain object detection are reflected in the following aspects: (1) There may be significant differences in the scene layout, number of objects, patterns between objects, and background of data in different domains. Therefore, in cross-domain object detection tasks, blindly directly aligning the feature distributions may not capture effective representations, leading to the problem of negative transfer, and causing the detection performance of the model on the target data to decrease rather than increase; (2) Some existing methods use generative techniques to introduce auxiliary data or adopt the student-teacher network model for training. In the former case, the training process often cannot achieve end-to-end, which limits the full-link optimization of the model from input to output, increasing the difficulty and operation steps of using the model. In the latter case, it increases the complexity and training difficulty of the model, making the training and deployment of the detection model more inconvenient; (3) The mainstream benchmark for cross-domain detection methods is the Faster RCNN detection model, whose efficiency and performance have fallen behind the current lightweight network architectures, which fundamentally limits the cross-domain detection ability of the model. While detectors such as the YOLO series can maintain high accuracy while having better real-time performance and computational efficiency, making them more suitable for actual application scenarios.

[0060] In view of this, referring to Figure 1 , this application provides an embodiment of a cross-domain object detection method based on hierarchical and multi-scale adversarial training. Abandoning the Faster RCNN detector, with the single-stage YOLOv5 detection model as the main body, a hierarchical multi-scale adversarial training strategy is proposed, aiming at efficient and high-precision cross-domain object detection. The main features of this application are reflected in: (1) According to the difference in the representation information of different-depth features, a hierarchical adversarial training strategy is designed in the backbone network part of YOLOv5, which fully promotes the alignment of cross-domain feature distributions and suppresses the negative transfer phenomenon caused by blind adaptation; (2) Based on the rich discriminative information of the feature map to be detected, a multi-scale adversarial training strategy is designed in the head network part of YOLOv5 to promote the adaptation of multi-scale object detection capabilities and ensure the discriminative and generalization capabilities of the model in cross-domain tasks; (3) The cross-domain adversarial training strategy is only used during the model training process. During the inference and detection stage, only the main body weights of YOLOv5 need to be loaded, which can fully improve the detection efficiency of the model while ensuring the above cross-domain detection ability and achieve real-time inference.

[0061] The technical solution of this embodiment includes:

[0062] (1) Obtain source domain images and target domain images: Among them, the source domain images contain rich annotation information, which can be existing local data related to the target domain or newly obtained data, and then be annotated; the target domain images can directly obtain the original data without further labeling. The samples are composed of source domain images and target domain images.

[0063] (2) Model and Environment: It is implemented using the coding and running environment of Docker + VScode + PyTorch. As shown in Figure 1 , the target detection model to be trained takes YOLOv5 as the main body and consists of a backbone network, a neck network, a head network, and an adversarial training adaptation strategy. Specifically, the backbone network uses the CSPDarknet53 structure for deep feature extraction; the neck network, as an intermediate link, further extracts features and performs feature fusion, including the Feature Pyramid Network structure and the Path Aggregation Network, etc.; the head network conducts multi-scale detection to handle targets of different scales respectively; the adversarial training adaptation strategy mainly consists of a gradient reversal layer and multiple domain discriminators. The gradient reversal layer can facilitate the unified adversarial learning process, and the domain discriminator distinguishes whether the features come from the source domain or the target domain.

[0064] (3) Network Training Process: Input the samples (source domain images and target domain images) into the network. In the backbone network stage, hierarchical adversarial training is carried out using the feature maps extracted by each C3 module. Specifically, the features of the 4th, 6th, and 9th layers of the backbone network are respectively used as shallow, middle, and deep features in the hierarchical adversarial training strategy, and the overall adaptation loss of the backbone network is calculated according to formulas (1), (2), (3), and (4).

[0065] Specifically, in this application, a hierarchical adversarial training strategy is designed in the YOLOv5 backbone network part according to the difference in the representation information of different depth features. Shallow features can capture a large amount of detailed information, which not only helps to improve the small target detection ability of the model but also helps to align the local features of cross-domain targets. This application proposes to use the least square loss for pixel-level cross-domain adaptation of shallow features, and its corresponding domain discriminator D sa The structure is as shown in Figure 2 . Suppose the shallow feature maps of the i-th source domain image and the j-th target domain image are f i s and , and u, v represent the pixel point coordinates in the shallow feature map. Then the shallow adaptation loss function is as shown in formula (1):

[0066]

[0067] Among them, D sa is the domain discriminator, is the shallow adaptation loss function;

[0068] The middle features contain local shapes and simple targets. This application proposes to use binary cross-entropy loss for image-level cross-domain adaptation of them, and its corresponding domain discriminator D ma The structure is as shown in Figure 3 . Suppose the middle feature map of the i-th sample is hi , and its corresponding domain label is d i (If the i-th sample belongs to the source domain image, the intermediate feature map h i The corresponding domain label d i is 0. If the i-th sample belongs to the target domain image, the intermediate feature map h i The corresponding domain label d i is 1), then the intermediate adaptation loss function is shown in Equation (2):

[0069]

[0070] where D ma is the domain discriminator, is the intermediate adaptation loss function;

[0071] There may be significant differences in the scene layout, number of targets, patterns between targets, and background of data from different domains. Blindly aligning the feature distributions directly may lead to negative transfer problems. This application proposes to use focal loss for image-level cross-domain adaptation of deep features, and by adjusting the weights of samples with different discrimination difficulties, the cross-domain generalization performance of the model can be more robustly improved. The corresponding domain discriminator D da is also as Figure 3 shown. Suppose the deep feature map of the i-th sample is z i , and its corresponding domain label is d i (If the i-th sample belongs to the source domain image, the deep feature map z i The corresponding domain label d i is 0. If the i-th sample belongs to the target domain image, the deep feature map z i The corresponding domain label d i is 1), γ is the weight control coefficient, then the deep adaptation loss function is shown in Equation (3):

[0072]

[0073] where, D da is the domain discriminator, is the deep adaptation loss function;

[0074] If a sample is easily classified by the domain discriminator, it means that the sample is far from the decision boundary and is not suitable for distribution adaptation. Therefore, it is hoped that the model assigns it a smaller loss to avoid negative transfer; on the contrary, if a sample is difficult to be classified by the domain discriminator, it means that the sample is close to the decision boundary. Therefore, it is hoped that the model assigns it a larger loss to achieve domain confusion. In summary, γ needs to take a value greater than 1. Based on the shallow, intermediate, and deep adaptation losses, the adaptation loss function of the backbone network part can be expressed as follows:

[0075]

[0076] Among them, is the adaptation loss function of the backbone network part.

[0077] Network training process: In the head network stage, multi-scale adversarial training is adopted using the feature map to be detected. Specifically, the features of the 17th, 20th, and 23rd layers of the head network are used as three different-scale feature maps for cross-domain adaptation, and the overall adaptation loss of the head network is calculated according to formulas (5) and (6);

[0078] Specifically, based on the rich multi-scale discriminant information of the feature map to be detected, this application designs a multi-scale adversarial training strategy in the YOLOv5 head network part. The features of the head network are related to the multi-scale detection ability of the model, so the pixel-by-pixel adaptation of the head network is efficient and beneficial to knowledge transfer. On the one hand, multi-scale adaptation can fully reduce local instance differences and ensure the multi-scale detection ability of the model; on the other hand, it can also reduce background noise and help the model distinguish between targets and backgrounds. Specifically, the domain discriminator D pixel has a structure as Figure 2 shown (three discriminators with the same structure but independent). Assume that the k-th scale feature map of the i-th sample is c i , and its corresponding domain label is d i (if the i-th sample belongs to the source domain image, the domain label d i corresponding to the k-th scale feature map c i is 0, if the i-th sample belongs to the target domain image, the domain label d i corresponding to the k-th scale feature map c i is 1), u, v represent specific pixel points, then the k-th scale adaptation loss function is as shown in formula (5):

[0079]

[0080] Among them, is the adaptation loss function of the k-th scale, is the domain discriminator of the k-th scale;

[0081] The multi-scale adaptation loss function of the head network is obtained by summing all scales:

[0082]

[0083] Among them, is the multi-scale adaptation loss function of the head network.

[0084] Network training process: The source domain labeled data has rich supervision information. Therefore, the discriminative ability of the model is further trained based on its class labels. This part directly uses the training loss (detection loss) of the YOLOv5 network, as shown in formula (7), where X s ,B s ,Y s respectively represent sets.

[0085]

[0086] Among them, is the training loss function of the YOLOv5 network, X s ,B s ,Y s respectively represent sets; is the localization loss, is the classification loss, is the confidence loss.

[0087] The overall optimization loss function of the network can be as shown in formula (8), and the optimization process is unified through the gradient reversal layer.

[0088]

[0089] Among them, is the training loss function of the YOLOv5 network, is the adaptation loss function of the backbone network part, is the multi-scale adaptation loss function of the head network, and λ is the trade-off parameter.

[0090] (4) Inference and detection stage: The network inference process is exactly the same as that of the YOLOv5 model, that is, the target domain image is input into the trained object detection model for detection to obtain the object detection results, which include the object bounding box and the class of the object. After training, the model does not need to load the weights of modules such as the domain discriminator, but only needs to load the YOLOv5 main body weights. Therefore, the method proposed in this application will not bring additional computational overhead and memory occupation to the model use, and can ensure the high-precision real-time detection ability of the model.

[0091] The present application provides a method that can reduce the domain distribution difference through a cross-domain adversarial training strategy without target domain data labels, and improve the detection performance of the model for target data. The main structure of the method proposed in the present application is the YOLOv5 detector. A hierarchical adversarial training strategy is designed in the backbone network stage. According to the difference in the characterization information of different hierarchical feature maps, hierarchical feature cross-domain distribution alignment is carried out to alleviate the negative transfer problem caused by blind adaptation and enhance the generalization ability of the detection model. At the same time, a multi-scale adversarial training strategy is designed in the head network stage. By using the rich discriminative information contained in the feature map to be detected, pixel-by-pixel adaptation is carried out to reduce the local instance difference, ensure the multi-scale detection ability of the detection model, and improve the detection accuracy of the model for the unannotated area of interest.

[0092] Detection task: In real life, the weather changes frequently, which is an important domain shift factor that causes the target detection model to fail to work properly. Therefore, based on the publicly available street view datasets Cityscapes (clear) and Foggy Cityscapes (foggy), a quantitative analysis of the cross-domain target detection results of the method proposed in the present application is carried out. As shown in Table 1, Faster RCNN, DAF (Domain adaptive Faster RCNN), MAF (Multi-adversarial adaptation faster rcnn), SWDA (Strong weak distribution alignment), ATF (Asymmetric tri-way faster rcnn), HTCN (Hierarchical Transferability Calibration Network), UMT (Unbiased mean teacher), TIA (Task-specific inconsistency alignment), and TDD (target-perceived dual-branch distillation) are all current mainstream cross-domain detection methods based on the Faster RCNN model; YOLOv5s, S-DAYOLO (stepwise domain adaptative YOLO), Confmix, and the method proposed in the present application are all cross-domain detection methods based on the YOLOv5 model.

[0093] As can be seen from Table 1, the method proposed in this application has achieved optimal performance. Moreover, compared with the main YOLOv5s, this application has significantly improved the detection accuracy in almost all categories, fully demonstrating the superior cross-domain detection ability of the proposed method. In addition, Faster RCNN is usually built based on the VGG16 network, with a parameter quantity of approximately 138.36M, and the model is relatively large. The detection speed of related methods can only reach 5-17 frames per second. On the contrary, for the method proposed in this application, the main YOLOv5s has only approximately 7.08M parameters, and the model is concise. During actual use, the detection speed can reach about 95 frames per second, fully demonstrating the timeliness of the detection in this application.

[0094] Table 1 Quantitative Results of Cross-Domain Detection in Cityscapes and Foggy Cityscapes

[0095]

[0096]

[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them. Although this application has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of this application can still be modified or equivalently replaced, and any modification or equivalent replacement that does not depart from the spirit and scope of this application shall be covered by the protection scope of the claims of this application.

Claims

1. A cross-domain target detection method based on hierarchical and multi-scale adversarial training, characterized in that: include: Input the sample into the target detection model to be trained to obtain a trained target detection model; The samples consist of source domain images and target domain images; The target domain image is input into the trained target detection model for detection to obtain the target detection result; the target detection result includes the target bounding box and the target category.

2. The method according to claim 1, characterized in that The source domain image is composed of labeled data; the labeled data includes a labeled target bounding box and the target category label; the target domain image is composed of raw data; the raw data is unlabeled data; the target bounding box is used to indicate the location of the target.

3. The method according to claim 1 or 2, characterized in that: The target detection model to be trained includes a backbone network, a neck network, a head network and an adversarial training adaptation strategy module; the adversarial training adaptation strategy module includes two groups of domain discriminators and a gradient flipping layer; The first group of domain discriminators is connected to the backbone network; the second group of domain discriminators is connected to the head network; the backbone network is connected to the head network through the neck network; the backbone network is used to extract shallow features, middle features and deep features of the input image; The head network is used to perform multi-scale detection on the input image to identify objects of different scales in the image; the first group of domain discriminators and the second group of domain discriminators are both used to distinguish whether the features come from the source domain or the target domain; the gradient flipping layer is used to unify the adversarial learning process.

4. The method according to claim 1 or 2, characterized in that: In the process of training the target detection model to be trained, the source domain image and the target domain image are input into the backbone network, and the feature map extracted by each C3 module is subjected to hierarchical adversarial training; During the head network training phase, the features of the three layers of the head network are used as three feature maps of different scales for cross-domain adaptation, and the multi-scale adaptation loss of the head network is calculated. Based on the category labels of the source domain annotated data, the training loss function of the YOLOv5 network is used to train the discrimination ability of the target detection model.

5. The method according to claim 4, characterized in that The hierarchical adversarial training of the feature graph extracted by each C3 module includes: The features of the three layers in the backbone network are used as shallow features, middle features, and deep features respectively, and the overall adaptation loss of the backbone network is calculated.

6. The method according to claim 5, characterized in that The calculating of the overall adaptation loss of the backbone network includes: Based on the least square loss, the shallow network features are adapted across domains at the pixel level to obtain the shallow adaptation loss function; Based on the binary cross entropy loss, the middle-layer network features are adapted across domains at the image level to obtain the middle-layer adaptation loss function. Based on the focal loss, the deep network features are adapted across domains at the image level to obtain the deep adaptation loss function; Based on the shallow adaptation loss function, the middle adaptation loss function and the deep adaptation loss function, the overall adaptation loss of the backbone network is obtained.

7. The method according to claim 6, characterized in that The pixel-level cross-domain adaptation of the shallow network features based on the least square loss is performed to obtain a shallow adaptation loss function, including: The least square loss is used to perform pixel-level cross-domain adaptation on the shallow network features. Assume that the shallow feature map of the i-th source domain image and the shallow feature map of the j-th target domain image are f i s and u,v represent the pixel coordinates in the shallow feature map, and the shallow adaptation loss function is: Among them, D sa is the domain discriminator, is the shallow adaptation loss function; The image-level cross-domain adaptation of the middle-layer network features based on the binary cross entropy loss is performed to obtain the middle-layer adaptation loss function, including: The middle-level features include local shapes and simple targets. The binary cross entropy loss is used to perform image-level cross-domain adaptation on the middle-level network features. Assume that the middle-level feature map of the i-th sample is h i , whose corresponding domain label is d i , then the middle layer adaptation loss function is: Among them, D ma is the domain discriminator, is the middle layer adaptation loss function; The image-level cross-domain adaptation of the deep network features based on the focal loss is performed to obtain a deep adaptation loss function, including: Assume that the deep feature map of the i-th sample is z i , whose corresponding domain label is d i , γ is the weight control coefficient, then the deep adaptation loss function is: Among them, D da is the domain discriminator, is the deep adaptation loss function; The adaptation loss function of the backbone network part of the YOLOv5 network is: in, is the adaptation loss function of the backbone network.

8. The method according to claim 4, characterized in that The multi-scale adaptation loss function of the calculation head network includes: Assume that the k-th scale feature map of the i-th sample is c i , whose corresponding domain label is d i , u, v represent the pixels of the feature map, then the adaptation loss function of the kth scale is: in, is the adaptation loss function of the k-th scale, is the domain discriminator of the kth scale; The multi-scale adaptation loss function of the head network is obtained by summing all scales: in, It is the multi-scale adaptation loss function of the head network.

9. The method according to claim 4, characterized in that The training loss function of the YOLOv5 network is: in, is the training loss function of the YOLOv5 network, X s ,B s ,Y s Respectively A collection of; For the positioning loss, is the classification loss, is the confidence loss.

10. The method according to claim 4, characterized in that The overall optimization loss function of the YOLOv5 network is: in, is the training loss function of the YOLOv5 network, is the adaptation loss function of the backbone network, is the multi-scale adaptation loss function of the head network, and λ is the trade-off parameter.