Multi-layer feature adaptive alignment cross-domain unmanned aerial vehicle target detection method and system
Through the cross-domain drone object detection method with multi-layer feature adaptive alignment, the domain offset problem caused by the difference in drone image feature is solved, and high-precision cross-domain object detection is achieved.
Patent Information
- Application Number
- CN202510145502.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, the characteristics of drone images are significant, making it difficult for traditional object detection algorithms to cope with the problem of domain offset.
A cross-domain drone target detection method with multi-layer feature adaptive alignment is adopted. Through the adaptive teacher-student mutual learning detection model, domain adversarial learning and strength and weakness data enhancement are combined with multi-layer weighted feature alignment, the difference between the source domain and the target domain is reduced.
Accurate detection can be carried out without border annotation, significantly reducing the cost of data annotation, improving cross-domain detection accuracy, and enhancing the generalization ability and application scope of the algorithm.
Smart Images

Figure CN120014241A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of unmanned aerial vehicle (UAV) image processing, and relates to a cross-domain UAV target detection method and system with multi-layer feature adaptive alignment. Background Art
[0002] As an emerging remote sensing platform, drones have been widely used in fields such as smart agriculture, environmental monitoring, and traffic management. Object detection in drone imagery is a key technology. In recent years, methods based on convolutional neural networks have achieved remarkable results on multiple drone imagery datasets. However, these methods rely on a large amount of labeled data with bounding boxes and assume that the training and test sets have the same distribution characteristics. This seriously affects the generalization ability of the detector when faced with new environments, resulting in performance degradation when facing domain shifts. An intuitive solution is to collect and annotate images of new environments, but this method is costly and time-consuming. Therefore, unsupervised domain adaptation (DA) methods have emerged to transfer knowledge from the source domain to the unlabeled target domain.
[0003] The core idea of domain-adaptive object detection (DAOD) is to learn domain-invariant feature representations by aligning source and target domain features. Existing alignment strategies can be divided into three categories: feature-level adaptation, pixel-level adaptation, and self-training methods. Feature-level adaptation reduces the feature differences between the source and target domains through adversarial learning or explicit metric learning, but existing methods often ignore the transferability differences between different feature layers. Pixel-level adaptation transforms the source domain image into an intermediate image with the target style and then trains the detector using supervised learning. However, generating an ideal image converter is very difficult and may exacerbate domain shift in some extreme cases. Self-training methods achieve category alignment by generating pseudo-labels, but generating high-quality pseudo-labels in the target domain remains challenging.
[0004] Most current domain adaptive object detection methods are based on ground-based image datasets, such as Cityscapes and BDD100k. However, instance-level (e.g., appearance, size, etc.) domain shift in drone imagery is more severe than in ground-based imagery because drone imagery experiences large variations in shooting angle, viewing distance, and altitude, and factors such as weather and lighting significantly impact image quality. Therefore, the application of domain adaptation in drone imagery faces greater challenges, primarily in the following ways: 1) significant image-level style differences make it more difficult to transfer source domain knowledge to the target domain; and 2) the generation of high-quality pseudo-labels remains an urgent problem. Summary of the Invention
[0005] The purpose of the present invention is to provide a cross-domain UAV target detection method and system with multi-layer feature adaptive alignment to solve the problem in the prior art that the traditional target detection algorithm is difficult to cope with domain offset due to significant differences in UAV image features in different data domains.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] In the first aspect, the present application discloses a cross-domain UAV target detection method with multi-layer feature adaptive alignment, comprising:
[0008] Acquire cross-domain drone target detection data across time, sensors, viewpoints, and weather conditions, and preprocess them to obtain source domain data and target domain data.
[0009] The source domain data and target domain data are input into the preset adaptive teacher-student mutual learning detection model for processing, and the detection results are output, including:
[0010] The source domain data is input into the student model to initialize the weight of the student model. After the weight is initialized, the weight of the teacher model is updated through the exponential moving average.
[0011] The target domain data is input into the updated teacher model for processing to obtain pseudo labels in the target domain;
[0012] The weight-initialized student model performs adaptive learning based on the source domain data and pseudo-labels in the target domain, iteratively updates the weights of the student model, and iteratively updates the weights of the teacher model through exponential moving average. This results in an adaptive teacher-student mutual learning detection model.
[0013] The target domain data is input into the student model of the adaptive teacher-student mutual learning detection model for detection, and the detection results are output.
[0014] Preferably, the four cross-domain drone target detection data across time, across sensors, across perspectives and across weather are produced based on two optical satellite remote sensing image datasets: the DIOR dataset and the DOTA dataset, and two drone remote sensing image datasets: the VisDrone dataset and the UAVDT dataset.
[0015] Preferably, the source domain data is input into the student model, and the target domain data is input into the teacher model and the student model respectively, specifically: the strongly enhanced images of the source domain data and the target domain data are input into the student model; the weakly enhanced images of the target domain data are input into the teacher model.
[0016] Preferably, the student model is adaptively learned and updated through a gradient reversal layer and a discriminator to obtain a student model with updated weights.
[0017] Preferably, the weight-initialized student model performs adaptive learning based on the source domain data and the pseudo labels in the target domain, iteratively updates the weight of the student model, and iteratively updates the weight of the teacher model through exponential moving average; the adaptive teacher-student mutual learning detection model is specifically obtained as follows:
[0018] S2031: The weight-initialized student model performs adaptive learning based on the source domain data and the pseudo labels in the target domain to obtain a weight-updated student model;
[0019] S2032: The weight-updated student model updates the weight of the teacher model through exponential moving average;
[0020] S2033: Repeat S2031 to S2032 to iteratively update the student model and the teacher model to obtain an adaptive teacher-student mutual learning detection model.
[0021] Preferably, the method for constructing the adaptive teacher-student mutual learning detection model includes:
[0022] The source domain data is used to initialize the student model and obtain the supervised loss.
[0023] The target domain data enters the teacher model to generate pseudo labels for the target domain images;
[0024] The pseudo labels of the target domain images are fed into the student model to obtain the unsupervised loss.
[0025] The source domain data and pseudo labels enter the student model, and adaptive learning is performed through the gradient reversal layer and the discriminator to obtain a multi-layer weighted feature alignment loss;
[0026] The total loss of the student model is calculated using a total loss function based on supervised loss, unsupervised loss and multi-layer weighted feature alignment loss. The feature encoder and detector in the student model are iteratively updated according to the supervised loss and unsupervised loss, and the feature encoder and discriminator in the student model are iteratively updated according to the multi-layer weighted feature alignment loss. The teacher model is iteratively updated following the student model through the exponential moving average of the student model. When the total loss of the student model converges, an adaptive teacher-student mutual learning detection model is obtained.
[0027] Preferably, the total loss function is:
[0028]
[0029] The supervised loss is:
[0030]
[0031] The unsupervised loss is:
[0032]
[0033] The multi-layer weighted feature alignment loss is:
[0034]
[0035] in, is the total loss; For supervised losses; is the unsupervised loss; is the multi-layer weighted feature alignment loss; unsup and λ dis is a hyperparameter used to control the corresponding loss weight, λ unsup =1.0;λ dis =0.05; represents the source domain image; Bounding box annotations representing source domain images; Indicates the corresponding category label; Show the target domain image; The classification loss for the region proposal network RPN; The regression loss for the region proposal network RPN; is the classification loss of the region of interest ROI; is the regression loss of the region of interest ROI; is the pseudo label generated by the teacher model on the target domain; w i is the quantitative value of the transferability of the i-th feature layer; is the discriminant loss of the output feature of the i-th feature layer, using binary cross entropy loss, where d = 0; is the adversarial optimization objective function of the output feature of the i-th feature layer; E i is the output value of the i-th layer of the feature encoder; D i is the output value of the discriminator layer i.
[0036] In a second aspect, the present application discloses a cross-domain UAV target detection system with multi-layer feature adaptive alignment, comprising:
[0037] The data acquisition and preprocessing unit is used to acquire UAV target detection data across four domains: time, sensor, viewing angle, and weather, and preprocess it to obtain source domain data and target domain data;
[0038] The data detection unit is used to input the source domain data and the target domain data into the preset adaptive teacher-student mutual learning detection model for processing and output the detection results, specifically including:
[0039] The source domain data is input into the student model to initialize the weight of the student model. After the weight is initialized, the weight of the teacher model is updated through the exponential moving average.
[0040] The target domain data is input into the updated teacher model for processing to obtain pseudo labels in the target domain;
[0041] The weight-initialized student model performs adaptive learning based on the source domain data and pseudo-labels in the target domain, iteratively updates the weights of the student model, and iteratively updates the weights of the teacher model through exponential moving average. This results in an adaptive teacher-student mutual learning detection model.
[0042] The target domain data is input into the student model of the adaptive teacher-student mutual learning detection model for detection, and the detection results are output.
[0043] In a third aspect, the present application discloses an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the cross-domain UAV target detection method with multi-layer feature adaptive alignment as described in any one of the above items are implemented.
[0044] In a fourth aspect, the present application discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the cross-domain UAV target detection method with multi-layer feature adaptive alignment described above.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] 1) Accurate detection without bounding box annotation: The present invention can achieve accurate target detection without bounding box annotation of new data domains, thereby significantly reducing the cost of data annotation.
[0047] 2) Multi-layer weighted feature adaptive alignment strategy: This paper designs a multi-layer weighted feature adaptive alignment strategy that differentiates the transferability of different feature layers. By adaptively selecting feature layers with high transferability for knowledge transfer, the quality of pseudo-labels is effectively improved, significantly enhancing the accuracy of cross-domain object detection in drone imagery.
[0048] 3) Improving cross-domain detection accuracy: The present invention can significantly improve target detection accuracy across time, sensors, viewing angles, and weather conditions, further enhancing the generalization capability and application scope of the algorithm.
[0049] In summary, this application adopts an adaptive teacher-student mutual learning detection model, which reduces the difference between the source domain and the target domain by combining domain adversarial learning with multi-layer weighted feature alignment and strong and weak data enhancement. In the student model, highly transferable feature layers are adaptively selected for knowledge transfer, thereby promoting the consistency of feature distribution between the source domain and the target domain. The teacher model acquires knowledge from the student model through a mutual learning strategy while avoiding over-reliance on source domain data. At the same time, the discriminator and gradient reversal layer are used for adaptive learning to reduce the domain offset of the student model and improve the pseudo-label accuracy of the teacher model. It effectively solves the domain offset problem caused by the difference in drone image features in different data domains, and can perform high-precision detection without bounding box annotation, which significantly improves the accuracy of cross-domain target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0051] Figure 1 Flowchart of the present invention;
[0052] Figure 2 This is the network structure diagram of the cross-domain UAV target detection method with multi-layer feature adaptive alignment in the present invention;
[0053] Figure 3 The detection results of the method of the present invention and the comparative method in four cross-domain modes: cross-time, cross-sensor, cross-viewing angle, and cross-weather;
[0054] Figure 4 These are the detection results of the method of the present invention in four cross-domain modes: cross-time, cross-sensor, cross-viewing angle, and cross-weather. DETAILED DESCRIPTION
[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.
[0056] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.
[0057] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0058] In the description of the embodiments of the present invention, it should be noted that if the terms "upper," "lower," "horizontal," "inner," etc. appear, the orientation or positional relationship indicated is based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the inventive product is typically placed when in use. These terms are merely for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or component referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on the present invention. In addition, the terms "first," "second," etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0059] In addition, if the term "horizontal" appears, it does not mean that the component must be absolutely horizontal, but can be slightly tilted. For example, "horizontal" only means that its direction is more horizontal than "vertical", and does not mean that the structure must be completely horizontal, but can be slightly tilted.
[0060] In the description of the embodiments of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "disposed," "installed," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to internal connections between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0061] The present invention is described in further detail below with reference to the accompanying drawings:
[0062] See also Figure 1 , this application discloses a cross-domain UAV target detection method with multi-layer feature adaptive alignment, including:
[0063] S1: Acquire UAV target detection data across four domains: time, sensor, view, and weather, and preprocess them to obtain source domain data and target domain data;
[0064] S2: Input the source domain data and target domain data into the preset adaptive teacher-student mutual learning detection model for processing and output the detection results, which specifically include:
[0065] S201: The source domain data is input into the student model to initialize the weight of the student model; after the weight is initialized, the student model updates the weight of the teacher model through the exponential moving average;
[0066] S202: The target domain data is input into the updated teacher model for processing to obtain pseudo labels in the target domain;
[0067] S203: The weight-initialized student model performs adaptive learning based on the source domain data and the pseudo labels in the target domain, iteratively updates the weights of the student model, and iteratively updates the weights of the teacher model through exponential moving average; thus, an adaptive teacher-student mutual learning detection model is obtained;
[0068] S204: The target domain data is input into the student model of the adaptive teacher-student mutual learning detection model for detection, and the detection results are output.
[0069] In some embodiments, a cross-domain UAV target detection method with multi-layer feature adaptive alignment includes:
[0070] S1: Acquire UAV target detection data across four domains: time, sensor, view, and weather, and preprocess them to obtain source domain data and target domain data;
[0071] S2: The source domain data is input into the student model, and the target domain data is input into the teacher model and the student model respectively;
[0072] S3: Use source domain data to train the student model and initialize the weights of the student model;
[0073] S4: The student model after weight initialization updates the weight of the teacher model through exponential moving average;
[0074] S5: The teacher model with updated weights processes the input target domain data to obtain pseudo labels in the target domain.
[0075] S6: The student model's weights are adaptively learned and updated based on the source domain data and the pseudo-labels generated by the teacher model in the target domain. After each update of the student model's weights, the teacher model's weights are updated using an exponential moving average. Through multiple iterative updates, the pseudo-labels generated by the teacher model in the target domain become more accurate. Furthermore, these pseudo-labels optimize the student model's weights. Through multiple training updates, the optimal student model is obtained.
[0076] S7: The best student model performs accurate detection in the target domain data and obtains the detection results.
[0077] In some embodiments, a cross-domain drone target detection method with multi-layer feature adaptive alignment includes:
[0078] Step 1: Collection and preprocessing of four cross-domain UAV target detection datasets across time, cross-sensor, cross-view and cross-weather.
[0079] Step 2: Network design of the cross-domain UAV target detection method with multi-layer weighted feature adaptive alignment, which includes the teacher-student model mutual learning strategy, the adversarial learning strategy of multi-layer weighted feature alignment, and the loss function design.
[0080] Step 3: Perform model training and test evaluation on four cross-domain drone target detection datasets across time, cross-sensor, cross-view and cross-weather.
[0081] Step 4: Input drone images of different data domains into the trained model to obtain detection results.
[0082] In some embodiments, this application uses four public datasets, including two optical satellite remote sensing image datasets: the DIOR dataset and the DOTA dataset, and two drone remote sensing image datasets: the VisDrone dataset and the UAVDT dataset. Based on these four datasets, four cross-domain drone target detection datasets are produced: cross-time, cross-sensor, cross-viewing angle, and cross-weather. Each dataset is divided into source domain data and target domain data.
[0083] In some embodiments, the weight-initialized student model performs adaptive learning based on the source domain data and the pseudo labels in the target domain, iteratively updates the weights of the student model, and iteratively updates the weights of the teacher model through exponential moving average; the adaptive teacher-student mutual learning detection model is obtained as follows:
[0084] S2031: The weight-initialized student model performs adaptive learning based on the source domain data and the pseudo labels in the target domain to obtain a weight-updated student model;
[0085] S2032: The weight-updated student model updates the weight of the teacher model through exponential moving average;
[0086] S2033: Repeat S2031 to S2032 to iteratively update the student model and the teacher model to obtain an adaptive teacher-student mutual learning detection model.
[0087] See also Figure 2 This application also discloses a method for constructing an adaptive teacher-student mutual learning detection model, including:
[0088] 1) Initialize the student model with source domain data to obtain supervised loss;
[0089] 2) The target domain data enters the teacher model to generate pseudo labels for the target domain images;
[0090] 3) The pseudo-label of the target domain image enters the student model to obtain the unsupervised loss;
[0091] 4) The source domain data and pseudo labels enter the student model, and adaptive learning is performed through the gradient reversal layer and the discriminator to obtain a multi-layer weighted feature alignment loss;
[0092] 5) The total loss of the student model is calculated using a total loss function based on the supervised loss, unsupervised loss, and multi-layer weighted feature alignment loss. The feature encoder and detector in the student model are iteratively updated according to the supervised loss and unsupervised loss, and the feature encoder and discriminator in the student model are iteratively updated according to the multi-layer weighted feature alignment loss. The teacher model is iteratively updated following the student model through the exponential moving average of the student model. When the total loss of the student model converges, an adaptive teacher-student mutual learning detection model is obtained.
[0093] The adaptive teacher-student mutual learning detection model disclosed in this paper reduces the discrepancy between the source and target domains by combining domain adversarial learning with multi-layer weighted feature alignment and strong and weak data augmentation. In the student model, highly transferable feature layers are adaptively selected for knowledge transfer, thereby promoting consistency in the feature distributions of the source and target domains. The teacher model then learns from the student model through a mutual learning strategy while avoiding over-reliance on source domain data.
[0094] In some embodiments, the student model undergoes adaptive learning and updating through a gradient reversal layer and a discriminator, resulting in a weighted student model. Adaptive learning using the discriminator and gradient reversal layer reduces domain shift in the student model and improves the pseudo-labeling accuracy of the teacher model. This invention significantly improves the accuracy of cross-domain drone target detection and the generalization capability of the model.
[0095] In some embodiments, a cross-domain UAV target detection method with multi-layer feature adaptive alignment includes:
[0096] Step 1: Collection and preprocessing of four cross-domain UAV target detection datasets across time, cross-sensor, cross-view and cross-weather.
[0097] Step 2: Network design of the cross-domain UAV target detection method with multi-layer weighted feature adaptive alignment, which includes the teacher-student model mutual learning strategy, the adversarial learning strategy of multi-layer weighted feature alignment, and the loss function design.
[0098] Step 2.1: Design of the overall network structure. The model designed by the present invention includes a target-specific teacher model and a cross-domain student model. The teacher model only accepts the target domain weakly enhanced image, the student model accepts the source domain and target domain Two training strategies are used: teacher-student mutual learning strategy and adversarial learning strategy with multi-layer weighted feature alignment.
[0099] First, an object detector is trained using labeled data from the source domain, and the feature encoder and detector are initialized. During the mutual learning phase, the object detector is replicated as a teacher model and a student model. The teacher model generates pseudo-labels to guide the student model's training, and the student model updates the teacher model's knowledge using an exponential moving average (EMA). Through iteration, the pseudo-labels are continuously improved. Simultaneously, adaptive learning is performed using a discriminator and gradient reversal layers to reduce domain shift in the student model and improve the accuracy of the teacher model's pseudo-labels.
[0100] Step 2.2: Teacher-student model mutual learning strategy. The model of the present invention consists of a student model and a teacher model with the same structure. The student model learns through standard gradient updates, while the teacher model is updated through the weighted exponential moving average (EMA) of the student model. In order to generate accurate target domain pseudo labels, the teacher model receives weakly enhanced images, while the student model receives strongly enhanced images. Specifically, the teacher model performs weak enhancement processing such as random horizontal flipping and cropping on the target samples; the student model performs strong enhancement such as random color jittering, grayscale processing, Gaussian blurring, and cropping image blocks. The mutual learning strategy includes three steps:
[0101] 1) Model initialization
[0102] The supervised loss for training and initializing the student model using labeled source data can be defined as:
[0103]
[0104] Among them, RPN loss is the loss used to learn the Region Proposal Network (RPN) to generate candidate proposals, while the Region of Interest (ROI) loss is the loss for the ROI prediction branch. Both RPN and ROI perform bounding box regression (reg) and classification (cls). The present invention uses binary cross entropy loss (BCEWithLogitsLoss) as and And use L1 loss as and
[0105] 2) Optimize the student model using target pseudo labels
[0106] After obtaining the pseudo labels of the target domain images from the teacher model, the student model can be updated using the following loss function:
[0107]
[0108] in represents the pseudo labels generated by the teacher model on the target domain, Represents the target domain image. In this process, unsupervised loss should not be applied to the bounding box regression task, because the confidence score of the bounding box predicted on the unlabeled data can only represent the confidence of each object category, rather than the location of the generated bounding box.
[0109] 3) Gradually update the teacher model from the student model.
[0110] The present invention adopts the exponential moving average (EMA) method to update the teacher model by gradually copying the weights of the student model. The update formula can be defined as:
[0111] θ t ←αθ t +(1-α)θ s
[0112] Among them, θ t and θ s denote the network parameters of the teacher model and the student model respectively, α = 0.9996.
[0113] Step 2.3: Adversarial learning strategy for multi-layer weighted feature alignment. Multiple domain classifiers (D) are placed after the feature encoder (Backbone, E) of the student model. Given the domain label d of each input image, the domain discriminator D is updated using binary cross entropy loss. Specifically, the image of the source domain is labeled as d = 0, and the image of the target domain is labeled as d = 1. The discriminator loss of the i-th feature layer (here including the three feature layers res2, res3, and res4 in the feature encoder) It can be expressed as:
[0114]
[0115] The adversarial optimization objective function of the output feature of the i-th feature layer can be defined as follows:
[0116]
[0117] In order to simplify the minimax optimization process, a gradient reversal layer (GRL) is added between the feature encoder and the discriminator to generate reverse gradients. The objective function of multi-layer feature alignment can be written as:
[0118]
[0119] This paper introduces a weighting mechanism to highlight the unbalanced feature confusion capabilities of different domain discriminators. The detailed steps of the implementation process are as follows:
[0120] (1) According to The calculation formula calculates the loss of each adversarial classifier. A larger loss indicates that the feature has poor discrimination between the source domain and the target domain, which means that the feature has high transferability, and vice versa.
[0121] (2) Quantify the transferability of the i-th feature layer. Transferability can be understood as the contribution of the knowledge contained in the feature to cross-domain adaptation. Features with high transferability contribute significantly to the improvement of domain adaptation performance. It is defined as follows:
[0122]
[0123] (3) The multi-layer weighted feature alignment loss is as follows:
[0124]
[0125] Step 2.4: Loss function design. Total loss function The summary is as follows:
[0126]
[0127] Among them, λ unsup and λ dis is a hyperparameter used to control the corresponding loss weight. unsup =1.0;λ dis =0.05; it should be noted that and is used to train the feature encoder and detector in the student model, while It is used to update the feature encoder and discriminator. The teacher model is only updated by exponential moving average.
[0128] Step 3: Perform model training and test evaluation on four cross-domain drone target detection datasets across time, cross-sensor, cross-view and cross-weather.
[0129] The comparison methods used in this paper include the baseline method Faster R-CNN, which is trained only on the source domain dataset and tested in the target domain; feature-level adaptive methods such as DA-Faster R-CNN, SWDA, and H2FA; and self-training methods such as UT, PT, AT, and CMT. Model training and testing evaluation were performed on four cross-domain drone target detection datasets: cross-time, cross-sensor, cross-viewpoint, and cross-weather.
[0130] Step 4: Input drone images of different data domains into the trained model to obtain detection results.
[0131] Visual comparison of the detection results of the method of the present invention and the comparison method on four cross-domain UAV target detection datasets across time, cross-sensor, cross-viewpoint and cross-weather.
[0132] The present application also discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the cross-domain UAV target detection method with multi-layer feature adaptive alignment as described in any one of the above items are implemented.
[0133] The present application also discloses a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the steps of the cross-domain UAV target detection method with multi-layer feature adaptive alignment are implemented as described in any one of the above items.
[0134] [Example]
[0135] The specific steps of the embodiment are as follows:
[0136] Step 1: Collection and preprocessing of four cross-domain drone target detection datasets: cross-time, cross-sensor, cross-viewpoint, and cross-weather. Specifically, the following steps are included:
[0137] Step 1.1: Cross-temporal UAV target detection dataset preparation
[0138] Due to the difference in lighting conditions, images taken during the day and at night will cause data domain offset. To solve this problem, the present invention constructed two cross-time target detection datasets: VisDrone_daytime_to_night and UAVDT_daytime_to_night. In the construction of the VisDrone_daytime_to_night dataset, 6,000 daytime images in the VisDrone dataset were selected as source domain data, and 1,200 night images were selected as target domain data. Since some categories are less common in night images (such as bicycles, tricycles, and awning tricycles), these categories are ignored, and "people" and "pedestrian" are merged into one "people" category. The dataset finally contains 6 target categories, and the specific details are shown in Table 1.
[0139] For the UAVDT_daytime_to_night dataset, 8,000 daytime images were selected from the UAVDT dataset as source domain data, and 2,000 nighttime images were selected as target domain data. This dataset contains three target categories: “car,” “truck,” and “bus.” See Table 2 for details.
[0140] Table 1. Cross-time drone target detection dataset VisDrone_daytime_to_night
[0141]
[0142] Table 2. Cross-time UAV target detection dataset UAVDT_daytime_to_night
[0143]
[0144] Step 1.2: Cross-sensor UAV object detection dataset preparation
[0145] Due to the significant differences in images captured by different devices, the present invention constructs a cross-sensor target detection dataset VisDrone_to_UAVDT. The dataset construction process is as follows: First, 6,000 images are selected from the VisDrone training set as source domain data, named VisDrone_to_UAVDT_source_training. Then, 6,000 images are randomly selected from the UAVDT training set as target domain training data, named VisDrone_to_UAVDT_target_training. Finally, 2,000 images are randomly selected from the UAVDT test set as the target domain test set, named VisDrone_to_UAVDT_target_testing. These three datasets contain three target categories in common: car, truck, and bus. See Table 3 for details.
[0146] Table 3. Cross-sensor drone target detection dataset VisDrone_to_UAVDT
[0147]
[0148] Step 1.3: Preparation of cross-view UAV target detection dataset.
[0149] The DIOR and DOTA datasets primarily contain orthophotos, while the UAVDT and VisDrone datasets primarily contain front and side views. Scale changes caused by differences in viewing angles can significantly impact detection performance. To address this, we constructed four datasets: DIOR to VisDrone, DIOR to UAVDT, DOTA to VisDrone, and DOTA to UAVDT.
[0150] First, preprocess the DIOR dataset. The DIOR dataset contains 20 categories of remote sensing targets, with all types of vehicles labeled as a single category, "car." Therefore, the three categories "car," "truck," and "bus" in the VisDrone and UAVDT datasets were merged into "car," and other irrelevant categories were discarded. Of the 23,463 images in the DIOR dataset, only 6,428 contain the "car" category. 6,000 of these images were selected as source domain data, named DIOR_source_training. Then, 6,000 images were selected from the VisDrone dataset as target domain training data, named VisDrone_target_training, and the VisDrone test set was used as target domain testing data, named VisDrone_target_testing. Similarly, 6,000 images were selected from the UAVDT dataset as target domain training data, named UAVDT_target_training, and 2,000 images were selected as target domain testing data, named UAVDT_target_testing.
[0151] When preprocessing the DOTA dataset, horizontal bounding boxes were used and the original image was segmented into 800×800 sub-images with a 100-pixel overlap. Regions smaller than 800×800 were padded with zero-pixel values. The DOTA dataset contains 15 categories, of which only "large-vehicle" and "small-vehicle" are related to vehicles, covering the three vehicle categories of "car," "truck," and "bus." Therefore, the cropped sub-images were screened and 4,357 images containing the "large-vehicle" and "small-vehicle" categories were selected. These two categories were merged into the "car" category, and the annotations of other irrelevant categories were discarded. Ultimately, these 4,357 images were used as the source domain data and named DOTA_source_training. The target domain data for the VisDrone and UAVDT datasets are the same as above. Detailed details of the datasets are shown in Table 4.
[0152] Table 4. Cross-view drone target detection dataset
[0153]
[0154] Step 1.4: Cross-weather drone target detection dataset preparation
[0155] In practical applications, detectors need to accurately detect targets under various weather conditions. Therefore, it is crucial to improve the model's generalization ability for different weather conditions. This paper selected 6,000 sunny images from the UAVDT dataset as source domain data, named UAVDT_sunny_to_rainy_training, and selected 1,500 rainy images as target domain data, named UAVDT_sunny_to_rainy_testing. Due to the relatively small number of "bus" and "truck" in the target domain, this paper only evaluates the "car" category. The specific details of the dataset are shown in Table 5.
[0156] Table 5. Cross-weather target detection dataset UAVDT_sunny_to_rainy
[0157]
[0158] Step 2: Adaptive teacher-student mutual learning detection model is constructed, and its network structure is shown in the following figure: Figure 2 As shown, the generation process includes the following sub-steps:
[0159] Step 2.1: Design the overall network structure
[0160] Given N containing annotations in the source domain data s images And given N in the target domain t Unlabeled images in, Represents the source domain image Bounding box annotation of Indicates the corresponding category label. Target domain image There is no annotation. The ultimate goal of cross-domain object detection is to use and Designing domain-invariant detectors.
[0161] The model designed by the present invention includes a target-specific teacher model and a cross-domain student model. The teacher model only accepts the target domain weakly enhanced image, the student model accepts the source domain and target domain Two training strategies are used: teacher-student mutual learning strategy and adversarial learning strategy with multi-layer weighted feature alignment.
[0162] First, an object detector is trained using labeled data from the source domain, and the feature encoder and detector are initialized. During the mutual learning phase, the object detector is replicated as a teacher model and a student model. The teacher model generates pseudo-labels to guide the student model's training, and the student model updates the teacher model's knowledge using an exponential moving average (EMA). Through iteration, the pseudo-labels are continuously improved. Simultaneously, adaptive learning is performed using a discriminator and gradient reversal layers to reduce domain shift in the student model and improve the accuracy of the teacher model's pseudo-labels.
[0163] Step 2.2: Teacher-Student Model Mutual Learning Strategy
[0164] The model of the present invention consists of a student model and a teacher model with the same structure. The student model learns through standard gradient updates, while the teacher model is updated through the weighted exponential moving average (EMA) of the student model. In order to generate accurate target domain pseudo-labels, the teacher model receives weakly enhanced images, while the student model receives strongly enhanced images. Specifically, the teacher model performs weak enhancement processing such as random horizontal flipping and cropping on the target samples; the student model performs strong enhancement such as random color jittering, grayscale processing, Gaussian blurring, and cropping image blocks. The mutual learning strategy includes three steps:
[0165] (1) Model initialization
[0166] In the self-training framework, initialization is of great significance because it relies on the teacher model to generate reliable pseudo labels to optimize the student model on the unlabeled target domain. To achieve this goal, we first use the available supervised source data Through supervised loss Optimize the model. Therefore, the supervised loss for training and initializing the student model using labeled source data can be defined as:
[0167]
[0168] Among them, RPN loss is the loss used to learn the Region Proposal Network (RPN) to generate candidate proposals, while the Region of Interest (ROI) loss is the loss for the ROI prediction branch. Both RPN and ROI perform bounding box regression (reg) and classification (cls). The present invention uses binary cross entropy loss (BCEWithLogitsLoss) as and And use L1 loss as and
[0169] (2) Optimizing the student model using target pseudo-labels
[0170] Since there are no available labels in the target domain, the present invention adopts a pseudo-labeling method to generate virtual labels on the target domain images to train the student model. To filter out noisy pseudo-labels, a confidence threshold δ is set on the bounding box predicted by the teacher model to remove false positives, where δ = 0.8. In addition, duplicate bounding box predictions are excluded by performing non-maximum suppression (NMS) on each category. Therefore, after obtaining the pseudo-labels of the target domain images from the teacher model, the following loss function can be used to update the student model:
[0171]
[0172] in represents the pseudo labels generated by the teacher model on the target domain, In this process, the unsupervised loss should not be used for the bounding box regression task, because the confidence scores of the bounding boxes predicted on unlabeled data can only represent the confidence of each object category, rather than the location of the generated bounding box.
[0173] (3) Gradually update the teacher model from the student model
[0174] In order to obtain high-quality pseudo labels from the target image, the present invention adopts the exponential moving average (EMA) method to update the teacher model by gradually copying the weights of the student model. The update formula can be defined as:
[0175] θ t ←αθ t +(1-α)θ s
[0176] Among them, θ t and θ s denote the network parameters of the teacher model and the student model respectively, α = 0.9996.
[0177] Step 2.3: Adversarial learning strategy for multi-layer weighted feature alignment.
[0178] In order to achieve multi-layer adversarial learning, multiple domain discriminators (Domain Classifier, D for short) are placed after the feature encoder (Backbone, E for short) of the student model. The discriminator is designed to distinguish which domain (source domain or target domain) the derived feature E(X) comes from. Then the probability that each input sample belongs to the target domain can be defined as D(E(X)), and the probability of belonging to the source domain can be defined as 1-D(E(X)). Given the domain label d of each input image, the binary cross entropy loss is used to update the domain discriminator D. Specifically, the image of the source domain is labeled d=0, and the image of the target domain is labeled d=1. The discriminator loss of the i-th feature layer (here including the three feature layers res2, res3, and res4 in the feature encoder) It can be expressed as:
[0179]
[0180] On the other hand, in the adversarial learning process, the feature encoder E is trained to produce features that can confuse the discriminator D, while the discriminator D tries to distinguish which domain these features come from. Therefore, the adversarial optimization objective function of the output features of the i-th feature layer can be defined as follows:
[0181]
[0182] In order to simplify the minimax optimization process, a gradient reversal layer is added between the feature encoder and the discriminator to generate reverse gradients. During the gradient calculation process, the gradient reversal layer reverses the gradients passed through, so that the gradients of the feature encoder E are calculated in the opposite direction. This helps to maximize the discriminator loss by learning E, while only minimizing the target
[0183] Furthermore, the objective function of multi-layer feature alignment can be written as:
[0184]
[0185] Because features extracted from different feature layers have specific properties, namely, their adaptability to different scenarios (such as lighting and viewing angle), this paper introduces a weighting mechanism to highlight the unbalanced feature confusion capabilities of classifiers in different domains. The detailed steps of the implementation process are as follows:
[0186] (1) According to The calculation formula calculates the loss of each adversarial classifier. A larger loss indicates that the feature has poor discrimination between the source domain and the target domain, which means that the feature has high transferability, and vice versa.
[0187] (2) Quantify the transferability of the i-th feature layer. Transferability can be understood as the contribution of the knowledge contained in the feature to cross-domain adaptation. Features with high transferability contribute significantly to the improvement of domain adaptation performance. It is defined as follows:
[0188]
[0189] (3) The multi-layer weighted feature alignment loss is as follows:
[0190]
[0191] Leveraging the aforementioned multi-layer weighted feature alignment loss, the student model can address domain bias in visual features and help the teacher model generate accurate pseudo labels after multiple exponential moving average updates.
[0192] Step 2.4: Loss function design.
[0193] Total loss function The summary is as follows:
[0194]
[0195] Among them, λ unsup and λ dis is a hyperparameter used to control the corresponding loss weight. It should be noted that and is used to train the feature encoder and detector in the student model, while It is used to update the feature encoder and discriminator. The teacher model is only updated by exponential moving average.
[0196] Step 3: Perform model training and test evaluation on four cross-domain drone target detection datasets
[0197] The comparison methods used in this paper include the baseline method Faster RCNN, which is trained only on the source domain dataset and tested in the target domain; feature-level adaptive methods such as DA-Faster RCNN, SWDA, and H2FA; and self-training-based methods such as UT, PT, AT, and CMT.
[0198] Step 3.1: Training and testing the adaptive object detection model across time domains
[0199] Due to changes in lighting conditions (such as daytime and nighttime), images collected at different time periods suffer from domain migration issues. To address this issue, we designed two sets of experiments: VisDrone cross-temporal object detection (VisDrone_daytime_to_night dataset) and UAVDT cross-temporal object detection (UAVDT_daytime_to_night dataset). In the tests, the trained models achieved state-of-the-art performance in both cross-temporal domain adaptation tasks, with mAP50 reaching 33.4% and 64.6%, respectively. Compared to using only source domain data, the detector performance improved by 5.0% and 16.9%, respectively. The test results are shown in Tables 6 and 7.
[0200] Table 6. Test results on the VisDrone cross-temporal drone target detection dataset
[0201]
[0202]
[0203] Step 3.2: Cross-sensor domain adaptive object detection model training and testing
[0204] Due to the differences in images captured by different devices, this paper uses the VisDrone and UAVDT datasets to study cross-camera adaptability. Specifically, the model is trained using the VisDrone training set as the source domain and the UAVDT training set as the target domain, and then evaluated on the UAVDT test set. The test results, shown in Table 8, show that our method achieves the best mAP50 of 27.5%. This demonstrates that our method is effective in domain transfer for cross-camera adaptation scenarios.
[0205] Table 7. Test results on the UAVDT cross-temporal drone target detection dataset
[0206] method mAP50 car truck truck Faster RCNN 47.7 78.8 29.5 29.5 DA-Faster RCNN 46.3 76.4 35.5 35.5 SWDA 57.5 69.7 57.4 57.4 UT 62.1 66.1 74.8 74.8 PT 42.8 83.0 20.2 20.2 H2FA 62.7 75.8 50.2 50.2 AT 63.0 81.7 53.8 53.8 CMT 61.7 85.9 67.1 67.1 Method of the present invention 64.6 81.0 60.7 60.7
[0207] Table 8. Test results on the cross-sensor drone target detection dataset
[0208]
[0209]
[0210] Step 3.3: Cross-view domain adaptive object detection model training and testing
[0211] The present invention designed four experiments to train the model. In Experiment 1 and Experiment 2, the source domain used the training set of the DIOR dataset, and the target domain used the VisDrone and UAVDT datasets respectively; in Experiment 3 and Experiment 4, the source domain used the training set of the DOTA dataset, and the target domain used the VisDrone and UAVDT datasets respectively. It should be noted that these experiments only evaluated the common category (car). The test results are shown in Table 9. In the four cross-view migration tasks, compared with Faster RCNN using only source domain data, the performance improvement of this method is 32.8%, 31.7%, 33.2% and 40.8% respectively. The results show that the perspective change significantly affects the generalization ability of the model. Compared with CMT, the performance improvement of this method on the four migration tasks is 6.9%, 5.8%, 12.4% and 6.1% respectively. The main performance improvement comes from the multi-layer weighted feature alignment loss, which effectively alleviates the scale change problem in cross-domain adaptation.
[0212] Table 9. Detection results on the cross-view drone target detection dataset
[0213]
[0214]
[0215] Step 3.4: Training and testing the adaptive target detection model across weather domains
[0216] In the real world, detectors need to detect objects in a variety of weather conditions, so improving the model's generalization capabilities under different weather conditions is crucial. To simulate this scenario, the present invention selected 6,000 sunny images from the UAVDT dataset as source domain data and 1,500 rainy images as target domain data for training. Due to the small number of buses and trucks in the target domain, the present invention only evaluated cars. The test results are shown in Table 10, showing that compared with the PT, AT, and CMT methods, the performance of the present invention's method improved by 23.7%, 1.9%, and 1.1%, respectively. This result fully verifies the effectiveness and applicability of the present invention's method under different weather conditions.
[0217] Table 10. Test results on the cross-weather drone target detection dataset
[0218] method car AP50 Faster RCNN 21.9 DA-Faster RCNN 33.2 SWDA 25.3 UT 38.4 PT 38.8 H2FA 37.4 AT 60.6 CMT 61.4 Method of the present invention 62.5
[0219] Step 4: Input drone images of different data domains into the trained model to obtain detection results.
[0220] Visual comparison of the detection results of the method of the present invention and the comparison method on four cross-domain UAV target detection datasets across time, cross-sensor, transmission angle and cross-weather. Figure 3 The results of the comparison methods are presented in various cross-domain adaptation scenarios. Due to domain shift, the Faster R-CNN model trained solely on source domain data suffers from severe missed detections and false detections. After domain adaptation, the detection performance of the SWDA, UT, and AT methods is significantly improved. Compared with these methods, the proposed method can further reduce the number of false positive and false negative samples. Figure 4 The test results of the proposed method are shown. In different domain adaptation tasks, the proposed method has achieved excellent performance.
[0221] In summary, this invention discloses a cross-domain drone target detection method using multi-layer adaptive feature alignment. First, based on four public datasets (VisDrone, UAVDT, DIOR, and DOTA), four cross-domain drone target detection datasets (cross-time, cross-sensor, cross-view, and cross-weather) are constructed and preprocessed. Then, they are processed in a pre-defined adaptive teacher-student mutual learning detection model, outputting detection results. This pre-defined adaptive teacher-student mutual learning detection model combines domain adversarial learning with multi-layer weighted feature alignment and strong and weak data augmentation to reduce the differences between the source and target domains. In the student model, highly transferable feature layers are adaptively selected for knowledge transfer, thereby promoting consistency in feature distribution between the source and target domains. The teacher model acquires knowledge from the student model through a mutual learning strategy while avoiding over-reliance on source domain data. Simultaneously, a discriminator and gradient reversal layer are used for adaptive learning to reduce domain shift in the student model and improve the pseudo-labeling accuracy of the teacher model. Subsequently, training and testing are performed on the four cross-time, cross-sensor, cross-view, and cross-weather datasets. Finally, drone images from different data domains are fed into the trained model to obtain detection results. This method effectively solves the domain shift problem caused by the differences in drone image features across different data domains, enabling high-precision detection without bounding box annotation, significantly improving the accuracy of cross-domain object detection.
[0222] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0223] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0224] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0225] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0226] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. A cross-domain UAV target detection method with multi-layer feature adaptive alignment, characterized in that: include: Acquire UAV target detection data across four domains: time, sensor, view and weather, and preprocess to obtain source domain data and target domain data; The source domain data and the target domain data are input into the preset adaptive teacher-student mutual learning detection model for processing, and the detection results are output, including: The source domain data is input into the student model to initialize the weight of the student model. After the weight is initialized, the weight of the teacher model is updated through the exponential moving average. The target domain data is input into the updated teacher model for processing to obtain pseudo labels in the target domain; The weight-initialized student model performs adaptive learning based on the source domain data and the pseudo-labels in the target domain, iteratively updates the weight of the student model, and iteratively updates the weight of the teacher model through exponential moving average; thus, an adaptive teacher-student mutual learning detection model is obtained; The target domain data is input into the student model of the adaptive teacher-student mutual learning detection model for detection, and the detection results are output.
2. According to the multi-layer feature adaptive alignment cross-domain UAV target detection method of claim 1, it is characterized in that: The four cross-domain drone target detection data across time, across sensors, across perspectives and across weather are produced based on two optical satellite remote sensing image datasets: the DIOR dataset and the DOTA dataset, and two drone remote sensing image datasets: the VisDrone dataset and the UAVDT dataset.
3. According to the cross-domain UAV target detection method with multi-layer feature adaptive alignment according to claim 1, it is characterized in that: The source domain data is input into the student model, and the target domain data is input into the teacher model and the student model respectively. Specifically, the strongly enhanced images of the source domain data and the target domain data are input into the student model; the weakly enhanced images of the target domain data are input into the teacher model.
4. According to claim 1, a cross-domain UAV target detection method with multi-layer feature adaptive alignment is characterized in that: The student model is adaptively learned and updated through a gradient reversal layer and a discriminator to obtain a weight-updated student model.
5. According to the cross-domain UAV target detection method with multi-layer feature adaptive alignment as described in claim 1, it is characterized in that: The weight-initialized student model performs adaptive learning based on the source domain data and the pseudo-labels in the target domain, iteratively updates the weight of the student model, and iteratively updates the weight of the teacher model through exponential moving average; the adaptive teacher-student mutual learning detection model is obtained as follows: S2031: The weight-initialized student model performs adaptive learning based on the source domain data and the pseudo labels in the target domain to obtain a weight-updated student model; S2032: The weight-updated student model updates the weight of the teacher model through exponential moving average; S2033: Repeat S2031 to S2032 to iteratively update the student model and the teacher model to obtain an adaptive teacher-student mutual learning detection model.
6. The cross-domain UAV target detection method with multi-layer feature adaptive alignment according to claim 1 is characterized in that: The method for constructing the adaptive teacher-student mutual learning detection model includes: The source domain data is used to initialize the student model and obtain the supervised loss. The target domain data enters the teacher model to generate pseudo labels for the target domain images; The pseudo-labels of the target domain images are fed into the student model to obtain the unsupervised loss. The source domain data and pseudo labels enter the student model, and adaptive learning is performed through the gradient reversal layer and the discriminator to obtain a multi-layer weighted feature alignment loss; Based on the supervised loss, unsupervised loss and multi-layer weighted feature alignment loss, the total loss of the student model is calculated using the total loss function. The feature encoder and detector in the student model are iteratively updated according to the supervised loss and unsupervised loss, and the feature encoder and discriminator in the student model are iteratively updated according to the multi-layer weighted feature alignment loss. The teacher model is iteratively updated following the student model through the exponential moving average of the student model. When the total loss of the student model converges, an adaptive teacher-student mutual learning detection model is obtained.
7. The cross-domain UAV target detection method with multi-layer feature adaptive alignment according to claim 6 is characterized in that: The total loss function is: The supervised loss is: The unsupervised loss is: The multi-layer weighted feature alignment loss is: in, is the total loss; For supervised losses; is the unsupervised loss; is the multi-layer weighted feature alignment loss; unsup and λ dis is a hyperparameter used to control the corresponding loss weight, λ unsup =1.0;λ dis =0.05; represents the source domain image; Represents the bounding box annotation of the source domain image; Indicates the corresponding category label; Show the target domain image; The classification loss for the region proposal network RPN; The regression loss for the region proposal network RPN; is the classification loss of the region of interest ROI; is the regression loss of the region of interest ROI; is the pseudo label generated by the teacher model on the target domain; w i is the transferability quantification value of the i-th feature layer; is the discriminant loss of the output feature of the i-th feature layer, using binary cross entropy loss, where d = 0; is the adversarial optimization objective function of the output feature of the i-th feature layer; E i is the output value of the i-th layer of the feature encoder; D i is the output value of the i-th layer of the discriminator.
8. A cross-domain UAV target detection system with multi-layer feature adaptive alignment, characterized in that: include: The data acquisition preprocessing unit is used to acquire UAV target detection data across four domains: time, sensor, view and weather, and preprocess the data to obtain source domain data and target domain data. The data detection unit is used to input the source domain data and the target domain data into the preset adaptive teacher-student mutual learning detection model for processing and output the detection results, specifically including: The source domain data is input into the student model to initialize the weight of the student model. After the weight is initialized, the weight of the teacher model is updated through the exponential moving average. The target domain data is input into the updated teacher model for processing to obtain pseudo labels in the target domain; The weight-initialized student model performs adaptive learning based on the source domain data and the pseudo-labels in the target domain, iteratively updates the weight of the student model, and iteratively updates the weight of the teacher model through exponential moving average; thus, an adaptive teacher-student mutual learning detection model is obtained; The target domain data is input into the student model of the adaptive teacher-student mutual learning detection model for detection, and the detection results are output.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the cross-domain UAV target detection method with multi-layer feature adaptive alignment as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the cross-domain UAV target detection method with multi-layer feature adaptive alignment as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Cross-domain target detection method based on multi-layer feature alignment
CN110363122A
Rubbish sentry box detection method based on teacher adaptive framework
CN116895015A
High-resolution remote sensing image unsupervised adaptive target detection method
CN117475295A
Model training method, cross-domain target detection method and electronic equipment
CN118038163A
Aerial remote sensing image cross-domain target detection method
CN118570666A