An Unmanned Aerial Vehicle Small Target Detection Method Based on a Multi-Channel Domain Generalization Network
By building a multi-channel domain generalization network and using the YOLO-v8 network for feature alignment and domain generalization learning, the problem of unstable performance of drone small target detection in different environments is solved, and the detection accuracy and robustness of small targets in low-resolution images are improved.
Patent Information
- Application Number
- CN202510286090.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-12
AI Technical Summary
The existing drone small-objective detection models perform poorly in different environments, especially in low-resolution images with insufficient small-object feature information, and domain differences lead to performance attenuation.
Build a multi-channel domain generalization network, including high-resolution and low-resolution object detection networks and low-resolution domain generalization networks, and connect strong constraints and adversarial constraints, use the YOLO-v8 network for feature alignment and domain generalization learning to enhance the robustness of small-object detection.
The detection accuracy and domain generalization capability of low-resolution networks for small targets is improved, ensuring high-performance object detection under different environments and conditions.
Smart Images

Figure CN119810607B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision, image processing, and object detection, and particularly relates to a method for detecting small UAV targets based on a multi-channel domain generalization network. Background Art
[0002] UAV target detection is a key branch in the field of computer vision. With the rapid development of UAV technology, its application fields are continuously expanding, including military, civilian, and commercial uses. Along with this, potential security threats of UAVs have begun to emerge, such as issues like illegal intrusion and malicious attacks. The main purpose of UAV target detection is to identify, locate, and classify UAVs in image or video materials. Currently, this technology has been widely applied in multiple scenarios in daily life, such as airport security, protection of sensitive areas, and video surveillance systems.
[0003] Early object detection technologies mainly relied on manually designed features and machine learning classifiers. Classic methods, such as the Viola-Jones algorithm using Haar features and methods based on HOG (Histogram of Oriented Gradients) features, performed poorly when dealing with complex backgrounds, multi-object recognition, and size variations. With the progress of the computer vision field, the introduction of deep learning technology marked the emergence of new algorithms. R-CNN (Regions with CNN features) was a pioneering work in this regard. It located targets in images through region proposals and used a convolutional neural network (CNN) to extract features and perform classification. Its subsequent improved version, Fast R-CNN, significantly improved speed and efficiency by integrating the detection process into a single network and introducing the RoI (Region of Interest) pooling layer to extract features from the shared feature map. On this basis, one-stage object detectors such as YOLO (You Only Look Once) and SSD (Single Shot MultiBox Detector) further revolutionized the field. They directly predicted the class and location of targets through a single forward propagation process, greatly accelerating the detection speed. The latest YOLO-v8 algorithm has further developed in this series, not only maintaining the high efficiency of the YOLO series but also showing significant improvements in accuracy and generalization ability, making it perform better in complex real-time detection scenarios.
[0004] In object detection, small object detection is a particularly challenging task, mainly because small objects occupy few pixels in the image and have insufficient feature information. What's more complex is that when the data of these small objects comes from different environments, there will be domain differences between the datasets, that is, different distribution characteristics. Such domain differences may be caused by different acquisition conditions (such as lighting, background complexity), different data acquisition devices, or different image processing techniques. These differences lead to the fact that an object detection model that performs well in one environment may perform poorly in another environment, because the model may not be able to effectively extract universal features applicable to all environments from the image. This performance degradation caused by domain differences is particularly obvious when dealing with small objects, because the recognizable features of these objects are limited. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for detecting small objects of drones based on a multi-channel domain generalization network to improve the detection accuracy and robustness of small objects.
[0006] The purpose of the present invention is achieved through the following technical solutions:
[0007] A method for detecting small objects of drones based on a multi-channel domain generalization network includes:
[0008] Construct and train a multi-channel domain generalization network;
[0009] Input a low-resolution image containing small objects of drones into the trained multi-channel domain generalization network for detection and recognition;
[0010] The construction and training of the multi-channel domain generalization network includes the steps of:
[0011] Collect high-resolution images of drones as the first training image set, use the images obtained by decimating the high-resolution images of drones by a factor of ten as the second training image set, and use the low-resolution images as the third training image set;
[0012] Construct a multi-channel domain generalization network, which includes a high-resolution object detection network, a low-resolution object detection network, a low-resolution domain generalization network, and a discriminator. The high-resolution object detection network and the low-resolution object detection network are connected by a strong constraint, and the low-resolution object detection network and the low-resolution domain generalization network are connected by an adversarial constraint;
[0013] The high-resolution object detection network adopts the YOLO-v8 network. During model training, input the first training image set into the high-resolution object detection network, and obtain the first intermediate features through the intermediate layer of the high-resolution object detection network;
[0014] The low-resolution object detection network uses the YOLO-v8 network. During model training, the second training image set is input into the low-resolution object detection network, and the second intermediate feature is obtained through the intermediate layer of the low-resolution object detection network;
[0015] The first intermediate feature and the second intermediate feature are strongly constrained. The loss used for the strong constraint is the feature alignment loss, and the mean square error loss is used as the feature alignment loss;
[0016] The low-resolution domain generalization network uses the YOLO-v8 network. During model training, the third training image set is input into the low-resolution domain generalization network, and the third intermediate feature is obtained through the intermediate layer of the low-resolution domain generalization network; The low-resolution domain generalization network performs adversarial learning based on the third intermediate feature and the second intermediate feature, and then connects the intermediate layer of the low-resolution object detection network and the intermediate layer of the low-resolution domain generalization network to the discriminator. The discriminator evaluates whether the low-resolution object detection network can effectively learn the domain features of the low-resolution domain generalization network and outputs a discrimination result; The loss used during the training of the discriminator is the domain generalization loss.
[0017] Furthermore, using the mean square error loss as the feature alignment loss is specifically expressed as:
[0018]
[0019] where represents the actual value, represents the predicted value, N is the number of samples.
[0020] Furthermore, the calculation formula of the domain generalization loss is specifically expressed as:
[0021]
[0022] where D is the discriminator, θ is the parameter of the low-resolution object detection network, is the adversarial loss function, which is used to measure the discrimination ability of the discriminator D for the feature source domain, represents the process in which the network parameter θ minimizes the adversarial loss function while the discriminator D maximizes the adversarial loss function ; x represents the image features extracted from the low-resolution domain generalization network, representing the distribution of data; is the output of the low-resolution object detection network, and respectively represent the data distributions of the low-resolution domain generalization network and the data dependent on the parameters of the low-resolution object detection network. E represents the mathematical expectation operation on the original multi-domain training data distribution.
[0023] Furthermore, the YOLO-v8 network includes three parts: Backbone, Neck, and Head. The Backbone is used for feature extraction, the Neck is used for multi-scale feature fusion, and the Head is used for object detection and classification tasks. The Head part includes a detection head and a classification head;
[0024] The classification loss function of the YOLO-v8 network uses binary cross-entropy loss:
[0025]
[0026] where N is the number of samples, is the true label of the i-th sample, is the probability that the model predicts the i-th sample as the positive class;
[0027] The regression loss function of the YOLO-v8 network uses DFL loss:
[0028]
[0029] where S i and S i+1 represent the model outputs or data at two different states or time points; k true is the true value of the regression target, k i and k i+1 are the adjacent interval endpoints after discretizing the continuous value k true and satisfy k i ≤ k true ≤ k i+1 .
[0030] Furthermore, inputting the low-resolution image containing small UAV targets into the trained multi-channel domain generalization network for detection and recognition, the specific process is as follows:
[0031] Input the low-resolution image to be recognized into the low-resolution object detection network. The Backbone part of the trained YOLO-v8 network extracts large-scale features, the Neck part fuses the feature maps from different stages of the Backbone, and the Head part outputs the predicted bounding boxes for small UAV targets.
[0032] Furthermore, the total loss of the multi-channel domain generalization network is:
[0033]
[0034] Among them, are the weight factors of the feature alignment loss MSE, domain generalization loss V , classification loss BCE and regression loss DFL respectively.
[0035] The beneficial effects of the present invention are as follows:
[0036] The present invention uses a high-resolution network to strongly constrain a low-resolution network, thereby providing more high-resolution information features to the low-resolution network and enhancing the low-resolution network's perception ability for small targets. At the same time, through adversarial learning with the low-resolution domain generalization network of the input low-resolution datasets of different scenarios, the domain generalization ability of the network for small target features is enhanced. In this way, the network can more accurately detect small targets under different environments and changing conditions, ensuring high performance regardless of any background or conditions, thus solving common problems in small target detection, such as insufficient feature information and challenges brought by inter-domain variations. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is the multi-channel domain generalization network architecture diagram provided by the embodiment of the present invention;
[0038] Figure 2 is the YOLO-v8 network structure diagram;
[0039] Figure 3 is the discriminator network structure diagram. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] Referring to Figures 1 - 3 , the present invention provides a technical solution:
[0042] A method for detecting small targets of drones based on a multi-channel domain generalization network, comprising:
[0043] Step S1: Construct and train a multi-channel domain generalization network;
[0044] Step S2: Input the low-resolution image containing small targets of drones into the trained multi-channel domain generalization network for detection and recognition;
[0045] The step S1 of constructing and training the multi-channel domain generalization network includes the following steps:
[0046] Collect high-resolution images of drones as the first training image set, use the images obtained by decimating the high-resolution images of drones by a factor of ten as the second training image set, and use the low-resolution images as the third training image set;
[0047] Construct a multi-channel domain generalization network, as Figure 1 shown. The multi-channel domain generalization network includes a high-resolution object detection network, a low-resolution object detection network, a low-resolution domain generalization network, and a discriminator. The high-resolution object detection network and the low-resolution object detection network are connected by a strong constraint, and the low-resolution object detection network and the low-resolution domain generalization network are connected by an adversarial constraint;
[0048] The high-resolution object detection network uses the YOLO-v8 network. During model training, input the first training image set into the high-resolution object detection network to obtain the first intermediate feature through the intermediate layer of the high-resolution object detection network;
[0049] The low-resolution object detection network uses the YOLO-v8 network. During model training, input the second training image set into the low-resolution object detection network to obtain the second intermediate feature through the intermediate layer of the low-resolution object detection network;
[0050] Perform a strong constraint on the first intermediate feature and the second intermediate feature. The loss used for the strong constraint is the feature alignment loss, and the mean square error loss is used as the feature alignment loss, which is specifically expressed as:
[0051]
[0052] where represents the actual value, represents the predicted value, N is the number of samples. The mean square error loss quantifies the error of the model prediction by calculating the average of the sum of the squares of the differences between the predicted value and the actual value.
[0053] For the low-resolution object detection network, the object detection channel uses the decimated images of the same dataset as the high-resolution object detection network by a factor of ten. By performing a strong constraint on the features obtained from the intermediate layer of the high-resolution object detection network, the intermediate layer that is more sensitive to small object detection can learn the high-resolution features of the high-resolution object detection network, thereby improving the accuracy of small object detection.
[0054] These three channels are connected through specific connection mechanisms - strong constraint and adversarial constraint. Such a setting enables each channel to maintain its independent feature processing while optimizing the overall feature extraction and information fusion process through mutual learning.
[0055] The low-resolution domain generalization network adopts the YOLO-v8 network. During model training, the third training image set is input into the low-resolution domain generalization network, and the third intermediate features are obtained through the intermediate layer of the low-resolution domain generalization network. The low-resolution domain generalization network performs adversarial learning based on the third intermediate features and the second intermediate features, and then connects the intermediate layer of the low-resolution object detection network and the intermediate layer of the low-resolution domain generalization network to the discriminator. The discriminator evaluates whether the low-resolution object detection network can effectively learn the domain features of the low-resolution domain generalization network and outputs a discrimination result. The network structure diagram of the discriminator is as shown in Figure 3 shown; through the feedback of the discriminator, the low-resolution object detection network can continuously adjust and optimize its parameters to more accurately adapt to and simulate the performance of the feature domain of the low-resolution domain generalization network.
[0056] The loss used during discriminator training is the domain generalization loss, and the specific calculation formula is expressed as:
[0057]
[0058] where D is the discriminator, θ is the parameters of the low-resolution object detection network, is the adversarial loss function, which is used to measure the discrimination ability of the discriminator D for the feature source domain, represents the process of minimizing the adversarial loss function while the network parameters θ maximize the adversarial loss function of the discriminator D; x represents the image features extracted from the low-resolution domain generalization network, representing the distribution of data; is the output of the low-resolution object detection network, and respectively represent the data distributions of the low-resolution domain generalization network and the low-resolution object detection network depending on the network parameters θ. E represents the mathematical expectation operation on the original multi-domain training data distribution.
[0059] In this configuration, the discriminator D is trained to maximize its ability to distinguish between two types of samples, while the parameters of the low-resolution object detection network are adjusted by minimizing the discriminator's recognition ability for its output. The purpose is to make the domain of the output feature image of the low-resolution object detection network better imitate the domain of the feature image of the low-resolution domain generalization network. Through adversarial learning, the domain adaptability of the low-resolution object detection network to the low-resolution domain generalization network is continuously improved, and the domain generalization ability of the object detection network is enhanced.
[0060] As shown in Figure 2, the YOLO-v8 network consists of three parts: Backbone, Neck, and Head. The Backbone is used for feature extraction, the Neck is used for multi-scale feature fusion, and the Head is used for object detection and classification tasks. The Head part includes a detection head and a classification head;
[0061] The classification loss function of the YOLO-v8 network uses binary cross-entropy loss:
[0062]
[0063] where N is the number of samples, is the true label of the i-th sample, is the probability that the model predicts the i-th sample as the positive class;
[0064] The regression loss function of the YOLO-v8 network uses DFL loss:
[0065]
[0066] where S i and S i+1 represent the model outputs or data at two different states or time points; k true is the true value of the regression target, k i and k i+1 are the adjacent interval endpoints after discretizing the continuous value k true and satisfy k i ≤ k true ≤ k i+1 .
[0067] The high-resolution object detection network mainly processes high-resolution images to extract large-scale features. In this channel, the model can identify objects in a large range, including some small targets, through convolutional layers and multi-scale feature extraction techniques because high-resolution images provide sufficient details to support complex feature recognition. The low-resolution object detection network uses the same dataset as the high-resolution object detection network but with significant downsampling. This design aims to enhance the network's sensitivity to small target detection through lower-resolution images. By the feature alignment loss, the features in this channel are made to approach those of the high-resolution object detection network, achieving effective fusion of the two features, reducing the performance degradation caused by information loss between high and low resolutions, and improving the ability to recognize small targets.
[0068] Furthermore, the total loss of the multi-channel domain generalization network is as follows:
[0069]
[0070] where are the weight factors of the feature alignment loss MSE, domain generalization loss V , classification loss BCE and regression loss DFL respectively. In the default configuration of this embodiment, the value range of each weight factor is 0.1 ≤ α, β, η, μ ≤ 2.0. For the small target detection scenario of drones, the initial values are set as α = 1.0, β = 0.5, η = 0.8, μ = 1.2 to strengthen feature alignment and regression accuracy.
[0071] The output of the entire network is calculated by combining the object detection result loss, feature alignment loss, and domain generalization loss of YOLO-v8 in the above ratio. This loss combination method can not only improve the accuracy of object detection but also strengthen the feature fusion and domain generalization ability between different resolutions, ensuring the high performance and robustness of the network in different environments.
[0072] In step S2, the low-resolution image containing small drone targets is input into the trained multi-channel domain generalization network for detection and recognition. The specific process is as follows:
[0073] The low-resolution image to be recognized is input into the low-resolution object detection network. The Backbone part of the trained YOLO-v8 network extracts large-scale features, the Neck part fuses the feature maps from different stages of the Backbone, and the Head part outputs the predicted bounding boxes for small drone targets.
[0074] The core advantage of the present invention is to use a high-resolution network to strongly constrain a low-resolution network, thereby providing more high-resolution information features to the low-resolution network and improving the low-resolution network's perception ability for small targets. At the same time, through adversarial learning with the low-resolution domain generalization network that inputs low-resolution datasets of different scenarios, the domain generalization ability of the network for small target features is enhanced. In this way, the network can more accurately detect small targets in different environments and changing conditions, ensuring high performance regardless of any background or conditions, thus solving common problems in small target detection, such as insufficient feature information and challenges brought by domain variations.
[0075] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in related fields. Any changes and variations made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.
Claims
1. A method for detecting small targets of unmanned aerial vehicles based on a multi-channel domain generalization network, characterized in that, Including: Construct and train a multi-channel domain generalization network; Input the low-resolution image containing small UAV targets into the trained multi-channel domain generalization network for detection and recognition; The constructing and training of the multi-channel domain generalization network includes the steps: Collect high-resolution UAV images as the first training image set, use the images obtained by decimating the high-resolution UAV images by a factor of ten as the second training image set, and use the low-resolution images as the third training image set; Construct a multi-channel domain generalization network, which includes a high-resolution object detection network, a low-resolution object detection network, a low-resolution domain generalization network, and a discriminator. The high-resolution object detection network and the low-resolution object detection network are connected by a strong constraint; The low-resolution object detection network and the low-resolution domain generalization network are connected by an adversarial constraint; where: The high-resolution object detection network uses the YOLO-v8 network. During model training, input the first training image set into the high-resolution object detection network, and obtain the first intermediate feature through the intermediate layer of the high-resolution object detection network; The low-resolution object detection network uses the YOLO-v8 network. During model training, input the second training image set into the low-resolution object detection network, and obtain the second intermediate feature through the intermediate layer of the low-resolution object detection network; Perform a strong constraint on the first intermediate feature and the second intermediate feature. The loss used for the strong constraint is the feature alignment loss. The mean square error loss is used as the feature alignment loss, which is specifically expressed as: wherein represents the actual value, represents the predicted value, N is the number of samples; The low-resolution domain generalization network uses the YOLO-v8 network. During model training, input the third training image set into the low-resolution domain generalization network, and obtain the third intermediate feature through the intermediate layer of the low-resolution domain generalization network; The low-resolution domain generalization network performs adversarial learning based on the third intermediate feature and the second intermediate feature, and then connects the intermediate layer of the low-resolution object detection network and the intermediate layer of the low-resolution domain generalization network to the discriminator. The discriminator evaluates whether the low-resolution object detection network can effectively learn the domain features of the low-resolution domain generalization network and outputs a discrimination result; The loss used during the training of the discriminator is the domain generalization loss, and the calculation formula of the domain generalization loss is specifically expressed as: Among them, D is the discriminator, and θ is the parameter of the low-resolution object detection network. is the adversarial loss function, which is used to measure the discriminator D's ability to distinguish the source domain of features. represents the process in which the network parameter θ minimizes the adversarial loss function when the discriminator D maximizes the adversarial loss function . x represents the image features extracted from the low-resolution domain generalization network, representing the data distribution. is the output of the low-resolution object detection network. and respectively represent the data distributions of the low-resolution domain generalization network and the data depending on the low-resolution object detection network parameters. E represents the mathematical expectation operation on the original multi-domain training data distribution. Through the feedback of the discriminator, the low-resolution object detection network continuously adjusts and optimizes its parameters.
2. The method for detecting small targets of an unmanned aerial vehicle based on a multi-channel domain generalization network according to claim 1, wherein: The YOLO-v8 network includes three parts: Backbone, Neck, and Head. The Backbone is used for feature extraction, the Neck is used for multi-scale feature fusion, and the Head is used for object detection and classification tasks. The Head part includes a detection head and a classification head; The classification loss function of the YOLO-v8 network uses binary cross-entropy loss: where N is the number of samples, is the true label of the i-th sample, is the probability that the model predicts the i-th sample as the positive class; The regression loss function of the YOLO-v8 network uses DFL loss: Among them S i and S i+1 represent the model outputs or data at two different states or time points; k true is the true value of the regression target, k i and k i+1 are the adjacent interval endpoints after discretizing the continuous value k true and satisfy k i ≤ k true ≤ k i+1 .
3. The method for detecting small targets of an unmanned aerial vehicle based on a multi-channel domain generalization network according to claim 2, wherein: The process of inputting the low-resolution image containing small UAV targets into the trained multi-channel domain generalization network for detection and recognition is specifically as follows: The low-resolution image to be recognized is input into the low-resolution object detection network. The Backbone part of the trained YOLO-v8 network extracts large-scale features, the Neck part fuses the feature maps from different stages of the Backbone, and the Head part outputs the predicted bounding boxes for generating small UAV targets.
4. A method for detecting small targets of unmanned aerial vehicles based on a multi-channel domain generalization network according to claim 1, characterized in that: The total loss of the multi-channel domain generalization network is as follows: Among them, are the weight factors of the feature alignment loss MSE, domain generalization loss V , classification loss BCE and regression loss DFL respectively.
Citation Information
Patent Citations
Vehicle target detection method for remote sensing application scene
CN111899172A
Hidden danger area intrusion detection method and device, equipment and storage medium
CN117197756A