Unmanned aerial vehicle target detection method based on DC GMA-YOLOv10 infrared and visible light fusion

By fusing infrared and visible light image features using the DC GMA-YOLOv10 model, the performance limitations of UAV detection under low light conditions were solved, enabling effective detection in all weather conditions and improving detection accuracy and robustness.

CN120877150APending Publication Date: 2025-10-31SHENYANG AEROSPACE UNIVERSITY
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510987197.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing drone detection methods have limited performance under low light conditions, resulting in frequent missed detections and false alarms. They cannot achieve all-weather detection, and a single sensor cannot meet the diverse detection needs in complex environments.

Method used

An infrared and visible light fusion method based on DC GMA-YOLOv10 is adopted. By constructing a dual-modal feature extraction network through an improved group hybrid attention module DC GMA, a cross-modal differential sensing fusion module CMDAF, and an information enhancement sampling module MSFS, infrared and visible light image features are fused to enhance detection performance.

Benefits of technology

This improves the robustness and performance of the drone detection system, ensuring effective detection of drones under various environmental conditions and enhancing detection accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877150A_ABST
    Figure CN120877150A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle target detection method based on DC GMA-YOLOv10 infrared and visible light fusion, and relates to the technical field of image detection. The method comprises the following steps: firstly, collecting and manufacturing infrared and visible light unmanned aerial vehicle image data sets of an unmanned aerial vehicle target in a complex environment; then, a DGM-YOLOv10 model is constructed, and the DGM-YOLOv10 model is trained according to the image data set; according to the DGM-YOLOv10 model, a feature extraction network of the YOLOv10 model is changed into two branches, and an improved group mixed attention module DC GMA is introduced to obtain a bimodal feature extraction network; a cross-modal differential perception fusion module CMDAF is introduced between the bimodal feature extraction networks; an information enhancement sampling module MSFS is introduced into the neck network; and on the basis of the trained DGM-YOLOv10 model, paired visible light and infrared unmanned aerial vehicle images are input for detection. According to the invention, based on the recognition of the convolutional neural network model, the robustness and performance of the unmanned aerial vehicle detection system can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image detection technology, and in particular to a method for detecting unmanned aerial vehicle targets based on the fusion of infrared and visible light using DC GMA-YOLOv10. Background Technology

[0002] With the rapid development of drone technology, its applications in military reconnaissance, border patrol, disaster relief, and urban security are becoming increasingly widespread. However, the actual application and detection environments are often dynamic and complex, with harsh weather conditions and complex terrain posing numerous challenges to the effective detection of drones.

[0003] Most commonly used drone detection methods rely on visible light images. Visible light images offer high resolution and rich texture details and color information, serving as a crucial foundation for image processing and target detection research. However, under low-light conditions, the performance of drone target detection based on visible light images is limited, with frequent missed detections and false alarms, making effective all-weather drone target detection impossible. Infrared sensing systems, by sensing the infrared radiation waves of target objects within a specific spectral range, eliminate the dependence of traditional optical imaging on the visible light environment. Infrared imaging technology leverages the difference in thermal conductivity between the target and the background, identifying the thermal features generated during drone flight at night or in low-light conditions to obtain target information that visible light sensors cannot capture, compensating for the susceptibility of visible light sensors to illumination effects. However, infrared images suffer from low resolution and poor texture, failing to provide a wealth of drone target details.

[0004] In summary, in complex and constantly changing flight environments, a single sensor often cannot meet the diverse needs of UAV target detection tasks. It is necessary to make comprehensive use of visible light sensors and infrared sensors to ensure effective detection of UAV targets in various environments. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a drone target detection method based on the fusion of infrared and visible light using DC GMA-YOLOv10, which addresses the shortcomings of the prior art and enables target detection of drones.

[0006] To solve the above-mentioned technical problems, the technical solution adopted by this invention is: a UAV target detection method based on DC GMA-YOLOv10 infrared and visible light fusion, comprising:

[0007] Collect and create datasets of infrared and visible light UAV images of UAV targets in complex environments;

[0008] A DC GMA-YOLOv10 model is constructed and trained based on the image dataset. The DC GMA-YOLOv10 model includes an improved group hybrid attention module (DC GMA) and a YOLOv10 model. The improved group hybrid attention module (DC GMA) replaces the ordinary convolution in the group hybrid attention mechanism (GMA) with dilated convolution. The feature extraction network of the YOLOv10 model is modified to have two branches. The improved group hybrid attention module (DC GMA) is introduced before the second convolutional layer of each branch to obtain a dual-modal feature extraction network for feature extraction of infrared and visible light images.

[0009] A cross-modal differential sensing fusion module (CMDAF) is introduced between the dual-modal feature extraction networks of the DC GMA-YOLOv10 model. This module uses a progressive approach to exchange complementary features between infrared and visible light images, thereby reducing feature differences between the two modalities.

[0010] Information Augmentation Sampling Module (MSFS) is introduced between the neck networks of the DGM-YOLOv10 model to perform upsampling operations, replacing the original upsampling module in YOLOv10.

[0011] Based on the trained DC GMA-YOLOv10 model, pairs of visible light and infrared UAV images are input for detection.

[0012] Furthermore, the specific method for collecting and creating the infrared and visible light UAV image dataset of UAV targets in complex environments is as follows:

[0013] First, the infrared and visible light cameras of the photoelectric turntable are used to capture drone flight videos under different weather and environmental conditions; the two modes of video are extracted frame by frame to obtain continuous images, and then the continuous images are sampled at fixed intervals to ensure the difference between the images; the sampled images are then screened.

[0014] Adaptive histogram partitioning and brightness correction are used to enhance the images in the infrared dataset to remove image noise and improve image details;

[0015] The TWMM method was used to register visible light and infrared images to form a dataset of infrared and visible light UAV flight images.

[0016] Furthermore, the specific method for image enhancement of infrared dataset images using adaptive histogram partitioning and brightness correction is as follows:

[0017] For adaptive segmentation of infrared grayscale histograms, Gaussian filtering and local weighted scatter smoothing algorithms are used to smooth the original grayscale histograms and remove peaks in the histograms.

[0018] The histogram is divided into multiple sub-histograms based on local minima. For each segment of the histogram, the gray density metric is used for discrimination. Low gray density values ​​correspond to foreground and high gray density values ​​correspond to background. An adaptive threshold is used to further classify each interval as foreground or background, maximizing the inter-class variance to optimize the classification effect.

[0019] For the foreground sub-histogram, a local contrast-weighted distribution is used to replace the traditional intensity distribution, thereby enhancing the local details of the image.

[0020] For the background sub-histogram, the mapping range is recalculated, contrast is improved, and noise is controlled to reduce background noise and prevent excessive enhancement of background noise;

[0021] A visual correction factor is used to adjust the size of the foreground histogram and reduce noise interference.

[0022] The particle swarm optimization algorithm is used to correct the average brightness of the enhanced image based on the reference image, so as to ensure that the overall brightness of the output image is moderate and natural.

[0023] Furthermore, the specific method by which the cross-modal differential sensing fusion module CMDAF progressively exchanges complementary features between infrared and visible light modal images is as follows:

[0024] (a) Acquire infrared and visible light features of the same size and number of channels;

[0025] (b) Perform channel alignment and normalization on the two features;

[0026] (c) Complementary features representing the differences between the two modes are obtained through difference calculation;

[0027] (d) Perform global average pooling on complementary features to generate channel weights and spatial weights respectively;

[0028] (e) Multiply each feature by the obtained channel weight and spatial weight to obtain the infrared and visible light fused features;

[0029] (f) Then, the infrared and visible light fused features are added to the original mode branches respectively;

[0030] Multiple cross-modal differential sensing fusion modules (CMDAF) are used for different feature levels, that is, the above process (a)-(f) is repeated multiple times to complete the feature exchange between the two modalities.

[0031] Furthermore, the specific method for the information enhancement sampling module MSFS to perform upsampling operation is as follows:

[0032] (1) Obtain the infrared and visible light fused feature map to be processed;

[0033] (2) The offset of the infrared and visible light fused feature map is adjusted, and after the offset is adjusted, the feature map is sampled by bilinear interpolation to obtain a feature map containing local detail enhancement information.

[0034] (3) Perform pixel rearrangement on the infrared and visible light fused feature map, and perform bilinear interpolation sampling and offset adjustment in sequence to obtain a feature map containing global context information;

[0035] (4) Fuse the two sets of feature maps, which contain local detail enhancement information and global context information respectively;

[0036] (5) Output the fused feature map;

[0037] Multiple information enhancement sampling modules (MSFS) are used for different feature levels, that is, the above process (1)-(5) is repeated multiple times to complete the upsampling operation.

[0038] Furthermore, the process of training the DC GMA-YOLOv10 model is as follows:

[0039] The prepared image dataset is divided into training, validation, and test sets, and the drones are labeled in YOLO format to generate corresponding label text files, including the bounding box coordinates and class labels of the targets. In each training iteration, a batch of infrared and visible light images is randomly selected from the training set, and these images are input in pairs into the DC GMA-YOLOv10 model to perform the following operations:

[0040] Step S1: First, initialize the parameters of the YOLOv10 model backbone network and the improved group hybrid attention module DC GMA, and load the pre-trained weights using transfer learning to accelerate the convergence speed; then, input the infrared image and the visible light image into the dual-branch feature extraction channel of the YOLOv10 model backbone network respectively.

[0041] Step S2: During the forward propagation, the infrared and visible light channels respectively pass through the group hybrid attention module DC GMA containing dilated convolution to extract deep semantic features, capturing the spatial structure and target features within each modality; then, the cross-modal differential perception fusion module CMDAF is used for interactive fusion of features, and the fused features are sent to the information enhancement sampling module MSFS for upsampling processing to further enhance the expressive power of the features; finally, the feature map is sent to the detection head of the YOLOv10 model to generate the target classification probability, bounding box coordinates, and confidence score; the difference between the output result and the true annotation is calculated to construct the loss function;

[0042] Step S3: Enter the backpropagation process, use the chain rule to propagate error information forward from the self-detector, calculate the gradients of each convolutional layer and the cross-modal differential sensing fusion module CMDAF and the information augmentation sampling module MSFS layer by layer, and update the parameters in the network through the stochastic gradient descent algorithm; repeat the above forward and backpropagation process until the loss function converges;

[0043] Repeat steps S1-S3 above, processing a new batch each time, until the entire training set has been traversed.

[0044] Furthermore, the specific method for detecting pairs of visible light and infrared UAV images based on the trained DC GMA-YOLOv10 model is as follows:

[0045] The registered infrared image and the visible light image to be detected are fed into the dual-modal feature extraction network of the DC GMA-YOLOv10 model; through the improved group hybrid attention module DC GMA, modal features of infrared and visible light images are extracted, capturing the structural features and semantic information that are different in the two modalities;

[0046] After feature extraction is completed in each branch, the output multi-scale feature map is sent to the cross-modal differential sensing fusion module CMDAF. The cross-modal differential sensing fusion module CMDAF adopts a stepwise fusion strategy. Through feature difference calculation and adaptive information guidance mechanism, it alternately guides the feature representations of infrared and visible light modes to perform difference enhancement and complementary fusion in channel and spatial dimensions to obtain fused features, thereby reducing redundancy and bias between modes.

[0047] The fused features are fed into the Information Enhancement Sampling Module (MSFS) for upsampling. The MSFS obtains a feature map containing local detail enhancement information and global context information by combining pixel rearrangement and offset adjustment operations. The feature map is then upsampled and fused to further enhance the expressive power of the features and improve the spatial resolution.

[0048] The obtained fused features are input into the YOLOv10 detection head. By introducing a task decoupling loss function and a dynamic label assignment strategy, the target category and bounding box coordinates are generated to obtain the final detection result.

[0049] The beneficial effects of adopting the above technical solution are as follows: The UAV target detection method based on DC GMA-YOLOv10 infrared and visible light fusion provided by this invention constructs a DC GMA-YOLOv10 model including an improved group hybrid attention module (DC GMA). This replaces the ordinary convolution in the group hybrid attention mechanism (GMA) with dilated convolution, improving the model's feature representation capability while maintaining inference speed. The DC GMA-YOLOv10 model includes a YOLOv10 model, with the feature extraction network of the YOLOv10 model modified into two branches, each incorporating the improved DC GMA module to extract features from infrared and visible light images. A cross-modal differential sensing fusion module (CMDAF) is introduced, progressively exchanging complementary features between infrared and visible light images to reduce feature differences between the two modalities. Based on the trained DC GMA-YOLOv10 model, pairs of visible light and infrared UAV images are input for detection. This invention, based on the recognition of a convolutional neural network model, can improve the robustness and performance of the UAV detection system, ensuring effective detection of UAVs under various environmental conditions. Attached Figure Description

[0050] Figure 1 A flowchart of a UAV target detection method based on DC GMA-YOLOv10 infrared and visible light fusion provided as an example of the present invention;

[0051] Figure 2 A schematic diagram of the structure of the group hybrid attention module DC GMA provided as an example of the present invention;

[0052] Figure 3 A schematic diagram of the cross-modal differential fusion sensing module provided as an example of the present invention;

[0053] Figure 4 A schematic diagram of the information enhancement sampling module provided as an example of the present invention. Detailed Implementation

[0054] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0055] In this embodiment, a UAV target detection method based on DC GMA-YOLOv10 infrared / visible light fusion is used, such as... Figure 1As shown, the process includes: collecting infrared and visible light UAV datasets for creating UAV targets in complex environments; constructing a DC GMA-YOLOv10 model and training it based on the dataset; the DC GMA-YOLOv10 model includes an improved group hybrid attention module (DC GMA), which replaces the ordinary convolution in the group hybrid attention mechanism (GMA) with dilated convolution, improving the model's feature representation capability while maintaining inference speed. The DC GMA-YOLOv10 model includes a YOLOv10 model, with the feature extraction network modified to have two branches. The improved group hybrid attention module (DC GMA) is introduced before the second convolutional layer of each of the two branches to extract features from infrared and visible light images. A cross-modal differential sensing fusion module (CMDAF) is introduced between the dual-modal feature extraction networks of the DC GMA-YOLOv10 model. This module progressively exchanges complementary features from infrared and visible light images, reducing feature differences between the two modalities. An information augmentation sampling module (MSFS) is introduced into the neck network of the DGM-YOLOv10 model to replace the original upsampling module in YOLOv10. MSFS obtains feature maps containing local detail enhancement information and global contextual information through pixel rearrangement and offset adjustment operations, and then upsamples these feature maps. The upsampled feature maps are then fused, enabling the network model to possess both detail enhancement and contextual fusion capabilities, achieving upsampling while improving model expressive power. Based on the trained DC GMA-YOLOv10 model, pairs of visible light and infrared UAV images are input for detection. This invention, based on a convolutional neural network model, can improve the robustness and performance of UAV detection systems, ensuring effective UAV detection under various environmental conditions.

[0056] like Figure 1 As shown, S101: Collect infrared / visible light UAV datasets of UAV targets in complex environments to train and validate the DC GMA-YOLOv10 model;

[0057] In this embodiment, a usable infrared / visible light UAV dataset was created using an HP-Z50 (DMA) three-band optoelectronic turntable and four different types of UAVs. First, UAV flight videos were captured using the infrared and visible light cameras of the turntable under different weather and environmental conditions. Frames of the two video modalities were extracted to obtain continuous images, and then these continuous images were sampled at 20-frame intervals to ensure image diversity. The sampled images were then filtered, removing overly blurry or significantly inaccurate images. To address the issue of image quality degradation caused by interference from focal plane array detectors in infrared imaging technology, adaptive histogram partitioning and brightness correction were used to enhance the infrared dataset images, removing image noise and improving image details.

[0058] Adaptive histogram partitioning and brightness correction address the problems of excessive background enhancement, noise amplification, and brightness distortion in traditional image enhancement algorithms when processing infrared images. First, adaptive segmentation of the infrared grayscale histogram is performed. Gaussian filtering and a locally weighted scatter smoothing algorithm are used to smooth the original grayscale histogram, removing spikes. Then, the histogram is divided into multiple sub-histograms based on local minima. For each segment of the histogram, grayscale density is used for discrimination: low grayscale density values ​​correspond to foreground, and high grayscale density values ​​correspond to background. An adaptive threshold is used to further classify each interval as either foreground or background, maximizing inter-class variance to optimize classification performance. For the foreground sub-histogram, a locally contrast-weighted distribution replaces the traditional intensity distribution. This new weighted distribution method considers the frequency of grayscale levels, effectively enhancing local image details. For the background sub-histogram, strategies such as recalculating the mapping range, increasing contrast, and controlling noise ensure the ratio of the background interval to the entire dynamic range, preventing excessive background noise enhancement and improving image clarity and visual experience. In the final stage of image enhancement, considering the characteristics of the human eye, a visual correction factor is used to adjust the size of the foreground histogram and reduce noise interference. A particle swarm optimization algorithm is employed to correct the average brightness of the enhanced image based on the reference image, ensuring that the overall brightness of the output image is moderate and natural.

[0059] To ensure the accuracy of image fusion, the TWMM method is used to strictly register the visible light and infrared images.

[0060] TWMM is an automated registration method for thermal infrared and visible light images of UAVs, achieving bimodal image registration through the following four steps. First, saliency detection and candidate block generation methods are used to select several small patches (atomic patches) in the image to be registered, which have the richest texture and most significant features. Then, the similarity map corresponding to each atomic patch is calculated. By extracting image features and using weighted template matching, the similarity map of all atomic patches is calculated. Template matching involves creating a template based on the target (in this invention, the UAV target in the infrared image is used as the benchmark), and searching in the image to be registered (in this invention, the visible light image) for the target most similar to the template image and the region whose mean or variance is closest to the template. Template matching not only considers the pixel values ​​within the atomic patches but also improves the algorithm's focus on key features by introducing specific region weights. Subsequently, multi-level local max pooling is used to generate pyramid similarity maps of different patch sizes. Larger patches have the ability to acquire global information, while smaller patches focus on local details. This hierarchical construction strategy enables the model to cope with local feature distortions. The TWMM method utilizes the maximum index backtracking algorithm, starting from the top level of the pyramid similarity map and deriving corresponding points level by level. Maximum index backtracking, by combining similarity maps at all levels, solves the problems of poor small patch discrimination and inaccurate large patch localization, enabling TWMM to obtain a sufficient number of accurate corresponding points. Finally, thresholding and distance metrics are used to filter out unqualified corresponding points, and then the homography between infrared and visible light images is calculated, ensuring the accuracy and stability of the final result.

[0061] like Figure 1 As shown, S102: Based on the dataset, train the DC GMA-YOLOv10 model. This model is based on the YOLOv10 model. The feature extraction network of the YOLOv10 model is divided into two branches, and an improved group hybrid attention module DC GMA is added to each branch to extract features from infrared and visible light images. A cross-modal differential sensing fusion module CMDAF is also introduced. An information augmentation sampling module MSFS is introduced into the neck network to replace the original upsampling module in YOLOv10.

[0062] Specifically, YOLOv10, an object detection model proposed in 2014, further improves the performance-efficiency boundary of the YOLO model in terms of both post-processing and model architecture. YOLOv10 is the first to propose a NMS-free training strategy, eliminating the NMS training dependency of the YOLO series through dual label allocation and consistent matching metrics. This improves the model's inference speed and solves the problem of missed detections caused by NMS errors suppressing low-confidence ground truth bounding boxes in complex environments. The dual label allocation strategy combines commonly used one-to-many and one-to-one allocation. During training, both allocation strategies are used simultaneously: one-to-many allocation assigns multiple positive samples to each ground truth bounding box, providing rich supervision information; simultaneously, one-to-one allocation assigns only one positive sample to each ground truth bounding box, ensuring the model can select the best predicted bounding box. During inference, only one-to-one allocation is used, ensuring that each target has only one positive sample, avoiding NMS post-processing. The consistent matching metric ensures that the two branches remain consistent during training, ensuring consistency between one-to-one and one-to-many allocations in the optimization direction, further improving model performance.

[0063] YOLOv10 employs spatial-channel separation downsampling to maximize information retention during the downsampling process while reducing computational cost and parameter count. It simplifies the computational burden of classification tasks by using lightweight classification heads, focusing on optimizing regression heads that significantly improve detection performance. By leveraging the redundancy in each stage of the intrinsic rank analysis model, it replaces redundant stages with the proposed Compact Inverted Block (CIB) structure, achieving higher detection efficiency without compromising performance. Furthermore, by using large-kernel convolutions in the CIBs of deeper stages, it avoids shallow feature contamination, expands the receptive field, and enhances the model's ability to capture complex scenes. An efficient partial self-attention module is designed to introduce global representation learning capabilities into the model at low computational cost, improving detection accuracy.

[0064] like Figure 2 To further enhance the model's feature extraction capabilities, an improved group hybrid attention module (DC GMA) was added to the two-way YOLOv10 feature extraction network. This allows the feature extraction network to extract features from individual pixels while also focusing on the correlation between individual pixels and groups of pixels, thus achieving higher feature representation capabilities. To further reduce parameters and maintain inference speed, the ordinary convolutions in the hybrid attention module were replaced with dilated convolutions, forming the DC GMA module.

[0065] Group Mix Attention (GMA) differs from traditional attention mechanisms in that it can simultaneously capture the correlations between tokens, between token groups, and between token groups.

[0066] For feature processing requirements of specific tasks, traditional convolutional layers, through local connectivity mechanisms, can effectively capture the spatial hierarchical structure of input data. However, some application scenarios still require the network to have directional focusing capabilities to achieve deep parsing of key regions and enhance semantic associations. Let X∈R N×d Given the input tokens, N is the number of tokens, and d is the dimension, the output Y of a standard self-attention system is:

[0067] Y = softmax(XX) T )X

[0068] By using the definition of matrix multiplication, and utilizing XX T The correlation between each token pair is calculated, and the output A∈R of the softmax function is... N×N For attention graphs, multiplication AX means linearly recombinating token groups based on the attention graph at each position. However, this self-attention only considers the correlation between tokens within a single pattern and cannot perform cross-pattern processing across different pattern groups.

[0069] The Group Hybrid Attention (GMA) mechanism simulates a group pattern, uniformly splitting the Query, Key, and Value into multiple segments and performing different group aggregations to generate group proxies. Unlike traditional attention mechanisms, GMA computes the attention graph based on tokens and group proxies, and then recombines tokens and token groups in the Value based on this. Specifically, Q, K, and V are evenly divided into n segments, using X (i∈[1,…,n]) to represent a segment (X can be Q, K, and V), and using Aggregators... i (X i Aggregation is represented by ) . When performing attention calculations, Aggregate is used. i (X i The groups i ∈ [1,…,n] are concatenated to produce X′, thus obtaining the group agents Q′, K′, and V′. Attention is then calculated on the group agents to obtain the output.

[0070] The Group Hybrid Attention (GMA) mechanism enhances the feature extraction capability of attention mechanisms without reducing spatial resolution. GMA uses depthwise convolutions with kernel sizes of 3×3, 5×5, and 7×7 to implement the aggregator, while associating K×K tokens (K represents the kernel size), ensuring more complete and comprehensive associations between models.

[0071] Dilated convolutions (DCs) are a commonly used convolutional operation in convolutional neural networks. By setting a dilation rate parameter, and keeping the kernel size constant, DCs expand the feature-receptive region through the spacing between kernel elements, thereby improving the network's ability to understand input data. Compared to ordinary convolutions with the same receptive field, the introduction of a dilation factor reduces the number of parameters while maintaining the resolution of the feature map.

[0072] Let the kernel size of dilated convolution be k, and the dilation factor be d. Then, the formula for calculating the equivalent kernel size k' is as follows:

[0073] k′=k+(k-1)(d-1)

[0074] The formula for calculating the receptive field of layer i+1 is as follows:

[0075] RF i+1 =RF i +(k′-1)×S i

[0076]

[0077] Among them, RF i+1 Represents the receptive field of the (i+1)th layer, RF i Let S represent the receptive field of the (i+1)th layer, k′ represent the size of the convolution kernel, and S i This represents the product of all previous step sizes (excluding the current layer).

[0078] Furthermore, based on the improved YOLOv10 model, the DC GMA-YOLOv10 model introduces a cross-modal differential sensing fusion module (CMDAF) between the dual-modal feature extraction networks of the DC GMA-YOLOv10 model.

[0079] like Figure 3 The cross-modal differential sensing fusion module CMDAF exchanges complementary features between infrared and visible light images in a progressive manner, reducing the differences between the two modal features and enhancing feature interaction between infrared and visible light images during the feature extraction stage, thereby improving the model's detection accuracy.

[0080] The definition of cross-modality differential aware fusion (CMDAF) is as follows:

[0081]

[0082] in, and The final output of the feature extraction part is the bimodal feature, namely the fused infrared image feature and the visible light image feature. δ(·) is the Sigmoid function, and GAP is the global average pooling. and Defined as complementary or common features of infrared and visible light images, as follows:

[0083]

[0084] The extracted infrared and visible light image features are processed to obtain complementary features representing the differences between the two modes. These features are then compressed into vectors using average pooling and processed by the Sigmoid function to generate channel weights. The complementary features are multiplied by the obtained channel weights, and the result is added to the original features as modality supplementary information.

[0085] Furthermore, based on the improved YOLOv10 model, the DGM-YOLOv10 model introduces the Information Augmentation Sampling Module (MSFS) into the neck network of the DGM-YOLOv10 model.

[0086] like Figure 4 The Information Augmentation Sampling Module (MSFS) uses pixel rearrangement and local perception operations to enable the network model to achieve upsampling while also possessing detail enhancement and context fusion capabilities, further improving the model's accuracy in identifying UAV targets.

[0087] The Multi-scale Feature Enhancement Sampling (MSFS) module is defined as follows:

[0088] M = Conv(Concat(LU,UL))

[0089] Where M is the feature map obtained after processing by the Information Augmentation Sampling (MSFS) module, i.e., an upsampled feature map containing local detail enhancement information and global context information; Conv is the convolution operation; Concat is the concatenation operation; and LU and UL are defined as the feature map containing local detail enhancement information and the feature map containing global context information, respectively, as follows:

[0090] LU = Upsample(Linear(N))

[0091] UL=Linear(Upsample(Shuffle(N)))

[0092] Where N represents the fused infrared and visible light feature maps to be processed, Upsample represents bilinear interpolation sampling, Linear represents offset adjustment, and Shuffle represents pixel rearrangement. By adjusting the offset of the fused infrared and visible light feature maps and then performing bilinear interpolation sampling on the feature maps after offset adjustment, a feature map LU containing local detail enhancement information is obtained. By performing pixel rearrangement on the fused infrared and visible light feature maps and then sequentially performing bilinear interpolation sampling and offset adjustment, a feature map UL containing global context information is obtained.

[0093] The specific method by which the information enhancement sampling module MSFS performs upsampling operation is as follows:

[0094] (1) Obtain the infrared and visible light fused feature map to be processed;

[0095] (2) The offset of the infrared and visible light fused feature map is adjusted, and after the offset is adjusted, the feature map is sampled by bilinear interpolation to obtain a feature map containing local detail enhancement information.

[0096] (3) Perform pixel rearrangement on the infrared and visible light fused feature map, and sequentially perform bilinear interpolation sampling and offset adjustment to obtain a feature map containing global context information;

[0097] (4) Fuse the two sets of feature maps, which contain local detail enhancement information and global context information respectively;

[0098] (5) Output the fused feature map;

[0099] Multiple information enhancement sampling modules (MSFS) are used for different feature levels, that is, the above process (1)-(5) is repeated multiple times to complete the upsampling operation.

[0100] In this embodiment, the training process for the DC GMA-YOLOv10 model is as follows:

[0101] The DC GMA-YOLOv10 model was trained using a server. The training environment consisted of Python 3.9, the deep learning framework PyTorch 2.0.1, CDUA 10.8, the operating system Ubuntu 20.04, 64GB of RAM, and an NVIDIA GeForce RTX4090 graphics card with 24GB of video memory.

[0102] The prepared dataset is divided into training, validation, and test sets, and corresponding label text files are generated, including the bounding box coordinates and class labels of the targets. In each training iteration, a batch of infrared and visible light images is randomly selected from the training set, and these images are input in pairs into the DC GMA-YOLOv10 model, and the following operations are performed:

[0103] First, the parameters of the YOLOv10 backbone network and the improved group hybrid attention module (DC GMA) are initialized, and pre-trained weights are loaded using transfer learning to accelerate convergence. Then, infrared and visible light images are input into the network's dual-modal feature extraction channels, respectively.

[0104] During the forward propagation, the infrared and visible light channels respectively pass through a group hybrid attention module (DC GMA) containing dilated convolutions to extract deep semantic features, capturing the spatial structure and target features within each modality. Then, the cross-modal differential perception fusion module (CMDAF) performs interactive feature fusion. An enhanced sampling module (MSFS) is introduced into the neck network to improve the model's ability to express the fused infrared and visible light features while achieving accurate upsampling. Finally, the feature information is fed into the detection head of the YOLOv10 model to generate target classification probabilities, bounding box coordinates, and confidence scores. The difference between the output results and the ground truth annotations is calculated to construct the loss function.

[0105] The process then proceeds to the backpropagation phase, where the chain rule is used to propagate error information forward from the self-detector. The gradients of each convolutional layer and fusion module are calculated layer by layer, and the network parameters are updated using stochastic gradient descent. These steps are repeated until the loss function converges.

[0106] Repeat the above steps, processing a new batch each time, until the entire training set has been traversed.

[0107] like Figure 1 As shown, S103: Based on the trained DC GMA-YOLOv10 model, input pairs of infrared and visible light images for detection.

[0108] In this embodiment, the registered infrared image and the visible light image to be detected are fed into the dual-branch feature extraction network of the model. Through the improved group hybrid attention module DC GMA, the modal features of the infrared and visible light images are extracted, capturing the structural features and semantic information that differ between the two modalities.

[0109] After feature extraction is completed in each branch, the output multi-scale feature map is sent to the cross-modal differential sensing fusion module CMDAF. The CMDAF module adopts a stepwise fusion strategy. Through feature difference calculation and adaptive information guidance mechanism, it alternately guides the feature representations of infrared and visible light modes to perform difference enhancement and complementary fusion in channel and spatial dimensions to obtain fused features, reduce redundancy and bias between modes, and improve the robustness and accuracy of fused features.

[0110] Information Augmentation Sampling (MSFS) is introduced into the neck network to replace the original upsampling module in YOLOv10. MSFS, through pixel rearrangement, offset adjustment, and feature sampling operations, enables the two processing branches of the network model to achieve upsampling while simultaneously possessing detail enhancement and context fusion capabilities. The two processing branches are then fused, reducing feature information loss during upsampling and improving the model's expressive power.

[0111] The obtained fused features are input into the YOLOv10 detection head. By introducing a task-decoupling loss function and a dynamic label assignment strategy, the target's category and bounding box coordinates are generated, resulting in the final detection result.

[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the present invention.

Claims

1. A method for UAV target detection based on DC GMA-YOLOv10 infrared and visible light fusion, characterized in that, include: Collect and create datasets of infrared and visible light UAV images of UAV targets in complex environments; A DGM-YOLOv10 model is constructed and trained based on the image dataset. The DGM-YOLOv10 model includes an improved group hybrid attention module (DC GMA) and a YOLOv10 model. The improved DC GMA replaces the ordinary convolution in the group hybrid attention mechanism (GMA) with dilated convolution. The feature extraction network of the YOLOv10 model is modified to have two branches. The improved DC GMA is introduced before the second convolutional layer of each branch to obtain a dual-modal feature extraction network for feature extraction of infrared and visible light images. A cross-modal differential sensing fusion module (CMDAF) is introduced between the dual-modal feature extraction networks of the DGM-YOLOv10 model. This module progressively exchanges complementary features between infrared and visible light images, reducing feature differences between the two modalities. Information Augmentation Sampling Module (MSFS) is introduced between the neck networks of the DGM-YOLOv10 model to perform upsampling operations, replacing the original upsampling module in YOLOv10. Based on the trained DGM-YOLOv10 model, pairs of visible light and infrared UAV images are input for detection.

2. The UAV target detection method based on DC GMA-YOLOv10 infrared and visible light fusion according to claim 1, characterized in that, The specific method for collecting and creating infrared and visible light UAV image datasets of UAV targets in complex environments is as follows: First, the infrared and visible light cameras of the photoelectric turntable are used to capture drone flight videos under different weather and environments; the two modes of video are extracted frame by frame to obtain continuous images, and then the continuous images are sampled at fixed intervals to ensure the difference between the images; Filter the sampled images; Adaptive histogram partitioning and brightness correction are used to enhance the images in the infrared dataset to remove image noise and improve image details; The TWMM method was used to register visible light and infrared images to form a dataset of infrared and visible light UAV flight images.

3. The UAV target detection method based on DC GMA-YOLOv10 infrared and visible light fusion according to claim 2, characterized in that, The specific method for image enhancement of infrared dataset images using adaptive histogram partitioning and brightness correction is as follows: For adaptive segmentation of infrared grayscale histograms, Gaussian filtering and local weighted scatter smoothing algorithms are used to smooth the original grayscale histograms and remove peaks in the histograms. The histogram is divided into multiple sub-histograms based on local minima. For each segment of the histogram, the gray density metric is used for discrimination. Low gray density values ​​correspond to foreground and high gray density values ​​correspond to background. An adaptive threshold is used to further classify each interval as foreground or background, maximizing the inter-class variance to optimize the classification effect. For the foreground sub-histogram, a local contrast-weighted distribution is used to replace the traditional intensity distribution, thereby enhancing the local details of the image. For the background sub-histogram, the mapping range is recalculated, contrast is improved, and noise is controlled to reduce background noise and prevent excessive enhancement of background noise; A visual correction factor is used to adjust the size of the foreground histogram and reduce noise interference. The particle swarm optimization algorithm is used to correct the average brightness of the enhanced image based on the reference image, so as to ensure that the overall brightness of the output image is moderate and natural.

4. The UAV target detection method based on DC GMA-YOLOv10 infrared and visible light fusion according to claim 1, characterized in that, The specific method by which the cross-modal differential sensing fusion module CMDAF progressively exchanges complementary features between infrared and visible light modal images is as follows: (a) Acquire infrared and visible light features of the same size and number of channels; (b) Perform channel alignment and normalization on the two features; (c) Complementary features representing the differences between the two modes are obtained through difference calculation; (d) Perform global average pooling on complementary features to generate channel weights and spatial weights respectively; (e) Multiply each feature by the obtained channel weight and spatial weight to obtain the infrared and visible light fused features; (f) Then, the infrared and visible light fused features are added to the original mode branches respectively; Multiple cross-modal differential sensing fusion modules (CMDAF) are used for different feature levels, that is, the above process (a)-(f) is repeated multiple times to complete the feature exchange between the two modalities.

5. The UAV target detection method based on DC GMA-YOLOv10 infrared and visible light fusion according to claim 1, characterized in that, The specific method by which the information enhancement sampling module MSFS performs upsampling operation is as follows: (1) Obtain the infrared and visible light fused feature map to be processed; (2) The offset of the infrared and visible light fused feature map is adjusted, and after the offset is adjusted, the feature map is sampled by bilinear interpolation to obtain a feature map containing local detail enhancement information. (3) Perform pixel rearrangement on the infrared and visible light fused feature map, and perform bilinear interpolation sampling and offset adjustment in sequence to obtain a feature map containing global context information; (4) Fuse the two sets of feature maps, which contain local detail enhancement information and global context information respectively; (5) Output the fused feature map; Multiple information enhancement sampling modules (MSFS) are used for different feature levels, that is, the above process (1)-(5) is repeated multiple times to complete the upsampling operation.

6. The UAV target detection method based on DC GMA-YOLOv10 infrared and visible light fusion according to claim 1, characterized in that, The process of training the DGM-YOLOv10 model is as follows: The prepared image dataset is divided into training, validation, and test sets. The drones are annotated in YOLO format to generate corresponding label text files, including the bounding box coordinates and class labels of the targets. In each training iteration, a batch of infrared and visible light images is randomly selected from the training set, and these images are input in pairs into the DGM-YOLOv10 model. The following operations are performed: Step S1: First, initialize the parameters of the YOLOv10 model backbone network and the improved group hybrid attention module DC GMA, and load the pre-trained weights using transfer learning to accelerate the convergence speed; then, input the infrared image and the visible light image into the dual-branch feature extraction channel of the YOLOv10 model backbone network respectively. Step S2: During the forward propagation process, the infrared channel and the visible light channel respectively pass through the group hybrid attention module DC GMA containing dilated convolution to extract deep semantic features, and each captures the spatial structure and target features within the modality. The features are then interactively fused using the cross-modal differential sensing fusion module CMDAF, and the fused features are then fed into the information enhancement sampling module MSFS for upsampling processing to further enhance the expressive power of the features. Finally, the feature map is fed into the detection head of the YOLOv10 model to generate the target classification probability, bounding box coordinates, and confidence score. The difference between the output results and the true annotations is calculated to construct the loss function. Step S3: Enter the backpropagation process, use the chain rule to propagate error information forward from the self-detector, calculate the gradients of each convolutional layer, the cross-modal differential sensing fusion module CMDAF and the information augmentation sampling module MSFS layer by layer, and update the parameters in the network through the stochastic gradient descent algorithm; repeat the above forward and backpropagation process until the loss function converges; Repeat steps S1-S3 above, processing a new batch each time, until the entire training set has been traversed.

7. The UAV target detection method based on DC GMA-YOLOv10 infrared and visible light fusion according to claim 6, characterized in that, The specific method for detecting pairs of visible light and infrared UAV images based on the trained DGM-YOLOv10 model is as follows: The registered infrared image and the visible light image to be detected are fed into the dual-modal feature extraction network of the DGM-YOLOv10 model; the modal features of the infrared and visible light images are extracted through the improved group hybrid attention module DC GMA, capturing the structural features and semantic information that are different in the two modalities; After feature extraction is completed in each branch, the output multi-scale feature map is sent to the cross-modal differential sensing fusion module CMDAF. The cross-modal differential sensing fusion module CMDAF adopts a stepwise fusion strategy. Through feature difference calculation and adaptive information guidance mechanism, it alternately guides the feature representations of infrared and visible light modes to perform difference enhancement and complementary fusion in channel and spatial dimensions to obtain fused features, thereby reducing redundancy and bias between modes. The fused features are fed into the Information Enhancement Sampling Module (MSFS) for upsampling processing; The Information Augmentation Sampling Module (MSFS) obtains feature maps containing local detail enhancement information and global context information by combining pixel rearrangement and local perception operations. It then upsamples the feature maps and fuses them to further enhance the expressive power of the features and improve spatial resolution. The obtained fused features are input into the YOLOv10 detection head. By introducing a task decoupling loss function and a dynamic label assignment strategy, the target category and bounding box coordinates are generated to obtain the final detection result.

Citation Information

Cited By

  • Weak target detection method based on infrared and visible light data fusion and related equipment

    CN121213899A

  • Wild animal detection method fusing unmanned aerial vehicle thermal infrared image and visible light image

    CN121305621A

  • Railway scene target detection and behavior identification method based on space-time double-flow characteristics

    CN121354053A

  • A railway scene target detection and behavior recognition method based on spatiotemporal dual-stream features

    CN121354053B

  • Underground coal mine personnel behavior detection method and system based on cross-modal fusion

    CN121838046A