All-weather unmanned aerial vehicle target detection method and system based on reflectivity guidance

Through a reflectivity-guided progressive feature alignment network (RGFNet), reflectivity features are used to align infrared and visible light features to solve the problem of performance degradation of drone target detection under different lighting conditions, and achieve efficient fusion and accuracy of all-weather target detection.

CN120298934APending Publication Date: 2025-07-11ANHUI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510469355.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing drone target detection methods have deteriorated performance under different lighting conditions, especially in extremely dark environments and uneven lighting conditions, making it difficult to effectively fuse the complementary information of visible light and infrared images, resulting in inaccurate feature extraction and high noise, affecting the effect of cross-modal fusion.

Method used

A reflectivity-guided progressive feature alignment network (RGFNet) is used to align infrared features through reflectivity features, and feature interactions are performed in the hidden state space to generate complementary features, and finally fuse them in the detection network to generate all-weather object detection results.

Benefits of technology

The robustness and accuracy of drone target detection under different lighting conditions are improved, the effectiveness of feature fusion is enhanced, and better detection performance is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298934A_ABST
    Figure CN120298934A_ABST
Patent Text Reader

Abstract

The invention discloses an all-weather unmanned aerial vehicle target detection method and system based on reflectivity guidance. The method comprises the following steps: extracting an RGB feature set and an IR feature set from input data; using a reflectivity characteristic with a dark light resistance characteristic as a reference characteristic to align the infrared characteristic, and then using the aligned infrared characteristic as the reference characteristic to further align the reflectivity and the visible light characteristic; the aligned visible light features, the aligned infrared features and the aligned reflectivity features are input into an illumination perception selective fusion module, feature interaction is carried out in a hidden state space, and corresponding complementary features are generated; and adding the alignment characteristic and the complementary characteristic to obtain an enhanced characteristic, then directly adding the enhanced characteristics of the RGB and IR modes to obtain a fusion characteristic, and sending the fusion characteristics of the last several stages to the neck and the head of the detection network to generate a final detection result. According to the invention, effective cross-modal fusion is realized to adapt to different degrees of darkness and illumination changes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and image processing, and particularly to an all-weather UAV target detection method based on reflectivity-guided progressive feature alignment. Background Art

[0002] Unmanned aerial vehicles (UAVs) have unique advantages such as being small, having a wide field of view, being flexible to operate, and having low safety risks. UAV-based target detection plays an important role in road vehicle monitoring, intelligent traffic management, search and rescue missions, and border monitoring. However, current algorithms mainly rely on visible light images containing rich texture and color information. This results in a sharp decline in performance when encountering different lighting conditions, such as extremely dark environments, uneven light distribution, overexposure, and varying degrees of darkness at night, essentially limiting the all-weather application of UAV detection systems.

[0003] To achieve all-weather practical applications, the use of multiple sensors on UAV platforms has received extensive attention from researchers. Specifically, infrared imaging can stably capture thermal information and object contours without being affected by visual conditions, effectively complementing the disadvantages of visible light images. Therefore, efficiently integrating complementary information from RGB (visible light) and IR (infrared) modalities can significantly enhance the robustness of the model and contribute to all-weather target detection in remote sensing scenarios. However, the widespread spatial misalignment between visible light-infrared image pairs severely degrades the performance of traditional cross-modal detection methods. This problem is particularly challenging under low-light conditions. At this time, the quality of visible light images drops sharply, resulting in severe degradation of target information. Although infrared images can provide supplementary clues, they cannot fully compensate for the appearance distortion or information loss in visible light images, which increases the difficulty of visible light-infrared image alignment and further weakens the effectiveness of cross-modal feature fusion. Therefore, traditional cross-modal target detection methods such as using UAVs for multi-modal object recognition, vehicle detection based on adaptive multi-modal feature fusion and cross-modal vehicle indexing (using RGB-T images), and precise 3D object recognition through multi-modal fusion from local to global (Logonet) still have significant performance degradation problems under challenging night conditions.

[0004] To address these challenges, recent research has improved feature consistency and detection performance by introducing an alignment module to predict and correct positional misalignment. However, this highly depends on selecting a single high-quality modality as a reference and cannot fully utilize the complementary information from RGB and IR modalities. Some methods, such as multi-spectral pedestrian detection using simultaneous detection and segmentation techniques, cross-modal interactive attention network for multi-spectral pedestrian detection, different modality learning (DML) for building semantic segmentation, etc., attempt to achieve cross-modal fusion through weighted summation based on feature connection or illumination-aware importance scoring. However, this simple fusion strategy cannot effectively utilize the rich complementary information between modalities.

[0005] Recent methods, such as cross-modal fusion transformers for multi-spectral object detection, unified transformer framework for object detection and segmentation (MaskDINO), calibration and complementary transformers (C2Former) for RGB-infrared object detection, etc., use transformers to achieve interactive feature fusion through cross-modal attention mechanisms. However, these methods overlook several key issues under different lighting conditions. In extremely dark conditions, there is severe information loss in visible images, making it almost impossible to extract RGB features. In weak or uneven light conditions, visible images exhibit inconsistent quality between different regions. Some regions may contain useful details while others are completely degraded, posing significant challenges to feature extraction. Even in moderately lit environments, the coexistence of overexposed and underexposed regions further complicates the problem. This lighting variation makes RGB feature extraction highly unreliable and incomplete, severely affecting the quality of the fused features generated by baseline methods. Under these complex lighting conditions, inaccurate feature extraction introduces a large amount of noise and inconsistency, making it extremely challenging to achieve effective cross-modal fusion to adapt to different degrees of darkness and lighting variations. Summary of the Invention

[0006] The technical problem to be solved by the present invention is how to achieve effective cross-modal fusion to adapt to different degrees of darkness and lighting variations.

[0007] The present invention solves the above technical problem by the following technical means: an all-weather unmanned aerial vehicle target detection method based on reflectivity guidance, including:

[0008] S1. Extract an RGB feature set and an IR feature set from the input data;

[0009] S2. Use the reflectivity feature with anti-dark-light characteristics as a reference feature to align the infrared features, and then use the aligned infrared features as a reference feature to further align the reflectivity and visible light features;

[0010] S3. Input the aligned visible light features, aligned infrared features, and aligned reflectivity features into the light perception selective fusion module, and perform feature interaction in the hidden state space to generate corresponding complementary features;

[0011] S4. Add the aligned features and complementary features to obtain enhanced features. Then, directly add the enhanced features of the RGB and IR modalities to get the fused features. Send the fused features of the last few stages to the neck and head of the detection network to generate the final detection results.

[0012] As a further optimized technical solution, step S1 specifically includes: First, use the pre-trained Retinex decomposition network to decompose the reflectivity feature F F and the keyword K indicating whether it is dark from the visible image. Then, use two DarkNet feature extraction networks with the same structure but different parameters to extract the RGB feature sets F R with spatial dimensions of 1 / 2, 1 / 4, ···, 1 / 32 and the infrared modality feature F I from the input data respectively.

[0013] As a further optimized technical solution, using the reflectivity feature with anti-low-light characteristics as the reference feature to align the infrared features in step S2 specifically includes: First, use the reflectivity feature F F as the reference feature and the infrared modality feature F I as the perception feature to perform the offset-guided deformable convolution operation. This process starts from the channel connection of these two features, then convolves to generate the initial offset feature, and then performs channel connection with the offset feature from the previous stage, and convolves again to generate the final offset feature. Finally, process the offset feature to output the offset Δn and the modulation scalar Δm. The offset Δn and the modulation scalar Δm perform deformable convolution with the perception feature to obtain the aligned infrared feature

[0014] As a further optimized technical solution, the first offset feature is directly obtained by performing one convolution on the concatenated perception feature and reference feature.

[0015] As a further optimized technical solution, using the aligned infrared feature as the reference feature to further align the reflectivity and visible light features in step S2 specifically includes: Using the aligned infrared feature as the reference feature and the reflectivity feature F F as the perception feature to perform the offset-guided deformable convolution operation. First, concatenate the two features along the channel dimension, then convolve to generate the initial offset feature, perform channel connection with the offset feature from the previous stage, convolve to generate the final offset feature, and then the reflectivity feature F FPerform deformable convolution with the final offset feature to obtain aligned reflectance features

[0016] Next, use the infrared features obtained previously as reference features, and directly perform deformable convolution with the final offset feature obtained using the reflectance feature F F as the perceptual feature and the visible light feature F R to obtain a new visible light feature. Then, sum the aligned reflectance features and the new visible light feature to obtain aligned visible light features

[0017] As a further optimized technical solution, step S3 specifically includes: First, concatenate the aligned visible light features aligned infrared features aligned reflectance features along the channel dimension to obtain an initially concatenated multimodal feature Multimodal feature After layer normalization, pass through the visual state space module to obtain processed intermediate features Intermediate features After layer normalization, pass through a multi-layer perceptron to obtain the final multimodal fusion feature Use different learnable scaling factors s1 and s2 before and after the visual state space module and the multi-layer perceptron to control the information of the skip connection. Finally, separate the final multimodal fusion feature along the channel dimension to obtain single-modal complementary features

[0018] As a further optimized technical solution, the process of passing the multimodal feature through the visual state space module after layer normalization to obtain processed intermediate features is as follows:

[0019] Multimodal feature First, after layer normalization, it is separated into x R , x I , x F for separate processing. Then, pass through a linear layer, a depthwise separable convolution, and the SiLU function respectively to obtain modal features At the same time, the multimodal feature also passes through a linear layer and the SiLU function respectively to obtain modal features x′ R , x′ I , x′ F ;

[0020] Subsequently, the processed modal features It is input into a shared SS2D for feature enhancement and fusion to obtain enhanced features

[0021] Enhanced features After layer normalization respectively, they are multiplied by the modal features x′ R 、x′ I 、x′ F and then linearized to obtain the modal features x″ R 、x″ I 、x″ F , and the modal features x″ R 、x″ I 、x″ F are concatenated along the channel dimension to obtain intermediate features

[0022] As a further optimized technical solution, in the SS2D, the input modal features are first processed by a linear layer to generate the state matrices: Δ R , B R , C R , Δ I , B I , C I , and Δ F , B F , C F . At the same time, the parameter matrices are initialized as: A R , D R , A I , D I , A F, D F . The continuous matrices A and B need to be discretized before being applied to the state space equation, where Δ is closely related to the discretization process of the state space model. Considering the unreliability of the visible light modality under low-light conditions, the keyword K obtained from the Retinex decomposition network is used to improve Δ R as follows:

[0023] Δ x =Δ R ×K + Δ F ,

[0024] where the value of K is 0 or 1, representing the dark scene and the bright scene respectively;

[0025] The discretization process of the continuous matrices A and B is expressed as:

[0026]

[0027] and Map their respective modal features to the hidden state hk ;

[0028] The mapping process is expressed as:

[0029]

[0030] The state matrix C is used to recover the output from the hidden state at each time step. To enable each modality to decode more useful information, C R, C I and C F are added together to obtain a global perception matrix C x , which is used to decode the hidden states of all three modalities, thereby generating a selective scan output and

[0031] As a further optimized technical solution, step S4 specifically includes: adding the aligned visible light features aligned infrared features aligned reflectivity features obtained in step S2 and the visible light complementary features infrared light complementary features reflectivity complementary features obtained in step S3 to obtain enhanced features. Then, the infrared-enhanced features and the visible light-enhanced features are added together to obtain the fusion P i . The fusion features P3, P4, and P5 of the last three stages are sent to the neck and head of the detection network YOLOv8 as inputs to generate the final detection results.

[0032] The present invention also provides an all-weather unmanned aerial vehicle target detection system based on reflectivity guidance, including:

[0033] An extraction module for extracting an RGB feature set and an IR feature set from input data;

[0034] An alignment module for using the reflectivity feature with anti-low-light characteristics as a reference feature to align the infrared features, and then using the aligned infrared features as a reference feature to further align the reflectivity and visible light features;

[0035] A complementary feature generation module for inputting the aligned visible light features aligned infrared features aligned reflectivity features into the light perception selective fusion module and performing feature interaction in the hidden state space to generate corresponding complementary features;

[0036] The detection module is used to add the alignment feature and the complementary feature to obtain an enhanced feature. Then, the enhanced features of the RGB and IR modalities are directly added to obtain the fused feature P. i The fused features of the last few stages are sent to the neck and head of the detection network to generate the final detection result.

[0037] The advantages of the present invention are as follows:

[0038] 1. The present invention proposes a new reflectance-guided progressive feature alignment network (RGFNet), a framework for infrared-visible remote sensing target detection. The proposed RGFNet includes an RCAM (reflectance collaborative alignment module), which uses the illumination invariance of reflectance to guide the bidirectional feature registration between the visible and infrared modalities, effectively reducing the modality differences under different illumination conditions. In addition, an LSFM (light perception selective fusion module) is designed to map the multi-modal features to a shared hidden state space through a selective state space mechanism, achieving effective feature interaction while maintaining a linear computational complexity.

[0039] 2. Extensive experiments conducted on two challenging datasets show that the present invention has superior performance compared to the current state-of-the-art visible-infrared target detection methods. Description of the Drawings

[0040] Figure 1 is a flowchart of the UAV target detection method according to an embodiment of the present invention;

[0041] Figure 2 is an overall framework diagram of RGFNet according to an embodiment of the present invention;

[0042] Figure 3 is a schematic structural diagram of RCAM according to an embodiment of the present invention;

[0043] Figure 4 is a schematic structural diagram of LSFM according to an embodiment of the present invention;

[0044] Figure 5 is a schematic structural diagram of the visual state space module (VSSM) according to an embodiment of the present invention;

[0045] Figure 6 is a schematic structural diagram of the 2D selective scan (SS2D) according to an embodiment of the present invention. Detailed Embodiments

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0047] The all-weather UAV target detection method of the present invention adopts a Reflectance Guided Progressive Feature Alignment Network (RGFNet), which mainly includes two key modules:

[0048] (1) A Reflectance-Guided Collaborative Alignment Module (RCAM) that uses the illumination-invariant characteristics of reflectance features to guide cross-modal feature alignment. This module adopts a two-way calibration strategy, in which reflectance features first guide the alignment of infrared features, and then the aligned infrared features are used to calibrate RGB features. This two-way process effectively reduces the position differences under different illumination conditions.

[0049] (2) A Light-Aware Selective Fusion Module (LSFM) that maps the aligned modalities to a shared hidden state space through a selective state space mechanism, further enhancing cross-modal interaction and achieving efficient adaptive feature fusion while maintaining linear computational complexity. This architecture enables the model of the present invention to maintain reliable detection performance under different illumination conditions.

[0050] Embodiment 1

[0051] Please refer to Figure 1 , the UAV target detection method of the present invention includes the following steps:

[0052] S1. Extract an RGB feature set and an IR feature set from the input data.

[0053] Refer to Figure 2 , first use a pre-trained Retinex decomposition network to decompose the reflectance feature F F and the keyword K indicating whether it is dark from the visible image. Then use two DarkNet feature extraction networks with the same structure but different parameters to extract RGB feature sets F with spatial dimensions of 1 / 2, 1 / 4, ···, 1 / 32 from the input data respectively Rand the infrared modal feature F I . These feature sets contain the key information of the image at different resolutions.

[0054] S2. Use the reflectivity feature with anti - dark - light characteristics as the reference feature to align the infrared feature, and then use the aligned infrared feature as the reference feature to further align the reflectivity feature and the visible - light feature.

[0055] See also Figure 3 As shown, RCAM uses modulated deformable convolution to bidirectionally align cross - modal features. First, use the reflectivity feature F F as the reference feature and the infrared modal feature F I as the perceptual feature to perform the offset - guided deformable convolution operation. This process starts from the channel concatenation of these two features, then convolution generates the initial offset feature, and then channel - concatenates with the offset feature from the previous - stage RCAM, and convolution generates the final offset feature. Note that for the first offset feature, it is directly obtained by convolving the concatenated perceptual feature and reference feature once. Finally, process the offset feature to output the offset Δn and the modulation scalar Δm. The offset Δn and the modulation scalar Δm perform deformable convolution with the perceptual feature to obtain the aligned infrared feature Aligned infrared feature is calculated as follows:

[0056] {Δn,Δm} = C off ([F F ,F I ),

[0057]

[0058] where [·,·] represents the concatenation operation along the channel dimension. C off (·) and C def (·) represent offset convolution and deformable convolution respectively. Specifically, each position n on the aligned infrared feature is calculated as:

[0059]

[0060] where K, n k and represent the number of kernel weights, the k - th fixed offset, and the k - th kernel weight respectively, Δn k is the offset at the k - th position obtained by learning, n k ∈{(-1, 1), (-1, 0),..., (1, 1)} is a regular grid of a 3×3 kernel defined by K = 9, and Δm within the range [0,1] corresponds to each n kThe contribution. Convolution is performed at non-uniform positions, where (n + n k + Δn k ) may be a decimal. To solve the fractional value, this operation is implemented using bilinear interpolation.

[0061] Next, using the aligned infrared feature as the reference feature and the reflectivity feature F F as the perception feature, a deformable convolution operation guided by offset is performed. First, the two features are concatenated along the channel dimension, and then convolution is performed to generate the initial offset feature, which is then channel-connected with the offset feature from the previous stage, and convolution is performed to generate the final offset feature. Then, the reflectivity feature F F and the final offset feature are used together to perform deformable convolution to obtain the aligned reflectivity feature

[0062] Then, the final offset feature obtained previously using the infrared feature as the reference feature and the reflectivity feature F F as the perception feature is directly used together with the visible light feature F R to perform deformable convolution to obtain a new visible light feature. Then, the aligned reflectivity feature and the new visible light feature are summed to obtain the aligned visible light feature

[0063] The formula of the entire RCAM is as follows:

[0064]

[0065] S3. Input the aligned visible light feature the aligned infrared feature the aligned reflectivity feature into LSFM. LSFM maps the aligned modalities to a shared hidden state space by selecting a state space mechanism and performs feature interaction in the hidden state space to generate corresponding complementary features.

[0066] Meanwhile, as shown in Figure 4 , the Mamba-based LSFM makes full use of the advantage of the illumination invariance of reflectivity to achieve global attention while avoiding additional computational overhead. First, the aligned visible light feature the aligned infrared feature the aligned reflectivity feature are concatenated along the channel dimension to obtain the initial concatenated multi-modal feature The multi-modal feature after layer normalization passes through the VSSM (Visual State Space Module) to obtain the intermediate feature after being processed by the VSSM The intermediate feature Add it to the feature after the first skip connection, then perform layer normalization and pass it through an MLP (Multi-Layer Perceptron), and then add it to the feature after the second skip connection to obtain the final multi-modal fusion feature Before and after the VSSM (Visual State Space Module) and the MLP (Multi-Layer Perceptron), different learnable scaling factors s1 and s2 are used to control the information of the skip connection. Finally, the final multi-modal fusion feature Separate along the channel dimension to obtain single-modal complementary features The entire process of LSFM can be expressed as:

[0067]

[0068] Among them, Concat represents the concatenation operation along the channel dimension, LayerNorm represents layer normalization, and Split represents the feature separation operation along the channel dimension.

[0069] Such as Figure 5 As shown, in the VSSM, the multi-modal feature First, after layer normalization, it is separated into x R , x I , x F For separate processing, and then respectively pass through a linear layer, a depthwise separable convolution, and a SiLU function to obtain the modal features At the same time, the multi-modal feature Also pass through a linear layer and a SiLU function respectively to obtain the modal features x′ R , x′ I , x′ F , and the process is described as follows:

[0070]

[0071] x′ R = SiLU(Linear(x R )),

[0072] x′ I = SiLU(Linear(x I )),

[0073] x′ F = SiLU(Linear(x F )),

[0074] Among them, DWConv represents the depthwise separable convolution, respectively represent the inputs of the visible light modality, the infrared modality, and the reflectance feature of SS2D (2D Selective Scanning) at time step t.

[0075] Subsequently, the processed modal features are input into a shared SS2D for feature enhancement and fusion to obtain enhanced features This process can be expressed as:

[0076]

[0077] Enhanced features After layer normalization respectively, they are multiplied by the modal features x′ R 、x′ I 、x′ F and then linearized to obtain the modal features x″ R 、x″ I 、x″ F ,The modal features x″ R 、x″ I 、x″ F are concatenated along the channel dimension to obtain intermediate features This process can be expressed as follows:

[0078]

[0079] where Concat represents the concatenation operation along the channel dimension, LayerNorm represents layer normalization, and Linear represents the linear layer.

[0080] As Figure 6 shown in the SS2D, the input modal features are first processed by a linear layer to generate the state matrices: Δ R ,B R ,C R ,Δ I ,B I ,C I ,and Δ F ,B F ,C F 。Meanwhile, the parameter matrices are initialized as: A R ,D R ,A I ,D I ,A F, D F 。The continuous matrices A and B need to be discretized before being applied to the state space equation, where Δ is closely related to the discretization process of the state space model. The magnitude of Δ affects the model's ability to capture information in the time series, thus affecting the effective information propagation or forgetting between different parts of the sequence. Considering the unreliability of the visible light modality under low light conditions, the present invention uses the keyword K obtained from the Retinex decomposition network to improve Δ R as follows:

[0081] Δx = Δ R × K + Δ F ,

[0082] Wherein, the value of K is 0 or 1, representing a dark scene and a bright scene respectively.

[0083] The discretization process of the continuous matrices A and B can be expressed as:

[0084]

[0085] and Map their respective modal features to the hidden state h k .

[0086] The mapping process can be expressed as:

[0087]

[0088] The state matrix C is used to recover the output from the hidden state at each time step. To enable each modality to decode more useful information, add C R, C I and C F to get a global perception matrix C X , which is used to decode the hidden states of all three modalities to generate a selective scan output and

[0089] C x = C R + C I + C F ,

[0090]

[0091] It should be noted that adding to and enhances the representation ability of the visible modal features and strengthens the interaction with the infrared modal features. This can be expressed as follows:

[0092]

[0093] S4. Add the aligned features and complementary features to obtain enhanced features. Then, directly add the enhanced features of the two modalities to get the fused feature P i , and only send the fused features of the last few stages to the neck and head of the detection network to generate the final detection result.

[0094] Such as Figure 2As shown, the aligned visible light features obtained in step S2 aligned infrared features aligned reflectance features and the visible light complementary features obtained in step S3 infrared light complementary features reflectance complementary features are added together to obtain enhanced features. Then, the infrared-enhanced features and the visible light-enhanced features are added together to obtain the fused P i . The fused features P3, P4, and P5 of the last three stages are sent to the neck and head of the detection network YOLOv8 as inputs to generate the final detection results, so as to achieve the best performance and efficient calculation.

[0095] The training objective of the entire network is optimized based on the following defined final loss function:

[0096] L = λ1L box + λ2L cls + λ3L dfl

[0097] where L box represents the bounding box loss, L cls represents the classification loss, and L dfl is the loss for optimizing the bounding box coordinates. The weight parameters of L box , L cls and L dfl are represented by λ1, λ2, and λ3 respectively.

[0098] The working principle of target detection guided by the visible light image and the thermal infrared image of the drone in the present invention is as follows:

[0099] The input of RGFNet is the visible light image and the thermal infrared image of the drone, where the thermal infrared image of the drone is obtained by downsampling the GT thermal infrared image, and the visible light image of the drone is acquired under the drone platform. The drone platform used in the present invention is the DJI Matrice 300 RTK. The drone is equipped with a six-way infrared TOF sensor and is also equipped with the Zenmuse H20 series, which integrates a 20-megapixel zoom camera, a 12-megapixel wide-angle camera, a 30 Hz high-frame-rate thermal imaging camera, and a laser rangefinder with a detection range of up to 1200 meters, and has a maximum 23x hybrid optical zoom.

[0100] The present invention innovatively extends the YOLOv8 (Ultralytics YOLO) framework to a dual-branch structure and designs a new detection architecture RGFNet based on it. As Figure 2As shown in the figure, the framework includes a pre-trained Retinex decomposition network, a two-stream feature extraction DarkNet network, and three RGFBs, while the detection network includes a neck and a head for cross-modal object detection. For simplicity, the details of the convolutional modules for feature extraction in the backbone network are omitted in the figure.

[0101] To address the challenge of misalignment under different lighting intensity and distribution conditions, the present invention introduces a reflectance-guided collaborative alignment module (RCAM), which uses the illumination-invariant property of reflectance features as reference features to align infrared features, and then uses the aligned infrared features as a reference to further align reflectance and visible light features. Then, the aligned reflectance features, infrared features, and visible light features are input into a light perception selective fusion module (LSFM), and feature interactions are performed in the hidden state space to generate corresponding complementary features. Next, the aligned features and complementary features are added together to obtain enhanced features. Then, the enhanced features of the two modalities are directly added together to obtain the fused feature P. i RGFBs are only added in the last three stages to generate fused features P3, P4, and P5, which are the inputs to the YOLOv8 neck and head for predicting the final detection results to achieve optimal performance and efficient computation.

[0102] After the processing of RGFBs is completed, the infrared-enhanced features and visible light-enhanced features are added together to generate the fused feature P. i After obtaining these fused features, the features of the last three stages are input into the detector to supervise the training process of the network. This detector consists of two main parts: the neck and the head.

[0103] Embodiment 2

[0104] The present invention also provides a reflectance-guided all-weather UAV target detection system corresponding to Embodiment 1, including:

[0105] An extraction module for extracting an RGB feature set and an IR feature set from the input data;

[0106] An alignment module for using reflectance features with anti-low-light characteristics as reference features to align infrared features, and then using the aligned infrared features as reference features to further align reflectance and visible light features;

[0107] A complementary feature generation module for inputting the aligned visible light features the aligned infrared features the aligned reflectance features into the light perception selective fusion module and performing feature interactions in the hidden state space to generate corresponding complementary features;

[0108] The detection module is used to add the alignment feature and the complementary feature to obtain an enhanced feature, and then directly add the enhanced features of the RGB and IR modalities to obtain the fused feature P. i Send the fused features of the last few stages to the neck and head of the detection network to generate the final detection result.

[0109] The operations performed by each module are the same as those in the first embodiment.

[0110] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An all-weather unmanned aerial vehicle target detection method guided by reflectivity, characterized in that, Including: S1. Extract the RGB feature set and the IR feature set from the input data; S2. Use the reflectivity feature with anti - low - light characteristics as the reference feature to align the infrared feature, and then use the aligned infrared feature as the reference feature to further align the reflectivity and visible - light features; S3. Input the aligned visible - light features, aligned infrared features, and aligned reflectivity features into the light - perception selective fusion module, and perform feature interaction in the hidden state space to generate corresponding complementary features; S4. Add the aligned features and complementary features to obtain enhanced features. Then, directly add the enhanced features of the RGB and IR modalities to get the fused features, and send the fused features of the last few stages to the neck and head of the detection network to generate the final detection result.

2. The all-weather unmanned aerial vehicle target detection method based on reflectivity guidance according to claim 1, wherein Step S1 specifically includes: First, use the pre-trained Retinex decomposition network to decompose the reflectance feature F from the visible image F and the keyword K representing whether it is dark. Then, use two DarkNet feature extraction networks with the same structure but different parameters to extract RGB feature sets F with spatial dimensions of 1 / 2, 1 / 4, ···, 1 / 32 from the input data R and the infrared modality feature F I .

3. The all-weather UAV target detection method based on reflectivity guidance according to claim 1, wherein In step S2, using the reflectivity feature with anti-low-light characteristics as the reference feature to align the infrared feature specifically includes: First, using the reflectivity feature F F as the reference feature and the infrared modal feature F I as the perception feature to perform the offset-guided deformable convolution operation. This process starts from the channel connection of these two features, then convolves to generate the initial offset feature, then performs channel connection with the offset feature from the previous stage, and convolves again to generate the final offset feature. Finally, the offset feature is processed to output the offset Δn and the modulation scalar Δm. The offset Δn and the modulation scalar Δm perform deformable convolution with the perception feature to obtain the aligned infrared feature 4. The all-weather UAV target detection method based on reflectivity guidance according to claim 3, wherein, For the first offset feature, it is directly obtained by performing a single convolution on the spliced perception feature and reference feature.

5. The all-weather UAV target detection method based on reflectivity guidance according to claim 3, characterized in that, In step S2, using the aligned infrared features as reference features to further align the reflectance and visible light features specifically includes: using the aligned infrared features as reference features, using the reflectance feature F F as the perception feature, performing a deformable convolution operation guided by offset. First, the two features are concatenated along the channel dimension, and then convolved to generate the initial offset feature, which is then channel-connected with the offset feature from the previous stage and convolved to generate the final offset feature. Then, the reflectance feature F F and the final offset feature are used to perform a deformable convolution together to obtain the aligned reflectance feature Next, the infrared features aligned previously are used as reference features, and the reflectivity feature F F is used as the perception feature, and the finally obtained offset feature is directly convolved with the visible light feature F R using deformable convolution to obtain a new visible light feature. Then, the aligned reflectivity feature and the new visible light feature are summed to obtain the aligned visible light feature 6. The all-weather UAV target detection method based on reflectivity guidance according to claim 1, characterized in that, Step S3 specifically includes: First, the aligned visible light features aligned infrared features aligned reflectivity features are concatenated along the channel dimension to obtain the initially concatenated multi-modal features multi-modal features After layer normalization, it passes through the visual state space module to obtain the processed intermediate features intermediate features After layer normalization, it passes through a multi-layer perceptron to obtain the final multi-modal fusion features Different learnable scale factors s1 and s2 are used before and after the visual state space module and the multi-layer perceptron respectively to control the information of the skip connection. Finally, the final multi-modal fusion features are separated along the channel dimension to obtain single-modal complementary features 7. The all-weather UAV target detection method based on reflectivity guidance according to claim 6, characterized in that, Multi-modal features After layer normalization, it passes through the visual state space module to obtain the processed intermediate features The process is as follows: Multi-modal features First, after layer normalization, it is separated into x R , x I , x F for separate processing. Then, it passes through a linear layer, a depthwise separable convolution, and the SiLU function respectively to obtain modal features At the same time, the multi-modal features also pass through a linear layer and the SiLU function respectively to obtain modal features x' R , x' I , x' F ; Subsequently, the processed modal features are input into a shared SS2D for feature enhancement and fusion to obtain enhanced features Enhanced feature After layer normalization respectively, they are multiplied by the modality features x′ R , x′ I , x′ F and then linearized to obtain the modality features x″ R , x″ I , x″ F . The modality features x″ R , x″ I , x″ F are concatenated along the channel dimension to obtain the intermediate feature 8. The all-weather unmanned aerial vehicle target detection method based on reflectivity guidance according to claim 7, characterized in that In SS2D, the input modal features are first processed by a linear layer to generate a state matrix: Δ R ,B R ,C R ,Δ I ,B I ,C I , and Δ F ,B F ,C F . Meanwhile, the parameter matrices are initialized as: A R ,D R ,A I ,D I ,A F, D F . The continuous matrices A and B need to be discretized before being applied to the state space equation, where Δ is closely related to the discretization process of the state space model. The keyword K is obtained from the visible image using the Retinex decomposition network to improve Δ R as follows: Δ x = Δ R × K + Δ F , Among them, the value of K is 0 or 1, representing the dark scene and the bright scene respectively; The discretization process of the continuous matrices A and B is expressed as: and map their respective modal features to the hidden state h k ; The mapping process is expressed as: The state matrix C is used to recover the output from the hidden state at each time step. To enable each modality to decode more useful information, C R, C I and C F are added together to obtain a global perception matrix C x , which is used to decode the hidden states of all three modalities, thereby generating a selective scan output and 9. The all-weather UAV target detection method based on reflectivity guidance according to claim 1, wherein Step S4 specifically includes: adding the aligned visible light features obtained in step S2 aligned infrared features aligned reflectivity features and the visible light complementary features obtained in step S3 infrared light complementary features reflectivity complementary features to obtain enhanced features. Then, add the infrared-enhanced features and the visible light-enhanced features to get the fused feature P i , and send the fused features P3, P4, and P5 in the last three stages to the neck and head of the detection network YOLOv8 as inputs to generate the final detection results.

10. An all-weather unmanned aerial vehicle target detection system based on reflectivity guidance, characterized in that, Including: An extraction module for extracting the RGB feature set and the IR feature set from the input data; An alignment module for using the reflectivity feature with anti - low - light characteristics as the reference feature to align the infrared feature, and then using the aligned infrared feature as the reference feature to further align the reflectivity and visible - light features; A complementary feature generation module for generating aligned visible light features Aligned infrared features Aligned reflectivity features Input to the light perception selective fusion module and perform feature interaction in the hidden state space to generate corresponding complementary features; The detection module is used to add the alignment feature and the complementary feature to obtain an enhanced feature. Then, the enhanced features of the RGB and IR modalities are directly added to obtain the fused feature P i , and the fused features of the last few stages are sent to the neck and head of the detection network to generate the final detection result.

Citation Information

Cited By

  • Unmanned aerial vehicle detection method and system based on generative multi-modal fusion

    CN121280957A