Unmanned aerial vehicle dual-mode image registration and target detection method and simulation system

By registering the dual-mode images of the drone and improving the YOLOv8 network, the problem of low target detection accuracy in complex environments is solved, and high-precision target detection effect is achieved.

CN120339346AInactive Publication Date: 2025-07-18HEBEI UNIV OF ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510298858.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In complex environments, the single-mode image detection accuracy is low, especially in low light or complex backgrounds. Traditional visible light vision and infrared light sensors each have significant limitations, making it difficult to detect targets with high accuracy.

Method used

The dual-mode image registration method is adopted to register visible and infrared light images through perspective transformation, and the YOLOv8 network is improved, and the DMFF feature fusion module, LSKA large convolution kernel attention mechanism and SPConv convolution structure are introduced to improve feature extraction and expression capabilities.

Benefits of technology

It improves the target detection accuracy of the drone in complex environments, makes up for the shortcomings of single-modal detection, and achieves efficient and accurate target recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339346A_ABST
    Figure CN120339346A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of target detection, in particular to an unmanned aerial vehicle dual-mode image registration method, a target detection method and a simulation system, and the method comprises the steps: carrying out the registration of a dual-mode image through a frame selection region and perspective transformation; a YOLOv8 network is used as a basic framework and is improved; the specific improvement comprises the following steps: setting two input streams in an input part, and inputting a bimodal image into a backbone network; in the backbone network part, feature extraction is carried out on the two images, feature fusion is carried out by a DMFF feature fusion module in different stages of feature extraction, and the features are input into a neck network for feature expression enhancement; modifying the SPPF module, and adding an LSKA large convolution kernel attention mechanism in the SPPF module; in the neck network part, a C2f module is improved, and an SPConv convolution structure is introduced into the C2f module; and an unmanned aerial vehicle simulation system is arranged and is used for simulating dual-mode target detection of the unmanned aerial vehicle. According to the invention, the target detection precision of the unmanned aerial vehicle in a complex environment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of target detection and image registration, and more particularly to a method and simulation system for dual-modal image registration and target detection of an unmanned aerial vehicle (UAV). Background Art

[0002] With the rapid development of computer vision and deep learning, as well as the upgrade of UAV hardware, significant progress has been made in UAV target detection technology, enabling UAVs to efficiently and accurately detect targets in complex environments and be widely applied in fields such as reconnaissance and detection. When a UAV performs tasks in a real environment, it needs to flexibly search for targets in the environment, accurately obtain target information, and quickly transmit the obtained information to the ground terminal.

[0003] When a UAV performs tasks in a complex environment, higher requirements are put forward for the high-precision detection and positioning capabilities of targets. However, traditional visible-light vision UAVs have significant limitations when dealing with situations such as low light, occlusion, or complex backgrounds. Infrared light sensors are difficult to distinguish targets when the temperature difference is not significant, while visible-light sensors perform poorly at night or under adverse weather conditions. Using dual-modal image fusion technology to improve the robustness and accuracy of target detection has become an important direction in UAV target detection research. Summary of the Invention

[0004] In view of this, the present invention provides a method and simulation system for dual-modal image registration and target detection of a UAV, which can improve the accuracy of UAV target detection.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] In a first aspect, the present invention provides an image registration method for visible-light images and infrared-light images captured by a binocular camera with inconsistent image frames, which is characterized by including the following steps:

[0007] Select corresponding regions to be registered on dual-modal images with different resolutions;

[0008] Adopt perspective transformation to map the selected region on the visible-light image to the selected region on the infrared-light image;

[0009] Further, the perspective transformation formula is:

[0010]

[0011]

[0012]

[0013] x′ = m 11 x + m12 y + m 13

[0014] y' = m 21 x + m 22 y + m 23

[0015] w' = m 31 x + m 32 y + 1

[0016] Among them, the perspective transformation is a projective transformation, which can be represented by a 3×3 matrix M. (x, y) are the pixel coordinates in the original image, (x', y') are the pixel coordinates after transformation, and w' is the normalization factor. The final coordinates need to be normalized.

[0017] In a second aspect, the present invention provides a method for detecting targets in dual-modal images of an unmanned aerial vehicle, which is characterized by including the following steps:

[0018] Based on the YOLOv8 network as the basic architecture, it is improved to construct a target detection model;

[0019] Set a dual-branch structure in the image input part, and input the visible light image and the infrared light image into the backbone network respectively;

[0020] In the backbone network part, feature extraction is performed on the two images respectively, and feature fusion is performed by the DMFF feature fusion module at different stages of feature extraction, and then input into the neck network for enhancing feature expression;

[0021] Modify the SPPF module and add the LSKA large convolution kernel attention mechanism to it;

[0022] In the neck network part, improve the C2f module and introduce the SPConv convolution structure into it.

[0023] Furthermore, the DMFF feature fusion module includes:

[0024] The DMFF feature fusion module consists of a dimensionality reduction module, four convolutional branches of different scales, a CBAM attention module, a channel adaptive weighting module, and an upsampling module; the fused feature map first passes through the dimensionality reduction module for channel number compression, and then is respectively input into the convolutional branches of three different scales of 1×1, 3×3, and 5×5. The obtained features are added and then enter the CBAM attention module for enhancing spatial and channel attention. The enhanced features are restored to the original channel number through the upsampling module; the original input features and the CBAM-enhanced features are respectively weighted adaptively by channels, and then concatenated with the fused features in the channel direction, and finally the output features are adjusted through a 1×1 convolution.

[0025] Furthermore, the formula of DMFF is:

[0026] X d = W d *concat(X0,X1)

[0027] X f = W 1×1 *X d + W 3×3 *X d + W 5×5 *X d

[0028] M c = σ(W2δ(W1MaxPool(X f ))+W2δ(W1AvgPool(X f )))

[0029] M s = σ(F7([MaxPool(X f );AvgPool(X f )]))

[0030] X cbam = M c ·X f ·M s

[0031] w0 = Softmax([GAP(X0),GAP(X1),GAP(X cbam )])

[0032] X out = w0X0+w1X1+w cbam X cbam

[0033] X final = W proj *X out

[0034] The input of DMFF consists of two feature maps X0 and X1. First, channel dimension reduction is required. Among them, W d is a 1×1 convolutional kernel, * represents the convolution operation, and concat(X0,X1) represents the concatenation operation of the channel dimension; the reduced feature X d is respectively input into three convolutional kernels of different scales, W k×kConvolution kernels \(k\in\{1,3,5\}\) representing different scales are used, and the finally obtained multi-scale features are fused; then, through the CBAM attention mechanism, the CBAM channel attention obtains global information through max pooling and average pooling, where \(W1\) and \(W2\) are the weights of the fully connected layers, \(\delta\) is the ReLU activation function, and \(\sigma\) is the Sigmoid function; the spatial attention obtains spatial information through max pooling and average pooling, \(F7\) represents a \(7\times7\) convolution, and \([;]\) represents the channel concatenation operation, and the finally obtained features after CBAM processing; then, the channel attention weights of the input features \(X0\), \(X1\), and the enhanced feature \(X\) are calculated. GAP represents global average pooling, and Softmax is used to normalize the weights. \(w0\), \(w1\), and \(w\) respectively represent the channel weights of \(X0\), \(X1\), and \(X\); the finally obtained fused feature \(X\), after passing through a \(1\times1\) convolution to adjust the number of channels, where \(W\) is the weight of the \(1\times1\) convolution, which is used to adjust the output number of channels. cbam The channel attention weights of cbam \(X0\), \(X1\), and \(X\) cbam are calculated respectively. GAP represents global average pooling, and Softmax is used to normalize the weights. \(w0\), \(w1\), and \(w\) out respectively represent the channel weights of proj \(X0\), \(X1\), and \(X\); the finally obtained fused feature \(X\)

[0035] Furthermore, the modifications to SPPF include:

[0036] An LSKA module is introduced after the concat of SPPF;

[0037] Furthermore, the modified SPPF consists of two Conv modules, three cascaded \(k\times k\) max pooling layers, and a concat; after the fused feature map is input into one of the Conv modules, it is connected to three cascaded \(k\times k\) max pooling layers. The output of this Conv module and the output of each \(k\times k\) max pooling layer are jointly input into LSKA for feature enhancement, and after LSKA processing, it is connected to another Conv module.

[0038] Furthermore, the feature expression ability and receptive field are improved by using variable-sized convolution kernels. 1D convolution with a row-by-row and column-by-column structure is used to extract direction features, reducing the computational complexity.

[0039] Furthermore, the formula of LSKA is:

[0040] A = C 1×3 (X) → C 3×1 (A) → C 1×k (A) → C k×1 (A)

[0041] Y = Xe sigmoid(A)

[0042] where C 1×3 , C 3×1 represents 1D convolution, which is used to extract features in the horizontal and vertical directions. C 1×k and C k×1Further extract long - distance dependency information, where k is a variable hyperparameter. After the calculated attention weights A are normalized by Sigmoid, they are multiplied element - by - element with the input X to achieve attention adjustment.

[0043] Furthermore, the modifications to C2f include:

[0044] Replace Conv with the SPConv lightweight grouped convolution module in the C2f module;

[0045] Improve the feature extraction ability and reduce the computational cost by dynamically adjusting the contributions of 3×3 convolution and 1×1 convolution through adaptive weights;

[0046] Furthermore, the formula of SPConv is:

[0047] Y GWC =X*W GWC , groups = 2

[0048] Y PWC =X*W PWC

[0049] Y3 = Y GWC +Y PWC

[0050] α3 = AdaptiveAvgPool(Y3)

[0051] α1 = AdaptiveAvgPool(Y1)

[0052] β = Softmax([α3,α1])

[0053] Y = β3·Y3 + β1·Y1

[0054] Furthermore, the input feature map X has C channels, and is sent to 3×3 grouped convolution (GWC) and point - wise convolution (PWC) for calculation, and is sent to 1×1 standard convolution for calculation, where W GWC is the grouped 3×3 convolution kernel, and W PWC is the 1×1 convolution kernel. Y3 is the final 3×3 feature, which obtains α3 and α1 through adaptive feature weighting, and then uses Softmax normalization. The final output Y is a lightweight feature - enhanced map.

[0055] In the third aspect, the present invention provides a simulation system for dual - mode target detection of unmanned aerial vehicles, including:

[0056] An acquisition module for acquiring visible - light images and infrared - light images captured by a simulated unmanned aerial vehicle;

[0057] An image registration module for registering the problem of mismatched dual-modal image frames captured by the drone as described above;

[0058] A target detection module for performing target detection on visible light images and infrared light images using the drone dual-modal target detection method as described above;

[0059] A drone simulation module equipped with a simulated drone to simulate flight in a virtual scenario.

[0060] From the above technical solutions, it can be seen that compared with the prior art, the present invention has the following beneficial effects:

[0061] The present invention takes into account the problems of mismatched image frames captured by the drone binocular camera and low accuracy of single-modal target detection in complex environments. It designs an image registration method to register the visible light images and infrared light images captured by the binocular camera; designs a feature fusion module to perform feature fusion by the DMFF feature fusion module at different stages of feature extraction and input it into the neck network for enhanced feature expression; modifies the SPPF module and adds the LSKA large convolution kernel attention mechanism to it; in the neck network part, improves the C2f module and introduces the SPConv convolution structure into it. Finally, target detection is performed to make up for the significant limitations of visible light vision drones in dealing with low-light environments and the problem that it is difficult for infrared light sensors to distinguish targets when the temperature difference is not significant, thereby facilitating the accuracy of target detection of drones in complex environments. Description of the Drawings

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0063] Figure 1 A schematic diagram of the result of the image registration method provided by the present invention;

[0064] Figure 2 A flowchart of the drone dual-modal target detection method provided by the present invention;

[0065] Figure 3 A schematic diagram of the structure of the target detection model provided by the present invention;

[0066] Figure 4 A schematic diagram of the target detection result provided by the present invention;

[0067] Figure 5 A schematic diagram of the drone dual-modal target detection simulation system provided by the present invention; Specific implementation method

[0068] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work shall fall within the protection scope of the present invention.

[0069] As Figure 1 shown, the embodiments of the present invention disclose an image registration method for inconsistent visible light images and infrared light images captured by a binocular camera, including the following steps:

[0070] Select the corresponding areas to be registered on dual-modal images with different resolutions;

[0071] Use perspective transformation to map the selected area on the visible light image to the selected area on the infrared light image;

[0072] Furthermore, the perspective transformation formula is:

[0073]

[0074]

[0075]

[0076] x′ = m 11 x + m 12 y + m 13

[0077] y′ = m 21 x + m 22 y + m 23

[0078] w′ = m 31 x + m 32 y + 1

[0079] Among them, the perspective transformation is a projection transformation, which can be represented by a 3×3 matrix M. (x, y) are the pixel coordinates in the original image, (x′, y′) are the pixel coordinates after transformation, and w′ is the normalization factor. The final coordinates need to be normalized.

[0080] Generally speaking, the present invention designs an image registration method for inconsistent visible light images and infrared light images captured by a binocular camera. By using the selected area and perspective transformation, image registration is performed on dual-modal images with different resolutions and inconsistent images. The finally registered images can meet the training requirements of the dual-modal target detection network of the unmanned aerial vehicle.

[0081] As shown Figure 2 in the figure, an embodiment of the present invention discloses a dual-modal target detection method for drones, including the following steps:

[0082] Based on the YOLOv8 network as the basic architecture, it is improved to construct a target detection model;

[0083] Set a dual-branch structure in the image input part, and input visible light images and infrared light images into the backbone network respectively;

[0084] In the backbone network part, feature extraction is performed on the two types of images respectively, and feature fusion is performed by the DMFF feature fusion module at different stages of feature extraction, and then input into the neck network to enhance feature expression;

[0085] Modify the SPPF module and add the LSKA large convolutional kernel attention mechanism to it;

[0086] In the neck network part, improve the C2f module and introduce the SPConv convolutional structure into it.

[0087] Furthermore, the DMFF feature fusion module includes:

[0088] The DMFF feature fusion module consists of a dimensionality reduction module, four convolutional branches of different scales, a CBAM attention module, a channel adaptive weighting module, and an upsampling module; the fused feature map first passes through the dimensionality reduction module for channel number compression, and then is respectively input into the convolutional branches of three different scales of 1×1, 3×3, and 5×5. The obtained features are added and then enter the CBAM attention module for spatial and channel attention enhancement. The enhanced features are restored to the channel number through the upsampling module; the original input features and the CBAM enhanced features are respectively weighted adaptively by channels, and then concatenated with the fused features in the channel direction. Finally, the output features are adjusted by 1×1 convolution.

[0089] Furthermore, the formula of DMFF is:

[0090] X d = W d *concat(X0,X1)

[0091] X f = W 1×1 *X d + W 3×3 *X d + W 5×5 *X d

[0092] M c = σ(W2δ(W1MaxPool(Xf )) + W2δ(W1AvgPool(X f )))

[0093] M s =σ(F7([MaxPool(X f );AvgPool(X f )]))

[0094] X cbam =M c ·X f ·M s

[0095] w0=Softmax([GAP(X0), GAP(X1), GAP(X cbam )])

[0096] X out =w0X0 + w1X1 + w cbam X cbam

[0097] X final =W proj *X out

[0098] The input of DMFF consists of two feature maps X0 and X1. First, channel dimension reduction is required. Among them, W d is a 1×1 convolutional kernel, * represents the convolution operation, and concat(X0, X1) represents the concatenation operation of the channel dimension; the reduced feature X d is respectively input into three convolutional kernels of different scales. W k×k represents convolutional kernels of different scales k ∈ {1, 3, 5}, and the finally obtained multi-scale features are fused; then through the CBAM attention mechanism, the CBAM channel attention obtains global information through max pooling and average pooling. Among them, W1 and W2 are the weights of the fully connected layer, δ is the ReLU activation function, and σ is the Sigmoid function; the spatial attention obtains spatial information through max pooling and average pooling. F7 represents a 7×7 convolution, and [;] represents the channel concatenation operation, and the finally obtained feature after CBAM processing; then calculate the channel attention weights of the input features X0, X1 and the enhanced feature X cbam , GAP represents global average pooling, and Softmax is used to normalize the weights. w0, w1 and w cbam respectively represent the channel weights of X0, X1 and X cbam ; the finally obtained fused feature X out , after passing through a 1×1 convolution to adjust the number of channels, where W proj is the 1×1 convolution weight for adjusting the output number of channels.

[0099] Furthermore, the modifications to SPPF include:

[0100] An LSKA module is introduced after the concat of SPPF;

[0101] Furthermore, the modified SPPF consists of two Conv modules, three cascaded k×k max pooling layers, and one concat; after the fused feature map is input into one of the Conv modules, it is connected to three cascaded k×k max pooling layers. The output of this Conv module and the output of each k×k max pooling layer are jointly input into LSKA for feature enhancement, and after being processed by LSKA, it is connected to another Conv module.

[0102] Furthermore, the feature expression ability and receptive field are improved by using variable large convolution kernels, and 1D convolution with a row-by-row and column-by-column structure is used to extract direction features, reducing the computational amount.

[0103] Furthermore, the formula of LSKA is:

[0104] A = C 1×3 (X) → C 3×1 (A) → C 1×k (A) → C k×1 (A)

[0105] Y = X * sigmoid(A)

[0106] where C 1×3 , C 3×1 represents 1D convolution, which is used to extract features in the horizontal and vertical directions. C 1×k and C k×1 further extract long-range dependence information. Among them, k is a variable hyperparameter. After the calculated attention weight A is normalized by Sigmoid, it is multiplied element-wise with the input X to achieve attention adjustment.

[0107] Furthermore, the modifications to C2f include:

[0108] In the C2f module, the Conv is replaced with the SPConv lightweight grouped convolution module;

[0109] The contributions of the 3×3 convolution and the 1×1 convolution are dynamically adjusted through adaptive weights, improving the feature extraction ability while reducing the computational amount;

[0110] Furthermore, the formula of SPConv is:

[0111] Y GWC = X * W GWC , groups = 2

[0112] Y PWC = X * WPWC

[0113] Y3 = Y GWC +Y PWC

[0114] α3 = AdaptiveAvgPool(Y3)

[0115] α1 = AdaptiveAvgPool(Y1)

[0116] β = Softmax([α3,α1])

[0117] Y = β3·Y3 + β1·Y1

[0118] Furthermore, the input feature map X has C channels. It is fed into 3×3 grouped convolution (GWC) and pointwise convolution (PWC) calculations. It is fed into 1×1 standard convolution calculation, where W GWC is the 3×3 grouped convolution kernel, and W PWC is the 1×1 convolution kernel. Y3 is the final 3×3 feature, which undergoes adaptive feature weighting to obtain α3 and α1, and then Softmax normalization is used. Finally, Y is output to obtain a lightweight feature expression enhancement map.

[0119] As Figure 3 shown, it is the architecture diagram of the entire object detection model, which is an end-to-end dual-modal single-stage object detection network. The visible light image and the infrared light image are input into the network at a resolution of 640×640. Feature extraction is performed on the images of the two modalities respectively. At different stages of feature extraction, the image features of the two modalities are fused by the DMFF module, and then input into the SPPF module with the LSKA module added. After that, the 320×320 feature maps of the two modalities obtained are concatenated along the channel dimension and input into the subsequent neck network and the detection head. In the neck network, the C2f module embedded with the SPConv module is used to lighten the computation. Finally, the image is input into the detection head, and the detection head predicts the object category and location information respectively.

[0120] As Figure 4 shown, it is the detection result obtained by using the model trained with the training set in the LLVIP dataset to detect the pictures in the LLVIP dataset. The detection result shows the detection target box, category information, and confidence level. There are no missed detections and false detections, the confidence level is relatively high, and the detection result is good.

[0121] As Figure 5As shown, it is a schematic diagram of a dual-modal target detection simulation system for an unmanned aerial vehicle (UAV). In the simulation system, the flight of the simulated UAV is controlled, and the images captured by the visible light and infrared light cameras of the simulated UAV are read. Image registration is performed on the dual-modal images, and then target detection is carried out on the registered images.

[0122] Those skilled in the art should understand that the implementation of the present invention can be provided in the form of a method, a computer program product, etc. Therefore, the present invention can adopt the form of a complete software embodiment, a complete hardware embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0123] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for dual - mode image registration and target detection of an unmanned aerial vehicle, characterized in that: Image registration is performed based on the inconsistency of the image frames caused by the visible - light image and the infrared - light image captured by a binocular camera; The image registration method uses perspective transformation, calculates the homography matrix using four pairs of corresponding points, and then aligns the source image to the target image using affine transformation to achieve geometric correction and registration of the dual - mode images; Based on the YOLOv8 network as the basic framework, its model is optimized and improved to construct a model: Based on the detected target, the registered sample images are collected for model training; Based on the trained detection target model, target detection is performed on the visible - light image and the infrared - light image captured from the perspective of the unmanned aerial vehicle; The system is designed based on an unmanned aerial vehicle simulation platform; Among them, the improvement of the YOLOv8 network includes: A dual - branch structure is set in the image input part, and the visible - light image and the infrared - light image are respectively input into the backbone network; In the backbone network part, feature extraction is performed on the two images respectively, and at different stages of feature extraction, the DMFF feature fusion module performs feature fusion and inputs it into the neck network for enhanced feature expression; Modify the SPPF module, and add the LSKA large - convolution - kernel attention mechanism to it; In the neck network part, improve the C2f module and introduce the SPConv convolution structure into it.

2. The method for dual-modal image registration and target detection of an unmanned aerial vehicle according to claim 1, wherein, By selecting the corresponding areas to be registered on the infrared - light image and the visible - light image with different resolutions, and transforming the area on the visible - light image to the corresponding area on the infrared - light image according to perspective transformation, the formula is: x′ = m 11 x + m 12 y + m 13 y′ = m 21 x + m 22 y + m 23 w′ = m 31 x + m 32 y + 1 Among them, perspective transformation is a projective transformation, which can be represented by a 3×3 matrix M, (x, y) is the pixel coordinate in the original image, (x′, y′) is the pixel coordinate after transformation, w′ is the normalization factor, and the final coordinates need to be normalized.

3. A method for dual-modal image registration and target detection of an unmanned aerial vehicle according to claim 1, characterized in that, For the model trained for the registered images, the target detection model is deployed to the unmanned aerial vehicle simulation platform for verification.

4. A method for dual-modal image registration and target detection of an unmanned aerial vehicle according to claim 1, characterized in that, Modify the SPPF module, and the modifications include: Introduce the LSKA module after the concat of SPPF.

5. A method for dual-modal image registration and target detection of an unmanned aerial vehicle according to claim 1, characterized in that, Modify the C2f module, and the modifications include: Replace the Conv convolution in C2f with SPConv.

6. The method for dual-modal image registration and target detection of an unmanned aerial vehicle according to claim 1, wherein The DMFF feature fusion module consists of a dimensionality - reduction module, four convolutional branches with different scales, a CBAM attention module, a channel - adaptive weighting module, and an up - sampling module; the fused feature map first passes through the dimensionality - reduction module for channel number compression, and then is respectively input into the convolutional branches with scales of 1×1, 3×3, and 5×5. The obtained features are added and then enter the CBAM attention module for spatial and channel attention enhancement. The enhanced features are restored to the original channel number through the up - sampling module; the original input features and the CBAM - enhanced features are respectively weighted adaptively by channels, and then are concatenated in the channel direction with the fused features, and finally the output features are adjusted through 1×1 convolution.

7. A method for dual-modal image registration and target detection of an unmanned aerial vehicle according to claim 1, characterized in that The formula of DMFF is: X d = W d *concat(X0,X1) X f = W 1×1 * X d + W 3×3 * X d + W 5×5 * X d M c = σ(W2δ(W1MaxPool(X f )) + W2δ(W1AvgPool(X f ))) M s = σ(F7([MaxPool(X f );AvgPool(X f )])) X cbam = M c · X f · M s w0 = Softmax([GAP(X0), GAP(X1), GAP(X cbam )]) X out = w0X0 + w1X1 + w cbam X cbam X final = W proj * X out The input of DMFF consists of two feature maps X0 and X1. First, channel dimension reduction is required. Here, W d is a 1×1 convolutional kernel, * represents the convolution operation, and concat(X0,X1) represents the concatenation operation in the channel dimension; the reduced feature X d is input into three convolutional kernels of different scales respectively. W k×k represents convolutional kernels of different scales k∈{1,3,5}, and the finally obtained multi-scale features are fused; then through the CBAM attention mechanism, the CBAM channel attention obtains global information through max pooling and average pooling. Here, W1 and W2 are the weights of the fully connected layers, δ is the ReLU activation function, and σ is the Sigmoid function; the spatial attention obtains spatial information through max pooling and average pooling. F7 represents a 7×7 convolution, [;] represents the channel concatenation operation, and the finally obtained feature after CBAM processing; then calculate the channel attention weights of the input features X0, X1 and the enhanced feature X cbam . GAP represents global average pooling, Softmax is used to normalize the weights, and w0, w1 and w cbam represent the channel weights of X0, X1 and X cbam respectively; the finally obtained fused feature X out , after adjusting the number of channels through a 1×1 convolution, where W proj is the 1×1 convolution weight for adjusting the output number of channels.

8. A method for dual-modal image registration and target detection of an unmanned aerial vehicle according to claim 1, characterized in that, The modified SPPF consists of two Conv modules, three consecutive k×k max pooling layers in series, and a concat; after the fused feature map is input into one of the Conv modules, it is connected to three consecutive k×k max pooling layers in series. The output of this Conv module and the output of each k×k max pooling layer are jointly input into LSKA for feature enhancement, and after being processed by LSKA, it is connected to another Conv module.

9. A method for dual-modal image registration and target detection of an unmanned aerial vehicle according to claim 1, characterized in that, Enhance the feature expression ability and receptive field by using variable large convolutional kernels, and adopt 1D convolution with a row-by-row plus column-by-column structure to extract directional features and reduce the computational amount. The formula is: A = C 1×3 (X) → C 3×1 (A) → C 1×k (A) → C k×1 (A) Y = Xe sigmoid(A) Among them, C 1×3 , C 3×1 represents 1D convolution, which is used to extract features in the horizontal and vertical directions. C 1×k and C k×1 further extract long-range dependence information. Here, k is a variable hyperparameter. After the calculated attention weight A is normalized by Sigmoid, it is multiplied element-wise with the input X to achieve attention adjustment; In the neck network part, replace Conv with the SPConv lightweight grouped convolution module in the C2f module, dynamically adjust the contributions of the 3×3 convolution and the 1×1 convolution through adaptive weights, improve the feature extraction ability, and at the same time reduce the computational amount. The formula is: Y GWC =X * W GWC , groups = 2 Y PWC = X * W PWC Y3 = Y GWC +Y PWC α3 = AdaptiveAvgPool(Y3) α1 = AdaptiveAvgPool(Y1) β = Softmax([α3,α1]) Y = β3·Y3 + β1·Y1 The input feature map X has C channels, and is fed into the 3×3 grouped convolution (GWC) and pointwise convolution (PWC) calculations, and is fed into the 1×1 standard convolution calculation, where W GWC is the grouped 3×3 convolution kernel, and W PWC is the 1×1 convolution kernel. Y3 is the final 3×3 feature, which is adaptively feature weighted to obtain α3 and α1, and then Softmax normalization is used. Finally, Y is output to obtain the lightweight feature expression enhancement map.

10. A simulation system for dual-modal image registration and target detection of an unmanned aerial vehicle, characterized in that, Including: Start the simulation drone, load the object detection model, and obtain visible light images and infrared light images from the perspective of the simulation drone; An image registration method that uses the algorithm described in claim 2 to register the visible light image and the infrared light image; An object detection model for performing object detection on the visible light image and the infrared light image by using the dual-modal image object detection method of the drone described in any one of claims 1-10.