Remote sensing target multi-modal fusion detection method combining atmospheric scattering removal and super-resolution

Through the multimodal fusion detection method of combined atmospheric scattering removal and super-resolution multimodal fusion detection method, the problem of image quality degradation caused by atmospheric scattering in remote sensing image target detection is solved, and the effect of improving detection accuracy and efficiency is achieved.

CN120219722APending Publication Date: 2025-06-27ANHUI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510420436.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-04
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Remote sensing image target detection faces interference from complex environmental factors such as atmospheric scattering, cloud occlusion, and light changes, resulting in a decline in image quality and low detection accuracy.

Method used

Using a multimodal fusion detection method with combined atmospheric scattering removal and super-resolution multimodal fusion detection method, the pixel-level fusion module and super-resolution-assisted target detection network combines visible light and infrared image data to remove atmospheric scattering effects and improve image resolution.

Benefits of technology

It effectively reduces the problem of image quality reduction caused by the target image being affected by atmospheric scattering, and improves the accuracy and efficiency of remote sensing image object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219722A_ABST
    Figure CN120219722A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing target multi-modal fusion detection method combining atmospheric scattering removal and super-resolution, which comprises the following steps: 1, collecting paired remote sensing infrared-visible images and carrying out preprocessing, 2, establishing a pixel-level fusion module, 3, constructing a super-resolution assisted target detection network, and 4, carrying out multi-modal fusion detection on the super-resolution assisted target detection network. 4, constructing a super-resolution branch composed of an encoder and a decoder; and 5, training to obtain an optimal target detection network model which is used for predicting a remote sensing image. According to the method, the characteristics of visible light and infrared images can be fully utilized, and the problem of image quality reduction caused by the influence of atmospheric scattering on the target image is reduced, so that the remote sensing image target detection precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer vision and image processing, and specifically relates to a remote sensing target multi-modal fusion detection method that combines atmospheric scattering removal and super-resolution. Background Art

[0002] With the continuous development of remote sensing technology, remote sensing image target detection provides an important detection means in the fields of military reconnaissance, geological disaster investigation and treatment, environmental monitoring, etc. In recent years, the breakthrough of deep learning technology has provided powerful technical support for the extraction of massive remote sensing image information, significantly improving the accuracy and efficiency of remote sensing image target detection. However, remote sensing image target detection still faces many challenges.

[0003] First of all, remote sensing images are easily interfered by complex environmental factors, such as atmospheric scattering, cloud and fog occlusion, light changes, etc. These factors will lead to a decline in image quality and blurred target features, thus affecting the detection accuracy. Especially in the visible light modality, it is more vulnerable to atmospheric scattering, resulting in a decline in image quality. Visible light images are formed by the reflection of sunlight by objects, and the wavelength range of the visible light band is between 400-760nm. Therefore, Rayleigh scattering is more likely to occur in the atmosphere, resulting in a decrease in the contrast and clarity of the image. Infrared images, due to their longer wavelengths, are relatively less affected by atmospheric scattering, and thus have better penetration and imaging effects under low light or adverse weather conditions.

[0004] In addition, the single-modal information of remote sensing target detection is insufficient: Although visible light images can provide rich color and detail information and have a high spatial resolution, which is suitable for observing details such as the shape, texture, and color of objects, infrared images reflect the thermal radiation characteristics of objects, that is, the information of the surface temperature distribution of objects, and can form images in an environment with no or very weak visible light at all. However, visible light images are blurred or ineffective at night, in haze, or in strong backlighting. Although infrared images have strong anti-interference ability, they have low resolution and lack texture.

[0005] In addition, remote sensing image target detection is limited by low resolution, small target blurring, and complex background interference, resulting in low detection accuracy and high missed detection rate. Summary of the Invention

[0006] The present invention is to solve the above-mentioned deficiencies existing in the prior art, and proposes a remote sensing target multi-modal fusion detection method that combines atmospheric scattering removal and super-resolution, aiming to make full use of the characteristics of visible light and infrared images, reduce the problem of image quality decline caused by the influence of atmospheric scattering on the target image, and thus improve the detection accuracy of remote sensing image targets.

[0007] The present invention adopts the following technical solutions to achieve the above-mentioned invention purpose:

[0008] The feature of a multi-modal fusion detection method for remote sensing targets that combines atmospheric scattering removal and super-resolution in the present invention is as follows: It is carried out according to the following steps:

[0009] Step 1: Collect paired remote sensing infrared-visible images and perform preprocessing to obtain the preprocessed infrared remote sensing image dataset T and the visible light dataset ;

[0010] Step 2: Establish a pixel-level fusion module, including: a parallel infrared branch and visible light branch, and an SE module; each branch includes an attention mask generation layer and a bottleneck feature extraction layer, and processes the nth preprocessed infrared remote sensing image and the nth visible light remote sensing image after atmospheric scattering removal to obtain the nth fusion image ;

[0011] Step 3: Construct a super-resolution assisted object detection network, including: a multi-dimensional feature extraction module, a PPA module, a multi-dimensional feature fusion module, and a result prediction module, and process to obtain the predicted detection box and the predicted class label;

[0012] Step 3.1: The multi-dimensional feature extraction module is composed of d cascaded Dark blocks, and processes in sequence, so as to use the feature output by the (d - 2)th Dark block as the nth one-dimensional target feature , use the feature output by the (d - 1)th Dark block as the nth two-dimensional target feature , and the feature output by the dth Dark block as the nth three-dimensional target feature ;

[0013] Step 3.2: Construct a PPA module and process to obtain the nth three-dimensional enhanced feature representation ;

[0014] Step 3.3 The multi-dimensional feature fusion module includes: 2 upsampling fusion blocks and 2 downsampling fusion block units, and processes , and to obtain the nth one-dimensional prediction feature map , the nth two-dimensional prediction feature map and the nth three-dimensional prediction feature map ;

[0015] Step 3.4: The prediction module processes , , Process it to generate a prediction tensor containing bounding box offsets, confidence, and class probabilities. After decoding the prediction tensor, obtain the predicted detection boxes and predicted class labels;

[0016] Step 4: Based on the predicted detection boxes, predicted class labels, ground truth detection boxes, and ground truth class labels, construct the object detection loss using Equation (1) :

[0017] (1)

[0018] In Equation (1), represents the localization loss between the predicted detection box of and the ground truth detection box, represents the confidence loss of the object existence of represents , and are the weight coefficients of the localization loss, confidence loss, and class classification loss respectively;

[0019] Step 5: Construct a super-resolution branch composed of an encoder and a decoder, and use the feature output by the 3rd Dark module in the object detection network as the nth low-level feature , and use the feature output by the 5th Dark module as the nth high-level feature , thereby and are processed to obtain the nth super-resolution reconstruction feature ;

[0020] Step 6: Based on and , construct the super-resolution loss using Equation (2) :

[0021] (2)

[0022] Step 7: Construct the total loss function using Equation (3)

[0023] (3)

[0024] In Equation (3), and are two balance coefficients;

[0025] Step 8: Optimize and train the object detection network model composed of a multi-dimensional feature extraction module, a PPA module, a multi-dimensional feature fusion module, a result prediction module, and a super-resolution branch using the backpropagation algorithm, and calculate the total loss function to update the model parameters until the total loss function converges, so as to obtain the optimal object detection network model after training, which is used to predict the remote sensing image and obtain the image with predicted detection frames and predicted class labels.

[0026] Another feature of the remote sensing target multi-modal fusion detection method combining atmospheric scattering removal and super-resolution according to the present invention is that the above-mentioned step 1 is carried out according to the following process:

[0027] Step 1.1: After obtaining the visible light remote sensing image dataset with real detection frames and real class labels and performing size normalization and preprocessing, the preprocessed visible light remote sensing image dataset S = { , ,…, ,…, } is obtained, where represents the nth preprocessed visible light remote sensing image, and N is the total number of visible light remote sensing images;

[0028] Step 1.2: After obtaining the infrared remote sensing image dataset with real detection frames and real class labels and performing size normalization and preprocessing, the preprocessed infrared remote sensing image dataset T = { , ,…, ,…, } is obtained, where represents the nth preprocessed infrared remote sensing image;

[0029] Step 1.3: Input S = { , ,…, ,…, } into the network model based on a convolutional neural network for processing to obtain the visible light dataset after removing the atmospheric scattering effect , ,…, ,…, , where represents the nth visible light remote sensing image after removing the atmospheric scattering.

[0030] Furthermore, the above-mentioned step 2 is carried out according to the following process:

[0031] The attention mask generation layer in the infrared branch processes to generate the nth infrared attention map , and after element-wise multiplication with , the nth infrared weighted feature is obtained ;

[0032] The bottleneck feature extraction layer in the infrared branch processes to generate the nth infrared feature map ;

[0033] The attention mask generation layer in the visible light branch processes to generate the nth visible light attention map , and after element-wise multiplication with , the nth visible light weighted feature is obtained ;

[0034] The bottleneck feature extraction layer in the visible light branch processes to generate the nth visible light feature map ;

[0035] Combine with along the channel dimension and generate the nth combined feature map , then input it into the SE module for processing to generate the nth channel weight map ;

[0036] Multiply with channel-wise and output the nth fused image .

[0037] Furthermore, the PPA module in step 3.2 includes: a multi-branch feature extraction module and an attention mechanism module; among them, the multi-branch feature extraction module includes: parallel local branches, global branches, and sequential convolutional branches;

[0038] Perform point convolution operations on and then input them into the three branches for processing. After obtaining three different levels of features and fusing them, and then through the processing of the attention mechanism module, the nth three-dimensional enhanced feature representation is obtained.

[0039] Furthermore, each upsampling fusion block in step 3.3 includes: an upsampling unit, a feature splicing unit, and a CSP unit; each downsampling fusion block includes: a downsampling unit, a feature splicing unit, and a CSP unit;

[0040] Input into the first upsampling fusion block, and after the upsampling operation of the first upsampling unit, it is combined with After being input into the first feature splicing unit for splicing operation, it is then input into the first CSP unit for processing to obtain the nth two-dimensional fusion feature ;

[0041] Input into the second upsampling fusion block. After the upsampling operation of the second upsampling unit, it is combined with and input into the second feature splicing unit for splicing operation, and then input into the second CSP unit for processing to obtain the nth one-dimensional prediction feature map ;

[0042] Input into the first downsampling fusion block. After the downsampling operation of the first downsampling unit, it is combined with and input into the third feature splicing unit for splicing operation, and then input into the third CSP unit for processing to obtain the nth two-dimensional prediction feature map ;

[0043] Input into the second downsampling fusion block. After the downsampling operation of the second downsampling unit, it is combined with and input into the fourth feature splicing unit for splicing operation, and then input into the fourth CSP unit for processing to obtain the nth three-dimensional prediction feature map .

[0044] Further, step 5 is carried out according to the following process:

[0045] Step 5.1: The encoder consists of 3 CR modules and an upsampling module, and processes and to obtain the nth fused feature map ;

[0046] The first CR module processes to obtain the nth processed low-level feature ;

[0047] The upsampling module performs bilinear upsampling on to align its spatial dimension with to obtain the nth processed high-level feature ;

[0048] Input and are concatenated along the channel dimension and then processed by two CR modules in sequence to obtain the nth fused feature map ;

[0049] Step 5.2: The decoder uses the EDSR module to perform step-by-step spatial dimension magnification to obtain the nth super-resolution reconstruction feature .

[0050] An electronic device according to the present invention includes a memory and a processor, characterized in that the memory is used to store a program for supporting the processor to execute the remote sensing target multi-modal fusion detection method, and the processor is configured to execute the program stored in the memory.

[0051] A computer-readable storage medium according to the present invention, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it executes the steps of the remote sensing target multi-modal fusion detection method.

[0052] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0053] 1. The present invention applies the dehazing technology through a deep learning algorithm to eliminate the problem of image quality degradation caused by the atmospheric scattering effect, so as to accurately detect whether the remote sensing image contains a target and locate the target position.

[0054] 2. The present invention realizes the complementarity of the two-modal features by establishing a pixel-level fusion module to process the visible light and infrared light two-modal data, and uses a deep learning-based network to realize the positioning and classification of small targets in the remote sensing image, which helps to improve the target detection accuracy.

[0055] 3. The present invention designs a simple and flexible super-resolution branch to obtain high-resolution features, which can distinguish small objects in the large background of low-resolution input, thereby further improving the detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is a schematic flow chart of the present invention;

[0057] Figure 2 is a structural diagram of the pixel-level fusion module of the present invention;

[0058] Figure 3 is a structural diagram of the target detection network of the present invention;

[0059] Figure 4 is a structural diagram of the PPA module of the present invention;

[0060] Figure 5 is a structural diagram of the super-resolution branch of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0061] In the present invention, a multi-modal remote sensing image target detection method combining atmospheric scattering removal and super-resolution utilizes super-resolution technology to enhance image details, restore small target features, enhance the model's ability to capture subtle information, and improve detection performance, thereby solving the problems of image degradation, modal fusion, and small target detection, and providing an efficient and robust solution for remote sensing target detection. Specifically, as Figure 1 shown, the method includes the following steps:

[0062] Step 1: Collect paired remote sensing infrared-visible images and perform preprocessing;

[0063] Step 1.1: After obtaining the visible light remote sensing image dataset with real detection frames and real class labels and performing size normalization and preprocessing, the preprocessed visible light remote sensing image dataset S = { , ,…, ,…, } is obtained, where represents the nth preprocessed visible light remote sensing image, and N is the total number of visible light remote sensing images;

[0064] Step 1.2: After obtaining the infrared remote sensing image dataset with real detection frames and real class labels and performing size normalization and preprocessing, the preprocessed infrared remote sensing image dataset T = { , ,…, ,…, } is obtained, where represents the nth preprocessed infrared remote sensing image;

[0065] The dataset used in the present invention is the publicly available remote sensing aerial image vehicle dataset VEDAI, which contains 1248 images, each with two modalities of infrared and visible light, and the images focus on different backgrounds, including grasslands, highways, mountains, and urban areas. The size of all images is 1024 × 1024.

[0066] Step 1.3: Input S = { , ,…, ,…, } into the network model based on the convolutional neural network for processing to obtain the visible light dataset after removing the atmospheric scattering effect , ,…, ,…, , where represents the nth visible light remote sensing image after removing the atmospheric scattering.

[0067] Step 2: AsFigure 2 As shown, a pixel-level fusion module is established, including: a parallel infrared branch and visible light branch, and an SE module; each branch contains an attention mask generation layer and a bottleneck feature extraction layer, and and are processed to obtain the nth fused image ;

[0068] The attention mask generation layer in the infrared branch processes to generate the nth infrared attention map , and after element-wise multiplication with , the nth infrared weighted feature is obtained;

[0069] The bottleneck feature extraction layer in the infrared branch processes to generate the nth infrared feature map ;

[0070] The attention mask generation layer in the visible light branch processes to generate the nth visible light attention map , and after element-wise multiplication with , the nth visible light weighted feature is obtained;

[0071] The bottleneck feature extraction layer in the visible light branch processes to generate the nth visible light feature map ;

[0072] Combine and along the channel dimension and generate the nth joint feature map , then input it into the SE module for processing to generate the nth channel weight map ;

[0073] Multiply and channel by channel, and output the nth fused image .

[0074] Step 3: As Figure 3 shown, construct a super-resolution assisted object detection network, including: a multi-dimensional feature extraction module, a PPA module, a multi-dimensional feature fusion module, and a result prediction module, and process to obtain the predicted detection box and predicted class label;

[0075] Step 3.1: The multi-dimensional feature extraction module consists of d cascaded Dark blocks, and processes Processed in sequence, so that the feature output by the (d - 2)-th Dark block is used as the n-th target one-dimensional feature , and the feature output by the (d - 1)-th Dark block is used as the n-th target two-dimensional feature , and the feature output by the d-th Dark block is used as the n-th target three-dimensional feature ;

[0076] In the present invention, the multi-dimensional feature extraction module is implemented using the darknet53 network, which is composed of d Dark blocks connected in series, where d is taken as 4 here.

[0077] Step 3.2: As Figure 4 shown, constructing the PPA module includes: a multi-branch feature extraction module and an attention mechanism module;

[0078] Among them, the multi-branch feature extraction module includes parallel local branches, global branches, and sequential convolutional branches;

[0079] After performing point convolution operations on , they are respectively input into three branches for processing, obtaining three features at different levels, fusing them, and then processing through the attention mechanism module to obtain the n-th three-dimensional enhanced feature representation ;

[0080] In the present invention, the local branch divides the input feature map into non-overlapping 2×2 small blocks and then uses 3×3 convolution operations. The global branch divides the input feature map into non-overlapping 4×4 small blocks and then uses 3×3 convolution operations. The serial convolution is 3 serial 3×3 convolution operations. Through the complementarity of the parallel branches; in the present invention, through the complementarity of the parallel branches, the PPA module can enhance the feature expression of small targets at different scales.

[0081] Step 3.3 The multi-dimensional feature fusion module includes: 2 downsampling fusion blocks and 2 downsampling fusion block units, and processes , and to obtain the n-th one-dimensional prediction feature map , the n-th two-dimensional prediction feature map and the n-th three-dimensional prediction feature map ;

[0082] Each upsampling fusion block in Step 3.3 includes: an upsampling unit, a feature splicing unit, and a CSP unit; each downsampling fusion block includes: a downsampling unit, a feature splicing unit, and a CSP unit;

[0083] Input into the first downsampling fusion block, and after the upsampling operation of the first upsampling unit, it is combined with After being input into the first feature splicing unit for splicing operation, it is then input into the first CSP unit for processing to obtain the nth two-dimensional fusion feature ;

[0084] The is input into the second downsampling fusion block, and after the upsampling operation of the second upsampling unit, it is input into the second feature splicing unit together with for splicing operation, and then input into the second CSP unit for processing to obtain the nth one-dimensional prediction feature map .

[0085] The is input into the first downsampling fusion block, and after the downsampling operation of the first downsampling unit, it is input into the third feature splicing unit together with for splicing operation, and then input into the third CSP unit for processing to obtain the nth two-dimensional prediction feature map ;

[0086] The is input into the second downsampling fusion block, and after the downsampling operation of the second downsampling unit, it is input into the fourth feature splicing unit together with for splicing operation, and then input into the fourth CSP unit for processing to obtain the nth three-dimensional prediction feature map .

[0087] Step 3.4: The prediction module processes , , to generate a prediction tensor containing bounding box offsets, confidences, and class probabilities, and after decoding the prediction tensor, prediction detection boxes and prediction class labels are obtained.

[0088] Step 4: Based on the prediction detection boxes, prediction class labels, true detection boxes, and true class labels, an object detection loss is constructed using Equation (1) :

[0089] (1)

[0090] In Equation (1), represents the localization loss between the predicted detection box of and the true detection box, represents the confidence loss of the object existence of , represents the class classification loss of , , and They are the weight coefficients of the localization loss, confidence loss, and class classification loss respectively.

[0091] Step 5: As Figure 5 shown, construct a super-resolution branch composed of an encoder and a decoder, and use the features output by the 3rd Dark module in the object detection network as the nth low-level feature and use the features output by the 5th Dark module as the nth high-level feature , so as to and are processed to obtain the nth super-resolution reconstruction feature ;

[0092] Step 5.1: The encoder consists of 3 CR modules and an upsampling module, and processes and to obtain a fused feature map ;

[0093] The first CR module processes to obtain the nth processed low-level feature ;

[0094] The CR module includes a convolution operation and a ReLU activation function;

[0095] The upsampling module performs a bilinear upsampling operation on to align its spatial dimension with to obtain the nth processed high-level feature ;

[0096] After and are concatenated along the channel dimension and then processed by two CR modules in sequence, the nth fused feature map is obtained.

[0097] Step 5.2: The decoder uses the EDSR module to gradually enlarge the spatial size of to obtain the nth super-resolution reconstruction feature ;

[0098] Step 6: Based on and , construct the super-resolution loss using Equation (2):

[0099] (2)

[0100] Step 7: Use Equation (3) to construct the total loss function

[0101] (3)

[0102] In formula (3), and are two balance coefficients.

[0103] Step 8: Use the backpropagation algorithm to optimize and train the object detection network model composed of the multi-dimensional feature extraction module, PPA module, multi-dimensional feature fusion module, result prediction module, and super-resolution branch, and calculate the total loss function to update the model parameters until the total loss function converges, so as to obtain the optimal object detection network model after training, which is used to predict the remote sensing image to obtain an image with predicted detection boxes and predicted class labels.

[0104] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.

[0105] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the steps of the above method.

Claims

1. A multi-modal fusion detection method for remote sensing targets combining atmospheric scattering removal and super-resolution, characterized in that: Follow these steps: Step 1: Collect paired remote sensing infrared-visible images and preprocess them to obtain the preprocessed infrared remote sensing image dataset T and visible light dataset T ; Step 2: Establish a pixel-level fusion module, which includes: parallel infrared branch and visible light branch and a SE module; each branch contains an attention mask generation layer and a bottleneck feature extraction layer, and performs pixel-level fusion on the nth infrared remote sensing image after preprocessing. and the nth visible light remote sensing image after atmospheric scattering is removed Process and get the nth fused image ; Step 3: Construct a super-resolution assisted target detection network, including: multi-dimensional feature extraction module, PPA module, multi-dimensional feature fusion module, result prediction module, and Processing is performed to obtain the predicted detection box and predicted category label; Step 3.1: The multi-dimensional feature extraction module consists of d series-connected Dark blocks. Process them sequentially, so that the features output by the d-2th Dark block are used as the nth target one-dimensional features , the features output by the d-1th Dark block are used as the nth target two-dimensional features , the features output by the dth Dark block are used as the nth target 3D features ; Step 3.2: Build the PPA module and install it Processing is performed to obtain the nth three-dimensional enhanced feature representation ; Step 3.3 Multi-dimensional feature fusion module, including: 2 upsampling fusion blocks and 2 downsampling fusion block units, and , and Processing is performed to obtain the nth one-dimensional prediction feature map , the nth two-dimensional prediction feature map And the nth three-dimensional prediction feature map ; Step 3.4: Prediction module pair , , Processing is performed to generate a prediction tensor containing bounding box offset, confidence, and category probability. After decoding the prediction tensor, the predicted detection box and predicted category label are obtained. Step 4: Based on the predicted detection box and predicted category label as well as the true detection box and true category label, use formula (1) to construct the target detection loss : (1) In formula (1), express The positioning loss between the predicted detection box and the real detection box, express The confidence loss of the target existence, express The category classification loss is , and They are the weight coefficients for positioning loss, confidence loss, and category classification loss respectively; Step 5: Construct a super-resolution branch consisting of an encoder and a decoder, and use the features output by the third Dark module in the target detection network as the nth low-level feature , the features output by the 5th Dark module are used as the nth high-level features , thus and Processing is performed to obtain the nth super-resolution reconstruction feature ; Step 6: Based on and , using formula (2) to construct super-resolution loss : (2) Step 7: Use formula (3) to construct the total loss function : (3) In formula (3), and are 2 balance coefficients; Step 8: Use the back propagation algorithm to optimize the training of the target detection network model composed of the multidimensional feature extraction module, PPA module, multidimensional feature fusion module, result prediction module and super-resolution branch, and calculate the total loss function. To update the model parameters until the total loss function The optimal target detection network model after training is obtained until convergence, which is used to predict remote sensing images and obtain images with predicted detection boxes and predicted category labels.

2. The multi-modal fusion detection method for remote sensing targets combining atmospheric scattering removal and super-resolution according to claim 1 is characterized in that: The step 1 is performed according to the following process: Step 1.1: Obtain a visible light remote sensing image dataset with real detection frames and real category labels, normalize the size and preprocess it, and obtain the preprocessed visible light remote sensing image dataset S = { , ,…, ,…, },in, represents the nth preprocessed visible light remote sensing image, where N is the total number of visible light remote sensing images; Step 1.2: Obtain an infrared remote sensing image dataset with real detection frames and real category labels, normalize the size and preprocess it, and obtain the preprocessed infrared remote sensing image dataset T = { , ,…, ,…, },in, Represents the nth infrared remote sensing image after preprocessing; Step 1.3: Set S = { , ,…, ,…, } Input into the network model based on convolutional neural network for processing to obtain the visible light data set with atmospheric scattering effect removed , ,…, ,…, ,in, Represents the nth visible light remote sensing image after atmospheric scattering is removed.

3. The multi-modal fusion detection method for remote sensing targets combining atmospheric scattering removal and super-resolution according to claim 2 is characterized in that: The step 2 is performed according to the following process: The attention mask generation layer in the infrared branch is Process and generate the nth infrared attention map , and After element-by-element multiplication, the nth infrared weighted feature is obtained ; The bottleneck feature extraction layer in the infrared branch is Process and generate the nth infrared feature map ; The attention mask generation layer in the visible light branch is Process and generate the nth visible light attention map , and After element-by-element multiplication, we get the nth visible light weighted feature ; The bottleneck feature extraction layer in the visible light branch is Processing is performed to generate the nth visible light feature map ; Will and Splice along the channel dimension and generate the nth joint feature map Then it is input into the SE module for processing to generate the nth channel weight map ; Will and After channel-by-channel multiplication, the nth fused image is output .

4. The multi-modal fusion detection method for remote sensing targets combining atmospheric scattering removal and super-resolution according to claim 3 is characterized in that: The PPA module in step 3.2 includes: a multi-branch feature extraction module and an attention mechanism module; wherein the multi-branch feature extraction module includes: a parallel local branch, a global branch and a sequence convolution branch; right After the point convolution operation, the three branches are input for processing respectively, and the features at three different levels are obtained and fused. Then, they are processed by the attention mechanism module to obtain the nth three-dimensional enhanced feature representation. .

5. The multi-modal fusion detection method for remote sensing targets combining atmospheric scattering removal and super-resolution according to claim 4 is characterized in that: In step 3.3, each upsampling fusion block includes: an upsampling unit, a feature splicing unit and a CSP unit; each downsampling fusion block includes: a downsampling unit, a feature splicing unit and a CSP unit; Will Input into the first upsampling fusion block, and after the upsampling operation of the first upsampling unit, After being input into the first feature splicing unit for splicing operation, they are then input into the first CSP unit for processing to obtain the nth two-dimensional fusion feature. ; Will Input into the second upsampling fusion block, and after the upsampling operation of the second upsampling unit, After being input into the second feature concatenation unit for concatenation, they are then input into the second CSP unit for processing to obtain the nth one-dimensional prediction feature map. ; Will Input into the first downsampling fusion block, and after the downsampling operation of the first downsampling unit, After being input into the third feature splicing unit for splicing operation, they are input into the third CSP unit for processing to obtain the nth two-dimensional prediction feature map. ; Will Input into the second downsampling fusion block, and after the downsampling operation of the second downsampling unit, After being input into the fourth feature splicing unit for splicing operation, they are input into the fourth CSP unit for processing to obtain the nth three-dimensional prediction feature map. .

6. The multi-modal fusion detection method for remote sensing targets combining atmospheric scattering removal and super-resolution according to claim 5 is characterized in that: The step 5 is performed according to the following process: Step 5.1: The encoder consists of three CR modules and an upsampling module. and Processing is performed to obtain the nth fused feature map ; The first CR module Processing is performed to obtain the nth processed low-level features ; Upsampling module pair Perform bilinear upsampling to make its spatial dimension equal to Align to get the nth processed high-level features ; Will and After splicing along the channel dimension, and then processed by two CR modules in sequence, the nth fused feature map is obtained. ; Step 5.2: The decoder uses the EDSR module to The spatial size is enlarged step by step to obtain the nth super-resolution reconstruction feature .

7. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the remote sensing target multimodal fusion detection method according to any one of claims 1 to 6, and the processor is configured to execute the program stored in the memory.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multimodal fusion detection method for remote sensing targets described in any one of claims 1 to 6 are executed.