Remote sensing target detection method and system for low-visibility image
Through the multimodal image dehazing enhancement and feature distillation mechanism, combined with the hybrid backbone structure of Transformer, Mamba and CNN, the clarity and detection accuracy problems of remote sensing target detection under complex weather conditions are solved, and efficient and stable target detection effects are achieved.
Patent Information
- Application Number
- CN202511249600.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-03
AI Technical Summary
Existing remote sensing target detection methods suffer from reduced image clarity, blurred target boundaries, and low detection accuracy under complex weather conditions such as haze, low illumination, rain and snow. In addition, the Transformer model has high computational complexity, making it difficult to meet real-time and deployment requirements in resource-constrained environments.
It adopts multimodal remote sensing image dehazing and enhancement processing, a hybrid backbone structure integrating Transformer, Mamba and CNN, combines two-dimensional wavelet transform and adaptive selection of key areas to generate HR feature maps, and performs detection through cross-scale aggregation.
It improves target recognition and robustness under low visibility conditions, reduces the computational complexity of the model inference stage, enhances the stability and adaptability of the model in complex weather scenarios, and is suitable for end-to-end deployment.
Smart Images

Figure CN120747487A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing monitoring, and in particular to a remote sensing target detection method and system for low-visibility images. Background Art
[0002] With the development of remote sensing technology, remote sensing images have been widely used in fields such as environmental monitoring, disaster assessment, and urban planning. Object detection, a core task in remote sensing image analysis, aims to automatically identify and accurately locate various ground objects in high-resolution imagery. Convolutional neural networks (CNNs) and Transformers are commonly used in object detection algorithms.
[0003] CNN-based detection methods automatically extract hierarchical features from images by building multi-layer convolutional structures. They possess excellent local feature modeling capabilities and perform well in small object detection and edge information extraction. The Transformer model, through its self-attention mechanism, enables global information exchange between any locations in an image, overcoming the local receptive field limitations of CNNs and demonstrating strong advantages in modeling long-range dependencies and recognizing objects in complex backgrounds.
[0004] While the aforementioned methods have made some progress, technical challenges persist in practical applications. Remote sensing image quality degrades significantly in adverse weather conditions such as haze, low illumination, rain, and snow. This can lead to issues such as lack of clarity, blurred object boundaries, and reduced contrast, making it difficult for detection algorithms to accurately identify ground objects. In particular, the accuracy of detecting small or low-contrast objects decreases significantly. When processing high-resolution remote sensing imagery, the Transformer model suffers from a quadratic computational complexity due to its self-attention mechanism, resulting in high video memory usage and slow inference speeds, making it difficult to meet the real-time and resource-constrained deployment requirements.
[0005] Therefore, it is necessary to develop more accurate and efficient remote sensing target detection methods to improve the accuracy and efficiency of remote sensing target detection. Summary of the Invention
[0006] In order to solve the problems of reduced image clarity, blurred target boundaries, low detection accuracy and poor model robustness in existing remote sensing target detection methods under complex weather conditions such as haze, low illumination, rain and snow, the present invention provides a remote sensing target detection method and system for low-visibility images.
[0007] In a first aspect, the present invention provides a remote sensing target detection method for low-visibility images, which adopts the following technical solutions: A remote sensing target detection method for low-visibility images, comprising: Acquire multimodal remote sensing image data; Perform dehazing and enhancement processing on low-visibility input images; Normalize the dehazed RGB and IR images and then merge them; A hybrid backbone structure integrating Transformer, Mamba, and CNN is used to encode multimodal fusion images layer by layer. Decompose the backbone output features in the frequency domain based on two-dimensional wavelet transform; Generate HR feature maps by adaptively selecting key regions; The detection target box is obtained based on cross-scale aggregation of HR feature maps.
[0008] Furthermore, the dehazing and enhancement processing of the low-visibility input image includes utilizing a teacher network pre-trained on a clear multimodal image to guide a student network to perform feature distillation on the foggy input image. Through gradient alignment and degradation response constraints, the student network learns robust representation capabilities under foggy conditions, thereby enhancing the clarity and contrast of the image and reducing the occlusion effect of haze on small target detection. The final output is an enhanced RGB and IR image for subsequent fusion processing, wherein a preliminary degradation map is obtained by calculating the pixel-level difference between the clear image and the foggy image, and then normalized to quantify the degree of degradation in different areas of the image. Finally, a degradation weighted response distillation method is used to assign higher weights to severely degraded areas according to the degree of degradation. In the response map distillation process of classification and positioning, the student network pays more attention to targets in severely degraded areas, thereby achieving targeted improvement of the student network's detection capability in different degraded areas.
[0009] Furthermore, the degenerate weighted response distillation method includes the following steps: Calculate the pixel difference between the fog image and the clear image, where D represents the difference map, It is a remote sensing image under clear conditions. is the remote sensing image input under low visibility conditions; according to the formula Normalize D, where max(D) represents the maximum pixel difference in the difference map D, and min(D) represents the minimum pixel difference in the difference map D. The degradation degree map obtained after normalization reflects the degradation intensity of each area of the image and is used in the distillation process of the classification and positioning response map to guide the student network to focus on the severely degraded areas. According to the formula, , calculate the weighted response distillation loss; where C, H, and W represent the number of channels, height, and width of the feature map, respectively. and is the activation value of the teacher network and the student network in the cth channel, hth row, and wth column in the response feature map; according to the formula, , we get the dehazing enhancement loss function, where Edge distillation loss is used to constrain the perceptual consistency of the teacher and student networks in terms of edge structure; A degradation-weighted response distillation loss is used to emphasize response consistency in severely degraded regions; General detection losses for the student network, including classification loss and bounding box regression loss; Represents the weight coefficient of edge distillation loss; represents the weight coefficient of the degradation weighted response distillation loss; Represents the weight coefficient of the conventional detection loss.
[0010] Furthermore, the defogging RGB and IR images are normalized and then stitched together, including stitching them into four-channel images (R, G, B, I) along the channel dimension, wherein, according to the formula Get the fused image, where , C is the number of channels, H and W are the height and width of the image, {R, G, B} and {I} represent RGB images and IR images, Concat() represents the connection operation along the channel axis; the dehazed RGB and IR images are normalized to [0, 1]. Specifically, X is subsampled to 1 / n of the original image size to complete the SR module and accelerate the training process. According to the formula The sampled image is obtained, where D() represents n downsampling operations using bilinear interpolation, and then the subsampling results are sent to the backbone to generate multi-level features.
[0011] Furthermore, the hybrid backbone structure that integrates Transformer, Mamba and CNN is used to encode the multimodal fusion image layer by layer. In the high-resolution stage, a lightweight CNN module is used to extract edge and texture local information to adapt to the small target detection task; in the medium-resolution stage, a Mamba-based state space modeling module is introduced, and a multi-directional scanning mechanism is used to capture the global structure of rotated and deformed targets; in the low-resolution stage, a Transformer module based on multi-scale sliding window attention is combined to achieve semantic modeling and regional attention of large-scale targets. The features of each stage are connected across layers and fused at multiple scales to form a complete feature pyramid, and a multi-scale feature map is output. Among them, the dynamic convolution module adaptively generates convolution kernel weights according to the input feature content to achieve personalized processing of different image regions, which is expressed as: , where is the weight coefficient generated according to the input feature x, is the corresponding dynamic convolution kernel, and * is the convolution operation.
[0012] Furthermore, the frequency domain decomposition of the backbone output features based on the two-dimensional wavelet transform includes first extracting preliminary features from the backbone output feature map through a convolution operation, and then applying a low-pass filter LPF and a high-pass filter HPF in the horizontal and vertical directions through a two-dimensional Haar wavelet transform to decompose it into four sub-band components: low-frequency component , horizontal high-frequency component , vertical high-frequency component and diagonal high-frequency components Apply Laplace filter to high frequency components to enhance edge details, and then further strengthen target features and suppress background noise through Squeeze-and-Excitation attention and pixel attention. Apply Gaussian filter to smooth the background and suppress high-frequency noise, combine SE attention adaptively to reduce background redundancy and enhance target contrast; according to the formula , the frequency domain decomposition result of the input feature map after convolution and two-dimensional Haar wavelet transform (HWT) is obtained, where F represents the final output feature map of the backbone network, that is, the multimodal fusion feature after encoding by the hybrid backbone structure of Transformer, Mamba and CNN; f(·) represents the preliminary convolution operation on the final output feature map F to extract basic features, and HWT represents two-dimensional wavelet transform.
[0013] Furthermore, the HR feature map is generated by adaptively selecting key areas, including selecting multi-scale features from the fusion backbone network, wherein the low-level features are first input into the channel recalibration module to enhance the key texture channel response related to reconstruction; the high-level features are upsampled to match the spatial resolution of the low-level features; then the two are concatenated and fused into a set of spatially aligned multi-scale features; the fused features are processed by two serially connected CR modules to enhance inter-channel dependencies and suppress redundant background information; wherein, in the training stage, according to the formula, Calculate the pixel difference between the reconstructed image and the original input; where S represents the reconstructed image generated by upsampling of the SR branch, and X represents the original input image; according to the formula, Get the target detection loss; where, 、 、 are the weights of different layers of the three loss functions, weights 、 、 Adjust the error emphasis between the frame coordinates, frame dimensions, and object, non-object, and classification; according to the formula, The total loss is obtained, where 、 is the loss weight factor.
[0014] Furthermore, the method of generating the HR feature map by adaptively selecting the key area also includes reducing the amount of calculation by adaptively selecting the key area, wherein the formula C , and get the coarse attention, where Q, K, and V are the query, key, and value generated by 1×1 convolution. The high-level features K and V perform cross-layer attention calculations on the bottom-level Q and calculate the fine attention, which is expressed as , after getting careful attention, where and It is a fine-grained token obtained by filtering out the top k most informative key-value pair indexes based on the global similarity s and mapping them to the underlying feature network. Through the formula, , The final output is obtained, where represents convolution, represents upsampling, Represents depthwise convolution.
[0015] Furthermore, the detection target frame is obtained based on cross-scale aggregation of HR feature maps, including semantically guided coarse-grained feature fusion between each layer of the feature pyramid through a cross-layer attention mechanism, using higher-level feature maps as keys and values to query adjacent low-level feature maps, performing cross-layer attention calculations, and generating coarse attention maps. The coarse attention stage is performed in a high-level semantic space, which can significantly reduce the amount of attention calculation while maintaining the context modeling capability; subsequently, key areas with larger response intensity are screened from the coarse attention results as candidate areas for fine attention operations, and standard self-attention operations are performed in this area to finely model and enhance features of local details; through this coarse-first-fine-later attention strategy, the excessive response of traditional global attention to non-critical background areas is effectively avoided, and the model's ability to focus on small targets and edge details is improved.
[0016] In a second aspect, a remote sensing target detection system for low-visibility images includes: The data acquisition module is configured to acquire medical images and multimodal remote sensing image data; The pre-processing module is configured to perform dehazing and enhancement processing on the low-visibility input image; The fusion module is configured to normalize the dehazed RGB and IR images and then perform splicing and fusion; The encoding module is configured to encode the multimodal fusion image layer by layer using a hybrid backbone structure that integrates Transformer, Mamba, and CNN; The decomposition module is configured to perform frequency domain decomposition of the backbone output features based on a two-dimensional wavelet transform; The selection module is configured to generate HR feature maps by adaptively selecting key regions; The detection module is configured to obtain a detection target box based on cross-scale aggregation of the HR feature map.
[0017] In a third aspect, the present invention provides a computer-readable storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded and executed by a processor of a terminal device, for a remote sensing target detection method for low-visibility images.
[0018] In a fourth aspect, the present invention provides a terminal device comprising a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, and the instructions are suitable for being loaded and executed by the processor to implement the remote sensing target detection method for low-visibility images.
[0019] In summary, the present invention has the following beneficial technical effects: Improve target recognizability and robustness in low-visibility conditions: This invention effectively enhances the clarity and contrast of remote sensing images in severe weather conditions such as haze and rain through multimodal image dehazing enhancement and feature distillation mechanisms, reduces the occlusion interference of environmental degradation on small target detection, and enhances the stability and adaptability of the model in complex weather scenarios.
[0020] Reduce the computational complexity of the model inference stage: Using the Mamba, CNN, and Transformer fusion module as the backbone, combined with dynamic convolution and sparse self-attention mechanisms, it can adaptively adjust the calculation area according to the importance of features, achieve key area modeling, reduce redundant computing overhead, improve detection efficiency, and be suitable for end-to-end deployment.
[0021] Strengthen the ability to separate and model target structures and semantic features: By introducing two-dimensional wavelet transform after feature extraction, image features are decoupled into high-frequency structures and low-frequency semantic components, and branch modeling is performed, which effectively improves the model's independent perception of edge contours and semantic areas, and enhances the detection head's ability to comprehensively utilize different types of information. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is a schematic diagram of a remote sensing target detection method for low-visibility images according to Example 1 of the present invention; Figure 2 is a structural diagram of a model according to embodiment 1 of the present invention; Figure 3 This is a bar chart comparing the core performance of the model of Example 1 of the present invention with other models; Figure 4 This is a bar chart comparing the computational performance of the model of Example 1 of the present invention and other models. DETAILED DESCRIPTION
[0023] The present invention will be further described in detail below with reference to the accompanying drawings.
[0024] Example 1 Reference Figure 1 , a remote sensing target detection method for low-visibility images of this embodiment includes: S1: Acquire multimodal remote sensing image data and perform dehazing and enhancement processing on low-visibility input images to improve the clarity and contrast of low-visibility images and reduce the occlusion of small targets by environmental factors such as haze; S2: Normalize the dehazed RGB and IR images to [0, 1] and concatenate them into a four-channel image (R, G, B, I) along the channel dimension to form the fused input; S3: A hybrid backbone architecture combining Transformer, Mamba, and CNN encodes multimodal fusion images layer by layer. This architecture combines the global modeling capabilities of the Transformer, the state-space modeling advantages of Mamba, and the fine-grained capture of local details by CNN. While maintaining efficient modeling, it achieves bottom-up multi-scale feature representation: low-level features focus on preserving texture and edges, while high-level features prioritize semantic relationships and scene structure modeling.
[0025] To enhance its ability to perceive spatial structural changes and small object textures in remote sensing imagery, a lightweight dynamic convolution module is introduced at each stage of the backbone network. This adaptively adjusts the convolution kernel weights based on the spatial distribution of input features, enabling more refined contextual adjustment and response. Through the synergistic effect of state modeling and dynamic convolution, the network can more accurately perceive feature differences at different scales and regions, effectively improving target representation capabilities in low-visibility conditions.
[0026] S4: To further enhance the model's ability to decouple structural and semantic information, a two-dimensional wavelet transform is introduced to perform frequency-domain decomposition of the backbone output features. High-frequency components are rich in texture details and edge contours, making them suitable for structural enhancement; low-frequency components carry scene structure and semantic backbone information, making them suitable for global understanding.
[0027] S5: During the training phase, the SR branch is enabled. The encoder fuses the backbone's multi-layer features. The decoder uses three layers of deconvolution to upsample to twice the size of the input image, generating HR feature maps that guide the backbone network to learn clearer object representations. During the inference phase, the SR branch is removed, and only the fused backbone is used to extract features. S6: Structural-semantic fusion features are aggregated across scales via the detection head. Sparse self-attention is introduced at each level of the feature pyramid, adaptively selecting key regions to reduce computational overhead and prevent global attention from overreacting to background noise. The detection head ultimately outputs object bounding box parameters, confidence scores, and class probabilities.
[0028] Specifically, In S1, the defogging enhancement process is specifically as follows: An object detector that performs well on clear image datasets is used as a teacher network to guide the student network in performing feature distillation on foggy input images. During feature distillation, the Sobel operator is used to calculate the horizontal and vertical gradients of the feature map to highlight edges and semantic information in the image and enhance feature contrast.
[0029] According to the formula, , , , Convolution is performed on the feature map to calculate the edge gradient, where is the horizontal gradient convolution kernel, is the vertical gradient convolution kernel; is the feature map output by a certain layer of the teacher network, Feature maps output by the corresponding layer of the student network.
[0030] According to the formula, , , and get the gradient amplitude.
[0031] By minimizing the difference between the gradient representations of the teacher network and the student network, the student network's ability to perceive the edge of the target is enhanced, helping the student network to extract clearer semantic information in foggy images and improving the quality of feature extraction.
[0032] In addition, a degradation-weighted response distillation method is used. A preliminary degradation map is first generated by calculating the pixel-level difference between the clear and foggy images. This map is then normalized to quantify the degree of degradation in different regions of the image. Finally, a degradation-weighted response distillation method is used to assign higher weights to severely degraded regions based on the degree of degradation. During the response map distillation process for classification and localization, the student network is more focused on targets in severely degraded regions, thereby improving the student network's detection capabilities in different degraded regions. The final output is enhanced RGB and IR images for subsequent fusion processing.
[0033] According to the formula, Calculate the pixel difference between the fog image and the clear image, where D represents the difference map, It is a remote sensing image under clear conditions. is the input remote sensing image under low visibility conditions: According to the formula, Normalize D to obtain a degradation degree map, which reflects the degradation intensity of each region of the image: it is used in the distillation process of classification and localization response maps to guide the student network to focus on severely degraded areas.
[0034] According to the formula, , calculate the weighted response distillation loss. Where C, H, W represent the number of channels (Channel), height (Height), and width (Width) of the feature map respectively. and is the activation value of the teacher network and the student network in the cth channel, hth row, and wth column in the response feature map.
[0035] According to the formula, , we get the dehazing enhancement loss function. Edge distillation loss is used to constrain the perceptual consistency of the teacher and student networks in terms of edge structure; A degradation-weighted response distillation loss is used to emphasize response consistency in severely degraded regions; General detection losses for the student network, including classification loss and bounding box regression loss; Represents the weight coefficient of edge distillation loss; represents the weight coefficient of the degradation weighted response distillation loss; Represents the weight coefficient of the conventional detection loss.
[0036] In S2, the dehazed RGB and IR images are normalized to [0,1], specifically, For RGB images, which are usually in 8-bit format, the pixel value range is between 0-2550, so according to the formula, , The result is a three-channel image with a shape of (H, W, 3) and a value range between 0 and 1.
[0037] For infrared images IR, they are usually single-channel images, and their bit depth may be 8-bit or 16-bit.
[0038] If it is an 8-bit image, according to the formula, , Get the IR normalized result.
[0039] If it is a 16-bit image, according to the formula,
[0040] Get the IR normalized result.
[0041] In S2, the four-channel image (R, G, B, I) is spliced along the channel dimension. Specifically, According to the formula, Get the fused image, Where, , C is the number of channels, H and W are the height and width of the image, {R,G,B} and {I} represent RGB images and IR images, and Concat() represents the connection operation along the channel axis.
[0042] The IR image needs to be expanded by one channel dimension first. According to the formula, , Its shape is converted from (H, W) to (H, W, 1), and then it is spliced with the RGB image by channel to obtain the fused image After stitching, the shape of the image is (H, W, 4), and the channel order is usually R, G, B, IR.
[0043] In S2, normalized to [0,1], specifically, In order to speed up the training process and improve the efficiency of the SR module, the original image is first downsampled to 1 / n the size of the original image to complete the SR module and speed up the training process.
[0044] According to the formula, Get the downsampled image.
[0045] Where D() represents the downsampling operation using bilinear interpolation, X represents the original image, and n represents the downsampling factor. Bilinear interpolation calculates the weighted average of the four surrounding pixels to generate the value of each pixel in the new image, achieving a smooth reduction effect. After downsampling, the image is passed through a multi-layer convolutional network to extract and encode features at different scales, forming a multi-level feature map from low-level to high-level layers.
[0046] S3. Multimodal Fusion A hybrid backbone structure integrating Transformer, Mamba and CNN is used to encode the multimodal fusion image layer by layer. Specifically, At the high-resolution stage, a lightweight CNN module is used to extract local information such as edges and textures, making it suitable for small object detection. The input image undergoes preliminary processing via convolutional layers to extract low-level features, such as edges, corners, and other details. Multiple convolutional layers are used to extract features layer by layer, ensuring that local texture and edge information are fully captured. Convolution operations with a stride of 1 ensure that the spatial resolution of the feature map does not decrease excessively, ensuring that detailed information about small objects is captured. Batch normalization and the ReLU activation function are used to improve feature representation and accelerate training.
[0047] According to the formula, , and obtain the high-resolution stage output feature map.
[0048] Among them, W is the convolution kernel, b is the bias, is the input feature map.
[0049] At the medium-resolution stage, a Mamba-based state-space modeling module is introduced, employing a multi-directional scanning mechanism to capture the global structure of rotated and deformed targets. This state-space model is particularly suitable for modeling the spatiotemporal characteristics of images, making it particularly suitable for handling complex situations such as target rotation and deformation. Convolution operations in different directions are used to scan the target in all directions to capture its global structure. Images are processed using rotation-invariant convolution kernels to ensure that target features are effectively captured at all angles.
[0050] According to the formula, , and obtain the output feature map of the medium-resolution stage.
[0051] in, Parameters that model the state space, is the input feature map.
[0052] At low resolution, the Transformer module, combined with a multi-scale sliding window attention algorithm, enables semantic modeling and regional attention for large-scale objects (such as airports and woodlands). A self-attention mechanism weights different regions within the image, with particular attention paid to large-scale objects. A multi-scale sliding window approach is used to partition the image into regions, ensuring comprehensive perception of objects of varying sizes. Based on the features within each window, the similarity between them and the global features is calculated. Global semantic modeling is achieved by weightedly combining features from different regions.
[0053] According to the formula, , and obtain the weighted output feature map of the low-resolution stage.
[0054] Among them, Attention represents the self-attention operation, the input is the three mappings of the same feature map: query, key, and value, and the output is the weighted feature map.
[0055] Features from each stage are integrated through cross-layer connections and multi-scale fusion to form a complete feature pyramid. In the feature pyramid network, low-resolution features are fused with high-resolution features to maintain a balance between detail and semantic information. During the feature fusion process, cross-scale connections are used to effectively represent the multi-level image information at multiple scales. Finally, the output multi-scale feature map is passed through the detection head to predict object bounding box parameters, estimate confidence, and output object category probabilities.
[0056] In S3, dynamic convolution can further refine the feature extraction process by adaptively adjusting the convolution kernel weights to achieve personalized processing in different regions and scales, thereby enhancing the model's sensitivity to local image features. Specifically, The dynamic convolution module adaptively generates convolution kernel weights according to the input feature content to achieve personalized processing of different image regions.
[0057] According to the formula, , Where, is the weight coefficient generated according to the input feature x, is the corresponding dynamic convolution kernel, and * is the convolution operation.
[0058] In S4, the two-dimensional wavelet transform is specifically, The feature map output by the backbone, according to the formula, Preliminary features are extracted through convolution operations. Here, F represents the original feature map output by the backbone, f(·) represents the convolution operation, W is the convolution kernel weight, b is the bias term, and * represents the convolution operation.
[0059] Then, the two-dimensional Haar wavelet transform is applied in the horizontal and vertical directions in sequence with low-pass filters (LPF) and high-pass filters (HPF), and the image is decomposed into four sub-band components: low-frequency component ( ), horizontal high-frequency components ( ), vertical high-frequency component ( ) and diagonal high-frequency components ( ).
[0060] According to the formula, , we get the frequency domain decomposition result of the input feature map after convolution and two-dimensional Haar wavelet transform (HWT). Here, F represents the final output feature map of the backbone network, i.e., the multimodal fusion feature after encoding by the hybrid backbone structure integrating Transformer, Mamba, and CNN; f(·) represents the preliminary convolution operation on the final output feature map F to extract basic features. HWT represents two-dimensional wavelet transform. Specifically, According to the formula, The low-frequency components are obtained, effectively preserving the main structural features of the image.
[0061] Apply a high-pass filter horizontally and a low-pass filter vertically, then downsample by a factor of 2 to obtain the horizontal high-frequency components. Apply a low-pass filter horizontally and a high-pass filter vertically, then downsample by a factor of 2 to obtain the vertical high-frequency components. Apply a high-pass filter horizontally and vertically, then downsample by a factor of 2 to obtain the diagonal high-frequency components.
[0062] The low-pass filter is mainly used to retain the low-frequency components in the image, which correspond to the main structural features and background information of the image. The input feature map is processed in the horizontal and vertical directions. Where D(u,v) represents the radial distance to the center of the spectrum in the frequency domain, is the cutoff frequency.
[0063] High-pass filter is used to extract high-frequency components in images. These components contain texture, edge and other details, which are crucial for target detection. High-pass filter is based on the formula Process the input feature map.
[0064] In S5, enabling the SR branch is specifically as follows: The SR branch guides the backbone network to learn clearer object representations during training, thereby improving the resolution and detection accuracy of small objects in low-visibility remote sensing images. This branch extracts multi-scale features from the fused backbone network, including a low-level feature and a high-level feature.
[0065] In the encoder, low-level features are first fed into a channel recalibration module (CR module) to enhance the responses of key texture channels relevant for reconstruction. High-level features are then upsampled to match the spatial resolution of the low-level features. The two are then concatenated and fused into a set of spatially aligned multi-scale features. The fused features are further processed through two serially connected CR modules to enhance inter-channel dependencies and suppress redundant background information, thereby improving the structural integrity and detail fidelity of the super-resolved representation.
[0066] In the decoder, the fused feature maps are progressively upsampled through three deconvolution layers to reconstruct a high-resolution feature map with a spatial size twice that of the input image. This high-resolution output serves as the image reconstruction result and is pixel-wise aligned with the original input image.
[0067] In the training phase, according to the formula, Calculate the pixel difference between the reconstructed image and the original input. Where S represents the reconstructed image generated by upsampling of the SR branch, and X represents the original input image.
[0068] According to the formula, Get the target detection loss. Among them, 、 、 are the weights of different layers of the three loss functions, weights 、 、 Adjusts box coordinates, box dimensions, and error emphasis between objectness, non-objectness, and classification.
[0069] According to the formula, The total loss is obtained, where 、 is the loss weight factor.
[0070] In S6, sparse self-attention is introduced at each level of the feature pyramid to reduce the amount of computation by adaptively selecting key areas. Specifically: According to the formula, C , we get the coarse attention, where Q, K, and V are the query, key, and value generated by 1×1 convolution. The high-level features K and V perform cross-layer attention calculation on the low-level Q.
[0071] By formula, , get fine attention, where, and It is a fine-grained token obtained by filtering out the top k most informative key-value pair indexes according to the global similarity s and mapping them to the underlying feature network.
[0072] By formula, , Get the final output.
[0073] Experimental verification: To systematically validate the proposed method's object detection capabilities in low-visibility multimodal remote sensing imagery, we constructed a custom multimodal low-visibility small target detection dataset based on the VEIDAI remote sensing dataset. This dataset comprises two modalities: ① RGB low-visibility modality: Using an atmospheric scattering model to synthesize low-visibility images from raw VEIDAI visible light images, simulating realistic haze weather conditions; and ② infrared modality: Using raw infrared imaging, which has a strong ability to penetrate haze. The dataset contains 1,246 images at either 1024×1024 or 512×512 resolution, encompassing typical remote sensing scenes such as grasslands, roads, mountains, and cities. Each image contains small targets of varying scales.
[0074] The MPE-YOLO and SuperYOLO methods were set up for performance comparison with the present invention. The complete framework proposed by the present invention integrates the Transformer, CNN and Mamba models, and adds the corresponding self-attention mechanism.
[0075] All methods were evaluated using the same training and test sets. Evaluation metrics include: ① Precision: The ratio of samples predicted by the model to samples that actually belong to a particular class. This metric is specific to a particular class and reflects the model's prediction accuracy for that class. ② mAP@0.5: The mAP value at an IoU threshold of 0.5. ③ Parameters: The number of model parameters, also used to measure algorithm / model complexity. ④ GFLOPs: The number of floating-point operations per second.
[0076] Table 1 Comparison of performance data under different methods mAP@0.5 Precision Prarams GFLOPs MPE-YOLO 0.5366 0.4023 4.2 17.3 SuperYOLO 0.7025 0.6735 7.7 56.2 The present invention 0.8511 0.8203 8.5 70.3 Because MPE-YOLO cannot directly process low-visibility remote sensing images and can only process RGB single-modal images, its performance is significantly poor when processing custom datasets, and its parameters and computational complexity are very small. Although SuperYOLO cannot directly process low-visibility images, it can compensate with IR infrared images, which will improve accuracy to a certain extent, but it is still insufficient for extremely low-visibility environments. This invention further deepens multimodal fusion and target detection, improving detection accuracy, but its model parameters and computational complexity are relatively high.
[0077] Example 2 This embodiment provides a remote sensing target detection system for low-visibility images, including: The data acquisition module is configured as follows: A computer-readable storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded and executed by a processor of a terminal device, for a remote sensing target detection method for low-visibility images.
[0078] A terminal device includes a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, wherein the instructions are suitable for being loaded and executed by the processor to perform a remote sensing target detection method for low-visibility images.
[0079] The above are all preferred embodiments of the present invention, and are not intended to limit the scope of protection of the present invention. Therefore, any equivalent changes made based on the structure, shape, and principle of the present invention should be included in the scope of protection of the present invention.
Claims
1. A remote sensing target detection method for low-visibility images, characterized in that: include: Acquire multimodal remote sensing image data; Perform dehazing and enhancement processing on low-visibility input images; Normalize the dehazed RGB and IR images and then merge them; A hybrid backbone structure integrating Transformer, Mamba, and CNN is used to encode multimodal fusion images layer by layer. Decompose the backbone output features in the frequency domain based on two-dimensional wavelet transform; Generate HR feature maps by adaptively selecting key regions; The detection target box is obtained based on cross-scale aggregation of HR feature maps.
2. The remote sensing target detection method for low-visibility images according to claim 1, characterized in that: The dehazing and enhancement processing of the low-visibility input image includes utilizing a teacher network pre-trained on a clear multimodal image to guide a student network to perform feature distillation on the foggy input image. Through gradient alignment and degradation response constraints, the student network learns robust representation capabilities under foggy conditions, thereby enhancing the clarity and contrast of the image and reducing the occlusion effect of haze on small target detection. The final output is an enhanced RGB and IR image for subsequent fusion processing, wherein a preliminary degradation map is obtained by calculating the pixel-level difference between the clear image and the foggy image, and then normalized to quantify the degree of degradation in different areas of the image. Finally, a degradation weighted response distillation method is used to assign higher weights to severely degraded areas according to the degree of degradation. In the response map distillation process of classification and positioning, the student network is made to pay more attention to targets in severely degraded areas, thereby achieving targeted improvement of the student network's detection capability in different degraded areas.
3. The remote sensing target detection method for low-visibility images according to claim 2, characterized in that: The degenerate weighted response distillation method includes the following formula: Calculate the pixel difference between the fog image and the clear image, where D represents the difference map, It is a remote sensing image under clear conditions. is the remote sensing image input under low visibility conditions; according to the formula Normalize D, where max(D) represents the maximum pixel difference in the difference map D, and min(D) represents the minimum pixel difference in the difference map D. The degradation degree map obtained after normalization reflects the degradation intensity of each area of the image and is used in the distillation process of the classification and positioning response map to guide the student network to focus on the severely degraded areas. According to the formula, , calculate the weighted response distillation loss; Where C, H, and W represent the number of channels, height, and width of the feature map, respectively. and is the activation value of the teacher network and the student network in the cth channel, hth row, and wth column in the response feature map; according to the formula, , we get the dehazing enhancement loss function, where Edge distillation loss is used to constrain the perceptual consistency of the teacher and student networks in terms of edge structure; A degradation-weighted response distillation loss is used to emphasize response consistency in severely degraded regions; General detection losses for the student network, including classification loss and bounding box regression loss; Represents the weight coefficient of edge distillation loss; represents the weight coefficient of the degradation weighted response distillation loss; Represents the weight coefficient of the conventional detection loss.
4. The remote sensing target detection method for low-visibility images according to claim 3, characterized in that: The dehazed RGB and IR images are normalized and then stitched together, including stitching them into four-channel images (R, G, B, I) along the channel dimension, where according to the formula Get the fused image, where , C is the number of channels, H and W are the height and width of the image, {R, G, B} and {I} represent RGB images and IR images, Concat() represents the connection operation along the channel axis; the dehazed RGB and IR images are normalized to [0, 1]. Specifically, X is subsampled to 1 / n of the original image size to complete the SR module and accelerate the training process. According to the formula The sampled image is obtained, where D() represents n downsampling operations using bilinear interpolation, and then the subsampling results are sent to the backbone to generate multi-level features.
5. The remote sensing target detection method for low-visibility images according to claim 4, characterized in that: The hybrid backbone structure that integrates Transformer, Mamba, and CNN is used to encode the multimodal fusion image layer by layer. In the high-resolution stage, a lightweight CNN module is used to extract local edge and texture information to adapt to the small target detection task. In the medium-resolution stage, a Mamba-based state space modeling module is introduced, and a multi-directional scanning mechanism is used to capture the global structure of rotated and deformed targets. In the low-resolution stage, a Transformer module based on multi-scale sliding window attention is combined to achieve semantic modeling and regional attention of large-scale targets. Features at each stage are connected across layers and fused at multiple scales to form a complete feature pyramid, and a multi-scale feature map is output. Among them, the dynamic convolution module adaptively generates convolution kernel weights according to the input feature content to achieve personalized processing of different image regions, which is expressed as: , where is the weight coefficient generated according to the input feature x, is the corresponding dynamic convolution kernel, and * is the convolution operation.
6. The remote sensing target detection method for low-visibility images according to claim 5, characterized in that: The frequency domain decomposition of the backbone output features based on the two-dimensional wavelet transform includes first extracting preliminary features from the backbone output feature map through a convolution operation, and then applying a low-pass filter LPF and a high-pass filter HPF in the horizontal and vertical directions through a two-dimensional Haar wavelet transform to decompose it into four sub-band components: low-frequency component , horizontal high-frequency component , vertical high-frequency component and diagonal high-frequency components Apply Laplace filter to high frequency components to enhance edge details, and then further strengthen target features and suppress background noise through Squeeze-and-Excitation attention and pixel attention. Apply Gaussian filter to smooth the background and suppress high-frequency noise, combine SE attention adaptively to reduce background redundancy and enhance target contrast; according to the formula , we get the frequency domain decomposition result of the input feature map after convolution and two-dimensional Haar wavelet transform (HWT). Here, F represents the final output feature map of the backbone network, that is, the multimodal fusion feature after encoding by the hybrid backbone structure of Transformer, Mamba and CNN; f(·) represents the preliminary convolution operation on the final output feature map F to extract basic features, and HWT represents two-dimensional wavelet transform.
7. The remote sensing target detection method for low-visibility images according to claim 6, characterized in that: The method of generating HR feature maps by adaptively selecting key regions includes selecting multi-scale features from a fusion backbone network, wherein low-level features are first input into a channel recalibration module to enhance the key texture channel response related to reconstruction; high-level features are upsampled to match the spatial resolution of low-level features; then the two are concatenated and fused into a set of spatially aligned multi-scale features; the fused features are processed by two serial CR modules to enhance inter-channel dependencies and suppress redundant background information; wherein, in the training phase, according to the formula, Calculate the pixel difference between the reconstructed image and the original input; where S represents the reconstructed image generated by upsampling of the SR branch, and X represents the original input image; according to the formula, Get the target detection loss; where, 、 、 are the different layers of the three loss functions The weight, weight 、 、 Adjust the error emphasis between the frame coordinates, frame dimensions, and object, non-object, and classification; according to the formula, The total loss is obtained, where 、 is the loss weight factor.
8. The remote sensing target detection method for low-visibility images according to claim 7, characterized in that: The method of generating the HR feature map by adaptively selecting the key area also includes reducing the amount of calculation by adaptively selecting the key area, wherein, by formula C Get coarse attention C , where Q, K, V are the query, key, and value generated by 1×1 convolution, T is the transpose, the dimension of the dk line vector K, and the high-level features K and V perform cross-layer attention calculations on the bottom-level Q and calculate the fine attention , expressed as , after getting careful attention, where and It is a fine-grained token obtained by filtering out the top k most informative key-value pair indexes based on the global similarity s and mapping them to the underlying feature network. Through the formula, , The final output is obtained, where represents convolution, represents upsampling, Represents depthwise convolution.
9. The remote sensing target detection method for low-visibility images according to claim 8, characterized in that: The detection target box is obtained by cross-scale aggregation of HR feature maps, including semantically guided coarse-grained feature fusion between each layer of the feature pyramid through a cross-layer attention mechanism, using higher-layer feature maps as keys and values to query adjacent lower-layer feature maps, performing cross-layer attention calculations, and generating coarse attention maps. This coarse attention stage is performed in a high-level semantic space, which can significantly reduce the amount of attention calculation while maintaining context modeling capabilities; Subsequently, key areas with greater response intensity are screened out from the coarse attention results as candidate areas for fine attention operations. Standard self-attention operations are performed in these areas to finely model local details and enhance features. Through this coarse-first-fine-later attention strategy, the excessive response of traditional global attention to non-critical background areas is effectively avoided, and the model's ability to pay attention to small targets and edge details is improved.
10. A remote sensing target detection system for low-visibility images, characterized in that: include: A data acquisition module is configured to acquire medical images; Acquire multimodal remote sensing image data; The pre-processing module is configured to perform dehazing and enhancement processing on the low-visibility input image; The fusion module is configured to normalize the dehazed RGB and IR images and then perform splicing and fusion; The encoding module is configured to encode the multimodal fusion image layer by layer using a hybrid backbone structure that integrates Transformer, Mamba, and CNN; The decomposition module is configured to perform frequency domain decomposition of the backbone output features based on a two-dimensional wavelet transform; The selection module is configured to generate HR feature maps by adaptively selecting key regions; The detection module is configured to obtain a detection target box based on cross-scale aggregation of the HR feature map.
Citation Information
Patent Citations
Remote sensing image space-time fusion method and system based on Swin Transform and CNN parallel interaction
CN118470479A
Remote sensing image vehicle detection method based on super-resolution and multi-modal fusion
CN118982800A
Multi-modal image matching method and system based on saliency graph structure enhancement
CN120543997A
System and Method for Compressing and Restoring Data Using Hierarchical Autoencoders and a Knowledge-Enhanced Correlation Network
US20250260417A1
Cited By
Self-adaptive image restoration method based on degradation type and degree joint perception
CN120953136A
Plasmodium infection stage detection method and system
CN121033053A
Unmanned aerial vehicle foggy day target detection method
CN121170657A
Unmanned aerial vehicle fog target detection method
CN121170657B
Intelligent water conservancy image processing method and system
CN121190772A