A remote sensing image cloud shadow detection method, system, terminal and storage medium

By employing wavelet transform and heterogeneous feature decoupling techniques, the problems of misjudgment and missed detection in cloud shadow detection in remote sensing images were solved, achieving physical separation and geometric alignment of cloud shadow features and improving detection performance.

CN122223013APending Publication Date: 2026-06-16NANJING UNIV OF INFORMATION SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610654241.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-13
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing remote sensing image cloud shadow detection methods lack modeling of the spatial correspondence between cloud shadows, which makes it easy for detection results to be misjudged and missed. Furthermore, cloud shadows are easily confused with low-reflectivity ground features, affecting the detection effect.

Method used

Wavelet transform is used to decompose optical remote sensing image data to generate low-frequency and high-frequency sub-bands. An enhanced signal is generated through a cloud shadow feature spatial frequency enhancement module. Combined with heterogeneous feature decoupling and texture matching, heterogeneous centroid offset vectors are calculated to generate pseudo-cloud feature maps and pseudo-cloud shadow feature maps. Finally, cloud shadow segmentation is performed through a dual-head decoder to achieve physical separation and geometric alignment of cloud shadows.

Benefits of technology

It effectively captures detailed information of cloud and shadow boundaries, alleviates the problems of shadow blurring and phase reversal, achieves physical separation of cloud and shadow features and cross-class geometric alignment, and ensures the cloud and shadow detection effect in complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122223013A_ABST
    Figure CN122223013A_ABST
Patent Text Reader

Abstract

The application discloses a kind of remote sensing image cloud shadow detection method, system, terminal and storage medium of remote sensing data processing technical field, to solve the lack of cloud shadow spatial correspondence of prior art, and cloud shadow and low reflection ground object are easily confused, leading to detection result is prone to misjudgment and missed detection problem.It includes obtaining optical remote sensing image data, inputting optical remote sensing image data into window attention module to extract multi-scale spatial features, using wavelet transform to decompose multi-scale spatial features in frequency domain, generate low-frequency subband and high-frequency subband;Cloud shadow feature space-frequency enhancement module is inputted to low-frequency subband and high-frequency subband, and cloud shadow feature enhancement signal is generated;The application realizes the collaborative modeling of cloud and cloud shadow in optical remote sensing image in space and frequency domain, effectively solves the problem that cloud shadow detection is difficult, pairing is inconsistent and geometric constraint is insufficient in complex scene, ensures the cloud shadow detection effect of the application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method, system, terminal, and storage medium for detecting cloud shadows in remote sensing images, belonging to the field of remote sensing data processing technology. Background Technology

[0002] In the application of optical remote sensing images, remote sensing images are widely used in meteorological monitoring, disaster early warning, and surface information extraction. However, in actual imaging processes, clouds and their shadows can severely obstruct and interfere with surface information, reducing the usability and interpretation accuracy of remote sensing data. Therefore, how to accurately detect and separate clouds and cloud shadows has become one of the key issues in remote sensing image processing. Existing methods mainly rely on spectral features, threshold segmentation, and deep learning models to identify clouds and cloud shadows. However, due to the similarity in spectral response between clouds and highly reflective ground objects, as well as the confusion between cloud shadows and low-reflective ground objects, the detection results are prone to misjudgment and missed detection. In addition, most methods rely only on local features or pixel-level classification, lacking the ability to model the spatial correspondence and physical causes between clouds and cloud shadows, and making it difficult to characterize the geometric offset between the two.

[0003] As can be seen from the above, existing cloud shadow detection methods mainly rely on spectral features, threshold segmentation, and deep learning models to identify clouds and cloud shadows. They lack spatial correspondence between cloud shadows and low-reflectivity ground objects, which can easily lead to misjudgments and missed detections, thus affecting the actual detection effect. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method, system, terminal, and storage medium for detecting cloud shadows in remote sensing images. This method effectively captures detailed information about cloud shadow boundaries, alleviates the problems of blurring of shadow areas and phase reversal in traditional methods, achieves physical separation of cloud shadow features, realizes cross-category geometric alignment and physical-driven completion of cloud shadow features, and realizes collaborative modeling of clouds and cloud shadows in the spatial and frequency domains in optical remote sensing images. This effectively solves the problems of difficult cloud shadow detection, inconsistent pairing, and insufficient geometric constraints in complex scenes, ensuring the cloud shadow detection effect of this invention.

[0005] To solve the above-mentioned technical problems, the present invention is implemented using the following technical solution:

[0006] In a first aspect, the present invention provides a method for detecting cloud shadows in remote sensing images, comprising:

[0007] Acquire optical remote sensing image data, input the optical remote sensing image data into the window attention module to extract multi-scale spatial features, and use wavelet transform to decompose the multi-scale spatial features in the frequency domain to generate low-frequency sub-bands and high-frequency sub-bands.

[0008] The low-frequency sub-band and high-frequency sub-band are input into the cloud shadow feature spatial frequency enhancement module to generate the cloud shadow feature enhancement signal;

[0009] Bottleneck features are obtained based on the cloud shadow feature enhancement signal. The bottleneck features are then input into the cloud shadow heterogeneous feature decoupling module to obtain cloud semantic feature map and cloud shadow semantic feature map. Based on the cloud semantic feature map and cloud shadow semantic feature map, clean cloud features and clean cloud shadow features are obtained.

[0010] Calculate the heterogeneous centroid offset vector based on the clean cloud features and clean cloud shadow features;

[0011] Based on clean cloud features, clean cloud shadow features, and heterogeneous centroid offset vectors, pseudo cloud feature maps and pseudo cloud shadow feature maps are obtained. Based on clean cloud features, clean cloud shadow features, pseudo cloud feature maps, and pseudo cloud shadow feature maps, bidirectional texture matching scores are calculated, and edge protection gating weights are obtained. Based on bidirectional texture matching scores and edge protection gating weights, adaptive enhancement factors are generated.

[0012] Based on the adaptive enhancement factor, cloud semantic feature map, and cloud shadow semantic feature map, enhanced cloud semantic feature map and enhanced cloud shadow semantic feature map are obtained. The enhanced cloud semantic feature map and enhanced cloud shadow semantic feature map are fed into the dual-head decoder to complete cloud shadow segmentation. The enhanced cloud semantic feature map and enhanced cloud shadow semantic feature map are reconstructed by the dual-head decoder under the constraint of geometric consistency loss. At the same time, the pixel-level mask of cloud and cloud shadow is output to complete the detection work.

[0013] Furthermore, the step of inputting the low-frequency sub-band and high-frequency sub-band into the cloud shadow feature spatial frequency enhancement module to generate the cloud shadow feature enhancement signal includes:

[0014] A spatial weight map is generated based on the low-frequency subband and dark perception gating, with the specific expression as follows:

[0015]

[0016] In the formula: For spatial weighting, It is the Sigmoid activation function. For a convolution operation with a kernel size of 3, For low-frequency sub-band;

[0017] The high-frequency subband is converted to the complex frequency domain, and the cross-correlation spectrum between the cloud and its shadow is calculated using a two-way cross-correlation compensation mechanism. The specific expression is as follows:

[0018]

[0019]

[0020] In the formula: For horizontal high-frequency sub-bands, For high-frequency sub-bands in the vertical direction, For Fast Fourier Transform, This is a conjugate operation. A learnable phase rotation compensation factor. To convert to the complex frequency domain , To convert to the complex frequency domain , For Hanning window, The cross-correlation spectrum;

[0021] Obtain dark sensing weights, use a dynamic spectrum codebook to filter the cross-correlation spectrum, retain the spectrum that meets the cloud-shadow co-occurrence condition, reconstruct and generate the cloud-shadow enhanced signal, combine the spatial weight map and dark sensing weights, perform physical consistency enhancement on the high-frequency sub-band to obtain the enhanced high-frequency sub-band, perform inverse wavelet transform based on the low-frequency sub-band and the enhanced high-frequency sub-band to generate the cloud-shadow feature enhancement signal, the specific expression is as follows:

[0022]

[0023]

[0024]

[0025]

[0026] In the formula: To enhance the signal of cloud shadow, This represents the inverse fast Fourier transform. Indicates a dynamic spectrum codebook. This is a preset constant used to prevent the denominator from being zero. For the enhanced , For the enhanced , To enhance the signal for cloud shadow features, High-frequency sub-bands in the diagonal direction. For inverse wavelet transform, For dark perception weights, It is a learnable enhancement function.

[0027] Furthermore, the step of obtaining bottleneck features based on cloud shadow feature enhancement signals includes:

[0028] The cloud shadow feature enhancement signal is flattened and transposed, and then input into a linear layer to achieve dimensionality compression. The compressed cloud shadow feature enhancement signal is then input into the next layer window attention module for repeated operations to obtain high-dimensional features. Bottleneck features are generated based on the high-dimensional features, as shown in the following expression:

[0029]

[0030] In the formula: Bottleneck characteristics It is the ReLU activation function. For group normalization, For a convolution operation with a kernel size of 3, High-dimensional features;

[0031] The step of inputting bottleneck features into the cloud-shadow heterogeneous feature decoupling module to obtain cloud semantic feature maps and cloud-shadow semantic feature maps includes:

[0032] The bottleneck features are input into the cloud-shadow heterogeneous feature decoupling module, and two independent sets of cloud prototype vectors and cloud shadow prototype vectors are formed through channel projection. Based on the cloud prototype vectors and cloud shadow prototype vectors, cloud semantic feature maps and cloud shadow semantic feature maps are obtained. The specific expressions are as follows:

[0033]

[0034] In the formula: For cloud semantic feature maps, This is a semantic feature map of cloud shadows. This represents the prototype vector of the k-th cloud. This represents the k-th cloud shadow prototype vector, where k is the number of cloud prototype vectors and cloud shadow prototype vectors.

[0035] The step of obtaining clean cloud features and clean cloud shadow features based on the cloud semantic feature map and cloud shadow semantic feature map includes:

[0036] Based on the cloud semantic feature map, cloud shadow semantic feature map, and bottleneck features, cloud similarity and cloud shadow similarity are calculated, and the specific expressions are as follows:

[0037]

[0038] In the formula: For cloud similarity, For cloud shadow similarity, For normalized exponential functions, For a learnable projection matrix, for The transpose form, for The transpose form, This is the scaling factor;

[0039] Based on cloud similarity and cloud shadow similarity, the cloud enhancement signal and cloud shadow enhancement signal are weighted and aggregated, and their residuals are mapped back to the original feature space. After reconstruction, clean cloud features and clean cloud shadow features are obtained. The specific expressions are as follows:

[0040]

[0041]

[0042] In the formula: Characteristics of clean clouds. The characteristics of clean cloud shadows, For gradient clipping operators, To enhance cloud signal, To enhance the signal of cloud shadow, for The transpose form, for The transpose of .

[0043] Furthermore, the step of calculating the heterogeneous centroid offset vector based on clean cloud features and clean cloud shadow features includes:

[0044] Clean cloud features and clean cloud shadow features are mapped to the same physical coordinate space. Using the activation heatmap as weights, a weighted average is applied to the preset coordinate matrix Z to calculate the physical centroid coordinates of the cloud and the cloud shadow. The specific expressions are as follows:

[0045]

[0046]

[0047]

[0048] In the formula: This represents a two-dimensional cloud feature heatmap that has undergone mean pooling along the channel dimension. This represents a two-dimensional cloud shadow feature heatmap after mean pooling across all channels, where C represents the number of feature channels. This represents the coordinates of the pixel with x-coordinate i and y-coordinate j in the coordinate matrix Z. This represents the width of the coordinate matrix. Represents the height of the coordinate matrix. This is a preset constant used to prevent the denominator from being zero. The physical centroid coordinates of the cloud, The physical centroid coordinates of the cloud shadow. express The coordinates of the pixel with x-coordinate i and y-coordinate j. express The coordinates of the pixel with x-coordinate i and y-coordinate j;

[0049] Based on the physical centroid coordinates of the cloud and its shadow, the heterogeneous centroid offset vector is calculated, as shown in the following expression:

[0050]

[0051] In the formula: This is the offset vector of the heterogeneous centroid.

[0052] Furthermore, the step of obtaining pseudo-cloud feature maps and pseudo-cloud shadow feature maps based on clean cloud features, clean cloud shadow features, and heterogeneous centroid offset vectors includes:

[0053] The clean cloud features, clean cloud shadow features, and heterogeneous centroid offset vector are input into the forward and backward affine transformation matrices. Based on the forward and backward affine transformation matrices, pseudo-cloud feature maps and pseudo-cloud shadow feature maps are obtained, as shown in the following expressions:

[0054]

[0055]

[0056]

[0057]

[0058] In the formula: It is the forward affine transformation matrix. This is the inverse affine transformation matrix. The horizontal distance between clean cloud features and false cloud features. The vertical distance between clean cloud features and false cloud features. For bilinear interpolation resampling, To generate the sampling network; This is a pseudo-cloud feature map. This is a pseudo-cloud shadow feature map. It is an affine transformation operator;

[0059] Based on the clean cloud features, clean cloud shadow features, pseudo-cloud feature map, and pseudo-cloud shadow feature map, the bidirectional texture matching score is calculated, as shown in the following expression:

[0060]

[0061]

[0062]

[0063]

[0064] In the formula: Texture components that represent clean cloud features. For clean cloud-like texture components, This represents the channel compression and texture feature extraction operator. The texture component of the pseudo-cloud feature map. For the texture components of the pseudo-cloud shadow feature map, This represents the matching score of the cloud mapping to the cloud shadow. This represents the matching score of the cloud shadow mapping to the cloud. Scoring for bidirectional texture matching;

[0065] The specific expression for obtaining the edge protection gating weight is as follows:

[0066]

[0067]

[0068] In the formula: For the predicted thickness features of clouds, The predicted thickness feature of cloud shadows. A thickness sensor representing the cloud. The thickness sensor of the cloud shadow. Indicates the edge protection gating weight;

[0069] The adaptive enhancement factor is generated based on the bidirectional texture matching score and the edge protection gating weight, and the specific expression is as follows:

[0070]

[0071] In the formula: This represents the adaptive enhancement factor.

[0072] Furthermore, the enhanced cloud semantic feature map and enhanced cloud shadow semantic feature map are obtained based on the adaptive enhancement factor, cloud semantic feature map, and cloud shadow semantic feature map, and the specific expression is as follows:

[0073]

[0074]

[0075] In the formula: To enhance the semantic feature map of the cloud, To enhance the semantic feature map of cloud shadows, To compensate for the scaling factor of the feature, □ indicates element-wise multiplication.

[0076] Furthermore, the geometric consistency loss is obtained using the following specific expression:

[0077]

[0078]

[0079]

[0080] In the formula: Represents geometric consistency loss. This represents the normalized physical centroid coordinates of the cloud. This represents the physical centroid coordinates of the cloud shadow after normalization. This represents the expected cloud shadow location corresponding to the predicted m-th cloud. This indicates the weight of geometric loss in the total loss function. This represents the number of cloud-shadow pairs successfully extracted in the entire training batch. This represents the predicted set of centroids of cloud connected components. Let B represent the set of centroids of the connected components of the predicted cloud shadow, B represent the total number of predictions, and b represent the b-th prediction. This represents the nearest neighbor matching operator.

[0081] Secondly, the present invention provides a remote sensing image cloud shadow detection system, comprising:

[0082] The first processing module is used to acquire optical remote sensing image data, input the optical remote sensing image data into the window attention module to extract multi-scale spatial features, and use wavelet transform to decompose the multi-scale spatial features in the frequency domain to generate low-frequency sub-bands and high-frequency sub-bands.

[0083] Enhancement module: Used to input the low-frequency sub-band and high-frequency sub-band into the cloud shadow feature spatial frequency enhancement module to generate cloud shadow feature enhancement signal;

[0084] The second processing module is used to obtain bottleneck features based on the cloud shadow feature enhancement signal, input the bottleneck features into the cloud shadow heterogeneous feature decoupling module to obtain cloud semantic feature map and cloud shadow semantic feature map; and obtain clean cloud features and clean cloud shadow features based on cloud semantic feature map and cloud shadow semantic feature map.

[0085] First calculation module: used to calculate the heterogeneous centroid offset vector based on clean cloud features and clean cloud shadow features;

[0086] The second calculation module is used to obtain pseudo-cloud feature maps and pseudo-cloud shadow feature maps based on clean cloud features, clean cloud shadow features, and heterogeneous centroid offset vectors; calculate bidirectional texture matching scores based on clean cloud features, clean cloud shadow features, pseudo-cloud feature maps, and pseudo-cloud shadow feature maps; and obtain edge protection gating weights based on bidirectional texture matching scores and edge protection gating weights; and generate adaptive enhancement factors based on bidirectional texture matching scores and edge protection gating weights.

[0087] The detection module is used to obtain enhanced cloud semantic feature maps and enhanced cloud shadow semantic feature maps based on the adaptive enhancement factor, cloud semantic feature map, and cloud shadow semantic feature map. The enhanced cloud semantic feature maps and enhanced cloud shadow semantic feature maps are then fed into a dual-head decoder to complete cloud shadow segmentation. The enhanced cloud semantic feature maps and enhanced cloud shadow semantic feature maps are reconstructed by the dual-head decoder under the constraint of geometric consistency loss. At the same time, pixel-level masks of clouds and cloud shadows are output to complete the detection work.

[0088] Thirdly, the present invention provides a terminal, including a processor and a storage medium;

[0089] The storage medium is used to store instructions;

[0090] The processor is configured to operate according to the instructions to perform the steps of the method according to the first aspect.

[0091] Fourthly, a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method described in the first aspect.

[0092] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0093] This remote sensing image cloud shadow detection method decomposes optical remote sensing image data into low-frequency and high-frequency sub-bands through wavelet transform and generates cloud shadow feature enhancement signals, effectively capturing detailed information of cloud shadow boundaries and alleviating the problems of shadow region blurring and phase reversal in traditional methods. By decoupling bottleneck features, clean cloud features and clean cloud shadow features are constructed separately, achieving physical separation of cloud shadow features. Furthermore, heterogeneous centroid offset vectors are calculated, and pseudo-cloud feature maps and pseudo-cloud shadow feature maps are generated, achieving cross-class geometric alignment and physical-driven completion of cloud shadow features. At the same time, by combining bidirectional texture matching scores and edge protection gating weights, physical consistency enhancement is performed on cloud semantic feature maps and cloud shadow semantic feature maps, realizing the collaborative modeling of clouds and cloud shadows in the spatial and frequency domains in optical remote sensing images. This effectively solves the problems of difficult cloud shadow detection, inconsistent pairing, and insufficient geometric constraints in complex scenes, ensuring the cloud shadow detection effect of this invention. Attached Figure Description

[0094] Figure 1 This is a flowchart illustrating a remote sensing image cloud shadow detection method according to an embodiment of the present invention;

[0095] Figure 2 This is a schematic diagram of the model framework of a remote sensing image cloud shadow detection method provided by an embodiment of the present invention;

[0096] Figure 3 This is a schematic diagram of the input image of the L8-Biome dataset provided in an embodiment of the present invention;

[0097] Figure 4 This is a schematic diagram of the actual label images of the L8-Biome dataset provided in an embodiment of the present invention;

[0098] Figure 5 This is a schematic diagram of the segmentation results of UNet on the L8-Biome dataset provided in an embodiment of the present invention;

[0099] Figure 6 This is a schematic diagram of the segmentation results of RD_UNet on the L8-Biome dataset provided in an embodiment of the present invention;

[0100] Figure 7 This is a schematic diagram of the segmentation results of the Swin Transformer on the L8-Biome dataset provided in an embodiment of the present invention;

[0101] Figure 8 This is a schematic diagram of the segmentation results of HR_CloudNet on the L8-Biome dataset according to an embodiment of the present invention;

[0102] Figure 9 This is a schematic diagram of the segmentation results of Segformer on the L8-Biome dataset provided in an embodiment of the present invention;

[0103] Figure 10 This is a schematic diagram of the segmentation results of WaveVIT on the L8-Biome dataset provided in an embodiment of the present invention;

[0104] Figure 11 This is a schematic diagram of the segmentation results of a remote sensing image cloud shadow detection method provided in an embodiment of the present invention on the L8-Biome dataset. Detailed Implementation

[0105] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0106] The term "and / or" simply describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0107] Example 1:

[0108] like Figures 1-2As shown, the present invention provides a method for detecting cloud shadows in remote sensing images, comprising:

[0109] Acquire optical remote sensing image data, input the optical remote sensing image data into the window attention module to extract multi-scale spatial features, and use wavelet transform to decompose the multi-scale spatial features in the frequency domain to generate low-frequency sub-bands and high-frequency sub-bands.

[0110] Specifically, optical remote sensing image data is acquired, and multi-scale spatial features are extracted from the input window attention module. Wavelet transform is then used to decompose the multi-scale spatial features in the frequency domain. The size of the optical remote sensing image data is H*W*3, where H represents the image height, W represents the image width, and the number of image channels is 3. The specific expression is as follows:

[0111]

[0112] In the formula: X represents the input feature, and DWT is the wavelet transform. For low-frequency sub-band, For horizontal high-frequency sub-bands, For high-frequency sub-bands in the vertical direction, It is a high-frequency sub-band in the diagonal direction.

[0113] The low-frequency sub-band and high-frequency sub-band are input into the cloud shadow feature spatial frequency enhancement module to generate the cloud shadow feature enhancement signal;

[0114] Heterogeneous prototype graph inference is performed on the cloud shadow feature enhancement signal. This inference includes cloud shadow heterogeneous feature decoupling, heterogeneous centroid offset vector calculation, and blind offset geometric completion. Blind offset geometric completion includes calculating bidirectional texture matching scores, calculating edge protection gating weights, and generating adaptive enhancement factors, as detailed below:

[0115] Bottleneck features are obtained based on the cloud shadow feature enhancement signal. The bottleneck features are then input into the cloud shadow heterogeneous feature decoupling module to obtain cloud semantic feature map and cloud shadow semantic feature map. Based on the cloud semantic feature map and cloud shadow semantic feature map, clean cloud features and clean cloud shadow features are obtained.

[0116] Calculate the heterogeneous centroid offset vector based on the clean cloud features and clean cloud shadow features;

[0117] Based on clean cloud features, clean cloud shadow features, and heterogeneous centroid offset vectors, pseudo cloud feature maps and pseudo cloud shadow feature maps are obtained. Based on clean cloud features, clean cloud shadow features, pseudo cloud feature maps, and pseudo cloud shadow feature maps, bidirectional texture matching scores are calculated, and edge protection gating weights are obtained. Based on bidirectional texture matching scores and edge protection gating weights, adaptive enhancement factors are generated.

[0118] Based on the adaptive enhancement factor, cloud semantic feature map, and cloud shadow semantic feature map, enhanced cloud semantic feature map and enhanced cloud shadow semantic feature map are obtained. The enhanced cloud semantic feature map and enhanced cloud shadow semantic feature map are fed into the dual-head decoder to complete cloud shadow segmentation. The enhanced cloud semantic feature map and enhanced cloud shadow semantic feature map are reconstructed by the dual-head decoder under the constraint of geometric consistency loss. At the same time, the pixel-level mask of cloud and cloud shadow is output to complete the detection work.

[0119] This invention decomposes optical remote sensing image data into low-frequency and high-frequency sub-bands using wavelet transform and generates cloud shadow feature enhancement signals, effectively capturing detailed information about cloud shadow boundaries and alleviating the problems of shadow region blurring and phase inversion in traditional methods. By decoupling bottleneck features, clean cloud features and clean cloud shadow features are constructed separately, achieving physical separation of cloud shadow features. Furthermore, heterogeneous centroid offset vectors are calculated, and pseudo-cloud feature maps and pseudo-cloud shadow feature maps are generated, achieving cross-class geometric alignment and physically driven completion of cloud shadow features. At the same time, by combining bidirectional texture matching scores and edge protection gating weights, physical consistency enhancement is performed on cloud semantic feature maps and cloud shadow semantic feature maps. By combining physical laws with deep neural networks, this invention improves the problems of weak spatial correlation and poor interpretability in cloud shadow detection of existing models, realizing the collaborative modeling of clouds and cloud shadows in the spatial and frequency domains in optical remote sensing images. It effectively solves the problems of difficult cloud shadow detection, inconsistent pairing, and insufficient geometric constraints in complex scenes, ensuring the cloud shadow detection effect of this invention.

[0120] In this embodiment, the step of inputting the low-frequency sub-band and high-frequency sub-band into the cloud shadow feature spatial frequency enhancement module to generate the cloud shadow feature enhancement signal includes:

[0121] A spatial weight map is generated based on the low-frequency subband and dark perception gating to determine the region where cloud shadows exist. The specific expression is as follows:

[0122]

[0123] In the formula: For spatial weighting, It is the Sigmoid activation function. For a convolution operation with a kernel size of 3, For low-frequency sub-band;

[0124] High-frequency subband and The phenomenon of phase shift in cloud and cloud shadow features exists. A two-way cross-correlation compensation mechanism is constructed to capture the relative phase shift of features in the frequency domain, and mean fusion processing is performed on it to calculate the cross-correlation spectrum between cloud and cloud shadow. This achieves phase correction of cloud and cloud shadow spectra and enhancement of cross-subband feature consistency, as detailed below:

[0125] High-frequency subband and Transforming to the complex frequency domain, the cross-correlation spectrum between the cloud and its shadow is calculated using a bidirectional cross-correlation compensation mechanism. Furthermore, a learnable phase compensation factor is introduced to correct the fused spectrum, achieving phase alignment of the cloud shadow signal in the frequency domain. The specific expression is as follows:

[0126]

[0127]

[0128] In the formula: For horizontal high-frequency sub-bands, For high-frequency sub-bands in the vertical direction, For Fast Fourier Transform, This is a conjugate operation. A learnable phase rotation compensation factor. To convert to the complex frequency domain , To convert to the complex frequency domain , For Hanning window, The cross-correlation spectrum;

[0129] Obtain dark perception weights, use a dynamic spectrum codebook to filter the cross-correlation spectrum, retain the spectrum that meets the cloud-shadow co-occurrence condition, reconstruct and generate a cloud-shadow enhancement signal that meets the physical driving condition, combine the spatial weight map and dark perception weights, perform physical consistency enhancement on the high-frequency sub-band to obtain the enhanced high-frequency sub-band, perform inverse wavelet transform based on the low-frequency sub-band and the enhanced high-frequency sub-band to generate the cloud-shadow feature enhancement signal, the specific expression is as follows:

[0130]

[0131]

[0132]

[0133]

[0134] In the formula: To enhance the signal of cloud shadow, This represents the inverse fast Fourier transform. Indicates a dynamic spectrum codebook. This is a preset constant used to prevent the denominator from being zero. For the enhanced , For the enhanced , To enhance the signal for cloud shadow features, High-frequency sub-bands in the diagonal direction. For inverse wavelet transform, For dark perception weights, It is a learnable enhancement function.

[0135] Specifically, the bidirectional cross-correlation mechanism should include cross-correlation spectrum construction, symmetry denoising, and learnable phase reference correction:

[0136] After being smoothed by Hanning window and Perform complex field conjugate multiplication and simultaneously calculate the positive cross-correlation spectrum. With reverse cross-correlation spectrum The specific calculation formula is as follows:

[0137]

[0138]

[0139] By using the arithmetic mean method and Complex domain fusion is performed, and random phase noise is canceled by using conjugate symmetry. Stable cross-energy spectrum signals reflecting the common topological structure of cloud and shadow boundaries are extracted. Nonlinear compression processing is performed on the fused cross-correlation spectrum, and amplitude normalization is performed by calculating the exponential power of its modulus to balance the spectral energy under different radiation intensities while preserving phase information. The fused spectrum is multiplied by a preset complex form learnable phase compensation factor. By performing a linear rotation transformation in complex space, the phase reference of cloud and shadow in frequency domain response is forcibly aligned, eliminating the spatial reconstruction destructive interference caused by optical polarity reversal.

[0140] The learnable phase rotation compensation factor includes: introducing a set of learnable phase compensation factors with physical alignment function in the complex frequency space. This factor, as an adaptive rotation operator defined in the complex space, is specifically designed to correct the phase deflection caused by the opposite radiation polarity between the cloud and its shadow. By using the learnable phase compensation factor to perform nonlinear complex multiplication on the bidirectional cross-correlation fusion spectrum, linear phase remapping and dynamic phase rotation of the eigenvector are achieved. The core formula is:

[0141]

[0142] In the formula: Represents the original cross-correlation spectrum. This represents the processed cross-correlation spectrum. The imaginary unit, The phase bias learned by the model itself is used to force the feature responses of the cloud and the shadow to be conjugate aligned on the phase reference by using the rotation transformation of the complex space. This eliminates the destructive interference caused by optical polarity reversal during the spatial reconstruction process, thus changing feature cancellation into constructive interference. Through this physically heuristic phase calibration, the feature recognition and geometric continuity of the shadow area against the dark background are further improved.

[0143] In this embodiment, the step of obtaining bottleneck features based on cloud shadow feature enhancement signals includes:

[0144] The cloud shadow feature enhancement signal is flattened and transposed, and then input into a linear layer to achieve dimensionality compression. The compressed cloud shadow feature enhancement signal is then input into the next layer window attention module for repeated operations to obtain high-dimensional features. The high-dimensional features generated by the backbone network are then input into a 3×3 convolutional layer for local spatial information extraction, i.e., bottleneck features are generated based on the high-dimensional features. The specific expression is as follows:

[0145]

[0146] In the formula: Bottleneck characteristics It is the ReLU activation function. For group normalization, For a convolution operation with a kernel size of 3, High-dimensional features;

[0147] The step of inputting bottleneck features into the cloud-shadow heterogeneous feature decoupling module to obtain cloud semantic feature maps and cloud-shadow semantic feature maps includes:

[0148] The bottleneck features are input into the cloud-shadow heterogeneous feature decoupling module, and two independent sets of cloud prototype vectors and cloud shadow prototype vectors are formed through channel projection. Based on the cloud prototype vectors and cloud shadow prototype vectors, cloud semantic feature maps and cloud shadow semantic feature maps are obtained. The specific expressions are as follows:

[0149]

[0150] In the formula: For cloud semantic feature maps, This is a semantic feature map of cloud shadows. This represents the prototype vector of the k-th cloud. This represents the k-th cloud shadow prototype vector, where k is the number of cloud prototype vectors and cloud shadow prototype vectors.

[0151] The step of obtaining clean cloud features and clean cloud shadow features based on the cloud semantic feature map and cloud shadow semantic feature map includes:

[0152] Based on the cloud semantic feature map, cloud shadow semantic feature map, and bottleneck features, cloud similarity and cloud shadow similarity are calculated, and the specific expressions are as follows:

[0153]

[0154] In the formula: For cloud similarity, For cloud shadow similarity, For normalized exponential functions, For a learnable projection matrix, for The transpose form, for The transpose form, This is the scaling factor;

[0155] Based on cloud similarity and cloud shadow similarity, the cloud enhancement signal and cloud shadow enhancement signal are weighted and aggregated, and their residuals are mapped back to the original feature space. Under the constraint of geometric consistency loss, clean cloud features and clean cloud shadow features are obtained through reconstruction. The specific expression is as follows:

[0156]

[0157]

[0158] In the formula: Characteristics of clean clouds. The characteristics of clean cloud shadows, For gradient clipping operators, To enhance cloud signal, To enhance the signal of cloud shadow, for The transpose form, for The transpose of .

[0159] In this embodiment, calculating the heterogeneous centroid offset vector based on clean cloud features and clean cloud shadow features includes:

[0160] A pre-defined two-dimensional normalized coordinate matrix Z, with the same size as the optical remote sensing image data, is used to map clean cloud features and clean cloud shadow features to the same physical coordinate space. An activation heatmap is used as the weight to perform a weighted average on the pre-defined coordinate matrix Z, where the coordinate matrix Z has two channels. The physical centroid coordinates of the cloud and the cloud shadow are calculated using the following expressions:

[0161]

[0162]

[0163]

[0164] In the formula: This represents a two-dimensional cloud feature heatmap that has undergone mean pooling along the channel dimension. This represents a two-dimensional cloud shadow feature heatmap after mean pooling across all channels, where C represents the number of feature channels. This represents the coordinates of the pixel with x-coordinate i and y-coordinate j in the coordinate matrix Z. This represents the width of the coordinate matrix. Represents the height of the coordinate matrix. This is a preset constant used to prevent the denominator from being zero. The physical centroid coordinates of the cloud, The physical centroid coordinates of the cloud shadow. express The coordinates of the pixel with x-coordinate i and y-coordinate j. express The coordinates of the pixel with x-coordinate i and y-coordinate j;

[0165] Based on the physical centroid coordinates of the cloud and its shadow, the heterogeneous centroid offset vector is calculated, as shown in the following expression:

[0166]

[0167] In the formula: This is the heterogeneous centroid offset vector, which directly represents the direction and distance of the cloud shadow's translation relative to the cloud layer.

[0168] Specifically, the 2D cloud feature heatmap and 2D cloud shadow feature heatmap generated by the heterogeneous prototype attention mechanism are normalized to eliminate the amplitude interference of uneven illumination on the target centroid positioning. A coordinate matrix Z with the same size as the optical remote sensing image data is constructed and mapped to this continuous coordinate interval to form a 2D spatial probability distribution with 2D spatial positional significance. The spatial probability distribution is used to perform point-by-point weighted product operation on the 2D position tensor, and the physical centroid coordinates of the cloud entity and cloud shadow entity are calculated based on the spatial first moment integral. The centroid coordinates are represented in the form of continuous floating-point vectors. Based on the continuous floating centroid coordinates, the Euclidean difference between the two types of centroids in the 2D coordinate system is calculated to explicitly capture the spatial position offset caused by the solar altitude angle and the sensor observation angle. Finally, this heterogeneous offset vector is used as a spatial prior operator to guide subsequent branches to perform geometric alignment and feature completion with physical consistency.

[0169] In this embodiment, obtaining the pseudo-cloud feature map and pseudo-cloud shadow feature map based on clean cloud features, clean cloud shadow features, and heterogeneous centroid offset vector includes:

[0170] Clean cloud features, clean cloud shadow features, and heterogeneous centroid offset vectors are input into the forward and reverse affine transformation matrices. Based on these matrices, a spatial geometric translation matrix is ​​constructed using the heterogeneous centroid offset vector. A blind offset geometric completion method is then performed, with forward and reverse geometric mappings: the forward method translates the clean cloud features along the offset vector to obtain a pseudo-cloud shadow feature map, and the reverse method translates the clean cloud shadow features along the negative offset vector to obtain a pseudo-cloud feature map. Cross-class feature transfer is achieved through the physical correspondence between clouds and cloud shadows. Texture-rich cloud features are used to perform physically driven feature completion on weak cloud shadow regions. Simultaneously, cloud shadow features are used to verify the symmetry of cloud regions, resulting in pseudo-cloud feature maps and pseudo-cloud shadow feature maps. The specific expressions are as follows:

[0171]

[0172]

[0173]

[0174]

[0175] In the formula: It is the forward affine transformation matrix. This is the inverse affine transformation matrix. The horizontal distance between clean cloud features and false cloud features. The vertical distance between clean cloud features and false cloud features. For bilinear interpolation resampling, To generate the sampling network; This is a pseudo-cloud feature map. This is a pseudo-cloud shadow feature map. It is an affine transformation operator;

[0176] A bidirectional texture similarity measure is performed based on clean cloud features, clean cloud shadow features, pseudo-cloud feature maps, and pseudo-cloud shadow feature maps, and a bidirectional texture matching score is calculated. Specifically, low-dimensional texture features of the four are extracted using a texture projection module. Then, the cosine similarity among the clean cloud features, clean cloud shadow features, pseudo-cloud feature maps, and pseudo-cloud shadow feature maps is calculated, and the average value is taken as the bidirectional texture matching score. The specific expression is as follows:

[0177]

[0178]

[0179]

[0180]

[0181] In the formula: Texture components that represent clean cloud features. For clean cloud-like texture components, This represents the channel compression and texture feature extraction operator. The texture component of the pseudo-cloud feature map. For the texture components of the pseudo-cloud shadow feature map, This represents the matching score of the cloud mapping to the cloud shadow. This represents the matching score of the cloud shadow mapping to the cloud. Scoring for bidirectional texture matching;

[0182] Simultaneously, thickness information of the corresponding region is extracted through cloud thickness sensor and cloud shadow thickness sensor, and a fused dual-thickness Gaussian edge protection gating weight is constructed, i.e., the edge protection gating weight is obtained, and the specific expression is as follows:

[0183]

[0184]

[0185]

[0186] In the formula: For the predicted thickness features of clouds, The predicted thickness feature of cloud shadows. A thickness sensor representing the cloud. The thickness sensor of the cloud shadow. This represents the edge protection gating weight of the Gaussian gating mechanism. The edge protection gating weight has a smaller value in the target center and background region, but reaches a peak value at the target edge, thereby enhancing the characteristics of the edge transition zone. For thickness sensing;

[0187] The adaptive enhancement factor is generated based on the bidirectional texture matching score and the edge protection gating weight, and the specific expression is as follows:

[0188]

[0189] In the formula: This represents the adaptive enhancement factor.

[0190] In this embodiment, based on the adaptive enhancement factor, the cloud semantic feature map, and the cloud shadow semantic feature map, physical consistency enhancement is further performed on the cloud semantic feature map and the cloud shadow semantic feature map respectively to obtain enhanced cloud semantic feature maps and enhanced cloud shadow semantic feature maps. Then, through a residual connection structure, the physically guided pseudo-cloud feature map and pseudo-cloud shadow feature map are weighted and fused with the cloud semantic feature map and the cloud shadow semantic feature map to correct the relative positional deviation between the cloud and the cloud shadow in the feature space. The specific expression is as follows:

[0191]

[0192]

[0193] In the formula: To enhance the semantic feature map of the cloud, To enhance the semantic feature map of cloud shadows, To compensate for the scaling factor of the feature, □ indicates element-wise multiplication.

[0194] In this embodiment, the model of the present invention should use connected domain geometric consistency constraints during the training phase. By extracting the centroid displacement of the predicted cloud shadow mask, it is verified whether it is consistent with the position of the deduced physical offset vector, thereby further enhancing the physical geometric consistency of the model output.

[0195] After thresholding the enhanced cloud semantic feature map and the enhanced cloud shadow semantic feature map, the connected component analysis algorithm is used to identify discrete clouds and shadows in the image and remove noise fragments with too small an area. The geometric centroid of each independent instance in the image space is calculated and mapped to a normalized coordinate system consistent with the heterogeneous offset vector. Using the heterogeneous feature offset vector, the spatial translation simulation of the centroid of all detected cloud instances is performed to deduce the corresponding theoretical cloud shadow position. By constructing the nearest neighbor matching cost function, the Euclidean distance deviation between the simulated position and the actual observed cloud shadow instance centroid is calculated and converted into geometric consistency loss.

[0196] The geometric consistency loss is obtained as follows: For a single image, the set of centroids of the cloud connected components predicted by this invention is: The predicted set of centroids of the cloud shadow connected domains is Based on this, the specific expression for the geometric consistency loss is as follows:

[0197]

[0198]

[0199]

[0200] In the formula: Represents geometric consistency loss. This represents the normalized physical centroid coordinates of the cloud. This represents the physical centroid coordinates of the cloud shadow after normalization. This represents the expected cloud shadow location corresponding to the predicted m-th cloud. This indicates the weight of geometric loss in the total loss function. This represents the number of cloud-shadow pairs successfully extracted in the entire training batch. This represents the predicted set of centroids of cloud connected components. Let B represent the set of centroids of the connected components of the predicted cloud shadow, B represent the total number of predictions, and b represent the b-th prediction. This represents the nearest neighbor matching operator. Since there may be multiple clouds and multiple shadows in the graph, this operator will automatically find the cloud shadow instance that is closest to the predicted position of the current cloud.

[0201] In this embodiment, the processed cloud shadow features can be mapped to the output space using a dual-head FPN (Feature Pyramid Network) decoder. Based on the probability distribution and the predicted label for each category, fine-grained segmentation results of clouds and cloud shadows are obtained, which can be applied to the field of cloud shadow detection and segmentation in remote sensing images. In addition, after outputting the cloud shadow mask, the output results can also be applied to various downstream tasks, including cloud shadow removal, thin cloud detection, change detection, and large-scale land cover analysis, all of which are existing technologies and will not be described in detail here.

[0202] In some possible embodiments, such as Figure 3 and Figure 4 As shown, the dataset used in this invention is the L8-Biome dataset, which consists of multiple biome scenes captured by the Landsat-8 satellite OLI sensor (Land Imager). The spatial resolution is 30 meters, and it includes various complex land cover types such as clouds, shadows, and backgrounds. This dataset covers a variety of typical biomes such as forests, grasslands, deserts, and cities, and includes multi-temporal images under different cloud cover and solar altitude angle conditions. The labeled samples are abundant and highly representative. Figure 3 The input images for the L8-Biome dataset are shown. Figure 4 The image shows the actual labeled images from the L8-Biome dataset.

[0203] The comparative experiments employed methods including UNet (U-shaped convolutional network), Swim Transformer (window attention U-shaped network), RD_UNet (residual dense U-shaped network), HR_CloudNet (high-resolution cloud detection network), Segformer (lightweight semantic segmentation attention network), WaveVIT (wavelet transform-based visual attention network), and other advanced cloud and shadow segmentation models in recent years. Among these, UNet, as a classic U-shaped encoder-decoder structure, possesses excellent local feature extraction capabilities; Swim Transformer uses global attention for global context modeling; RD_UNet enhances feature reuse through residual dense connections; HR_CloudNet employs a high-resolution network structure to preserve fine-grained boundary information of cloud and shadow structures; Segformer utilizes an efficient Transformer structure to achieve lightweight multi-scale feature fusion; and WaveVIT combines wavelet transform with vision... Transformer (visual transformer) enhances the representation of frequency domain details; through comparison of various representative methods including traditional CNN (convolutional neural network), Transformer and WaveVIT hybrid architecture, the comprehensive advantages of the remote sensing image cloud shadow detection method based on physical centroid guidance and heterogeneous prototype graph inference proposed in this invention are verified in complex remote sensing cloud shadow segmentation tasks.

[0204] The hyperparameter settings are as follows: Haar wavelet time-frequency co-decomposition is used, the input image resolution is 256×256, the number of training epochs is set to 50, the initial learning rate is 0.0005, the AdamW optimizer (adaptive moment estimation with decoupled weight decay) is used, and the weight decay is... The learning rate scheduler uses a cosine annealing strategy with a maximum period of 100 and a minimum learning rate of [missing information]. The model is adjusted by setting the number of GNN (Graph Neural Network) prototypes to 16, the initial value of gamma in the bidirectional geometric completion module to 2.0, the initial value of cloud enhancement weight to 0.6, and the weight of geometric consistency constraint to 0.1. These hyperparameters work together to optimize the model training process and improve cloud segmentation performance.

[0205] Under these conditions, the experiments were repeated for all methods, and the average crossover ratio (miou) of clouds and cloud shadows and the overall accuracy (OA) are shown in Table 1. Individual crossover ratios for clouds and cloud shadows were also added.

[0206] Table 1. Cloud and shadow segmentation results on the L8-Biome dataset.

[0207]

[0208] In the table: OURs represents the scheme provided by this invention, mean B-iou is the boundary consistency, Cloud Iou is the crossover ratio of clouds, and Shadow Iou is the crossover ratio of cloud shadows.

[0209] As shown in Table 1, the remote sensing image cloud shadow detection network based on physical centroid guidance and heterogeneous prototype graph inference provided in this invention exhibits the best overall performance among all compared algorithms. Experimental results show that this method achieves the highest values ​​in most core metrics, especially in the most challenging cloud shadow detection task, where it reaches 18.03%, significantly higher than WaveVIT and Swin Transformer. In the cloud detection task (cloud intersection-union ratio), this method leads other models with an accuracy of 89.70%. In terms of average intersection-union ratio and boundary consistency, which reflect the overall segmentation ability, this method achieves 62.39% and 30.76%, respectively, both superior to existing state-of-the-art methods. This indicates that this method can effectively enhance the model's ability to capture weak shadow signals through heterogeneous feature decoupling and a physically driven geometric completion mechanism.

[0210] To present the classification results visually, Figures 5 to 10 The segmentation results of UNet, RD_UNet, SwinTransformer, HR_CloudNet, Segformer, WaveVIT, and a remote sensing image cloud shadow detection network based on physical centroid guidance and heterogeneous prototype graph inference are presented respectively. Figure 11 The segmentation results proposed in this invention are shown, and it can be intuitively seen that the method proposed in this invention enhances the correspondence between the cloud and cloud shadow segmentation results, and the two exhibit a high degree of consistency in shape and structure in the same scene.

[0211] This invention achieves collaborative modeling of clouds and cloud shadows in optical remote sensing images in both the spatial and frequency domains by combining physical centroid guidance, heterogeneous prototype graph reasoning, and bidirectional geometric completion. It effectively solves the problems of difficult cloud shadow detection, inconsistent pairing, and insufficient geometric constraints in complex scenes. Experiments show that the method achieves an overall segmentation accuracy of 63.04% on publicly available remote sensing cloud shadow detection datasets, which is significantly better than existing mainstream methods.

[0212] Example 2:

[0213] Based on the same inventive concept as Embodiment 1, the present invention provides a remote sensing image cloud shadow detection system, comprising:

[0214] The first processing module is used to acquire optical remote sensing image data, input the optical remote sensing image data into the window attention module to extract multi-scale spatial features, and use wavelet transform to decompose the multi-scale spatial features in the frequency domain to generate low-frequency sub-bands and high-frequency sub-bands.

[0215] Enhancement module: Used to input the low-frequency sub-band and high-frequency sub-band into the cloud shadow feature spatial frequency enhancement module to generate cloud shadow feature enhancement signal;

[0216] The second processing module is used to obtain bottleneck features based on the cloud shadow feature enhancement signal, input the bottleneck features into the cloud shadow heterogeneous feature decoupling module to obtain cloud semantic feature map and cloud shadow semantic feature map; and obtain clean cloud features and clean cloud shadow features based on cloud semantic feature map and cloud shadow semantic feature map.

[0217] First calculation module: used to calculate the heterogeneous centroid offset vector based on clean cloud features and clean cloud shadow features;

[0218] The second calculation module is used to obtain pseudo-cloud feature maps and pseudo-cloud shadow feature maps based on clean cloud features, clean cloud shadow features, and heterogeneous centroid offset vectors; calculate bidirectional texture matching scores based on clean cloud features, clean cloud shadow features, pseudo-cloud feature maps, and pseudo-cloud shadow feature maps; and obtain edge protection gating weights based on bidirectional texture matching scores and edge protection gating weights; and generate adaptive enhancement factors based on bidirectional texture matching scores and edge protection gating weights.

[0219] The detection module is used to obtain enhanced cloud semantic feature maps and enhanced cloud shadow semantic feature maps based on the adaptive enhancement factor, cloud semantic feature map, and cloud shadow semantic feature map. These enhanced cloud semantic feature maps and enhanced cloud shadow semantic feature maps are then fed into a dual-head decoder to complete cloud shadow segmentation. The dual-head decoder reconstructs the enhanced cloud semantic feature maps and enhanced cloud shadow semantic feature maps under the constraint of geometric consistency loss. Simultaneously, it outputs pixel-level masks of the cloud and cloud shadow, completing the detection process.

[0220] The specific functions of each module described above are explained in the relevant content of the method in Embodiment 1, and will not be repeated here.

[0221] Example 3:

[0222] This invention also provides a terminal, including a processor and a storage medium;

[0223] The storage medium is used to store instructions;

[0224] The processor is configured to operate according to the instructions to execute the steps of the method described in Embodiment 1.

[0225] Example 4:

[0226] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method described in Embodiment 1.

[0227] Since the storage medium provided in this embodiment of the invention can execute the method provided in Embodiment 1 of the invention, it has the corresponding functional modules and beneficial effects for executing the method.

[0228] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0229] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0230] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0231] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0232] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for detecting cloud shadows in remote sensing images, characterized in that, include: Acquire optical remote sensing image data, input the optical remote sensing image data into the window attention module to extract multi-scale spatial features, and use wavelet transform to decompose the multi-scale spatial features in the frequency domain to generate low-frequency sub-bands and high-frequency sub-bands. The low-frequency sub-band and high-frequency sub-band are input into the cloud shadow feature spatial frequency enhancement module to generate the cloud shadow feature enhancement signal; Bottleneck features are obtained based on the cloud shadow feature enhancement signal. The bottleneck features are then input into the cloud shadow heterogeneous feature decoupling module to obtain cloud semantic feature map and cloud shadow semantic feature map. Based on the cloud semantic feature map and cloud shadow semantic feature map, clean cloud features and clean cloud shadow features are obtained. Calculate the heterogeneous centroid offset vector based on the clean cloud features and clean cloud shadow features; Based on the clean cloud features, clean cloud shadow features, and heterogeneous centroid offset vector, obtain pseudo cloud feature maps and pseudo cloud shadow feature maps. Based on the clean cloud features, clean cloud shadow features, pseudo cloud feature maps, and pseudo cloud shadow feature maps, calculate the bidirectional texture matching score and obtain the edge protection gating weights. An adaptive enhancement factor is generated based on the bidirectional texture matching score and the edge protection gating weight; Based on the adaptive enhancement factor, cloud semantic feature map, and cloud shadow semantic feature map, enhanced cloud semantic feature map and enhanced cloud shadow semantic feature map are obtained. The enhanced cloud semantic feature map and enhanced cloud shadow semantic feature map are fed into the dual-head decoder to complete cloud shadow segmentation. The enhanced cloud semantic feature map and enhanced cloud shadow semantic feature map are reconstructed by the dual-head decoder under the constraint of geometric consistency loss. At the same time, the pixel-level mask of cloud and cloud shadow is output to complete the detection work.

2. The method for detecting cloud shadows in remote sensing images according to claim 1, characterized in that, The step of inputting the low-frequency sub-band and high-frequency sub-band into the cloud shadow feature spatial frequency enhancement module to generate the cloud shadow feature enhancement signal includes: A spatial weight map is generated based on the low-frequency subband and dark perception gating, with the specific expression as follows: In the formula: For spatial weighting, It is the Sigmoid activation function. For a convolution operation with a kernel size of 3, For low-frequency sub-band; The high-frequency subband is converted to the complex frequency domain, and the cross-correlation spectrum between the cloud and its shadow is calculated using a two-way cross-correlation compensation mechanism. The specific expression is as follows: In the formula: For horizontal high-frequency sub-bands, For high-frequency sub-bands in the vertical direction, For Fast Fourier Transform, This is a conjugate operation. A learnable phase rotation compensation factor. To convert to the complex frequency domain , To convert to the complex frequency domain , For Hanning window, The cross-correlation spectrum; Obtain dark sensing weights, use a dynamic spectrum codebook to filter the cross-correlation spectrum, retain the spectrum that meets the cloud-shadow co-occurrence condition, reconstruct and generate the cloud-shadow enhanced signal, combine the spatial weight map and dark sensing weights, perform physical consistency enhancement on the high-frequency sub-band to obtain the enhanced high-frequency sub-band, perform inverse wavelet transform based on the low-frequency sub-band and the enhanced high-frequency sub-band to generate the cloud-shadow feature enhancement signal, the specific expression is as follows: In the formula: To enhance the signal of cloud shadow, This represents the inverse fast Fourier transform. Indicates a dynamic spectrum codebook. This is a preset constant used to prevent the denominator from being zero. For the enhanced , For the enhanced , To enhance the signal for cloud shadow features, High-frequency sub-bands in the diagonal direction. For inverse wavelet transform, For dark perception weights, It is a learnable enhancement function.

3. The method for detecting cloud shadows in remote sensing images according to claim 1, characterized in that, The process of obtaining bottleneck features based on cloud shadow feature enhancement signals includes: The cloud shadow feature enhancement signal is flattened and transposed, and then input into a linear layer to achieve dimensionality compression. The compressed cloud shadow feature enhancement signal is then input into the next layer window attention module for repeated operations to obtain high-dimensional features. Bottleneck features are generated based on the high-dimensional features, as shown in the following expression: In the formula: Bottleneck characteristics It is the ReLU activation function. For group normalization, For a convolution operation with a kernel size of 3, High-dimensional features; The step of inputting bottleneck features into the cloud-shadow heterogeneous feature decoupling module to obtain cloud semantic feature maps and cloud-shadow semantic feature maps includes: The bottleneck features are input into the cloud-shadow heterogeneous feature decoupling module, and two independent sets of cloud prototype vectors and cloud shadow prototype vectors are formed through channel projection. Based on the cloud prototype vectors and cloud shadow prototype vectors, cloud semantic feature maps and cloud shadow semantic feature maps are obtained. The specific expressions are as follows: In the formula: For cloud semantic feature maps, This is a semantic feature map of cloud shadows. This represents the prototype vector of the k-th cloud. This represents the k-th cloud shadow prototype vector, where k is the number of cloud prototype vectors and cloud shadow prototype vectors. The step of obtaining clean cloud features and clean cloud shadow features based on the cloud semantic feature map and cloud shadow semantic feature map includes: Based on the cloud semantic feature map, cloud shadow semantic feature map, and bottleneck features, cloud similarity and cloud shadow similarity are calculated, and the specific expressions are as follows: In the formula: For cloud similarity, For cloud shadow similarity, For normalized exponential functions, For a learnable projection matrix, for The transpose form, for The transpose form, This is the scaling factor; Based on cloud similarity and cloud shadow similarity, the cloud enhancement signal and cloud shadow enhancement signal are weighted and aggregated, and their residuals are mapped back to the original feature space. After reconstruction, clean cloud features and clean cloud shadow features are obtained. The specific expressions are as follows: In the formula: Characteristics of clean clouds. The characteristics of clean cloud shadows, For gradient clipping operators, To enhance cloud signal, To enhance the signal of cloud shadow, for The transpose form, for The transpose of .

4. The method for detecting cloud shadows in remote sensing images according to claim 3, characterized in that, The step of calculating the heterogeneous centroid offset vector based on clean cloud features and clean cloud shadow features includes: Clean cloud features and clean cloud shadow features are mapped to the same physical coordinate space. Using the activation heatmap as weights, a weighted average is applied to the preset coordinate matrix Z to calculate the physical centroid coordinates of the cloud and the cloud shadow. The specific expressions are as follows: In the formula: This represents a two-dimensional cloud feature heatmap that has undergone mean pooling along the channel dimension. This represents a two-dimensional cloud shadow feature heatmap after mean pooling across all channels, where C represents the number of feature channels. This represents the coordinates of the pixel with x-coordinate i and y-coordinate j in the coordinate matrix Z. This represents the width of the coordinate matrix. Represents the height of the coordinate matrix. This is a preset constant used to prevent the denominator from being zero. The physical centroid coordinates of the cloud, The physical centroid coordinates of the cloud shadow. express The coordinates of the pixel with x-coordinate i and y-coordinate j. express The coordinates of the pixel with x-coordinate i and y-coordinate j; Based on the physical centroid coordinates of the cloud and its shadow, the heterogeneous centroid offset vector is calculated, as shown in the following expression: In the formula: This is the offset vector of the heterogeneous centroid.

5. The method for detecting cloud shadows in remote sensing images according to claim 4, characterized in that, The process of obtaining pseudo-cloud feature maps and pseudo-cloud shadow feature maps based on clean cloud features, clean cloud shadow features, and heterogeneous centroid offset vectors includes: The clean cloud features, clean cloud shadow features, and heterogeneous centroid offset vector are input into the forward and backward affine transformation matrices. Based on the forward and backward affine transformation matrices, pseudo-cloud feature maps and pseudo-cloud shadow feature maps are obtained, as shown in the following expressions: In the formula: It is the forward affine transformation matrix. This is the inverse affine transformation matrix. The horizontal distance between clean cloud features and false cloud features. The vertical distance between clean cloud features and false cloud features. For bilinear interpolation resampling, To generate the sampling network; This is a pseudo-cloud feature map. This is a pseudo-cloud shadow feature map. It is an affine transformation operator; Based on the clean cloud features, clean cloud shadow features, pseudo-cloud feature map, and pseudo-cloud shadow feature map, the bidirectional texture matching score is calculated, as shown in the following expression: In the formula: Texture components that represent clean cloud features. For clean cloud-like texture components, This represents the channel compression and texture feature extraction operator. The texture component of the pseudo-cloud feature map. For the texture components of the pseudo-cloud shadow feature map, This represents the matching score of the cloud mapping to the cloud shadow. This represents the matching score of the cloud shadow mapping to the cloud. Scoring for bidirectional texture matching; The specific expression for obtaining the edge protection gating weight is as follows: In the formula: For the predicted thickness features of clouds, The predicted thickness feature of cloud shadows. A thickness sensor representing the cloud. The thickness sensor of the cloud shadow. Indicates the edge protection gating weight; The adaptive enhancement factor is generated based on the bidirectional texture matching score and the edge protection gating weight, and the specific expression is as follows: In the formula: This represents the adaptive enhancement factor.

6. The method for detecting cloud shadows in remote sensing images according to claim 5, characterized in that, The enhanced cloud semantic feature map and enhanced cloud shadow semantic feature map are obtained based on the adaptive enhancement factor, cloud semantic feature map, and cloud shadow semantic feature map. The specific expression is as follows: In the formula: To enhance the semantic feature map of the cloud, To enhance the semantic feature map of cloud shadows, To compensate for the scaling factor of the feature, □ indicates element-wise multiplication.

7. The method for detecting cloud shadows in remote sensing images according to claim 4, characterized in that, The geometric consistency loss is obtained using the following specific expression: In the formula: Represents geometric consistency loss. This represents the normalized physical centroid coordinates of the cloud. This represents the physical centroid coordinates of the cloud shadow after normalization. This represents the expected cloud shadow location corresponding to the predicted m-th cloud. This indicates the weight of geometric loss in the total loss function. This represents the number of cloud-shadow pairs successfully extracted in the entire training batch. This represents the predicted set of centroids of cloud connected components. Let B represent the set of centroids of the connected components of the predicted cloud shadow, B represent the total number of predictions, and b represent the b-th prediction. This represents the nearest neighbor matching operator.

8. A remote sensing image cloud shadow detection system, characterized in that, include: The first processing module is used to acquire optical remote sensing image data, input the optical remote sensing image data into the window attention module to extract multi-scale spatial features, and use wavelet transform to decompose the multi-scale spatial features in the frequency domain to generate low-frequency sub-bands and high-frequency sub-bands. Enhancement module: Used to input the low-frequency sub-band and high-frequency sub-band into the cloud shadow feature spatial frequency enhancement module to generate cloud shadow feature enhancement signal; The second processing module is used to obtain bottleneck features based on the cloud shadow feature enhancement signal, input the bottleneck features into the cloud shadow heterogeneous feature decoupling module to obtain cloud semantic feature map and cloud shadow semantic feature map; and obtain clean cloud features and clean cloud shadow features based on cloud semantic feature map and cloud shadow semantic feature map. First calculation module: used to calculate the heterogeneous centroid offset vector based on clean cloud features and clean cloud shadow features; The second calculation module is used to obtain pseudo-cloud feature maps and pseudo-cloud shadow feature maps based on clean cloud features, clean cloud shadow features, and heterogeneous centroid offset vectors. Based on clean cloud features, clean cloud shadow features, pseudo-cloud feature maps, and pseudo-cloud shadow feature maps, it calculates bidirectional texture matching scores and obtains edge protection gating weights. An adaptive enhancement factor is generated based on the bidirectional texture matching score and the edge protection gating weight; The detection module is used to obtain enhanced cloud semantic feature maps and enhanced cloud shadow semantic feature maps based on the adaptive enhancement factor, cloud semantic feature map, and cloud shadow semantic feature map. The enhanced cloud semantic feature maps and enhanced cloud shadow semantic feature maps are then fed into a dual-head decoder to complete cloud shadow segmentation. The enhanced cloud semantic feature maps and enhanced cloud shadow semantic feature maps are reconstructed by the dual-head decoder under the constraint of geometric consistency loss. At the same time, pixel-level masks of clouds and cloud shadows are output to complete the detection work.

9. A terminal, characterized in that, Including processor and storage media; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to perform the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method according to any one of claims 1 to 7.