Cross-scale semantic segmentation method, system and device for cloud and cloud shadow
By integrating the state space model and the dual-branch encoder of the convolutional neural network, combined with the Mamba-convolution fusion module, the problem of global context and local features coordination in the semantic segmentation of cloud and cloud shadow in the existing technology is solved, and high-precision cloud and cloud shadow segmentation is achieved.
Patent Information
- Application Number
- CN202510535749.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to effectively coordinate global context and local features in semantic segmentation of clouds and cloud shadows, resulting in insufficient segmentation accuracy, especially when dealing with complex structures such as thin clouds and broken cloud shadows.
By integrating VSSM branches and convolutional branches based on the state space model, the dual-branch encoder is used to extract the global long-range dependency features and local detail features of the remote sensing image, and cross-scale fusion is performed through the Mamba-convolution fusion module to generate the segmentation results of cloud and cloud shadow.
Cross-scale feature interaction is realized, feature expression ability is improved, model analysis ability of complex clouds and cloud shadow features is enhanced, small-area clouds and cloud shadows can be accurately identified, misjudgment and misjudgment cases are reduced, and it is better than other models in terms of edge accuracy and local details.
Smart Images

Figure CN120070462A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a cross-scale semantic segmentation method, system and device for clouds and cloud shadows, belonging to the technical field of atmospheric science. Background Art
[0002] The semantic segmentation of clouds and cloud shadows has important application value in the fields of remote sensing and atmospheric science, directly affecting the accuracy of key tasks such as weather forecasting and climate modeling. Traditional deep learning methods face significant challenges in dealing with the complex shapes, long-range dependencies and local features of clouds and cloud shadows: Although the Transformer architecture (a neural network architecture consisting of an encoder and a decoder) can effectively model long-range dependencies, its computational complexity grows quadratically with the number of image patches, resulting in extremely high requirements for computing power and making it difficult to balance model performance and computational efficiency.
[0003] Although the computational complexity of the state space model (such as the Mamba model, i.e., the selective state space model) grows linearly, there are few inventions in the field of semantic segmentation of remote sensing images. Moreover, existing methods overly focus on long-range dependency modeling and ignore the importance of local features, which are crucial for pixel-level segmentation tasks.
[0004] Existing hybrid architectures fail to effectively coordinate global context and local features, resulting in insufficient segmentation accuracy for the complex structures of clouds and cloud shadows (such as thin clouds and fragmented cloud shadows), restricting the reliability of meteorological applications. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a cross-scale semantic segmentation method, system and device for clouds and cloud shadows. By integrating the global long-range dependency features of the VSSM branch and the local detail features of the convolutional branch, cross-scale feature interaction is achieved, further enhancing the feature expression ability and the model's ability to analyze complex cloud and cloud shadow features.
[0006] To achieve the above purpose, the present invention is implemented by the following technical solutions: In the first aspect, the present invention provides a cross-scale semantic segmentation method for clouds and cloud shadows, including: Obtain remote sensing images of clouds and cloud shadows; Extract the global long-range dependency features and local detail features of the remote sensing images through a dual-branch encoder including a VSSM branch and a convolutional branch, where: The VSSM branch is based on the state space model, and performs multi-directional scanning on the remote sensing images through a preset SS2D module to capture global long-range dependency features; The convolutional branch extracts local detail features based on the convolutional neural network; Fuse the global long-range dependence features and local detail features through a pre-constructed Mamba-convolution fusion module to obtain the fused features; Transfer the fused features to the decoder, and reconstruct the fused features through the decoder to generate the segmentation result of clouds and cloud shadows.
[0007] Furthermore, the SS2D module processes the remote sensing image through the following steps: Unfold the input remote sensing image into a one-dimensional sequence along four scanning paths, including from top left to bottom right, from bottom right to top left, from top right to bottom left, and from bottom left to top right; Each path sequentially passes through a preset S6 block for dynamic parameter adjustment, and the S6 block adaptively adjusts the parameter matrix of the state space model based on the input content; Fuse the sequences output by the four scanning paths through the Hadamard product and reorganize them into a two-dimensional feature map.
[0008] Furthermore, the parameter matrix of the S6 block is dynamically updated in the following manner: The input remote sensing image features are linearly projected to generate a parameter vector Δ and a dynamic weight matrix 、 denoted as, where: The parameter vector Δ is adjusted by the activation function SiLU for discretization; The dynamic weight matrix is initialized by the HIPPO matrix and optimized through backpropagation during training; The dynamic weight matrix is denoted as directly generated by linear projection; Use the zero-order hold method to discretize continuous parameters, and the formula is as follows: ; ; where, and are the parameter matrices for approximately discretizing the continuous equation by the zero-order hold method, I is the identity matrix; represents the matrix exponential; the parameter vector Δ is the time step discretization parameter, the dynamic weight matrices A and B are the state transition matrices, ΔA represents scaling the state transition matrix A by the time step discretization parameter Δ, and ΔB represents weighting the state transition matrix B by the time step discretization parameter Δ; Through the above formula for discretizing continuous parameters using the zero-order hold method, dynamically adjust the response weights of the state space model to the input features to achieve long-range dependence modeling.
[0009] Further, fusing the global long-range dependence features and local detail features through a pre-constructed Mamba-convolution fusion module to obtain the fused features includes: Performing channel attention weighting and spatial attention weighting on the local detail features to obtain the weighted local detail features; Performing multi-scale convolution operations on the global long-range dependence features to obtain the convolved global long-range dependence features; Fusing the weighted global long-range dependence features and local detail features to obtain the fused features.
[0010] Further, the decoder gradually restores the resolution through upsampling operations and integrates the fused features of each stage of the encoder through skip connections to output the pixel-level segmentation results.
[0011] In a second aspect, the present invention provides a cross-scale semantic segmentation system for clouds and cloud shadows, including: A dual-branch encoder, including a VSSM branch and a convolution branch. Among them, the VSSM branch is based on a state space model, and performs multi-directional scanning on the remote sensing image through a preset SS2D module to capture global long-range dependence features; the convolution branch extracts local detail features based on a convolutional neural network; A Mamba-convolution fusion module for cross-scale fusion of the global long-range dependence features and local detail features; A decoder that receives the fused features through skip connections and generates segmentation results.
[0012] Further, the VSSM branch is composed of multiple stacked VSSM blocks, and each VSSM block includes layer normalization, depth convolution, SiLU activation function, and an SS2D module.
[0013] Further, the system uses a cross-entropy loss function and an AdamW optimizer during training, and the learning rate is dynamically adjusted through the Poly strategy.
[0014] In a third aspect, the present invention provides a cross-scale semantic segmentation device for clouds and cloud shadows, including: A memory for storing computer programs / instructions; A processor for executing the computer programs / instructions to implement the steps of the method described in any one of the foregoing.
[0015] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in any one of the foregoing are implemented.
[0016] Compared with the prior art, the beneficial effects achieved by the present invention: The present invention provides a cross-scale semantic segmentation method, system and device for clouds and cloud shadows. By integrating the global long-range dependence features of the VSSM branch and the local detail features of the convolutional branch, cross-scale feature interaction is achieved, further enhancing the feature expression ability and the model's parsing ability for complex cloud and cloud shadow features. In the segmentation task of clouds and cloud shadows, it can not only accurately detect the target area, but also retain rich boundary details, reducing the situations of misjudgment and missed judgment. Its segmentation results are superior to other models in terms of edge accuracy and local details, can accurately identify small-area clouds and cloud shadows, and reduce misjudgment situations. It has significant advantages in feature extraction, fusion and multi-scale information processing, and can effectively improve the accuracy and robustness of cloud and cloud shadow segmentation. It not only performs excellently in overall performance, but also has high accuracy and stability when facing different types of land cover classifications. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is the overall structural diagram of the MCloud module provided by an embodiment of the present invention; Figure 2 is the structural diagram of the VSSM branch provided by an embodiment of the present invention; Figure 3 is the structural diagram of the SS2D provided by an embodiment of the present invention; Figure 4 is the structural diagram of the MCloud module provided by an embodiment of the present invention; Figure 5 is the comparison schematic diagram of different models on the 38-Cloud dataset provided by an embodiment of the present invention; Figure 6 is the comparison schematic diagram of different models on the SPARCS-Val dataset provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the protection scope of the present invention.
[0019] Embodiment 1. This embodiment introduces a cross-scale semantic segmentation method for clouds and cloud shadows, including: Obtain remote sensing images of clouds and cloud shadows; Extract the global long-range dependence features and local detail features of the remote sensing images through a dual-branch encoder including a VSSM branch and a convolutional branch, where: The VSSM branch is based on a state space model, and performs multi-directional scanning on the remote sensing images through a preset SS2D module to capture global long-range dependence features; The convolutional branch extracts local detail features based on a convolutional neural network; Fuse the global long-range dependence features and local detail features through a pre-constructed Mamba-convolution fusion module to obtain the fused features; Transfer the fused features to the decoder, and reconstruct the fused features through the decoder to generate the segmentation result of clouds and cloud shadows.
[0020] The cross-scale semantic segmentation method of clouds and cloud shadows provided in this embodiment is applied to the MCloud (cloud and cloud shadow semantic segmentation proposed in this invention) module. The MCloud module adopts a dual-branch encoder-decoder architecture, and its structure is as Figure 1 shown.
[0021] The encoder part is composed of two parallel branches: one is the Mamba architecture branch based on the state space model, which we call the VSSM (Visual State Space Model) branch; the other is the traditional convolution architecture branch. These two branches are responsible for modeling the long-range dependence relationship and local detail features of clouds and cloud shadows in the remote sensing image respectively, so as to capture and understand the complex information in the image from different perspectives.
[0022] The VSSM branch uses the VSSM block as its core building unit, and the specific structure of the VSSM block is as Figure 2 shown. After the input image features enter the VSSM block, they are first input into the layer normalization module, and the output data has been adjusted to the same scale, and then input into the linear layer module, and the output data has completed the feature channel expansion. Then the expanded data is successively input into the depth convolution operation module, the non-linear transformation of the SiLU activation function, and the SS2D (Selective Scan 2D selective scanning function in the two-dimensional space) module and the layer normalization module, and the input features after complete processing are output. Then, the output result and the input features directly mapped through the linear layer in the previous branch are aggregated to further fuse the information. Finally, the features after the aggregation operation and the original input features are integrated through the residual connection mechanism, and this result is used as the output of the VSSM block to achieve efficient feature transfer and stable network training.
[0023] Among them, the SS2D module is the core part of the VSSM module, and its structure is as Figure 3 shown, and it mainly divides the input feature image according to Figure 3Divided into 9 parts from 1 to 9, the SS2D module processes the image through four different scanning paths: 123456789 (from top left to bottom right), 987654321 (from bottom right to top left), 147258369 (from top right to bottom left), and 963852741 (from bottom left to top right). Each path unfolds the two-dimensional image into a one-dimensional sequence, enabling the model to capture the context information of the image from multiple directions. This dynamic feature allows the SS2D module to flexibly adjust its focus on image features when processing remote sensing images in different scenarios, improving the generalization ability and segmentation accuracy of the model. Through this design, the SS2D module can not only efficiently capture the global and local features in the image but also achieve a good balance between computational complexity and performance, providing strong technical support for the accurate segmentation of clouds and cloud shadows in remote sensing images. The calculation formula of the SS2D module is as follows: ; ; ; where represents the discretized calculation of multi-path scanning similar to VMamba (Visual Combined Selective Structured State Space Model), is the final two-dimensional feature map merged by the scans in four directions. The two-dimensional image is unfolded into a one-dimensional sequence x through four directions and along the sequence direction d. represents the d-th two-dimensional image, is the vector form of this image. y is the output result. The S6 block is a type of ssm parameterization method that combines a selection mechanism (called S6) dependent on the input. The S6 block has the ability to adjust dynamic weights and can adaptively adjust the model parameters according to the input data, thus capturing the complex patterns and features in the image more effectively. The S6 block is an improvement of the Mamba architecture based on the state space model and can selectively retain or filter information according to the input content, thus maintaining a linear time complexity when processing long sequences. The state space model is a mathematical model originating from modern control theory for describing the dynamic behavior of a system. It represents the internal state of the system through a set of state variables and describes the change of state variables over continuous time and the relationship between state variables and output variables through state equations and output equations. The ssm parameterization method assumes that the state of the dynamic system can be predicted by the following two mathematical equations: ; ; where, and represent the input and output of the system at time represents The internal state of the time system denotes the internal state of the time system is the parameter matrix of the model used to capture information from previous states and utilize this information to construct new states indicating the degree of influence of the input on the system defines how the state is transformed into an output similar to a residual connection providing a direct signal from the input to the output
[0024] However, in traditional ssm parameterization methods, the parameter matrix is static and does not change with the input content, which causes the model to be unable to dynamically adjust the information retention strategy according to the input content and lose context relevance. The S6 block of the Mamba architecture associates the parameter matrices and with the input through linear layer projection, enabling the model to selectively retain or ignore context information and adjust the state update rate according to the input content is the time step discretization parameter, controlling the state update rate. At the same time, the parameter matrix is initialized as a HIPPO matrix (Recurrent Memory with Optimal Polynomial Projections, a general framework for online compression of continuous signals and discrete time series), and is then dynamically adjusted during training for long-range dependence modeling. These improvements enhance the content awareness ability of the model while maintaining a low computational complexity. The dynamic parameter matrices and are dynamically updated as follows: ; ; ; where SiLU is the SiLU activation function is the weight matrix is the bias term , , The definitions of the three parameters are learnable weight matrices corresponding to the subscripts, used to project the input features into the parameter space , , are the bias terms corresponding to the subscripts, and together with the weight matrix they constitute an important part of the linear transformation
[0025] The calculation formula of the S6 block is as follows: ; ; Among them, is the state transition matrix, and are the input / output mapping matrices, is the direct transfer matrix, represents the internal state of the system at time represents the internal state of the system at time and represent the input and output of the system at time ; ; ; Among them, the solution principle of the initial step matrix is a matrix exponential operation, is the identity matrix, ΔA represents scaling the state transition matrix A by the time step discretization parameter Δ, and ΔB represents weighting the state transition matrix B by the time step discretization parameter Δ.
[0026] After being processed by block S6, the sequences in the four scanning directions are recombined. Through reverse operations, the one-dimensional sequences are restored to two-dimensional image blocks, which contain the context feature information integrated from four different directions. Then these feature information are fused through the Hadamard product (element-wise product). As an element-wise multiplication operation, the Hadamard product can effectively integrate the features in different directions, enhancing the model's understanding of the global information of the image. This fusion method not only retains the uniqueness of the features in each direction but also highlights their common features, thus improving the model's adaptability to complex scenarios.
[0027] Mamba-Convolution Fusion Module The present invention proposes a Mamba-Convolution fusion module, whose core design goal is to effectively integrate the long-range dependent feature information from the VSSM branch and the local feature information from the convolution branch, so as to achieve a more comprehensive and rich feature representation, and to improve the performance of the model on specific tasks. Its structure is as Figure 4 shown.
[0028] For the features extracted by the convolution branch, since they are essentially obtained by convolving the convolutional kernel with local regions of the input image, they mainly contain local feature information of the image. To enable these local features to be more effectively integrated into the global semantic understanding framework, a channel attention mechanism is first adopted to optimize them. The channel attention mechanism can highlight the channel features that have higher discriminability and contribution to the current specific task by evaluating and weighting the importance of each channel feature, thereby enhancing the discriminability and expressive ability of the features to a certain extent. On this basis, a spatial attention mechanism is further introduced. This mechanism can assign attention weights to different spatial positions in the feature map at a relatively low computational complexity, thereby capturing global feature information with important semantic value, realizing the effective combination of local features and global features, and enhancing the model's overall grasp and understanding ability of the features.
[0029] For the features from the VSSM branch, their advantage lies in being obtained through the VSSM block with excellent long-range dependence feature learning ability, which can capture long-range dependence relationships in images or data. This is crucial for semantic understanding and context modeling in many complex tasks. However, relying solely on long-range dependence features may lead to insufficient description of local detailed information. Therefore, convolution operations of different scales are used to further learn the detailed features of local regions. Convolution operations of different scales can extract features from local regions from multiple angles and granularities, thus enriching the expressive ability of the VSSM branch features. On the basis of retaining long-range dependence information, they can also more accurately depict the details of local features, further enhancing the integrity and richness of the features.
[0030] Finally, the features of the convolution branch and the VSSM branch after the above processing are fused and fed into the decoder through skip connections. This fusion method can not only retain the richness of the original features while maintaining a relatively low computational complexity, but also transfer feature information at different levels and semantic levels to the decoder through skip connections. Skip connections can effectively avoid the loss and degradation of feature information during transmission, enabling the decoder to obtain a more comprehensive and accurate feature representation, thereby significantly enhancing the decoder's ability to reconstruct features and its adaptability to complex tasks, providing strong support for the model to achieve excellent performance in various application scenarios.
[0031] This embodiment has the following beneficial effects: (1) The MCloud module realizes cross-scale feature interaction by integrating the global features of the VSSM branch and the local features of the convolution branch, further enhancing the feature expression ability and the model's ability to analyze complex cloud and cloud shadow features.
[0032] (2) It has stronger feature extraction and fusion capabilities compared to other models.
[0033] (3) In the cloud and cloud shadow segmentation task, it can not only accurately detect the target area, but also retain rich boundary details, reducing the situations of misjudgment and missed judgment. Its segmentation results are superior to other models in terms of edge accuracy and local details, and can accurately identify small-area clouds and cloud shadows, reducing the situation of misjudgment.
[0034] (4) It has significant advantages in feature extraction, fusion, and multi-scale information processing, and can effectively improve the accuracy and robustness of cloud and cloud shadow segmentation. It not only performs excellently in overall performance, but also has high accuracy and stability when facing different types of land cover classification.
[0035] (5) The Mamba architecture achieves a breakthrough balance between long-range dependence modeling and computational efficiency through the co-design of the state space model and convolution. The MCloud module significantly improves the computational efficiency while maintaining near-top segmentation performance, and is more suitable for actual deployment applications.
[0036] Specific verification (supporting materials): The experiment was conducted using PyTorch (an open-source deep learning framework developed by the Facebook AI Research team) on an NVIDIA RTX 4090 GPU (a graphics card from NVIDIA). Since most models in this experiment converge after 250 iterations, this study (the number of epochs for one training) was fixed at 300, and the batch size was 16. This study used the cross-entropy loss function and adopted the AdamW optimizer (an optimization algorithm in deep learning) with a weight decay coefficient of 0.001; during training, the Poly learning rate strategy (updating the learning rate once per deep learning update) was used, with the initial learning rate set to 0.001 and the Poly Power (the power exponent in the deep learning formula) set to 2. The learning rate (LearningRate, LR) for each round of training is described as follows: .
[0037] (1) The ablation experiment verification was conducted on the CloudSEN-12 (CloudSEN-12 is a large dataset for cloud semantic understanding) dataset, aiming to evaluate the contribution of different components in the model to the final segmentation performance. The following are the different combinations and their corresponding MIoU (mean intersection over union) metric values as shown in Table 1: Table 1 Ablation experiments of different modules in the network
[0038] Convolutional branch: The baseline model only uses the convolutional branch, with a MIoU of 73.22%. The convolutional branch can effectively extract local features but has limitations in dealing with long-range dependencies.
[0039] Convolutional branch + VSSM branch: After adding the VSSM branch to the baseline model, the MIoU is improved to 76.30%, an increase of 3.08%. The VSSM branch captures long-range dependencies through a state space model, significantly enhancing the model's perception of global information and thus greatly improving the segmentation performance.
[0040] Convolutional branch + VSSM branch + MCloud module adopted in the present invention: After further adding the MCloud module, the MIoU is improved to 78.19%, an increase of 1.89%. The MCloud module realizes cross-scale feature interaction by integrating the global features of the VSSM branch and the local features of the convolutional branch, further enhancing the feature expression ability and the model's ability to analyze complex cloud and cloud shadow features.
[0041] (2) Comparative experiments are conducted with current excellent models, which are mainly divided into three categories according to the architecture: based on convolutional structures, such as FCN (Fully Convolutional Network), DeepLab (Atrous Convolutional Architecture Convolution), OCRNet (Explicitly Modeling with Object Context), etc.; based on Transformer structures, such as SETR (Converting Image Segmentation to Sequence Analysis, i.e., Semantic Segmentation Transformer), PVT (Pyramid Vision Transformer), SwinUNet (Hierarchical U-shaped Network Structure); and convolutional-Transformer hybrid architectures, such as CVT (Convolutional Vision), MPViT (Multi-Path Vision), DBNet (Dual-Branch Network Model). And CCViM (Fusing Context Clustering and Visual State Space Model), VM-UNet (U-shaped Architecture Model Introducing Visual State Space Blocks), RS3Mamba (Visual State Space Model) of the Mamba architecture.
[0042] Tables 2 and 3 show the comparison of evaluation metrics of different models on the CloudSEN-12 dataset (a large dataset for cloud semantic understanding). From the overall ranking of the MIoU metric, the MCloud module proposed in the present invention performs the best, leading in the Pixel Accuracy (PA), Mean Pixel Accuracy (MPA), and Mean Intersection over Union (MIoU) metrics compared to traditional Convolutional Neural Networks (CNNs), Transformers, hybrid architectures, and networks based on the Mamba architecture, reaching 78.19%, 90.13%, and 88.85% respectively. This result indicates that the MCloud module has significant advantages in cloud and cloud shadow semantic segmentation tasks.
[0043] Comparison of overall evaluation metrics of different models on the CloudSEN-12 dataset
[0044] In comparison with models of other architectures, the MCloud module not only outperforms most models based on CNN and Transformer architectures but also significantly outperforms models based on the CNN-Transformer hybrid architecture. For example, compared with SegNet (a deep fully convolutional neural network architecture for image semantic segmentation) with a relatively good performance in the CNN architecture (MIoU of 77.01%) and SwinUNet with a relatively good performance in the Transformer architecture (MIoU of 77.53%), the MIoU of the MCloud module is 1.18% and 0.66% higher respectively. In addition, the performance of the MCloud module is also better than that of DBNet with a relatively good performance in the CNN-Transformer hybrid architecture (MIoU of 77.37%) and HyCloud (another theoretical hybrid architecture for comparison in the present invention) proposed in the previous invention of cloud and cloud shadow semantic segmentation based on attention mechanism multi-scale feature extraction (MIoU of 77.85%), which indicates that the MCloud module has stronger capabilities in feature extraction and fusion. In the comparison of models within the Mamba architecture, the performance of the MCloud module is also prominent. Compared with CCViM (MIoU of 74.5%), VM-Unet (MIoU of 77.13%), and RS3Mamba (MIoU of 77.91%), the MIoU of the MCloud module is 3.69%, 1.06%, and 0.28% higher respectively. This shows that the MCloud module based on the Mamba architecture, combined with the network structure and MCloud module designed in the present invention, has stronger feature extraction and fusion capabilities and can more effectively handle complex cloud and cloud shadow segmentation tasks.
[0045] Table 3 Comparison of classification evaluation indicators of different models on the CloudSEN-12 dataset
[0046] From the classification indicators, the MCloud module performs well in both cloud and cloud shadow segmentation tasks. In cloud segmentation, the P, R, and F1 of the MCloud module reach 92.15%, 92.50%, and 92.32%, respectively, which is significantly better than other models. In cloud shadow segmentation, the P, R, and F1 of the MCloud module reach 83.00%, 82.50%, and 82.75%, respectively, which is also better than other models. This shows that the MCloud module can not only accurately detect the target area in the cloud and cloud shadow segmentation tasks, but also retain rich boundary details and reduce misjudgments and missed judgments.
[0047] In order to further verify the performance of the MCloud module, this paper randomly selected five images in different scenes such as cities, villages, open spaces and waters, and used several models with the highest MIoU index to compare the segmentation results. Figure 5 The following are the comparison results, including (a) test image; (b) label image; (c) MCloud; (d) HyCloud; (e) DBNet; (f) SwinUNet; (g) OCRNet; (h) CDUNet. From the visualization results, the MCloud module performs best in the cloud and cloud shadow segmentation task. Its segmentation results are superior to other models in terms of edge accuracy and local details, and can accurately identify small-area clouds and cloud shadows, reducing misjudgment.
[0048] In summary, the performance of the MCloud module on the CloudSEN-12 dataset is significantly better than other Mamba architecture models, and also better than most CNN, Transformer, and CNN-Transformer hybrid architecture models. This shows that the MCloud module has significant advantages in feature extraction, fusion, and multi-scale information processing, and can effectively improve the accuracy and robustness of cloud and cloud shadow segmentation.
[0049] (3) Generalization Experiment: In order to evaluate the segmentation performance and generalization ability of the proposed MCloud module network, we conducted a generalization experiment on the 38-Cloud dataset. Tables 4 and 5 show the comparison between our network and the current advanced models on the 38-Cloud dataset.
[0050] In terms of network architecture categories, models based on the Mamba architecture have shown significant advantages in cloud detection tasks, with their performance comprehensively surpassing that of traditional CNN, Transformer, and hybrid architecture models. In terms of comprehensive performance, traditional CNN models (such as the Dual Attention Network DANet and the Fully Convolutional Network FCN-32s) and pure Transformer models (such as SETR) perform the weakest, with their MIoU both below 90%. Hybrid architecture models (such as DBNet and HyCloud) are better than single architectures, but still inferior to Mamba series models. Specifically, the MCloud module of the Mamba architecture ranks first with an MIoU of 94.60%, a PA of 97.58%, and an MPA of 97.62%, significantly leading other models. The algorithm based on Mamba shows significant advantages in cloud segmentation tasks, and its comprehensive performance comprehensively surpasses that of traditional convolution, Transformer, and hybrid architecture models. Traditional convolutional networks, such as DANet and BiSeNetV2 (a bilateral network structure for real-time semantic segmentation), and pure Transformer models (such as SETR) have obvious deficiencies in segmentation accuracy in complex scenarios due to the limitations of their feature modeling capabilities; although hybrid architecture models (such as DBNet and HyCloud) improve performance by fusing multiple types of features, they are still limited by computational complexity and local-global information interaction efficiency. In contrast, the Mamba architecture achieves a breakthrough balance between long-range dependence modeling and computational efficiency through the co-design of the state space model and convolution.
[0051] Table 4 Comparison of overall evaluation indicators of different models on the 38-Cloud dataset
[0052] The MCloud module network we proposed achieves a balance between global context awareness and local detail extraction through the co-design of the state space branch and the convolution branch, while abandoning the dependence on multi-band inputs to reduce computational complexity, and achieving leading performance with only visible light data. Its MIoU is increased by 1.33% compared with DBNet, and the number of parameters is reduced by 43%, providing an efficient and reliable solution for real-time remote sensing image processing. This result verifies the feasibility and potential of the Mamba architecture in remote sensing tasks of clouds and cloud shadows.
[0053] Figure 6Shows the comparison of the segmentation results of the MCloud module with models such as HyCloud, DBNet, and SwinUNet in complex background scenarios such as cloudless and multi-cloud. Among them, (a) test image; (b) label; (c) MCloud; (d) RS3Mamba; (e) VM-UNet; (f) HyCloudX; (g) SwinUNet; (h) DBNet; (i) DeepLab V3. It can be seen from the visualization results that the segmentation results of the MCloud module are significantly better than other models in terms of edge continuity and detail restoration ability. In the prediction results of DBNet and CDUNet (deep learning cloud detection), there are obvious serrated breaks in the cloud boundary, especially in the thin cloud area, local misjudgment is likely to occur (such as Figure 5 the middle cloud edge); Although OCRNet (Object Context Representation Network) improves the cloud body detection accuracy through multi-scale feature extraction, it is insufficient in adapting to the internal texture changes of the cloud layer, resulting in over-smoothing of the segmentation results in the thick cloud area. SwinUNet improves the coherence of the cloud contour based on the global modeling ability of Transformer, but there are still missed detection phenomena in the detection of small-scale cloud blocks. HyCloud fuses multi-scale context features through a convolutional-Transformer dual-branch structure, and its segmentation boundary accuracy has been improved compared with the above models, but in the scenario of cloudless and high surface reflectivity (such as Figure 6 shown), there are still a small amount of misdetection noises caused by snow interference.
[0054] Table 5 Comparison of classification evaluation indicators of different models on the 38-Cloud dataset
[0055] The segmentation results of the MCloud module we proposed show significant robustness and accuracy. Through the long-range dependence relationship modeled by the state space branch, it effectively captures the continuous characteristics of the cloud layer distribution and avoids the edge break problem; in the multi-cloud dense area, the local texture information extracted by the convolution branch and the global semantic guidance of the state space branch work together to achieve fine distinction between thick and thin areas inside the cloud layer; and under the interference of high-reflection background (such as Figure 6 shown), the MCloud module significantly suppresses misdetection noises by dynamically screening cross-scale context features, while completely retaining the boundary details of clouds and ground objects. It is worth noting that the MCloud module only relies on visible light band input, and its performance has approached the segmentation accuracy of the multi-band fusion model HyCloudX (dual-branch structure model), verifying the potential of the state space architecture in complex feature modeling.
[0056] To further evaluate the segmentation performance and generalization ability of the MCloud module network, the present invention also conducted a comparative experiment on the SPARCS-Val dataset with more classifications and more scenarios (the SPARCS-Val dataset was developed by Oregon State University in the United States, and its original intention was to verify the accuracy of cloud and cloud shadow masks derived from spatial cloud and cloud shadow removal algorithms). The experimental results are shown in Tables 6 and 7, where Table 6 shows the overall metrics and Table 7 shows the pixel accuracy of each class for different models. CS refers to cloud shadow classification, CSOW refers to cloud shadow classification over water, W refers to water classification, I / S refers to ice and snow classification, L refers to land classification, C refers to cloud classification, and F refers to flood classification.
[0057] The experimental results show that the MCloud module performs better than most other architecture models on the SPARCS-Val dataset. Specifically, the MIoU, PA, MPA, R, and F1 metrics of the MCloud module reached 77.47%, 93.77%, 88.23%, 85.06%, and 86.5% respectively, standing out among all models. This indicates that the MCloud module has strong generalization ability when dealing with complex datasets. From the perspective of class pixel accuracy, the MCloud module performs excellently in cloud (C) and land (L) classifications, with pixel accuracies of 92.06% and 95.86% respectively. In cloud shadow (CS) and cloud shadow over water (CSOW) classifications, the MCloud module also achieved good results, 81.95% and 74.38% respectively. In addition, in water (W) and ice / snow (I / S) classifications, the pixel accuracies of the MCloud module are 95.34% and 94.31% respectively, also showing excellent performance. Specifically, the high pixel accuracies of the MCloud module in cloud (C) and land (L) classifications indicate its high precision in distinguishing these two common land cover classes. For cloud shadow (CS) and cloud shadow over water (CSOW) classifications, although the accuracies are relatively low, it still shows good recognition ability, which may be related to the complexity and diversity of cloud shadows. In water (W) and ice / snow (I / S) classifications, the high pixel accuracies of the MCloud module further prove its effectiveness in dealing with these land cover features with different spectral and spatial characteristics. The segmentation performance of the MCloud module is relatively balanced across different classes, and it can effectively handle the segmentation tasks of various complex land cover features. This shows that the MCloud module not only performs excellently in overall performance but also has high accuracy and stability when facing different types of land cover classifications.
[0058] Table 6 Comparison of overall evaluation metrics of different models on the SPARCS-Val dataset
[0059] Tables 6 and 7 show the segmentation results of multiple models in different scenarios of the SPARCS-Val dataset. From the data in the tables, it can be seen that the segmentation results of the traditional convolutional structure network DeepLab V3 (a deep convolutional neural network architecture for semantic image segmentation) have problems of rough edges and misdetection, especially a large number of misdetections in water area classification. Although the segmentation results of the SwinUNet with a Transformer structure and the DBNet with a hybrid structure are relatively good, there are still misdetections in a certain range. In contrast, the segmentation results of the MCloud module are excellent. In the segmentation task of clouds and cloud shadows, the MCloud module can accurately segment the boundaries of clouds and cloud shadows, retain rich boundary details, and reduce the occurrence of misdetections. This is mainly due to the fact that the MCloud module introduces the Mamba architecture based on the state space model, and through the collaborative work of the state space architecture branch and the convolutional architecture branch, effective modeling of long-range dependencies and local features is achieved. In addition, the MCloud module designed further enhances the model's ability to analyze complex ground object features, enabling the model to more accurately segment clouds and cloud shadows.
[0060] Table 7 Comparison of classification evaluation metrics of different models on the SPARCS-Val dataset
[0061] In complex scenarios, such as in the case of ice and snow noise interference, the MCloud module can still maintain good segmentation performance. Compared with other networks based on the Mamba architecture, such as RS3Mamba and VM-UNet, the MCloud module also shows obvious advantages in segmentation results. Although RS3Mamba also has good performance in the segmentation of clouds and cloud shadows, there are still certain misdetections when dealing with complex scenarios. Similar problems also exist in the segmentation results of VM-UNet, especially in the case of ice and snow noise interference, the misdetection phenomenon is more obvious. In contrast, by introducing the MCloud module, the MCloud module further enhances the model's ability to analyze complex ground object features, thus showing excellent segmentation performance in different scenarios.
[0062] Performance improvement: Compared with models based on the convolutional-Transformer hybrid architecture, the MCloud module significantly reduces the computational complexity and the number of parameters while maintaining a high segmentation accuracy. Especially compared with the HyCloudX model proposed in the first invention, although the MIoU of the MCloud module is slightly lower, its computational complexity is only about one-third of that of HyCloudX, and the number of parameters is reduced by about 43%. This huge efficiency improvement stems from the linear computational complexity characteristics of the Mamba architecture, demonstrating the feasibility of introducing the Mamba architecture into the semantic segmentation task of remote sensing image clouds and cloud shadows.
[0063] From a broader perspective, models of different architecture series exhibit different balance characteristics between efficiency and performance. Convolutional models such as LinkNet have the highest computational efficiency but limited performance; Transformer models such as SwinUNet have better performance but heavier computational burdens; convolutional-Transformer hybrid architectures such as HyCloudX have the best performance but the highest computational complexity; while Mamba series models achieve a good balance between performance and efficiency. In particular, the MCloud module proposed in this invention achieves a segmentation performance close to or even exceeding that of most convolutional-Transformer hybrid architecture models while the number of parameters and computational complexity are only slightly higher than those of some convolutional models.
[0064] Example 2. This example provides a cross-scale semantic segmentation system for clouds and cloud shadows, including: A dual-branch encoder, including a VSSM branch and a convolutional branch. Among them, the VSSM branch is based on the state space model, and performs multi-directional scanning on the remote sensing image through a preset SS2D module to capture global long-range dependence features; the convolutional branch extracts local detail features based on the convolutional neural network; A Mamba-convolution fusion module for cross-scale fusion of the global long-range dependence features and local detail features; A decoder that receives the fused features through skip connections and generates segmentation results.
[0065] The VSSM branch is composed of multiple stacked VSSM blocks, and each VSSM block includes layer normalization, depth convolution, SiLU activation function, and SS2D module.
[0066] When training the system, the cross-entropy loss function and AdamW optimizer are adopted, and the learning rate is dynamically adjusted through the Poly strategy.
[0067] Example 3. This example provides a cross-scale semantic segmentation device for clouds and cloud shadows, including: A memory for storing computer programs / instructions; A processor for executing the computer program / instructions to implement the steps of the method according to any one of Embodiment 1.
[0068] Embodiment 4 provides a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, it implements the steps of the method according to any one of Embodiment 1.
[0069] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
[0070] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0071] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0072] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0073] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide for implementing in the flowFigure 1 one process or multiple processes and / or boxes Figure 1 steps of the functions specified in one box or multiple boxes.
[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure rather than to limit the scope of its protection. Although the present disclosure has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that after reading the present disclosure, various changes, modifications or equivalent replacements can still be made to the specific implementation manners of the invention. However, these changes, modifications or equivalent replacements are all within the scope of protection of the claims pending for publication.
Claims
1. A cross-scale semantic segmentation method for clouds and cloud shadows, characterized in that: include: Acquire remote sensing images of clouds and cloud shadows; The global long-range dependency features and local detail features of the remote sensing image are extracted through a dual-branch encoder consisting of a VSSM branch and a convolution branch, where: The VSSM branch is based on the state space model and uses the preset SS2D module to perform multi-directional scanning of remote sensing images to capture global long-range dependency features. The convolution branch extracts local detail features based on convolutional neural networks; The global long-range dependency features and local detail features are fused through the pre-built Mamba-convolution fusion module to obtain the fused features; The fused features are passed to the decoder, which reconstructs the fused features and generates the segmentation results of clouds and cloud shadows.
2. The cross-scale semantic segmentation method for clouds and cloud shadows according to claim 1, characterized in that: The SS2D module processes remote sensing images through the following steps: Expand the input remote sensing image into a one-dimensional sequence according to four scanning paths, including from upper left to lower right, from lower right to upper left, from upper right to lower left, and from lower left to upper right; Each path is sequentially subjected to dynamic parameter adjustment through a preset S6 block, wherein the S6 block adaptively adjusts the parameter matrix of the state space model based on the input content; The sequences output by the four scanning paths are fused through the Hadamard product and reorganized into a two-dimensional feature map.
3. The cross-scale semantic segmentation method for clouds and cloud shadows according to claim 1, characterized in that: The parameter matrix of the S6 block is dynamically updated in the following way: Input remote sensing image features generate parameter vector Δ and dynamic weight matrix through linear projection layer , represents, where: The parameter vector Δ is adjusted by the activation function SiLU for discretization; Dynamic Weight Matrix Initialized by the HIPPO matrix and optimized by back-propagation during training; Dynamic Weight Matrix It means that it is directly generated by linear projection; The continuous parameters are discretized using the zero-order hold method, and the formula is as follows: ; ; in, and is the parameter matrix of the discretized continuous equation approximated by the zero-order hold method, I is the identity matrix; Represents the matrix index; the parameter vector Δ is the time step discretization parameter, the dynamic weight matrix A, B is the state transfer matrix, ΔA means scaling the state transfer matrix A according to the time step discretization parameter Δ, and ΔB means weighting the state transfer matrix B according to the time step discretization parameter Δ; Through the above formula for discretizing continuous parameters using the zero-order hold method, the response weights of the state space model to the input features are dynamically adjusted to achieve long-range dependency modeling.
4. The cross-scale semantic segmentation method for clouds and cloud shadows according to claim 1, characterized in that: The pre-built Mamba-convolution fusion module is used to fuse the global long-range dependency features and the local detail features to obtain the fused features, including: Perform channel attention weighting and spatial attention weighting on the local detail features to obtain weighted local detail features; Perform multi-scale convolution operations on the global long-range dependency features to obtain the global long-range dependency features after convolution; The weighted global long-range dependency features and local detail features are fused to obtain fused features.
5. The cross-scale semantic segmentation method for clouds and cloud shadows according to claim 1, characterized in that: The decoder uses upsampling operations to gradually restore the resolution, integrates the fusion features of each stage of the encoder through skip connections, and outputs pixel-level segmentation results.
6. A cross-scale semantic segmentation system for clouds and cloud shadows, characterized in that: include: The dual-branch encoder includes a VSSM branch and a convolution branch. The VSSM branch is based on a state-space model and uses a preset SS2D module to perform multi-directional scanning of remote sensing images to capture global long-range dependency features. The convolution branch extracts local detail features based on a convolutional neural network. Mamba-convolution fusion module, used to fuse the global long-range dependency features and local detail features across scales; The decoder receives the fused features through skip connections and generates segmentation results.
7. The cloud and cloud shadow cross-scale semantic segmentation system according to claim 6, characterized in that: The VSSM branch is composed of a plurality of stacked VSSM blocks, each of which includes layer normalization, depth convolution, SiLU activation function and SS2D module.
8. The cloud and cloud shadow cross-scale semantic segmentation system according to claim 6, characterized in that: The system uses a cross entropy loss function and an AdamW optimizer during training, and the learning rate is dynamically adjusted through a Poly strategy.
9. A cloud and cloud shadow cross-scale semantic segmentation device, characterized in that: include: Memory, for storing computer programs / instructions; A processor, configured to execute the computer program / instructions to implement the steps of the method according to any one of claims 1 to 5.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Remote sensing image cloud and cloud shadow segmentation method based on double-branch fusion network
CN114943963A
Remote sensing image multi-dimensional attention semantic segmentation method and system
CN118470328A
Three-dimensional oral hard palate image segmentation method based on multidirectional state space model
CN118941585A
Cross-view binocular image super-resolution reconstruction method and system based on Mama
CN119273546A
Unmanned aerial vehicle remote sensing RGB-D image semantic segmentation method, device and equipment
CN119540560A
Cited By
Pathological image colon gland segmentation method, device, equipment, medium and program product
CN120339630A
Cross-wind-field adaptive wind power prediction method
CN120579672A
Remote sensing image segmentation method and system based on convolution-state space fusion and position trigger
CN120635462A
A Remote Sensing Image Segmentation Method and System Based on Convolution-State Space Fusion and Position Triggers
CN120635462B
Image segmentation method fusing state modeling and convolution perception mechanism
CN121010761A