Remote sensing image semantic segmentation model, segmentation system and segmentation method

Through the combination of MC-Unet architecture and feature coding, optimization and decoding modules, the problem of insufficient feature capture and adaptability in semantic segmentation of remote sensing images is solved, segmentation accuracy and operation speed are improved, and more efficient information extraction and land object recognition are achieved.

CN120279264APending Publication Date: 2025-07-08AEROSPACE DONGFANGHONG SATELLITE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510236370.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing remote sensing image semantic segmentation technology has insufficient feature capture and adaptability, resulting in limited improvement in segmentation accuracy and slow network operation speed.

Method used

The MC-Unet architecture is adopted, combined with the MobileNetV2 network and the circular cross-cross attention network, and the feature encoding, optimization and decoding modules are used to enhance the feature extraction and fusion capabilities, and lightweight and sparse feature optimization units are used to improve the network's feature capture and operation speed.

Benefits of technology

The segmentation accuracy and operation speed of semantic segmentation of remote sensing images are improved, and more efficient information extraction and land object recognition are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279264A_ABST
    Figure CN120279264A_ABST
Patent Text Reader

Abstract

The invention provides a remote sensing image semantic segmentation model, system and method. The segmentation model adopts an MC-Unet architecture and comprises a feature coding module used for receiving an original RGB input image and outputting feature maps of different levels according to the scales of the feature maps; the feature optimization module is used for receiving the feature map extracted by the feature coding module and outputting an optimized high-order feature map; and the feature decoding module is used for receiving the optimized high-order feature map processed by the feature optimization module and outputting a semantic segmentation result according to a semantic category. According to the method, efficient semantic segmentation of the remote sensing image can be realized: on one hand, the feature information abundance of the to-be-segmented pixel is enhanced by enhancing feature extraction and fusion capabilities, and the information extraction precision is improved; and on the other hand, the operation speed of the network is improved by adopting a lightweight feature extraction unit and a sparse feature optimization unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of remote sensing image technology, and particularly relates to an efficient semantic segmentation model, segmentation system, and segmentation method for remote sensing images. Background Art

[0002] With the development of remote sensing technology, remote sensing images have become an important way to obtain surface information. However, the massive data and complex scenes of remote sensing images pose huge challenges for information extraction, especially in the fine-grained recognition and analysis of ground objects. As a key means to solve this problem, semantic segmentation technology can divide remote sensing images at the pixel level into regions with specific semantic information, greatly improving the accuracy and efficiency of geographic information extraction.

[0003] In recent years, remote sensing image segmentation technology based on deep convolutional neural network (DCNN) has shown excellent performance. It can adaptively learn and extract feature information at different levels, including the underlying object detail features and the high-level abstract semantic features, becoming the preferred means for remote sensing image semantic segmentation tasks. Networks such as U-Net, SegNet, and DeepLab have become classic semantic segmentation networks. Among them, the more advanced DeepLab series models have achieved further optimization of multi-scale features by adding an Atrous Spatial Pyramid Pooling (ASPP) module on the basis of feature extraction, and have obtained more significant segmentation effects. However, although ASPP realizes feature optimization by adopting the "multi-scale" idea, it still cannot get rid of the inherent problems of long-distance feature capture difficulty, poor feature scale continuity, and weak adaptability in convolutional operations, which is the main reason restricting the further improvement of the accuracy of the network model.

[0004] Application Content

[0005] To overcome the deficiencies of the prior art, this application proposes a more efficient remote sensing image semantic segmentation model and semantic segmentation method. On the one hand, by strengthening the feature extraction and fusion capabilities, the abundance of feature information of the pixels to be segmented is enhanced, and the information extraction accuracy is improved; on the other hand, by adopting lightweight feature extraction and sparse feature optimization units, the network operation speed is increased.

[0006] This application provides a remote sensing image semantic segmentation model that adopts the MC-Unet architecture. The model includes:

[0007] A feature encoding module, which has a feature extraction backbone network for receiving the original RGB input image and outputting feature maps of different levels according to the feature map scale;

[0008] A feature optimization module, configured to receive the feature map extracted by the feature encoding module and output an optimized high-order feature map;

[0009] A feature decoding module, configured to receive the optimized high-order feature map processed by the feature optimization module and output a semantic segmentation result according to semantic categories,

[0010] Wherein, the feature optimization module captures long-range dependent features by using a cyclic cross attention network, models the dependence relationship between each pixel position in the feature map and other spatial positions, extracts global context information, so as to realize feature optimization;

[0011] Wherein, the feature decoding module adopts a skip connection means to fuse the high-order feature map processed by the feature optimization module with the relatively lower-order features captured by the feature encoding module, so as to enrich the information capacity of the pixels to be segmented.

[0012] According to the model provided by an embodiment of the present application, wherein,

[0013] The feature extraction backbone network of the feature encoding module is a MobileNetV2 network, which consists of 1 downsampling convolutional layer and 17 inverted residual blocks,

[0014] Wherein, the 17 inverted residual blocks are distributed in 7 different bottleneck blocks according to different parameter settings, and 2-fold downsampling of the feature map scale is completed at the positions of the 2nd, 4th, 7th and 14th residual blocks respectively, and the 2-fold downsampling at the positions of the 7th and 14th residual blocks is modified to keep the feature map scale unchanged.

[0015] According to the model provided by an embodiment of the present application, wherein,

[0016] The layer orders of the feature maps output by the feature encoding module include a low layer order L1, a middle layer order L2, and a high layer order L3;

[0017] The spatial scales corresponding to the feature maps of each layer order are respectively: 1 / 2, 1 / 4 and 1 / 8 of the spatial scale of the input image, and the corresponding number of channels are 32, 24 and 320 respectively.

[0018] According to the model provided by an embodiment of the present application, wherein the processing process of the feature optimization module includes:

[0019] S1: Dimension reduction of the features in the input feature map;

[0020] S2: Through the cyclic cross attention network, use the self-attention mechanism to capture the spatial dependence relationship of each pixel position in the vertical and horizontal directions, obtain the long-range spatial context information based on the sparse attention expression, and obtain the preliminarily optimized feature map;

[0021] S3: Adopt the residual idea to fuse the preliminarily optimized feature map with the feature map of the original input of the feature optimization module to obtain the finally output optimized high-order feature map, which is a high-order feature map containing context information.

[0022] According to the model provided by an embodiment of the present application, wherein, step S2 includes:

[0023] S2.1: Perform multiple convolution operations on the dimension-reduced input features respectively to obtain the query unit Q, the key-value unit K, and the representation unit V; reduce the dimensions of the query unit Q and the key-value unit K again while keeping the number of channels of the representation unit V unchanged;

[0024] S2.2: Perform an affinity operation on the query unit Q and the key-value unit K, and then generate the attention map A at each pixel position through softmax;

[0025] S2.3: Perform an aggregation calculation on the representation unit V and the sparse attention map A to obtain the aggregated new feature;

[0026] S2.4: Adopt the residual idea to add the obtained aggregated new feature to the dimension-reduced input feature to obtain the preliminarily optimized feature;

[0027] S2.5: Keep the network parameters unchanged, and repeat steps S2.1 to S2.4 to obtain the preliminarily optimized feature map, and the relevant information of all spatial positions of the feature map is aggregated at each feature position in the preliminarily optimized feature map.

[0028] According to the model provided by an embodiment of the present application, wherein, the processing process of the feature decoding module includes:

[0029] S1: Upsample the optimized high-order feature map L3' by a factor of 2, then fuse it with the middle-order feature of the same scale, and the fusion result goes through two layers of neuron operations to output the first high-low order fusion feature map;

[0030] S2: Upsample the first high-low order fusion feature map by a factor of 2, then fuse it with the low-order feature of the same scale, and the fusion result goes through two layers of neuron operations to output the second high-low order fusion feature map;

[0031] S3: Map the second high-low order fusion feature map into a feature map with the number of channels of the number of categories according to the semantic category to obtain the preliminary semantic segmentation result;

[0032] S4: Upsample the preliminary semantic segmentation result by a factor of 2 to obtain the final semantic segmentation result and output it.

[0033] According to the model provided by an embodiment of the present application, wherein, the upsampling in steps S1, S2, and S4 all adopts the bilinear neighborhood interpolation method.

[0034] A model provided according to an embodiment of the present application, wherein the feature fusion in steps S1 and S2 is performed in the channel dimension.

[0035] The present application also provides a remote sensing image semantic segmentation system, including:

[0036] A dataset preparation module for generating model training data and test data;

[0037] A lightweight model construction module for constructing a remote sensing image semantic segmentation model;

[0038] A model training module for training the remote sensing image semantic segmentation model to optimize the model network parameters;

[0039] A model inference module for generating semantic segmentation results for the test data;

[0040] A segmentation result visualization module for visualizing the output semantic segmentation results.

[0041] The present application also provides a remote sensing image semantic segmentation method for performing semantic segmentation using the remote sensing image semantic segmentation system, and the method includes:

[0042] S101: Prepare a dataset and complete the division of the training dataset and test data;

[0043] S102: Construct a remote sensing image semantic segmentation model according to the encoding-decoding architecture;

[0044] S103: Input the training data into the remote sensing image semantic segmentation model for model training to obtain the optimized model parameters;

[0045] S104: Input the test image into the trained model for model inference to obtain the semantic segmentation results;

[0046] S105: Perform label-to-RGB processing on the semantic segmentation results to obtain the visualized semantic segmentation results.

[0047] The efficiency of the remote sensing image semantic segmentation model, system and method disclosed in the present application is reflected in two aspects: segmentation accuracy and running speed. First, by strengthening the feature extraction and fusion capabilities, the abundance of feature information of the pixels to be segmented is enhanced, and the information extraction accuracy is improved. Second, by adopting lightweight feature extraction and sparse feature optimization units, the network running speed is increased. Description of the Drawings

[0048] The above characteristics, technical features, advantages and their implementation manners of the present application will be further described below in a clear and understandable manner by explaining the preferred embodiments and in combination with the accompanying drawings. The following drawings are only intended to illustrate and explain the present application schematically and do not limit the scope of the present application. Among them:

[0049] Figure 1 is a remote sensing image semantic segmentation model architecture according to an embodiment of the present application;

[0050] Figure 2 is a feature encoding module of a remote sensing image semantic segmentation model according to an embodiment of the present application;

[0051] Figure 3 is a feature optimization module of a remote sensing image semantic segmentation model according to an embodiment of the present application;

[0052] Figure 4 is a feature decoding module of a remote sensing image semantic segmentation model according to an embodiment of the present application;

[0053] Figure 5 is a schematic diagram of a remote sensing image semantic segmentation system according to an embodiment of the present application;

[0054] Figure 6 is a flowchart of the steps of a remote sensing image semantic segmentation method according to an embodiment of the present application. Specific Embodiments

[0055] In order to have a clearer understanding of the technical features, objectives and effects of the present application, the specific embodiments of the present application will now be described with reference to the accompanying drawings.

[0056] Aiming at the problems in the prior art, the present application proposes a more efficient remote sensing image semantic segmentation model and semantic segmentation method, including two aspects:

[0057] 1) In terms of accuracy, by strengthening the feature extraction and fusion capabilities, the abundance of feature information of the pixels to be segmented is enhanced. The new method breaks through the ability limitation of the ASPP network in feature optimization, and effectively improves the local-global feature capture ability of the network. The new method also combines the high-order-low-order skip connection idea of U-Net to further strengthen the fusion of features at different scales, and obtains the feature information to be segmented with more information capacity, which is beneficial to improving the segmentation accuracy of the network model;

[0058] 2) In terms of speed, by adopting a lightweight feature encoding structure based on depthwise separable convolution and a spatial context feature optimization module based on sparse attention expression, while taking into account the model accuracy, the network parameters are minimized, which helps to improve the speed of the network model. In addition, a remote sensing image semantic segmentation system is designed based on the above framework to significantly improve the accuracy and efficiency of ground object information extraction.

[0059] Example 1

[0060] According to an embodiment of the present application, a remote sensing image semantic segmentation model is provided. Compared with previous models, the semantic segmentation efficiency is effectively improved. The high efficiency of this model is reflected in two aspects: segmentation accuracy and running speed. By strengthening the feature extraction and fusion capabilities, the abundance of feature information of the pixels to be segmented is enhanced, and the information extraction accuracy is improved; then, by adopting lightweight feature extraction and sparse feature optimization units, the network running speed is increased.

[0061] As Figure 1 shown, the remote sensing image semantic segmentation model according to an embodiment of the present application adopts the MC-Unet architecture. The main body of the model includes a feature encoding module, a feature optimization module, and a feature decoding module.

[0062] 1. Feature Encoding Module

[0063] The feature encoding module is used to receive the original RGB input image and output feature maps of different levels (including levels L1 to L3) according to the feature map scale.

[0064] The structure of the feature encoding module is as Figure 2 shown. The MobileNetV2 network is used as the backbone network for feature extraction and is adaptively modified. Among them, the feature extraction part of the MobileNetV2 network consists of 1 downsampling convolutional layer and 17 inverted residual blocks. Among them, these 17 inverted residual blocks are distributed in 7 different bottleneck blocks according to different parameter settings, and 2-fold downsampling of the feature map scale is completed at the positions of the 2nd, 4th, 7th, and 14th residual blocks respectively. In this embodiment, the MobileNetV2 network is adaptively modified, and the 2-fold downsampling at the positions of the 7th and 14th residual blocks is modified to keep the feature map scale unchanged. By removing the last two downsamplings, the corresponding L1, L2, and L3 features are obtained.

[0065] In the extracted feature maps, the L1 level is the low level, the L2 level is the middle level, and the L3 level is the high level. The spatial scales corresponding to the feature maps of the above respective levels are: 1 / 2, 1 / 4, and 1 / 8 of the spatial scale of the input image, and the corresponding number of channels are 32, 24, and 320 respectively. In this embodiment, the spatial scale of the input image is 256×256, and the number of channels is 3.

[0066] The MobileNetV2 network adopted by the feature encoding module of this embodiment uses means such as depthwise separable convolution, inverted residual, and linear bottleneck, and can effectively reduce the amount of calculation and storage without significantly reducing the accuracy, having a significant lightweight advantage.

[0067] 2. Feature Optimization Module

[0068] The feature optimization module is used to receive the feature map extracted by the feature encoding module and output an optimized high-order feature map.

[0069] The structure of the feature optimization module is as Figure 3 shown. The Recurrent Criss-Cross Attention (RCCA) network is used to capture long-range dependent features, model the dependence relationship between each pixel position and other spatial positions in the feature map, and extract global context information to achieve feature optimization.

[0070] The processing process of the feature optimization module specifically includes:

[0071] S1: Reduce the dimension of the features in the input feature map. The purpose of reducing the feature dimension is to reduce the network calculation amount. In this embodiment, the feature map of the L3 layer extracted by the feature encoding module is used as the input of the feature optimization module, and its number of feature channels is 320. After the feature optimization module reduces the dimension of the L3 layer feature map by a factor of 4, a feature map with 80 feature channels is obtained;

[0072] S2: Through the recurrent criss-cross attention network, use the self-attention mechanism to capture the spatial dependence relationship of each pixel position in the vertical and horizontal directions, obtain the long-range spatial context information based on the sparse attention expression, and obtain the preliminarily optimized feature map. Specifically include:[[]]

[0073] S2.1: For the input features after dimension reduction, use 3 1×1 convolution operations respectively to obtain a query unit (Query, denoted as Q), a key-value unit (Key, denoted as K), and a representation unit (Value, denoted as V). To further reduce the calculation amount, the query unit Q and the key-value unit K are dimensionally reduced again, and the number of channels of the representation unit V remains unchanged. In this embodiment, the channels of Q and K are dimensionally reduced by a factor of 8 to obtain a feature map with 10 channels;

[0074] S2.2: Perform an affinity operation on the query unit Q and the key-value unit K, and then generate an attention map A for each pixel position through softmax.

[0075] Among them, the specific steps of the affinity calculation are: for any position u in the query unit Q, extract the channel vector Qu at this position, and calculate the correlation (or similarity) between the channel vectors of all feature points at the corresponding position in the key-value unit K in the vertical and horizontal directions:

[0076] d i,u = Q u × Ω i,u T ···(1)

[0077] Among them, d i,u represents the correlation degree between vectors, and Ω i,u represents the channel vectors corresponding to the feature points in the horizontal and vertical directions with u as the intersection point in K. For each position, through the above calculation, an attention vector composed of the correlation degrees of the corresponding feature points in the horizontal and vertical directions can be obtained, and then the corresponding attention weights at the corresponding positions can be obtained through softmax calculation, that is, the sparse attention map A.

[0078] S2.3: Perform an aggregation calculation on the representation unit V and the sparse attention map A to obtain a new aggregated feature.

[0079] Among them, the specific steps of the aggregation calculation are: for each position u in V, extract the channel vectors Φ i,u of all feature points in the horizontal and vertical directions at this position, and then perform a vector multiplication with the sparse attention map A at the corresponding position to obtain a new feature value at this position:

[0080] V u ' = ∑A i,u Φ i,u ···(2)

[0081] Among them, V′ u represents the new aggregated feature at position u, and A i,u represents the sparse attention map corresponding to position u.

[0082] S2.4: Adopt the residual idea, add the obtained new aggregated feature to the input feature H after dimensionality reduction to obtain the preliminarily optimized feature H’:

[0083] H′ u = V′ u + H u ···(3)

[0084] Among them, H' u represents the final feature at position u, and V u ' represents the new aggregated feature at position u. After the above steps, each feature position in H' has aggregated the context information in the vertical and horizontal directions.

[0085] S2.5: Keep the network parameters unchanged, repeat S2.1~S2.4 to obtain the preliminarily optimized feature map H”. At this time, each feature position in H” has aggregated the relevant information of all spatial positions of the feature map.

[0086] S3: Adopt the residual idea, fuse the preliminarily optimized feature map H” with the feature map of the original input of the feature optimization module to obtain the finally output optimized high-order feature map L3', and this feature map L3' contains rich context information.

[0087] 3. Feature Decoding Module

[0088] The feature decoding module is used to receive the optimized high - order feature map processed by the feature optimization module and output the semantic segmentation result according to the semantic category.

[0089] The structure of the feature decoding module is as Figure 4 shown. By using the skip connection method, the high - order feature map processed by the feature optimization module is fully fused with the relatively lower - order features in the different - level features captured by the feature encoding module, further enriching the information capacity of the pixels to be segmented.

[0090] The processing process of the feature decoding module specifically includes:

[0091] S1: Upsample the optimized high - order feature map L3' by a factor of 2, then fuse it with the middle - order feature of the same scale. The fusion result undergoes two - layer neuron operations to output the first high - low - order fusion feature map;

[0092] S2: Similar to the processing process of step S1 above, upsample the first high - low - order fusion feature map by a factor of 2, then fuse it with the low - order feature of the same scale. The fusion result undergoes two - layer neuron operations to output the second high - low - order fusion feature map;

[0093] S3: Map the second high - low - order fusion feature map into a feature map with the number of channels of the category according to the semantic category, that is, the preliminary semantic segmentation result S;

[0094] S4: Upsample the preliminary semantic segmentation result S by a factor of 2 to obtain the final semantic segmentation result and output it.

[0095] The above - mentioned feature fusion is performed in the channel dimension.

[0096] The above - mentioned upsampling all adopts the bilinear neighborhood interpolation method.

[0097] Embodiment 2

[0098] Based on the above - mentioned embodiment, the present application also discloses a remote - sensing image semantic segmentation system, the structure of which is as Figure 5 shown, including:

[0099] A dataset preparation module, used to generate model training data and test data;

[0100] A lightweight model construction module, used to construct the remote - sensing image semantic segmentation model provided in Embodiment 1 above;

[0101] A model training module, used to train the remote - sensing image semantic segmentation model to optimize the model network parameters;

[0102] A model inference module for generating semantic segmentation results for test data;

[0103] A segmentation result visualization module for visualizing the output semantic segmentation results.

[0104] Embodiment 3

[0105] This application also discloses a method for semantic segmentation of remote sensing images, which uses the remote sensing image semantic segmentation system provided in the above Embodiment 2 for semantic segmentation, and its process is as Figure 6 shown, including:

[0106] S101: Prepare a data set and complete the division of the training data set and test data.

[0107] According to a preset ratio, the data is divided into two parts: training data and test data. Among them, the training data is used for model training, and the test data is used for model inference. In particular, the data all include two parts: remote sensing images and their labels.

[0108] 1) In order or randomly, crop the remote sensing images and their corresponding labels in the training data into training samples of a fixed size. To ensure the robustness of the network model, a training sample expansion strategy is usually adopted, such as randomly cropping in the original image and combining data augmentation for the slices to obtain the training data set. In this embodiment, the cropping size is set to 256×256.

[0109] 2) Randomly split the training data set into two parts: a training sample set (Training) and a validation sample set (Validation) according to a set ratio. Among them, the training sample set is used for network parameter training, while the validation sample set is used to assist in adjusting network parameters. In this embodiment, the ratio is set to 0.25, that is, the validation sample set accounts for 0.25, and the rest is the training sample set. It should be noted that the two parts of the sample sets are not fixed. To prevent overfitting in model training, a cross-validation method is usually adopted, that is, in each round of model training process, the training sample set and the validation sample set are re-divided according to this ratio.

[0110] S102: Construct a lightweight network model MC-Unet according to the encoding-decoding architecture.

[0111] S103: Input the training data into the model for model training to obtain the optimal model parameters.

[0112] 1) To reduce the network learning time and improve robustness, a transfer learning strategy is adopted, and the initial parameters of the feature extraction backbone network MobileNetV2 are replaced with the pre-trained network parameters.

[0113] 2) Set appropriate network training hyperparameters and input the training dataset in batches to train the model parameters. When the input data is the training sample set, the model is in the training mode, and the model performs backpropagation and parameter adjustment and optimization; while when the input data is the validation sample set, the model is in the inference mode, and the model does not perform backpropagation and parameter adjustment;

[0114] 3) Finally, output the trained model parameters.

[0115] S104: Input the test image into the trained model to perform model inference and obtain the semantic segmentation result.

[0116] 1) Import the model framework and its trained model parameters, and set the model to the inference mode.

[0117] 2) Input the test image into the model to perform model inference and output the semantic segmentation result.

[0118] S105: Perform label-to-RGB processing on the semantic segmentation result to obtain the visual semantic segmentation result.

[0119] According to the color mapping relationship, map the label values in the semantic segmentation result to RGB values to generate the visual color semantic segmentation result.

[0120] Although this application has been disclosed above with preferred embodiments, it is not used to limit this application. Any person skilled in the art can make possible changes and modifications to the technical solution of this application by using the methods and technical contents disclosed above without departing from the spirit and scope of this application. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of this application without departing from the technical solution of this application shall fall within the protection scope of the technical solution of this application.

[0121] The content not described in detail in the specification of this application belongs to the well-known technology of those skilled in the art.

Claims

1. A remote sensing image semantic segmentation model that adopts the MC-Unet architecture. The model includes: A feature encoding module with a feature extraction backbone network for receiving an original RGB input image and outputting feature maps of different levels according to the feature map scale; A feature optimization module for receiving the feature maps extracted by the feature encoding module and outputting optimized high-order feature maps; A feature decoding module for receiving the optimized high-order feature maps processed by the feature optimization module and outputting a semantic segmentation result according to semantic categories. Among them, the feature optimization module uses a cyclic cross attention network to capture long-range dependent features, models the dependence relationship between each pixel position and other spatial positions in the feature map, and extracts global context information to achieve feature optimization. Among them, the feature decoding module uses a skip connection method to fuse the high-order feature maps processed by the feature optimization module with the relatively lower-order features captured by the feature encoding module to enrich the information capacity of the pixels to be segmented.

2. The model according to claim 1, wherein The feature extraction backbone network of the feature encoding module is a MobileNetV2 network, which consists of 1 downsampling convolutional layer and 17 inverted residual blocks. Among them, the 17 inverted residual blocks are distributed in 7 different bottleneck blocks according to different parameter settings, and 2-fold downsampling of the feature map scale is completed at the positions of the 2nd, 4th, 7th, and 14th residual blocks respectively, and the 2-fold downsampling at the positions of the 7th and 14th residual blocks is modified to keep the feature map scale unchanged.

3. The model according to claim 1, wherein The levels of the feature maps output by the feature encoding module include low-level L1, middle-level L2, and high-level L3; The corresponding spatial scales of the feature maps at each level are: 1 / 2, 1 / 4, and 1 / 8 of the spatial scale of the input image, and the corresponding number of channels are 32, 24, and 320 respectively.

4. The model according to claim 1, wherein, The processing process of the feature optimization module includes: S1: Dimension reduction of the features in the input feature map; S2: Through the cyclic cross attention network, use the self-attention mechanism to capture the spatial dependence relationship of each pixel position in the vertical and horizontal directions, obtain long-range spatial context information based on the sparse attention expression, and obtain a preliminarily optimized feature map; S3: Adopt the residual idea to fuse the preliminarily optimized feature map with the feature map of the original input of the feature optimization module to obtain the finally output optimized high-order feature map, and this final feature map is a high-order feature map containing context information.

5. The model according to claim 4, wherein Step S2 includes: S2.1: Perform multiple convolution operations on the dimension-reduced input features respectively to obtain a query unit Q, a key-value unit K, and a representation unit V; reduce the dimensions of the query unit Q and the key-value unit K again, and keep the number of channels of the representation unit V unchanged; S2.2: Perform an affinity operation on the query unit Q and the key-value unit K, and then generate an attention map A for each pixel position through softmax; S2.3: Perform an aggregation calculation on the representation unit V and the sparse attention map A to obtain a new aggregated feature. S2.4: Adopt the residual idea, add the newly aggregated features to the input features after dimensionality reduction to obtain the preliminarily optimized features; S2.5: Keep the network parameters unchanged, repeat steps S2.1 to S2.4 to obtain the preliminarily optimized feature map. Each feature position in the preliminarily optimized feature map aggregates the relevant information of all spatial positions of the feature map.

6. The model according to claim 1, wherein, The processing process of the feature decoding module includes: S1: Upsample the optimized high-order feature map L3' by a factor of 2, then fuse it with the mid-order feature of the same scale. The fusion result undergoes two layers of neuron operations to output the first high-low order fusion feature map; S2: Upsample the first high-low order fusion feature map by a factor of 2, then fuse it with the low-order feature of the same scale. The fusion result undergoes two layers of neuron operations to output the second high-low order fusion feature map; S3: Map the second high-low order fusion feature map into a feature map with the number of channels equal to the number of categories according to the semantic categories to obtain the preliminary semantic segmentation result; S4: Upsample the preliminary semantic segmentation result by a factor of 2 to obtain the final semantic segmentation result and output it.

7. The model according to claim 6, wherein The upsampling in steps S1, S2, and S4 all adopts the bilinear neighborhood interpolation method.

8. The model according to claim 6, wherein, The feature fusion in steps S1 and S2 is performed in the channel dimension.

9. A remote sensing image semantic segmentation system, comprising: A dataset preparation module for generating model training data and test data; A lightweight model construction module for constructing the remote sensing image semantic segmentation model according to claim 1; A model training module for training the remote sensing image semantic segmentation model to optimize the model network parameters; A model inference module for generating semantic segmentation results for test data; A segmentation result visualization module for visualizing the output semantic segmentation results.

10. A remote sensing image semantic segmentation method, using the remote sensing image semantic segmentation system according to claim 9 for semantic segmentation, the method comprising: S101: Prepare the dataset and complete the division of the training dataset and test data; S102: Construct the remote sensing image semantic segmentation model according to claim 1 according to the encoding-decoding architecture; S103: Input the training data into the remote sensing image semantic segmentation model for model training to obtain the optimized model parameters; S104: Input the test image into the trained model for model inference to obtain the semantic segmentation result; S105: Perform label-to-RGB processing on the semantic segmentation result to obtain the visualized semantic segmentation result.