Land ecological remote sensing image segmentation method for monitoring returning farmland to forest area
By constructing a multi-level remote sensing image segmentation model and combining semantic and attention mechanisms, the problems of low segmentation accuracy and weak generalization ability of remote sensing images in areas where farmland has been converted back to forest in existing technologies have been solved, and higher accuracy remote sensing image segmentation has been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QINGDAO HAOHAI NETWORK TECH
- Filing Date
- 2025-04-27
- Publication Date
- 2026-04-21
AI Technical Summary
Existing land remote sensing image segmentation methods suffer from low accuracy and weak generalization ability when monitoring areas where farmland has been converted back to forest. They are difficult to identify small plots or complex terrains and are easily affected by factors such as cloud cover, seasonal changes, and similar spectral features, leading to missegmentation or omission.
A land ecology remote sensing image segmentation method for monitoring areas where farmland has been converted back to forest is adopted. By constructing a segmentation model that includes a remote sensing image feature enhancement module, a multi-scale semantic fusion module, a local information capture module, and an edge information enhancement module, and combining semantic spatial attention, channel attention, and spatiotemporal attention mechanisms, the image feature extraction and edge refinement capabilities are improved.
It significantly improved the segmentation accuracy and boundary refinement effect of remote sensing images of land ecology in areas where farmland has been converted back to forest, and enhanced the ability to identify and segment complex scenes.
Smart Images

Figure CN120451185B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of land ecological remote sensing image segmentation technology, specifically relating to a method for segmenting land ecological remote sensing images in areas where farmland has been converted back to forest. Background Technology
[0002] Monitoring the land area of areas designated for reforestation is a crucial research topic for agricultural and forestry researchers. In the past, lacking satellite remote sensing imagery, monitoring personnel primarily relied on manual field surveys combined with traditional mapping methods. Specific methods included regularly organizing visits to these areas, visually observing, comparing topographic maps, manually recording vegetation cover changes, and using aerial photographs to roughly assess the area's extent. However, this approach was inefficient, had a long monitoring cycle, and the accuracy of the data heavily depended on the experience and sense of responsibility of the monitoring personnel.
[0003] With the development of satellite remote sensing image acquisition technology, monitoring the extent of land converted from farmland to forest using satellite remote sensing has become an important research topic. Satellite remote sensing technology has improved the real-time monitoring of land converted from farmland to forest. However, the segmentation technology based on remote sensing images also has obvious drawbacks in application. For example, remote sensing image segmentation algorithms are limited by image resolution, making it difficult to accurately identify small plots or complex terrain areas. Remote sensing image segmentation algorithms are easily affected by factors such as cloud cover, seasonal changes, shadows, or the similarity of spectral characteristics between crops and forests and grasslands, leading to missegmentation or omission. In addition, traditional threshold segmentation algorithms have problems such as insufficient adaptability to land use types with blurred boundaries and weak generalization ability.
[0004] Therefore, this invention proposes a land ecological remote sensing image segmentation method for monitoring areas where farmland has been converted back to forest, in order to solve the problems of weak generalization ability and low segmentation accuracy of existing land remote sensing image segmentation methods. Summary of the Invention
[0005] To address the aforementioned problems, this invention provides a method for segmenting land ecology remote sensing images in areas where farmland has been converted back to forest.
[0006] To achieve the above objectives, the present invention employs the following technical solution:
[0007] This invention provides a method for segmenting land ecology remote sensing images in areas where farmland has been converted back to forest, comprising the following steps:
[0008] S1. Obtain remote sensing images of the land ecology in the areas where farmland has been converted back to forest;
[0009] S2. Construct a land ecology remote sensing image segmentation model, the model including a remote sensing image feature enhancement module, a remote sensing image multi-scale semantic fusion module, a local information capture module, an edge information enhancement module, and a remote sensing image segmentation module;
[0010] S3. Input the land ecology remote sensing image into the remote sensing image feature enhancement module to perform feature enhancement operations and obtain the enhanced remote sensing image features;
[0011] S4. Input the enhanced remote sensing image features into the semantic space attention module in the remote sensing image multi-scale semantic fusion module to capture image spatial features and obtain semantic space attention module features. Input the semantic space attention module features into the multi-scale texture fusion module in the remote sensing image multi-scale semantic fusion module to obtain multi-scale semantic fusion features.
[0012] S5. Input the enhanced remote sensing image features into the remote sensing image channel attention module of the local information capture module to obtain the remote sensing image channel attention features. Input the remote sensing image channel attention features into the remote sensing image spatiotemporal attention module of the local information capture module. After the spatiotemporal feature extraction of the remote sensing image spatiotemporal attention module, the remote sensing image spatiotemporal attention features are obtained.
[0013] S6. Input the spatiotemporal attention features of the remote sensing image into the multi-scale spatial feature module of the edge information enhancement module to obtain multi-scale spatial features. Input the multi-scale spatial features into the edge grouping attention module of the edge information enhancement module to obtain edge grouping attention features.
[0014] S7. The multi-scale semantic fusion features and edge grouping attention features are fused together and then concatenated to obtain the fused features. The fused features are then input into the remote sensing image segmentation module, which outputs the corresponding predicted mask matrix to obtain the segmentation result.
[0015] Furthermore, step S3 specifically includes:
[0016] Land ecological remote sensing images The image is input into the remote sensing image feature enhancement module. First, the kernel size of the remote sensing image feature enhancement module is... In the convolutional layer, initial features are obtained after initial feature extraction. Initial features The kernels are processed by two parallel convolutions, each with a kernel size of [size missing]. and Depthwise separable convolutions yield two distinct depthwise separable features. and Two depth-separable features and By splicing, depth-separable splicing features are obtained. ,in, Indicates the height of the remote sensing image. Indicates the width of the remote sensing image. The number of channels in a remotely sensed image is expressed by the following formula:
[0017] ,
[0018] ,
[0019] in, Indicates the kernel size as Convolutional layers, This indicates a splicing operation. and These represent the kernel size as follows: and Depthwise separable convolution; concatenating depthwise separable features The input to the convolution kernel size is The convolutional layers expand their channel count to obtain depth-separable extended features. ,use The activation function is used for processing, and finally the convolution kernel size is... The convolutional layer recovers its channel count, resulting in enhanced remote sensing image features. The formula is expressed as follows:
[0020] ,
[0021] in, This indicates the activation function operation.
[0022] Furthermore, step S4 specifically includes:
[0023] Enhanced remote sensing image features The input is fed into the semantic space attention module, and then passes through a global average pooling layer with a convolutional kernel size of [missing value]. Convolutional layers and Activation function to obtain the first intermediate feature First intermediate feature The input to the convolution kernel size is Convolutional layers Activation function to obtain the second intermediate feature Second intermediate feature Features of enhanced remote sensing images Channel-weighted fusion is performed to obtain semantic space attention module features. The formula is expressed as follows:
[0024] ,
[0025] ,
[0026] in, This indicates a global average pooling operation. This represents the activation function. This indicates a channel weighted fusion operation. express Activation function;
[0027] Semantic space attention module features The input is fed into the multi-scale texture fusion module, and after passing through a convolution kernel with a size of [missing value], it is processed. The first convolutional feature is obtained from the convolutional layer. First convolutional feature The inputs are fed into three parallel, separable convolutional sub-networks of different sizes. First, the first convolutional feature is input into the first sub-network to obtain the first sub-network feature. The first sub-network includes convolutional kernels with a size of The depthwise separable convolution and the kernel size are Convolutional layers; first convolutional features The input is fed into the second sub-network to obtain the features of the second sub-network. The second sub-network includes convolutional kernels with a size of The depthwise separable convolution and the kernel size are The convolutional layers; finally, the first convolutional feature... The input is fed into the third sub-network to obtain the features of the third sub-network. The third sub-network includes convolutional kernels with a size of The depthwise separable convolution and the kernel size are The convolutional layer will incorporate the features of the first sub-network. Second sub-network characteristics and third sub-network features By splicing and blending, multi-scale fusion features are obtained. The formula is expressed as follows:
[0028] ,
[0029] ,
[0030] ,
[0031] ,
[0032] ,
[0033] in, , and These represent the kernel size as follows: , and Depth-separable convolution, Indicates the kernel size as convolutional layers, This indicates a feature concatenation operation; it fuses features across multiple scales. The input to the convolution kernel size is In the convolutional layer, the second convolutional feature is obtained. ;
[0034] The second convolutional features Features of the semantic space attention module Channel-weighted fusion is performed and then passed through an adaptive pooling layer to obtain multi-scale semantic fusion features. The formula is expressed as follows:
[0035] ,
[0036] ,
[0037] in, This indicates an adaptive pooling operation.
[0038] Furthermore, step S5 specifically includes:
[0039] Enhanced remote sensing image features The remote sensing image channel attention module, input to the local information capture module, passes through an adaptive max-pooling layer and a convolution kernel size of [missing information]. Point convolution, compressing the number of channels to half of the original number of channels. Output the first compression feature First compression feature go through The activation function performs a non-linear operation on the features, and then a convolution kernel of size is applied. The point convolution restores the number of channels in the feature to the original number of channels, resulting in the first restored feature. The formula is expressed as follows:
[0040] ,
[0041] ,
[0042] in, This indicates an adaptive max-pooling layer. Indicates the kernel size as Point convolution;
[0043] Enhanced remote sensing image features After adaptive average pooling layers and convolutional kernels of size [missing information], Point convolutions compress the number of channels to half the original number of channels. Output the second compression feature Second compression feature go through The activation function performs a non-linear operation on the features, and then a convolution kernel of size is applied. The point convolution restores the number of channels in the feature to the original number of channels, resulting in the second restored feature. The formula is expressed as follows:
[0044] ,
[0045] ,
[0046] in, Indicates the kernel size as Point convolution;
[0047] The first restored feature Second reduction feature splicing and fusion and after Activation function to obtain attention features of remote sensing image channels The formula is expressed as follows:
[0048] ,
[0049] in, express Activation function;
[0050] Attention features of remote sensing image channels The remote sensing image spatiotemporal attention module, which inputs to the local information capture module, processes the data through max pooling and average pooling layers. The two processed features are then concatenated to obtain the pooled feature. Pooling features The input to the convolution kernel size is Step size is Point convolution, then through Activation function to obtain spatiotemporal attention features of remote sensing images The formula is expressed as follows:
[0051] ,
[0052] ,
[0053] in, and These represent the max pooling layer and the average pooling layer, respectively. Indicates the kernel size as Step size is Point convolution.
[0054] Furthermore, step S6 specifically includes:
[0055] Spatiotemporal attention features of remote sensing images The spatiotemporal attention features of the remote sensing image are input into the multi-scale spatial feature module of the edge information enhancement module. The kernel size after the first multi-scale spatial feature sub-network is... Depth-separable convolution, batch normalization, and The activation function yields the first multi-scale spatial features. Spatiotemporal attention features of remote sensing images The kernel size after the second multi-scale spatial feature subnetwork is... Depth-separable convolution, batch normalization, and The activation function yields the second multi-scale spatial features. Spatiotemporal attention features of remote sensing images The kernel size after the third multi-scale spatial feature sub-network is Depth-separable convolution, batch normalization, and Activation function to obtain third-scale spatial features ; to include the first multi-scale spatial features Second multi-scale spatial features and third-scale spatial features Perform feature concatenation to obtain multi-scale spatial features. The formula is expressed as follows:
[0056]
[0057]
[0058]
[0059]
[0060] in, This indicates a batch of standardized operations. and These represent the kernel size as follows: as well as The depth of the separable convolutional layer;
[0061] Multi-scale spatial features The multi-scale spatial features are input into the edge grouping attention module of the edge information enhancement module. After convolution kernel size is The first group convolutional features are obtained by performing group convolution and batch normalization operations. Multi-scale spatial features After convolution kernel size is The second group convolution feature is obtained by performing group convolution and batch normalization operations. ; Convolve the first group of features Second group convolution features The convolutional fusion features of the first group are obtained by splicing and merging. First group convolutional fusion features go through Activation function, kernel size is Convolutional layers, batch normalization operations, and Activation function to obtain edge group attention features The formula is expressed as follows:
[0062] ,
[0063] ,
[0064] ,
[0065] ,
[0066] in, and These represent the kernel size as follows: The grouped convolution and the kernel size are Grouped convolution.
[0067] Furthermore, step S7 specifically includes:
[0068] Multi-scale semantic fusion features and edge grouping attention features After splicing, the fused features are obtained. ; to integrate features The image is input into the remote sensing image segmentation module and processed by a convolution kernel with a size of [missing information]. The point convolution, the kernel size is The point convolution is applied, and the features are restored to the resolution of the original input image through upsampling, outputting image features. The formula is expressed as follows:
[0069] ,
[0070] ,
[0071] in, Indicates an upsampling operation;
[0072] Image features go through The operation determines the category to which each pixel belongs, and finally... Color mapping operations generate the final mask matrix. The formula is expressed as follows:
[0073] ,
[0074] in, express operate, This indicates a color mapping operation.
[0075] Furthermore, during model training, the cross-entropy loss function is used as the objective function for model optimization. The Adam optimizer is used during model training, and the initial learning rate is set to 1e-4.
[0076] The advantages of this invention are:
[0077] This invention discloses a method for segmenting land ecology remote sensing images in areas where farmland has been converted to forest. It utilizes a remote sensing image feature enhancement module to improve the semantic expressive power of the original image, and combines a semantic spatial attention mechanism and a multi-scale texture fusion strategy in a multi-scale semantic fusion module to achieve deep fusion of information at different scales. Simultaneously, a local information capture module is introduced, which jointly models local and dynamic features in the remote sensing image through channel attention and spatiotemporal attention mechanisms, enhancing the model's ability to perceive regional heterogeneous changes. Regarding edge information enhancement, multi-scale spatial feature extraction and edge grouping attention mechanisms effectively highlight detailed information about ecological boundary areas in the remote sensing image. Finally, the multi-scale semantic fusion features and edge features are fused, and the corresponding mask matrix is output through the remote sensing image segmentation module, significantly improving the accuracy and boundary refinement effect of land ecology remote sensing image segmentation in areas where farmland has been converted to forest. Attached Figure Description
[0078] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0079] Figure 1 This is a flowchart of the steps of the method of the present invention;
[0080] Figure 2 This image shows a comparison of the land remote sensing image segmentation results of the method of this invention with other models. Detailed Implementation
[0081] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0082] Example 1
[0083] In this embodiment, as Figure 1 As shown, this invention provides a method for segmenting land ecology remote sensing images in areas where farmland has been converted back to forest. The specific steps include:
[0084] S1. Obtain remote sensing images of the land ecology in the areas where farmland has been converted back to forest;
[0085] S2. Construct a land ecology remote sensing image segmentation model, the model including a remote sensing image feature enhancement module, a remote sensing image multi-scale semantic fusion module, a local information capture module, an edge information enhancement module, and a remote sensing image segmentation module;
[0086] S3. Input the land ecology remote sensing image into the remote sensing image feature enhancement module to perform feature enhancement operations and obtain the enhanced remote sensing image features;
[0087] Specifically, land ecological remote sensing images The image is input into the remote sensing image feature enhancement module. First, the kernel size of the remote sensing image feature enhancement module is... In the convolutional layer, initial features are obtained after initial feature extraction. Initial features The kernels are processed by two parallel convolutions, each with a kernel size of [size missing]. and Depthwise separable convolutions yield two distinct depthwise separable features. and Two depth-separable features and By splicing, depth-separable splicing features are obtained. ,in, Indicates the height of the remote sensing image. Indicates the width of the remote sensing image. The number of channels in a remotely sensed image is expressed by the following formula:
[0088] ,
[0089] ,
[0090] in, Indicates the kernel size as convolutional layers, This indicates a splicing operation. and These represent the kernel size as follows: and Depthwise separable convolution; concatenating depthwise separable features The input to the convolution kernel size is The convolutional layers expand their channel count to obtain depth-separable extended features. ,use The activation function is used for processing, and finally the convolution kernel size is... The convolutional layer recovers its channel count, resulting in enhanced remote sensing image features. The formula is expressed as follows:
[0091] ,
[0092] in, This indicates the activation function operation.
[0093] S4. Input the enhanced remote sensing image features into the semantic space attention module in the remote sensing image multi-scale semantic fusion module to capture image spatial features and obtain semantic space attention module features. Input the semantic space attention module features into the multi-scale texture fusion module in the remote sensing image multi-scale semantic fusion module to obtain multi-scale semantic fusion features.
[0094] Specifically, the enhanced remote sensing image features The input is fed into the semantic space attention module, and then passes through a global average pooling layer with a convolutional kernel size of [missing value]. Convolutional layers and Activation function to obtain the first intermediate feature First intermediate feature The input to the convolution kernel size is Convolutional layers Activation function to obtain the second intermediate feature Second intermediate feature Features of enhanced remote sensing images Channel-weighted fusion is performed to obtain semantic space attention module features. The formula is expressed as follows:
[0095] ,
[0096] ,
[0097] in, This indicates a global average pooling operation. This represents the activation function. This indicates a channel weighted fusion operation. express Activation function;
[0098] Semantic space attention module features The input is fed into the multi-scale texture fusion module, and after passing through a convolution kernel with a size of [missing value], it is processed. The first convolutional feature is obtained from the convolutional layer. First convolutional feature The inputs are fed into three parallel, separable convolutional sub-networks of different sizes. First, the first convolutional feature is input into the first sub-network to obtain the first sub-network feature. The first sub-network includes convolutional kernels with a size of The depthwise separable convolution and the kernel size are Convolutional layers; first convolutional features The input is fed into the second sub-network to obtain the features of the second sub-network. The second sub-network includes convolutional kernels with a size of The depthwise separable convolution and the kernel size are The convolutional layers; finally, the first convolutional feature... The input is fed into the third sub-network to obtain the features of the third sub-network. The third sub-network includes convolutional kernels with a size of The depthwise separable convolution and the kernel size are The convolutional layer will incorporate the features of the first sub-network. Second sub-network characteristics and third sub-network features By splicing and blending, multi-scale fusion features are obtained. The formula is expressed as follows:
[0099] ,
[0100] ,
[0101] ,
[0102] ,
[0103] ,
[0104] in, , and These represent the kernel size as follows: , and Depth-separable convolution, Indicates the kernel size as convolutional layers, This indicates a feature concatenation operation; it fuses features across multiple scales. The input to the convolution kernel size is In the convolutional layer, the second convolutional feature is obtained. ;
[0105] The second convolutional features Features of the semantic space attention module Channel-weighted fusion is performed and then passed through an adaptive pooling layer to obtain multi-scale semantic fusion features. The formula is expressed as follows:
[0106] ,
[0107] ,
[0108] in, This indicates an adaptive pooling operation.
[0109] S5. Input the enhanced remote sensing image features into the remote sensing image channel attention module of the local information capture module to obtain the remote sensing image channel attention features. Input the remote sensing image channel attention features into the remote sensing image spatiotemporal attention module of the local information capture module. After the spatiotemporal feature extraction of the remote sensing image spatiotemporal attention module, the remote sensing image spatiotemporal attention features are obtained.
[0110] Specifically, the enhanced remote sensing image features The remote sensing image channel attention module, input to the local information capture module, passes through an adaptive max-pooling layer and a convolution kernel size of [missing information]. Point convolution, compressing the number of channels to half of the original number of channels. Output the first compression feature First compression feature go through The activation function performs a non-linear operation on the features, and then a convolution kernel of size is applied. The point convolution restores the number of channels in the feature to the original number of channels, resulting in the first restored feature. The formula is expressed as follows:
[0111] ,
[0112] ,
[0113] in, This indicates an adaptive max-pooling layer. Indicates the kernel size as Point convolution;
[0114] Enhanced remote sensing image features After adaptive average pooling layers and convolutional kernels of size [missing information], Point convolutions compress the number of channels to half the original number of channels. Output the second compression feature Second compression feature go through The activation function performs a non-linear operation on the features, and then a convolution kernel of size is applied. The point convolution restores the number of channels in the feature to the original number of channels, resulting in the second restored feature. The formula is expressed as follows:
[0115] ,
[0116] ,
[0117] in, Indicates the kernel size as Point convolution;
[0118] The first restored feature Second reduction feature splicing and fusion and after Activation function to obtain attention features of remote sensing image channels The formula is expressed as follows:
[0119] ,
[0120] in, express Activation function;
[0121] Attention features of remote sensing image channels The remote sensing image spatiotemporal attention module, which inputs to the local information capture module, processes the data through max pooling and average pooling layers. The two processed features are then concatenated to obtain the pooled feature. Pooling features The input to the convolution kernel size is Step size is Point convolution, then through Activation function to obtain spatiotemporal attention features of remote sensing images The formula is expressed as follows:
[0122] ,
[0123] ,
[0124] in, and These represent the max pooling layer and the average pooling layer, respectively. Indicates the kernel size as Step size is Point convolution.
[0125] S6. Input the spatiotemporal attention features of the remote sensing image into the multi-scale spatial feature module of the edge information enhancement module to obtain multi-scale spatial features. Input the multi-scale spatial features into the edge grouping attention module of the edge information enhancement module to obtain edge grouping attention features.
[0126] Specifically, the spatiotemporal attention features of remote sensing images The spatiotemporal attention features of the remote sensing image are input into the multi-scale spatial feature module of the edge information enhancement module. The kernel size after the first multi-scale spatial feature sub-network is... Depth-separable convolution, batch normalization, and The activation function yields the first multi-scale spatial features. Spatiotemporal attention features of remote sensing images The kernel size after the second multi-scale spatial feature subnetwork is... Depth-separable convolution, batch normalization, and The activation function yields the second multi-scale spatial features. Spatiotemporal attention features of remote sensing images The kernel size after the third multi-scale spatial feature sub-network is Depth-separable convolution, batch normalization, and Activation function to obtain third-scale spatial features ; to include the first multi-scale spatial features Second multi-scale spatial features and third-scale spatial features Perform feature concatenation to obtain multi-scale spatial features. The formula is expressed as follows:
[0127]
[0128]
[0129]
[0130]
[0131] in, This indicates a batch of standardized operations. and These represent the kernel size as follows: as well as The depth of the separable convolutional layer;
[0132] Multi-scale spatial features The multi-scale spatial features are input into the edge grouping attention module of the edge information enhancement module. After convolution kernel size is The first group convolutional features are obtained by performing group convolution and batch normalization operations. Multi-scale spatial features After convolution kernel size is The second group convolution feature is obtained by performing group convolution and batch normalization operations. ; Convolve the first group of features Second group convolution features The convolutional fusion features of the first group are obtained by splicing and merging. First group convolutional fusion features go through Activation function, kernel size is Convolutional layers, batch normalization operations, and Activation function to obtain edge group attention features The formula is expressed as follows:
[0133] ,
[0134] ,
[0135] ,
[0136] ,
[0137] in, and These represent the kernel size as follows: The grouped convolution and the kernel size are Grouped convolution.
[0138] S7. The multi-scale semantic fusion features and edge grouping attention features are fused together and then concatenated to obtain the fused features. The fused features are then input into the remote sensing image segmentation module, which outputs the corresponding predicted mask matrix to obtain the segmentation result.
[0139] Specifically, multi-scale semantic fusion features and edge grouping attention features After splicing, the fused features are obtained. ; to integrate features The image is input into the remote sensing image segmentation module and processed by a convolution kernel with a size of [missing information]. The point convolution, the kernel size is The point convolution is applied, and the features are restored to the resolution of the original input image through upsampling, outputting image features. The formula is expressed as follows:
[0140] ,
[0141] ,
[0142] in, Indicates an upsampling operation;
[0143] Image features go through The operation determines the category to which each pixel belongs, and finally... Color mapping operations generate the final mask matrix. The formula is expressed as follows:
[0144] ,
[0145] in, express operate, This indicates a color mapping operation.
[0146] Specifically, during model training, the cross-entropy loss function is used as the model's objective function. The Adam optimizer is used during model training, and the initial learning rate is set to 1e-4.
[0147] Example 2
[0148] In this embodiment, to verify the effectiveness of the land ecological remote sensing image segmentation method for monitoring areas of land converted from farmland to forest proposed in the field of land remote sensing image recognition and segmentation, the proposed land remote sensing image segmentation method and existing remote sensing image segmentation methods are experimentally verified under the same experimental configuration conditions. The obtained remote sensing image segmentation results are analyzed in detail to verify the effectiveness of the method proposed in the field of land remote sensing image segmentation.
[0149] In the comparative experiment between the proposed method and other existing remote sensing image segmentation methods, three mainstream remote sensing image segmentation methods were selected as comparison models. The experimental results were analyzed in detail to verify the effectiveness of the proposed method in the field of land remote sensing image segmentation. The three comparison models used in the experiment are: the MANet model, based on the U-Net architecture, can effectively capture the correlation of segmentation targets in complex scenes in remote sensing images, but its ability to recognize local image details is weak; the DC-Swin model uses the Swing Transformer as the backbone network, which is suitable for large-scale scene segmentation tasks, but its computational complexity is high, and the model requires a large amount of training data; the CTMFNet model combines CNN and Transformer architectures, which performs well in terms of overall performance, but the model has a large number of parameters, and the training and inference processes of the model have high requirements for computational resources.
[0150] In a comparative experiment of the proposed land remote sensing image segmentation method for monitoring areas where farmland has been converted back to forest, three experimental metrics were used to evaluate the effectiveness of each model: F1-score, which balances precision and recall to comprehensively evaluate the model's ability to segment and identify land in areas where farmland has been converted back to forest; mean intersection-union ratio (mIoU), which measures the model's accuracy in locating target boundaries by calculating the ratio of the intersection and union of the predicted segmentation results and the true labels; and overall accuracy (OA), which represents the proportion of correctly classified pixels out of the total number of pixels, providing a rapid evaluation of the global segmentation effect.
[0151] The dataset used in the experiment is DeepGlobe, which uses high-resolution land remote sensing images acquired through satellite remote sensing technology. This dataset covers the geographical environment of multiple regions and includes seven land cover types (forest, farmland, water bodies, urban areas, bare land, grassland, and shrubs). It also includes images from different seasons and years to reflect dynamic changes in the land surface, specifically for scenarios such as reforestation and urbanization. The samples in each category of this dataset exhibit a long-tailed distribution. In the experiment, the land remote sensing image dataset was divided into training, validation, and test sets in a 7:1:2 ratio to verify the model's generalization ability and segmentation accuracy in complex scenarios.
[0152] The experimental results of the proposed land remote sensing image segmentation method and the comparison model on the DeepGlobe dataset are shown in Table 1. The proposed method achieves F1 score of 89.40%, mean Intersection over Union (mIoU) of 84.25%, and overall accuracy (OA) of 91.37%, which are significantly better than the other three comparison models. Compared with the suboptimal model CTMFNet (F1-score 87.33%, mIoU 81.64%, OA 88.52%), the proposed method improves F1-score by 2.07%, mIoU by 2.61%, and OA by 2.85%, verifying the advantages of the proposed method in small-class recognition and complex boundary segmentation.
[0153] Table 1 shows the comparison results of validating the proposed method on the dataset.
[0154]
[0155] The actual segmentation results of the proposed method and other comparative models in the experiment are as follows: Figure 2 As shown, the red areas in the image represent the colors of the separated houses, while the gray-white areas represent the colors used to identify the woodland. Figure 2 As can be seen, the proposed method's segmentation and recognition effect is more consistent with the forest area in the input image. The other three models are far less accurate in segmenting and recognizing forest areas in complex scenes in the image than the proposed method. This shows that the proposed method effectively solves the problem of low image segmentation accuracy and weak generalization ability caused by complex scenes and blurred boundaries in land remote sensing images, and verifies the effectiveness of the proposed method.
[0156] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A land ecological remote sensing image segmentation method for monitoring a land-for-forest region, characterized in that, The method comprises the following steps: S1. Obtain a land ecological remote sensing image of a land-for-forest area; S2. Construct a land ecological remote sensing image segmentation model, the model comprising a remote sensing image feature enhancement module, a remote sensing image multi-scale semantic fusion module, a local information capturing module, an edge information enhancement module, and a remote sensing image segmentation module; S3. Input the land ecological remote sensing image into the remote sensing image feature enhancement module for feature enhancement operation to obtain enhanced remote sensing image features; S4. Input the enhanced remote sensing image features into a semantic space attention module in the remote sensing image multi-scale semantic fusion module for image spatial feature capturing to obtain semantic space attention module features, and input the semantic space attention module features into a multi-scale texture fusion module in the remote sensing image multi-scale semantic fusion module to obtain multi-scale semantic fusion features; S5. Input the enhanced remote sensing image features into a remote sensing image channel attention module of the local information capturing module to obtain remote sensing image channel attention features, and input the remote sensing image channel attention features into a remote sensing image space-time attention module of the local information capturing module to obtain remote sensing image space-time attention features through space-time feature extraction of the remote sensing image space-time attention module; The specific steps comprise: Enhanced remote sensing image features The remote sensing image channel attention module, input to the local information capture module, passes through an adaptive max-pooling layer and a convolution kernel size of [missing information]. Point convolution, compressing the number of channels to half of the original number of channels. Output the first compression feature First compression feature go through The activation function performs a non-linear operation on the features, and then a convolution kernel of size is applied. The point convolution restores the number of channels in the feature to the original number of channels, resulting in the first restored feature. ; Enhanced remote sensing image features After adaptive average pooling layers and convolutional kernels of size [missing information], Point convolutions compress the number of channels to half the original number of channels. Output the second compression feature Second compression feature go through The activation function performs a non-linear operation on the features, and then a convolution kernel of size is applied. The point convolution restores the number of channels in the feature to the original number of channels, resulting in the second restored feature. ; The first reduction feature And the second reduction feature Splice fusion and pass Activation function, get remote sensing image channel attention feature ; S6. Input the remote sensing image space-time attention features into a multi-scale spatial feature module of the edge information enhancement module to obtain multi-scale spatial features, input the multi-scale spatial features into an edge grouping attention module of the edge information enhancement module to obtain edge grouping attention features; S7. Fuse the multi-scale semantic fusion features and the edge grouping attention features, obtain fusion features through a splicing operation, input the fusion features into the remote sensing image segmentation module, output a final predicted corresponding mask matrix, and obtain a segmentation result.
2. The land ecological remote sensing image segmentation method for monitoring the area of returning farmland to forest according to claim 1, characterized in that, Step S3 specifically comprises: The land ecological remote sensing image is input to a remote sensing image feature enhancement module, and the image is processed by the remote sensing image feature enhancement module. The initial feature is obtained by performing initial feature extraction on the remote sensing image. The initial feature is input to two parallel depth separable convolutions with kernel sizes of respectively to obtain two different depth separable features and . The two depth separable features are spliced to obtain a depth separable splicing feature , wherein represents the height of the remote sensing image, represents the width of the remote sensing image, represents the channel number of the remote sensing image, and the formula is as follows: , , wherein, represents a convolutional layer with a convolution kernel size of , represents a concatenation operation, and represent deep separable convolutions with a convolution kernel size of and , respectively; the deep separable concatenated features are input to a convolutional layer with a convolution kernel size of to expand the number of channels thereof to obtain deep separable expanded features , which are processed using an activation function, and finally a convolutional layer with a convolution kernel size of is used to restore the number of channels thereof to obtain an enhanced remote sensing image feature , which is expressed by the following formula: , wherein, denotes an activation function operation.
3. The land ecological remote sensing image segmentation method for monitoring the area of returning farmland to forest according to claim 2, characterized in that, Step S4 specifically comprises: Enhanced remote sensing image features The input is fed into the semantic space attention module, and then passes through a global average pooling layer with a convolutional kernel size of [missing value]. Convolutional layers and Activation function to obtain the first intermediate feature First intermediate feature The input to the convolution kernel size is Convolutional layers Activation function to obtain the second intermediate feature Second intermediate feature Features of enhanced remote sensing images Channel-weighted fusion is performed to obtain semantic space attention module features. The formula is expressed as follows: , , wherein, represents a global average pooling operation, represents an activation function, represents a channel-weighted fusion operation, represents an activation function; The semantic space attention module features are input into a multi-scale texture fusion module, and after a convolutional layer with a convolution kernel size of , first convolutional features are obtained; the first convolutional features are input into three parallel separable convolution sub-networks of different sizes, first, the first convolutional features are input into a first sub-network to obtain first sub-network features , the first sub-network includes a depth separable convolution with a convolution kernel size of and a convolutional layer with a convolution kernel size of ; the first convolutional features are input into a second sub-network to obtain second sub-network features , the second sub-network includes a depth separable convolution with a convolution kernel size of and a convolutional layer with a convolution kernel size of ; finally, the first convolutional features are input into a third sub-network to obtain third sub-network features , the third sub-network includes a depth separable convolution with a convolution kernel size of and a convolutional layer with a convolution kernel size of , the first sub-network features , the second sub-network features and the third sub-network features are spliced and fused to obtain multi-scale fusion features , and the formula is as follows: , , , , , wherein, , and respectively represent depth separable convolutions with kernel size , and , represent a convolution layer with kernel size , represents a feature concatenation operation; the multi-scale fusion features are input into a convolution layer with kernel size to obtain second convolution features ; The second convolutional feature with the semantic space attention module feature The channel weighting fusion operation is performed and the adaptive pooling layer is passed to obtain a multi-scale semantic fusion feature , which is expressed as follows: , , wherein, denotes an adaptive pooling operation.
4. The land ecological remote sensing image segmentation method for monitoring the area of returning farmland to forest according to claim 3, characterized in that, In step S5, the first reduction feature is obtained The formula is expressed as follows: , , wherein, denotes an adaptive max pooling layer, denotes a point convolution with a kernel size of ; obtaining a second reduction characteristic which is expressed by the following equation: , , wherein, represents a point convolution with a kernel size of ; Obtaining channel attention features of a remote sensing image The formula is expressed as follows: , wherein represents activation function; Attention features of remote sensing image channels The remote sensing image spatiotemporal attention module, which inputs to the local information capture module, processes the data through max pooling and average pooling layers. The two processed features are then concatenated to obtain the pooled feature. Pooling features The input to the convolution kernel size is Step size is Point convolution, then through Activation function to obtain spatiotemporal attention features of remote sensing images The formula is expressed as follows: , , wherein, and respectively represent a max-pooling layer and an average-pooling layer, represents a dot convolution with a convolution kernel size of , and a step size of .
5. The land ecological remote sensing image segmentation method for monitoring the area of returning farmland to forest according to claim 4, characterized in that, Step S6 specifically comprises: Remote sensing image spatio-temporal attention feature The multi-scale spatial feature module input to the edge information enhancement module, the remote sensing image spatio-temporal attention feature The convolution kernel size of the first multi-scale spatial feature subnetwork is The depth separable convolution, batch normalization operation and Activation function, get the first multi-scale spatial feature Remote sensing image spatio-temporal attention feature The convolution kernel size of the second multi-scale spatial feature subnetwork is The depth separable convolution, batch normalization operation and Activation function, get the second multi-scale spatial feature Remote sensing image spatio-temporal attention feature The convolution kernel size of the third multi-scale spatial feature subnetwork is The depth separable convolution, batch normalization operation and Activation function, get the third multi-scale spatial feature The first multi-scale spatial feature , the second multi-scale spatial feature And the third multi-scale spatial feature Feature splicing operation, get multi-scale spatial feature , the formula is as follows: wherein, denotes a batch normalization operation, and denotes a depthwise separable convolution layer with kernel size and respectively. Multi-scale spatial features The multi-scale spatial features are input into the edge grouping attention module of the edge information enhancement module. After convolution kernel size is The first group convolutional features are obtained by performing group convolution and batch normalization operations. Multi-scale spatial features After convolution kernel size is The second group convolution feature is obtained by performing group convolution and batch normalization operations. ; Convolve the first group of features Second group convolution features The convolutional fusion features of the first group are obtained by splicing and merging. First group convolutional fusion features go through Activation function, kernel size is Convolutional layers, batch normalization operations, and Activation function to obtain edge group attention features The formula is expressed as follows: , , , , wherein, and denote grouped convolutions with kernel size and grouped convolutions with kernel size respectively.
6. The land ecological remote sensing image segmentation method for monitoring the area of returning farmland to forest according to claim 5, characterized in that, Step S7 specifically comprises: Multi-scale semantic fusion features and edge grouping attention features After splicing, the fused features are obtained. ; to integrate features The image is input into the remote sensing image segmentation module and processed by a convolution kernel with a size of [missing information]. The point convolution, the kernel size is The point convolution is applied, and the features are restored to the resolution of the original input image through upsampling, outputting image features. The formula is expressed as follows: , , wherein denotes an up-sampling operation; The image features are extracted After The operation determines the category to which each pixel belongs, and finally generates the final mask matrix through Color mapping operation The formula is as follows: , wherein represents operates, represents a color mapping operation.
7. The land ecological remote sensing image segmentation method for monitoring the area of returning farmland to forest according to claim 6, characterized in that, In the model training process, a cross-entropy loss function cross-entropy is used as the optimization objective function of the model, and an Adam optimizer is used for model training, with an initial learning rate of 1e-4.