Land ecology remote sensing image segmentation method for monitoring farmland-to-forest returning area
Through the land ecological remote sensing image segmentation method, feature enhancement, multi-scale semantic fusion and edge information enhancement modules are used to solve the problems of low segmentation accuracy and weak generalization ability of remote sensing image, and high-precision segmentation and boundary refinement of the returning farmland to forest areas are achieved.
Patent Information
- Application Number
- CN202510540415.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-04-27
AI Technical Summary
When monitoring the area where farmland is returned to forest, the existing land remote sensing image segmentation method has low segmentation accuracy and weak generalization ability, it is difficult to identify small-area plots and complex terrain, and is susceptible to cloud occlusion, seasonal changes and spectral feature similarity.
The land ecological remote sensing image segmentation method is adopted, and the feature enhancement module, multi-scale semantic fusion module, local information capture module and edge information enhancement module are combined with the semantic spatial attention mechanism, channel attention and space-time attention mechanism to improve the image feature expression and edge refinement capabilities, and the output mask matrix is segmented.
It significantly improves the segmentation accuracy and boundary refinement effect of land ecological remote sensing images in areas where farmland is returned to forest, and improves the recognition and generalization ability of complex scenes.
Smart Images

Figure CN120451185A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of land ecological remote sensing image segmentation, and in particular relates to a land ecological remote sensing image segmentation method for monitoring areas where farmland is being converted to forest. Background Art
[0002] Monitoring the extent of land in areas undergoing farmland-to-forest conversion is a key topic for researchers in agriculture and forestry. In the past, in the absence of satellite remote sensing imagery, monitoring efforts primarily relied on manual field surveys combined with traditional mapping techniques. These methods included regularly organizing monitoring visits to areas undergoing farmland-to-forest conversion, visually observing, comparing topographic maps, manually recording vegetation cover changes, and using aerial photography to roughly assess the extent of the area. However, this method is inefficient, requires a long monitoring cycle, and the accuracy of the data is highly dependent on the experience and commitment of the monitoring personnel.
[0003] With the development of satellite remote sensing image capture technology, using satellite remote sensing to monitor the scope of land converted from farmland to forest has become an important topic for researchers. Satellite remote sensing technology has improved the real-time performance of land monitoring for returning farmland to forest. However, segmentation technology based on remote sensing images also has obvious drawbacks in its application. For example, remote sensing image segmentation algorithms are limited by image resolution and are difficult to accurately identify small plots or complex terrain areas. Remote sensing image segmentation algorithms are easily affected by factors such as cloud occlusion, seasonal changes, shadows, or similar spectral characteristics of crops and forests and grasses, resulting in mis-segmentation or missed segmentation. In addition, traditional threshold segmentation algorithms have problems such as insufficient adaptability to land use types with fuzzy boundaries and weak generalization ability.
[0004] Therefore, the present invention proposes a land ecological remote sensing image segmentation method for monitoring the area of returning farmland to forest, so as to solve the problems of weak generalization ability and low segmentation accuracy of existing land remote sensing image segmentation methods. Summary of the Invention
[0005] In order to solve the above problems, the present invention provides a land ecological remote sensing image segmentation method for monitoring the area of returning farmland to forest.
[0006] To achieve the above-mentioned purpose, the present invention is implemented through the following technical solutions: The present invention provides a land ecological remote sensing image segmentation method for monitoring a land reclamation and restoration area, comprising the following steps: S1. Obtaining remote sensing images of land ecology in the area of returning farmland to forest; S2. Constructing a land ecological remote sensing image segmentation model, the model includes a remote sensing image feature enhancement module, a remote sensing image multi-scale semantic fusion module, a local information capture module, an edge information enhancement module, and a remote sensing image segmentation module; S3. Inputting the land ecological remote sensing image into the remote sensing image feature enhancement module to perform feature enhancement operation to obtain enhanced remote sensing image features; S4. Input the enhanced remote sensing image features into the semantic spatial attention module in the remote sensing image multi-scale semantic fusion module to capture image spatial features, thereby obtaining semantic spatial attention module features. Input the semantic spatial attention module features into the multi-scale texture fusion module in the remote sensing image multi-scale semantic fusion module to obtain multi-scale semantic fusion features. S5. Input the enhanced remote sensing image features into the remote sensing image channel attention module of the local information capture module to obtain remote sensing image channel attention features. Input the remote sensing image channel attention features into the remote sensing image spatiotemporal attention module of the local information capture module to extract spatiotemporal features of the remote sensing image spatiotemporal attention module to obtain remote sensing image spatiotemporal attention features. S6. Input the spatiotemporal attention features of the remote sensing image into the multi-scale spatial feature module of the edge information enhancement module to obtain a multi-scale spatial feature, and input the multi-scale spatial feature into the edge grouping attention module of the edge information enhancement module to obtain an edge grouping attention feature; S7. Fuse the multi-scale semantic fusion features and edge grouping attention features, obtain fusion features through splicing operations, input the fusion features into the remote sensing image segmentation module, output the final predicted corresponding mask matrix, and obtain the segmentation result.
[0007] Furthermore, step S3 specifically includes: Land ecological remote sensing images Input to the remote sensing image feature enhancement module, the image First, the convolution kernel size of the remote sensing image feature enhancement module is In the convolution layer, the initial features are obtained after initial feature extraction ; Initial features After two parallel convolution kernels of size and Depth-wise separable convolution, resulting in two different depth-wise separable features and , two depth-separable features and Splicing to obtain depth-separable splicing features ,in, Indicates the height of the remote sensing image, Indicates the width of the remote sensing image, Represents the number of channels of the remote sensing image. The formula is as follows: , , in, Indicates that the convolution kernel size is The convolutional layer, Represents a splicing operation, and They represent the convolution kernel size respectively. and Depthwise separable convolution; depthwise separable concatenation features The input to the convolution kernel size is The convolution layer expands its channel number to obtain depth-separable extended features ,use Activation function processing, and finally the convolution kernel size is The convolution layer restores the number of channels and obtains the enhanced remote sensing image features , the formula is as follows: , in, Represents the activation function operation.
[0008] Furthermore, step S4 specifically includes: The enhanced remote sensing image features Input into the semantic space attention module, after the global average pooling layer and the convolution kernel size is The convolutional layer and Activation function, get the first intermediate feature ; First intermediate feature The input to the convolution kernel size is The convolutional layer, Activation function to obtain the second intermediate feature ; Second intermediate feature and enhanced remote sensing image features Perform channel weighted fusion to obtain semantic space attention module features , the formula is as follows: , , in, represents the global average pooling operation, represents the activation function, represents the channel weighted fusion operation, express Activation function; The semantic space attention module features Input into the multi-scale texture fusion module, after the convolution kernel size is The convolution layer obtains the first convolution feature ; The first convolution feature Input into three parallel separable convolutional sub-networks of different sizes. First, the first convolutional feature is input into the first sub-network to obtain the first sub-network feature. , the first sub-network includes a convolution kernel size of The depth-wise separable convolution and the convolution kernel size are Convolutional layer; the first convolution feature Input to the second sub-network to obtain the second sub-network features , the second sub-network includes a convolution kernel size of The depth-wise separable convolution and the convolution kernel size are Convolutional layer; Finally, the first convolution feature Input to the third sub-network to obtain the third sub-network features , the third sub-network includes a convolution kernel size of The depth-wise separable convolution and the convolution kernel size are The convolution layer of the first sub-network features , the second sub-network characteristics and the third subnetwork characteristics Splicing and fusion to obtain multi-scale fusion features , the formula is as follows: , , , , , in, 、 and They represent the convolution kernel size respectively. 、 and Depthwise separable convolution, Indicates that the convolution kernel size is The convolutional layer, Represents feature splicing operation; multi-scale fusion features The input to the convolution kernel size is In the convolution layer, the second convolution feature is obtained ; The second convolution feature Features of semantic space attention module Perform channel weighted fusion operation and pass through the adaptive pooling layer to obtain multi-scale semantic fusion features , the formula is as follows: , , in, Represents an adaptive pooling operation.
[0009] Furthermore, step S5 specifically includes: The enhanced remote sensing image features The remote sensing image channel attention module input to the local information capture module passes through the adaptive maximum pooling layer and the convolution kernel size is Point convolution, the number of channels is compressed to the original number of channels , output the first compressed feature ; First compression feature go through The activation function performs nonlinear operations on the features, and then uses the convolution kernel size of The point convolution restores the number of channels of the feature to the original number of channels and obtains the first restored feature , the formula is as follows: , , in, represents the adaptive maximum pooling layer, Indicates that the convolution kernel size is Point convolution; Enhanced remote sensing image features After the adaptive average pooling layer and the convolution kernel size is Point convolution compresses the number of channels to the original number of channels , output the second compressed feature ; Second compression feature go through The activation function performs nonlinear operations on the features, and then uses the convolution kernel size of The point convolution restores the number of channels of the feature to the original number of channels, and obtains the second restored feature , the formula is as follows: , in, Indicates that the convolution kernel size is Point convolution; The first reduction feature and the second reduction feature Splicing and fusion Activation function to obtain the channel attention features of remote sensing images , the formula is as follows: , in, express Activation function; Remote sensing image channel attention features The remote sensing image spatiotemporal attention module input to the local information capture module is processed by the maximum pooling layer and the average pooling layer respectively, and the two processed features are spliced to obtain the pooled feature , the pooled features The input to the convolution kernel size is , the step length is Point convolution, then through Activation function to obtain the spatiotemporal attention features of remote sensing images , the formula is as follows: , , in, and Represent the maximum pooling layer and the average pooling layer respectively, Indicates that the convolution kernel size is , the step length is Point convolution.
[0010] Furthermore, step S6 specifically includes: The spatiotemporal attention features of remote sensing images Input to the multi-scale spatial feature module of the edge information enhancement module, remote sensing image spatiotemporal attention features The convolution kernel size after the first multi-scale spatial feature sub-network is Depthwise separable convolution, batch normalization operations, and Activation function to obtain the first multi-scale spatial feature ; Spatial-temporal attention features of remote sensing images The convolution kernel size after the second multi-scale spatial feature sub-network is Depthwise separable convolution, batch normalization operations, and Activation function to obtain the second multi-scale spatial feature ; Spatial-temporal attention features of remote sensing images The convolution kernel size after the third multi-scale spatial feature sub-network is Depthwise separable convolution, batch normalization operations, and Activation function to obtain the third multi-scale spatial feature ; The first multi-scale spatial feature , the second multi-scale spatial feature And the third multi-scale spatial feature Perform feature splicing operations to obtain multi-scale spatial features , the formula is as follows: in, represents the batch normalization operation, and They represent the convolution kernel size respectively. as well as Depthwise separable convolutional layers; Multi-scale spatial features Input to the edge grouping attention module of the edge information enhancement module, the multi-scale spatial features After the convolution kernel size is The group convolution and batch normalization operations are performed to obtain the first group convolution feature ; Multi-scale spatial features After the convolution kernel size is The grouped convolution and batch normalization operations are performed to obtain the second grouped convolution feature ; Convolution features of the first group and the second group convolutional features Perform splicing and fusion to obtain the first group convolution fusion feature ; The first group convolution fusion feature go through The activation function and convolution kernel size are The convolutional layers, batch normalization operations, and Activation function to obtain edge group attention features , the formula is as follows: , , , , in, and They represent the convolution kernel size respectively. The group convolution and convolution kernel size are Grouped convolution.
[0011] Furthermore, step S7 specifically includes: Multi-scale semantic fusion features and edge grouping attention features After the splicing operation, the fusion features are obtained ; The fusion features Input into the remote sensing image segmentation module, after the convolution kernel size is Point convolution, convolution kernel size is The point convolution is performed, and the features are restored to the resolution of the original input image through upsampling operation, and the output image features are , the formula is as follows: , , in, Represents an upsampling operation; The image features go through Operation determines the category to which each pixel belongs, and finally Color mapping operation to generate the final mask matrix , the formula is as follows: , in, express operate, Represents a color mapping operation.
[0012] Furthermore, during the model training process, the cross-entropy loss function is used as the optimization objective function of the model. The Adam optimizer is used for model training, and the initial learning rate is set to 1e-4.
[0013] The advantages of the present invention are: This invention discloses a land ecology remote sensing image segmentation method for monitoring areas where farmland has been converted to forest. This method utilizes a remote sensing image feature enhancement module to enhance the semantic expression capabilities of the original image. It combines the semantic spatial attention mechanism and multiscale texture fusion strategy in the remote sensing image multiscale semantic fusion module to achieve deep fusion of information at different scales. Furthermore, a local information capture module is introduced to jointly model local and dynamic features in remote sensing images through channel attention and spatiotemporal attention mechanisms, enhancing the model's ability to perceive regional heterogeneous changes. Regarding edge information enhancement, multiscale spatial feature extraction and edge grouping attention mechanisms are used to effectively highlight detailed information about ecological boundary areas in remote sensing images. Finally, the multiscale semantic fusion features are fused with edge features, and the corresponding mask matrix is output by the remote sensing image segmentation module. This significantly improves the accuracy and boundary refinement of land ecology remote sensing image segmentation in areas where farmland has been converted to forest. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0015] Figure 1 is a flow chart of the steps of the method of the present invention; Figure 2This is a comparison diagram of land remote sensing image segmentation results using the method of the present invention and other models. DETAILED DESCRIPTION
[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments derived by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0017] Example 1 In this embodiment, Figure 1 As shown, the present invention provides a land ecological remote sensing image segmentation method for monitoring the area of returning farmland to forest, and the specific steps include: S1. Obtaining remote sensing images of land ecology in the area of returning farmland to forest; S2. Constructing a land ecological remote sensing image segmentation model, the model includes a remote sensing image feature enhancement module, a remote sensing image multi-scale semantic fusion module, a local information capture module, an edge information enhancement module, and a remote sensing image segmentation module; S3. Inputting the land ecological remote sensing image into the remote sensing image feature enhancement module to perform feature enhancement operation to obtain enhanced remote sensing image features; Specifically, land ecological remote sensing images Input to the remote sensing image feature enhancement module, the image First, the convolution kernel size of the remote sensing image feature enhancement module is In the convolution layer, the initial features are obtained after initial feature extraction ; Initial features After two parallel convolution kernels of size and Depth-wise separable convolution, resulting in two different depth-wise separable features and , two depth-separable features and Splicing to obtain depth-separable splicing features ,in, Indicates the height of the remote sensing image, Indicates the width of the remote sensing image, Represents the number of channels of the remote sensing image. The formula is as follows: , , in, Indicates that the convolution kernel size is The convolutional layer, Represents a splicing operation, and They represent the convolution kernel size respectively. and Depthwise separable convolution; depthwise separable concatenation features The input to the convolution kernel size is The convolution layer expands its channel number to obtain depth-separable extended features ,use Activation function processing, and finally the convolution kernel size is The convolution layer restores the number of channels and obtains the enhanced remote sensing image features , the formula is as follows: , in, Represents the activation function operation.
[0018] S4. Input the enhanced remote sensing image features into the semantic spatial attention module in the remote sensing image multi-scale semantic fusion module to capture image spatial features, thereby obtaining semantic spatial attention module features. Input the semantic spatial attention module features into the multi-scale texture fusion module in the remote sensing image multi-scale semantic fusion module to obtain multi-scale semantic fusion features. Specifically, the enhanced remote sensing image features Input into the semantic space attention module, after the global average pooling layer and the convolution kernel size is The convolutional layer and Activation function to obtain the first intermediate feature ; First intermediate feature The input to the convolution kernel size is The convolutional layer, Activation function to obtain the second intermediate feature ; Second intermediate feature and enhanced remote sensing image features Perform channel weighted fusion to obtain semantic space attention module features , the formula is as follows: , , in, represents the global average pooling operation, represents the activation function, represents the channel weighted fusion operation, express Activation function; The semantic space attention module features Input into the multi-scale texture fusion module, after the convolution kernel size is The convolution layer obtains the first convolution feature ; The first convolution feature Input into three parallel separable convolutional sub-networks of different sizes. First, the first convolutional feature is input into the first sub-network to obtain the first sub-network feature. , the first sub-network includes a convolution kernel size of The depth-wise separable convolution and the convolution kernel size are Convolutional layer; the first convolution feature Input to the second sub-network to obtain the second sub-network features , the second sub-network includes a convolution kernel size of The depth-wise separable convolution and the convolution kernel size are Convolutional layer; Finally, the first convolution feature Input to the third sub-network to obtain the third sub-network features , the third sub-network includes a convolution kernel size of The depth-wise separable convolution and the convolution kernel size are The convolution layer of the first sub-network features , the second sub-network characteristics and the third subnetwork characteristics Splicing and fusion to obtain multi-scale fusion features , the formula is as follows: , , , , , in, 、 and They represent the convolution kernel size respectively. 、 and Depthwise separable convolution, Indicates that the convolution kernel size is The convolutional layer, Represents feature splicing operation; multi-scale fusion features The input to the convolution kernel size is In the convolution layer, the second convolution feature is obtained ; The second convolution feature Features of semantic space attention module Perform channel weighted fusion operation and pass through the adaptive pooling layer to obtain multi-scale semantic fusion features , the formula is as follows: , , in, Represents an adaptive pooling operation.
[0019] S5. Input the enhanced remote sensing image features into the remote sensing image channel attention module of the local information capture module to obtain remote sensing image channel attention features. Input the remote sensing image channel attention features into the remote sensing image spatiotemporal attention module of the local information capture module to extract spatiotemporal features of the remote sensing image spatiotemporal attention module to obtain remote sensing image spatiotemporal attention features. Specifically, the enhanced remote sensing image features The remote sensing image channel attention module input to the local information capture module passes through the adaptive maximum pooling layer and the convolution kernel size is Point convolution, the number of channels is compressed to the original number of channels , output the first compressed feature ; First compression feature go through The activation function performs nonlinear operations on the features, and then uses the convolution kernel size of The point convolution restores the number of channels of the feature to the original number of channels and obtains the first restored feature , the formula is as follows: , , in, represents the adaptive maximum pooling layer, Indicates that the convolution kernel size is Point convolution; Enhanced remote sensing image features After the adaptive average pooling layer and the convolution kernel size is Point convolution compresses the number of channels to the original number of channels , output the second compressed feature ; Second compression feature go through The activation function performs nonlinear operations on the features, and then uses the convolution kernel size of The point convolution restores the number of channels of the feature to the original number of channels, and obtains the second restored feature , the formula is as follows: , in, Indicates that the convolution kernel size is Point convolution; The first reduction feature and the second reduction feature Splicing and fusion Activation function to obtain the channel attention features of remote sensing images , the formula is as follows: , in, express Activation function; Remote sensing image channel attention features The remote sensing image spatiotemporal attention module input to the local information capture module is processed by the maximum pooling layer and the average pooling layer respectively, and the two processed features are spliced to obtain the pooled feature , the pooling features The input to the convolution kernel size is , the step length is Point convolution, then through Activation function to obtain the spatiotemporal attention features of remote sensing images , the formula is as follows: , , in, and Represent the maximum pooling layer and the average pooling layer respectively, Indicates that the convolution kernel size is , the step length is Point convolution.
[0020] S6. Input the spatiotemporal attention features of the remote sensing image into the multi-scale spatial feature module of the edge information enhancement module to obtain a multi-scale spatial feature, and input the multi-scale spatial feature into the edge grouping attention module of the edge information enhancement module to obtain an edge grouping attention feature; Specifically, the spatiotemporal attention features of remote sensing images Input to the multi-scale spatial feature module of the edge information enhancement module, remote sensing image spatiotemporal attention features The convolution kernel size after the first multi-scale spatial feature sub-network is Depthwise separable convolution, batch normalization operations, and Activation function to obtain the first multi-scale spatial feature ; Spatial-temporal attention features of remote sensing images The convolution kernel size after the second multi-scale spatial feature sub-network is Depthwise separable convolution, batch normalization operations, and Activation function to obtain the second multi-scale spatial feature ; Spatial-temporal attention features of remote sensing images The convolution kernel size after the third multi-scale spatial feature sub-network is Depthwise separable convolution, batch normalization operations, and Activation function to obtain the third multi-scale spatial feature ; The first multi-scale spatial feature , the second multi-scale spatial feature And the third multi-scale spatial feature Perform feature splicing operations to obtain multi-scale spatial features , the formula is as follows: in, represents the batch normalization operation, and They represent the convolution kernel size respectively. as well as Depthwise separable convolutional layers; Multi-scale spatial features Input to the edge grouping attention module of the edge information enhancement module, the multi-scale spatial features After the convolution kernel size is The group convolution and batch normalization operations are performed to obtain the first group convolution feature ; Multi-scale spatial features After the convolution kernel size is The grouped convolution and batch normalization operations are performed to obtain the second grouped convolution feature ; Convolution features of the first group and the second group convolutional features Perform splicing and fusion to obtain the first group convolution fusion feature ; The first group convolution fusion feature go through The activation function and convolution kernel size are The convolutional layers, batch normalization operations, and Activation function to obtain edge group attention features , the formula is as follows: , , , , in, and They represent the convolution kernel size respectively. The group convolution and convolution kernel size are Grouped convolution.
[0021] S7. Fuse the multi-scale semantic fusion features and edge grouping attention features, obtain fusion features through splicing operations, input the fusion features into the remote sensing image segmentation module, output the final predicted corresponding mask matrix, and obtain the segmentation result.
[0022] Specifically, the multi-scale semantic fusion features and edge grouping attention features After the splicing operation, the fusion features are obtained ; The fusion features Input into the remote sensing image segmentation module, after the convolution kernel size is Point convolution, convolution kernel size is The point convolution is performed, and the features are restored to the resolution of the original input image through upsampling operation, and the output image features are , the formula is as follows: , , in, Represents an upsampling operation; The image features go through Operation determines the category to which each pixel belongs, and finally Color mapping operation to generate the final mask matrix , the formula is as follows: , in, express operate, Represents a color mapping operation.
[0023] Specifically, during the model training process, the cross-entropy loss function is used as the optimization objective function of the model. The Adam optimizer is used during model training, and the initial learning rate is set to 1e-4.
[0024] Example 2 In this embodiment, in order to verify the effectiveness of the land ecological remote sensing image segmentation method for monitoring the returning farmland to forest area proposed in the present invention in the field of land remote sensing image recognition and segmentation, the proposed land remote sensing image segmentation method and the existing remote sensing image segmentation method are experimentally verified under the same experimental configuration conditions, and the obtained remote sensing image segmentation results are subjected to detailed demonstration and analysis, thereby verifying the effectiveness of the method proposed in the present invention in the field of land remote sensing image segmentation.
[0025] In a comparative experiment between the proposed method and other existing remote sensing image segmentation methods, three currently mainstream remote sensing image segmentation methods were selected as comparison models, and the obtained experimental results were analyzed in detail to verify the effectiveness of the proposed method in the field of land remote sensing image segmentation. The three comparison models used in the experiment are: the MANet model is based on the U-Net architecture and can effectively capture the correlation of segmentation targets in complex scenes in remote sensing images, but the model has weak recognition ability for local image details; the DC-Swin model uses the Swin Transformer as the backbone network and is suitable for large-scale scene segmentation tasks, but the computational complexity of this structure is high and the model requires a large amount of training data support; the CTMFNet model combines the CNN and Transformer architectures and performs well in terms of comprehensive performance, but the model has a large number of model parameters, and the model training and inference process require high computing resources.
[0026] In a comparative experiment on a proposed land remote sensing image segmentation method for monitoring areas where farmland has been converted to forest, three experimental metrics were used to evaluate the effectiveness of each model: the F1-score comprehensively evaluates the model's ability to segment and identify land in areas where farmland has been converted to forest by balancing precision and recall; the mean intersection over union (mIoU) measures the model's accuracy in locating target boundaries by calculating the ratio of the intersection and union of the predicted segmentation result to the true label; and the overall accuracy (OA) represents the proportion of correctly classified pixels to the total pixels, providing a quick assessment of the global segmentation effect.
[0027] The DeepGlobe dataset used in the experiment uses satellite remote sensing technology to collect high-resolution land remote sensing images. This dataset covers the geographical environment of multiple regions and includes seven land cover types (forest, farmland, water bodies, urban areas, bare land, grassland, and shrubs). It also includes images from different seasons and years to reflect dynamic changes in the land surface, such as those for returning farmland to forest and urban expansion. The samples in each category in this dataset exhibit a long-tail distribution. In the experiment, the land remote sensing image dataset was divided into training, validation, and test sets in a ratio of 7:1:2 to verify the model's generalization ability and segmentation accuracy in complex scenarios.
[0028] Table 1 shows the experimental results of the proposed land remote sensing image segmentation method and the comparison models on the DeepGlobe dataset. The proposed method achieves 89.40%, 84.25%, and 91.37% in the three experimental metrics of F1 score, mean intersection over union (mIoU), and overall accuracy (OA), respectively, significantly outperforming the other three comparison models. Compared to the second-best model CTMFNet (F1-score 87.33%, mIoU 81.64%, and OA 88.52%), the proposed method improves F1-score by 2.07%, mIoU by 2.61%, and OA by 2.85%, validating the proposed method's advantages in small category recognition and complex boundary segmentation.
[0029] Table 1 Comparison results of the proposed method in the dataset The actual segmentation effect of the proposed method and other comparison models in the experiment is as follows: Figure 2 As shown in the figure, the red part represents the color of the separated houses, and the grayish-white part represents the identification color of the forest. Figure 2 It can be seen that the segmentation and recognition effect of the proposed method is more consistent with the forest range in the input image. The other three models are far less accurate than the proposed method in segmenting and recognizing complex forest scenes in the image. This shows that the proposed method effectively solves the problems of low image segmentation accuracy and weak generalization ability caused by complex scenes and blurred boundaries in land remote sensing images, verifying the effectiveness of the method proposed in this invention.
[0030] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A land ecological remote sensing image segmentation method for monitoring the area of returning farmland to forest, characterized in that: The following steps are involved: S1. Obtaining remote sensing images of land ecology in the area of returning farmland to forest; S2. Constructing a land ecological remote sensing image segmentation model, the model includes a remote sensing image feature enhancement module, a remote sensing image multi-scale semantic fusion module, a local information capture module, an edge information enhancement module, and a remote sensing image segmentation module; S3. Inputting the land ecological remote sensing image into the remote sensing image feature enhancement module to perform feature enhancement operation to obtain enhanced remote sensing image features; S4. Input the enhanced remote sensing image features into the semantic spatial attention module in the remote sensing image multi-scale semantic fusion module to capture image spatial features, thereby obtaining semantic spatial attention module features. Input the semantic spatial attention module features into the multi-scale texture fusion module in the remote sensing image multi-scale semantic fusion module to obtain multi-scale semantic fusion features. S5. Input the enhanced remote sensing image features into the remote sensing image channel attention module of the local information capture module to obtain remote sensing image channel attention features. Input the remote sensing image channel attention features into the remote sensing image spatiotemporal attention module of the local information capture module to extract spatiotemporal features of the remote sensing image spatiotemporal attention module to obtain remote sensing image spatiotemporal attention features. S6. Input the spatiotemporal attention features of the remote sensing image into the multi-scale spatial feature module of the edge information enhancement module to obtain a multi-scale spatial feature, and input the multi-scale spatial feature into the edge grouping attention module of the edge information enhancement module to obtain an edge grouping attention feature; S7. Fuse the multi-scale semantic fusion features and edge grouping attention features, obtain fusion features through splicing operations, input the fusion features into the remote sensing image segmentation module, output the final predicted corresponding mask matrix, and obtain the segmentation result.
2. The land ecological remote sensing image segmentation method for monitoring the returning farmland to forest area according to claim 1 is characterized in that: Step S3 specifically includes: Land ecological remote sensing images Input to the remote sensing image feature enhancement module, the image First, the convolution kernel size of the remote sensing image feature enhancement module is In the convolution layer, the initial features are obtained after initial feature extraction ; Initial features After two parallel convolution kernels of size and Depth-wise separable convolution, resulting in two different depth-wise separable features and , two depth-separable features and Splicing to obtain depth-separable splicing features ,in, Indicates the height of the remote sensing image, Indicates the width of the remote sensing image, Represents the number of channels of the remote sensing image. The formula is as follows: , , in, Indicates that the convolution kernel size is The convolutional layer, Represents a splicing operation, and They represent the convolution kernel size respectively. and Depthwise separable convolution; depthwise separable concatenation features The input to the convolution kernel size is The convolution layer expands its channel number to obtain depth-separable extended features ,use Activation function processing, and finally the convolution kernel size is The convolution layer restores the number of channels and obtains the enhanced remote sensing image features , the formula is as follows: , in, Represents the activation function operation.
3. The land ecological remote sensing image segmentation method for monitoring the returning farmland to forest area according to claim 2 is characterized in that: Step S4 specifically includes: The enhanced remote sensing image features Input into the semantic space attention module, after the global average pooling layer and the convolution kernel size is The convolutional layer and Activation function to obtain the first intermediate feature ; First intermediate feature The input to the convolution kernel size is The convolutional layer, Activation function to obtain the second intermediate feature ; Second intermediate feature and enhanced remote sensing image features Perform channel weighted fusion to obtain semantic space attention module features , the formula is as follows: , , in, represents the global average pooling operation, represents the activation function, represents the channel weighted fusion operation, express Activation function; The semantic space attention module features Input into the multi-scale texture fusion module, after the convolution kernel size is The convolution layer obtains the first convolution feature ; The first convolution feature Input into three parallel separable convolutional sub-networks of different sizes. First, the first convolutional feature is input into the first sub-network to obtain the first sub-network feature. , the first sub-network includes a convolution kernel size of The depth-wise separable convolution and the convolution kernel size are Convolutional layer; the first convolution feature Input to the second sub-network to obtain the second sub-network features , the second sub-network includes a convolution kernel size of The depth-wise separable convolution and the convolution kernel size are Convolutional layer; Finally, the first convolution feature Input to the third sub-network to obtain the third sub-network features , the third sub-network includes a convolution kernel size of The depth-wise separable convolution and the convolution kernel size are The convolution layer of the first sub-network features , the second sub-network characteristics and the third subnetwork characteristics Splicing and fusion to obtain multi-scale fusion features , the formula is as follows: , , , , , in, 、 and They represent the convolution kernel size respectively. 、 and Depthwise separable convolution, Indicates that the convolution kernel size is The convolutional layer, Represents feature splicing operation; multi-scale fusion features The input to the convolution kernel size is In the convolution layer, the second convolution feature is obtained ; The second convolution feature Features of semantic space attention module Perform channel weighted fusion operation and pass through the adaptive pooling layer to obtain multi-scale semantic fusion features , the formula is as follows: , , in, Represents an adaptive pooling operation.
4. The land ecological remote sensing image segmentation method for monitoring the returning farmland to forest area according to claim 3 is characterized in that: Step S5 specifically includes: The enhanced remote sensing image features The remote sensing image channel attention module input to the local information capture module passes through the adaptive maximum pooling layer and the convolution kernel size is Point convolution, the number of channels is compressed to the original number of channels , output the first compressed feature ; First compression feature go through The activation function performs nonlinear operations on the features, and then uses the convolution kernel size of The point convolution restores the number of channels of the feature to the original number of channels and obtains the first restored feature , the formula is as follows: , , in, represents the adaptive maximum pooling layer, Indicates that the convolution kernel size is Point convolution; Enhanced remote sensing image features After the adaptive average pooling layer and the convolution kernel size is Point convolution compresses the number of channels to the original number of channels , output the second compressed feature ; Second compression feature go through The activation function performs nonlinear operations on the features, and then uses the convolution kernel size of The point convolution restores the number of channels of the feature to the original number of channels, and obtains the second restored feature , the formula is as follows: , in, Indicates that the convolution kernel size is Point convolution; The first reduction feature and the second reduction feature Splicing and fusion Activation function to obtain the channel attention features of remote sensing images , the formula is as follows: , in, express Activation function; Remote sensing image channel attention features The remote sensing image spatiotemporal attention module input to the local information capture module is processed by the maximum pooling layer and the average pooling layer respectively, and the two processed features are spliced to obtain the pooled feature , the pooled features The input to the convolution kernel size is , the step length is Point convolution, then through Activation function to obtain the spatiotemporal attention features of remote sensing images , the formula is as follows: , , in, and Represent the maximum pooling layer and the average pooling layer respectively, Indicates that the convolution kernel size is , the step length is Point convolution.
5. The land ecological remote sensing image segmentation method for monitoring the returning farmland to forest area according to claim 4 is characterized in that: Step S6 specifically includes: The spatiotemporal attention features of remote sensing images Input to the multi-scale spatial feature module of the edge information enhancement module, remote sensing image spatiotemporal attention features The convolution kernel size after the first multi-scale spatial feature sub-network is Depthwise separable convolution, batch normalization operations, and Activation function to obtain the first multi-scale spatial feature ; Spatial-temporal attention features of remote sensing images The convolution kernel size after the second multi-scale spatial feature sub-network is Depthwise separable convolution, batch normalization operations, and Activation function to obtain the second multi-scale spatial feature ; Spatial-temporal attention features of remote sensing images The convolution kernel size after the third multi-scale spatial feature sub-network is Depthwise separable convolution, batch normalization operations, and Activation function to obtain the third multi-scale spatial feature ; The first multi-scale spatial feature , the second multi-scale spatial feature And the third multi-scale spatial feature Perform feature splicing operations to obtain multi-scale spatial features , the formula is as follows: in, represents the batch normalization operation, and They represent the convolution kernel size respectively. as well as Depthwise separable convolutional layers; Multi-scale spatial features Input to the edge grouping attention module of the edge information enhancement module, the multi-scale spatial features After the convolution kernel size is The group convolution and batch normalization operations are performed to obtain the first group convolution feature ; Multi-scale spatial features After the convolution kernel size is The grouped convolution and batch normalization operations are performed to obtain the second grouped convolution feature ; Convolution features of the first group and the second group convolutional features Perform splicing and fusion to obtain the first group convolution fusion feature ; The first group convolution fusion feature go through The activation function and convolution kernel size are The convolutional layers, batch normalization operations, and Activation function to obtain edge group attention features , the formula is as follows: , , , , in, and They represent the convolution kernel size respectively. The group convolution and convolution kernel size are Grouped convolution.
6. The land ecological remote sensing image segmentation method for monitoring the returning farmland to forest area according to claim 5 is characterized in that: Step S7 specifically includes: Multi-scale semantic fusion features and edge grouping attention features After the splicing operation, the fusion features are obtained ; The fusion features Input into the remote sensing image segmentation module, after the convolution kernel size is Point convolution, convolution kernel size is The point convolution is performed, and the features are restored to the resolution of the original input image through upsampling operation, and the output image features are , the formula is as follows: , , in, Represents an upsampling operation; The image features go through Operation determines the category to which each pixel belongs, and finally Color mapping operation to generate the final mask matrix , the formula is as follows: , in, express operate, Represents a color mapping operation.
7. The land ecological remote sensing image segmentation method for monitoring the area of returning farmland to forest according to claim 6 is characterized in that: During the model training process, the cross-entropy loss function is used as the optimization objective function of the model. The Adam optimizer is used for model training, and the initial learning rate is set to 1e-4.
Citation Information
Patent Citations
Semantic image segmentation method and system based on edge enhancement
CN111462126A
Landslide image segmentation method based on high and low semantic information fusion
CN118887406A
Semantic detail fusion and context enhancement remote sensing image segmentation method based on DeepLabv3 +
CN119672340A
Three-dimensional point-cloud semantic segmentation method based on multi-level boundary enhancement for unstructured environment
WO2024230038A1