Image segmentation method based on deep learning remote sensing image

Through multi-branch depth collaborative modeling and adaptive attention fusion, combined with residual inverse MLP structure, the multi-scale representation problem in remote sensing images is solved, the boundary clarity and small-object recognition capabilities of remote sensing image segmentation are improved, and the robustness and efficiency of remote sensing image segmentation model are achieved.

CN120580436AActive Publication Date: 2025-09-02JILIN AGRICULTURAL UNIV

Patent Information

Application Number
CN202510742144.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-02
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

When existing deep learning is combined with remote sensing images, it is difficult to effectively integrate multi-scale representations, resulting in blurred boundaries, inaccurate recognition of small targets, and difficult to balance calculation efficiency and segmentation accuracy when complex land objects are targeted.

Method used

Multi-branch deep collaborative modeling is adopted, combining multi-scale path segmentation, adaptive attention fusion and residual inverse MLP structure, and dynamic interactive modeling of local and global information is realized through convolutional neural network and SwiftFormer encoder, and region-aware gating mechanism and residual reconstruction learning are used for self-repair.

Benefits of technology

It significantly improves the boundary clarity and small object recognition capabilities of remote sensing image segmentation, dynamically balances the relationship between global and local modeling, enhances the robustness and segmentation accuracy of the model, and is suitable for complex remote sensing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580436A_ABST
    Figure CN120580436A_ABST
Patent Text Reader

Abstract

The invention provides an image segmentation method based on a deep learning remote sensing image, and relates to the field of image processing, and the method comprises the steps: S1, collecting a remote sensing image, carrying out the construction of a data set through the remote sensing image, and completing the data preprocessing; s2, constructing a remote sensing image semantic segmentation model; and S3, training, verifying and optimizing the remote sensing image semantic segmentation model constructed in the step S2 by adopting the data set in the step S1 to complete model construction. According to the method, segmentation precision and expression consistency are improved through multi-branch deep collaborative modeling, so that robustness of the model in boundary fuzzy, small target and label defect areas is improved; meanwhile, the method adopts multi-scale path segmentation, adaptive attention fusion and a residual error inverse MLP structure, more efficient multi-level feature representation is realized, and dynamic balance of global and local modeling relations is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to an image segmentation method based on deep learning remote sensing images. Background Art

[0002] Remote sensing imagery is a digital record of electromagnetic wave reflection or radiation from the Earth's surface, acquired through non-contact sensors (such as those carried by satellites, aircraft, and drones). Using pixels as the basic unit, it reflects the spectral characteristics of ground objects through combinations of different wavelengths. It is widely used in agriculture, environmental monitoring, urban planning, disaster management, and other fields. Currently, remote sensing image detection not only suffers from reduced accuracy due to illumination fluctuations (i.e., the same scene exhibits different characteristics under different lighting conditions), but also faces the challenges of complex regional changes and irregularities in their size and number. Deep learning, with its powerful feature extraction methods and nonlinear representation capabilities, has been widely applied to change detection tasks, demonstrating excellent performance. At present, there are some documents that apply deep learning to remote sensing images. For example, Chinese patent document CN119625553A discloses a vegetation cover estimation method and device based on remote sensing semantic segmentation. By adopting a lightweight design, the network significantly reduces the computational complexity while ensuring segmentation accuracy, making it suitable for equipment with limited resources and improving the flexibility of practical applications; secondly, the Block module combines depthwise separable convolution and residual connection, which not only reduces the number of model parameters and computational cost, but also enhances the feature extraction capability, and improves the accuracy and detail retention of the segmentation results; through the introduction of multi-path feature fusion and SiLU activation function, the network exhibits good nonlinear expression ability when processing multi-scale remote sensing images, effectively improving the estimation accuracy of complex landforms and vegetation cover; in addition, the application of layer normalization significantly enhances the stability of the model, ensuring that it can converge efficiently under different data distributions. However, remote sensing scenes naturally have objects with huge scale differences, making it difficult for single-scale models to effectively capture diverse target features; at the same time, complex surface cover structures and transitions often lead to blurred target boundaries, which in turn restricts segmentation performance and the discrimination of image boundary areas; in addition, the semantic relationships in remote sensing images often span a wide spatial range, requiring long-range dependencies to be ensured and contextual information to be avoided; in the above-mentioned applications of deep learning and remote sensing images, although it can improve recognition accuracy and reduce computational complexity to a certain extent, there are still problems such as the inability to efficiently and coherently fuse multi-scale representations, the inability to ensure boundary clarity in high-resolution images rich in noise and details, and the inability to balance computational efficiency and segmentation accuracy when processing large-scale data and complex-shaped targets. Summary of the Invention

[0003] In response to the problems existing in the above-mentioned prior art, the purpose of the present invention is to provide an image segmentation method based on deep learning remote sensing images. This method improves the robustness of the model in areas with blurred boundaries, small targets, and label defects through multi-branch deep collaborative modeling, improving segmentation accuracy and expression consistency; at the same time, this method adopts multi-scale path segmentation, adaptive attention fusion and residual inverse MLP structure to achieve more efficient multi-level feature representation, thereby dynamically balancing the global and local modeling relationship to solve the problems existing in the combination of deep learning and remote sensing images in the prior art.

[0004] The purpose of the present invention is achieved through the following technical solutions: An image segmentation method based on deep learning remote sensing images, comprising: Step S1: collect remote sensing images, use remote sensing images to construct data sets, and complete data preprocessing; Step S2: Construct a remote sensing image semantic segmentation model: First, construct a convolutional neural network (CNN) and a SwiftFormer encoder. Then, construct a three-branch channel and use a three-branch mutual guidance fusion mechanism to achieve dynamic interaction modeling of local and global information. After that, use a region-aware gating mechanism to achieve region-adaptive feature information flow regulation. Then, use a hierarchical feature fusion module to achieve multi-scale semantic information fusion. Finally, use residual reconstruction learning and self-repair mechanism to complete model self-identification, prediction errors, and self-repair. Step S3: Use the data set in step S1 to train, verify and optimize the remote sensing image semantic segmentation model constructed in step S2 to complete model construction.

[0005] Based on further optimization of the above scheme, the remote sensing images in step S1 use the LoveDA dataset, which contains large-scale, high-resolution remote sensing images, providing a total of 5,987 remote sensing images with a resolution of 1024×1024 and their corresponding pixel-level semantic labels. The dataset covers multiple regions in China, including urban, rural, and mixed areas, with significant urban-rural distribution differences, and can truly reflect the common inter-domain offset problem in actual remote sensing applications. The dataset is constructed using remote sensing images as follows: the dataset is cropped with a step size of 512 pixels to obtain images of 512×512 pixels each, and the images are preprocessed to ignore classes with pixel values ​​of 0. The remaining seven ground objects include background, building, road, water, barren soil, agriculture, and forest. The images are randomly rotated, translated, flipped, cropped, and other operations are performed to further increase the data volume to improve the generalization ability of the model. The dataset is then divided into a training set, a validation set, and a test set.

[0006] Based on the further optimization of the above solution, the "construction of convolutional neural network and SwiftFormer encoder" in step S2 is specifically as follows: The cascaded SwiftFormer encoder architecture uses an efficient attention mechanism to gradually extract multi-scale features at different network stages while maintaining the ability to model global context. This architecture enables the joint modeling of local detail features and long-range dependencies. For a given input feature , the SwiftFormer encoder first adopts a 3x3 depthwise separable convolution ( DWConv ) captures the spatial structure and is then passed through a 1x1 convolution ( Conv ) transforms the channel dimension, thereby compressing features while retaining local structural information, specifically:

[0007] At the same time, the SwiftFormer encoder's efficient additive attention mechanism (EAA) significantly reduces computational costs while maintaining the ability to capture global context information. The specific process is as follows: First, for the input features , obtain its query matrix through two linear projections Q and bond matrix K :

[0008] Where: , represents the learning parameters; nrepresents the sequence length, d represents the embedding dimension; Then, each query matrix is ​​calculated Q and learning vectors The scaled dot product between them generates a global attention score for each position :

[0009] Afterwards, the global attention score is used to aggregate the query matrix to obtain the global query vector q :

[0010] The global query vector is then fused with the key matrix through element-wise multiplication to encode the interaction between all spatial locations:

[0011] Where: Linear represents a linear transformation; Represents element-wise multiplication; Finally, the global context features are combined with the normalized query representation through a residual connection to generate the final output of the encoder:

[0012] Where: Norm () represents the normalization function.

[0013] Based on the further optimization of the above solution, the step S2 of "constructing a three-branch channel and using a three-branch mutual guidance fusion mechanism to realize dynamic interactive modeling of local and global information" is specifically as follows: First, a unified feature map is extracted from the backbone network of the convolutional neural network and the SwiftFormer encoder. F shared , and input it into three branches respectively: Detail branch ( Detail Branch ): , used to extract texture details; Boundary branch ( Boundary Branch ): , used to extract edge line information; Global branch ( Global Branch ): , used to extract context; Where: The feature mapping function representing the detail branch is composed of convolution, ReLU 、 BN etc. local extraction modules; Represents the boundary feature extraction module, including SobelConvolution, edge enhancement layer, etc.; express Transformer ,The global modeling module composed of attention mechanism, is biased towards semantic grouping and regional perception; Afterwards, the mutual guidance in the three branches is established through the guidance graph:

[0014] Where: 、 、 、 、 、 They represent the global branch guiding the detail branch graph, the boundary branch guiding the detail branch graph, the global branch guiding the boundary branch graph, the detail branch guiding the boundary branch graph, the detail branch guiding the global branch graph, and the boundary branch guiding the global branch graph respectively; Represents the Sigmoid activation function, whose range is [0,1]; Indicates channel dimension splicing; Finally, the outputs of the detail branch, boundary branch, and global branch are updated respectively through the guidance graph:

[0015] Where: 、 、 Represent the updated detail branch feature map, boundary branch feature map, and global branch feature map respectively; 、 、 、 、 、 They represent the learning parameters respectively.

[0016] Based on further optimization of the above solution, the step S2 of "using the region-aware gating mechanism to achieve region-adaptive feature information flow control" is specifically as follows: First, set up the regional structure feature extraction module. Each pixel in the region has a set of values, and the structural complexity S(x,y) To reflect whether the current pixel position is at an edge, corner, or area with drastic texture changes:

[0017]

[0018] Where: G x Indicates a point I(x,y) exist x Gradient in the axial direction; G y Indicates a pointI(x,y) exist y Gradient in the axial direction; By the semantic confidence of the current region C(x,y) To determine the accuracy of the prediction:

[0019] Where: Softmax() express Softmax function; F pred (x,y) Indicates that the model is at pixel point (x,y) The predicted feature vector at ; By local grayscale standard deviation T(x,y) Determine the texture intensity of the current area:

[0020] Where: I i Indicates (x,y) The pixel value of the small window centered at represents the pixel mean; N Indicates the number of pixels in the current area; The region-aware gating modules are deployed in detail branches, boundary branches, and global branches, and the structural complexity is adaptively selected. S(x,y) , semantic confidence C(x,y) and texture intensity T(x,y) , and dynamically adjust the information flow weight through the gating function, so as to complete the regional hierarchical difference modeling before feature fusion, improve the fusion effect and segmentation accuracy, specifically: Detail branch:

[0021]

[0022] Boundary branches:

[0023]

[0024] Global branch:

[0025]

[0026] Where: W d 、 W b 、 W gRepresent the weight parameters corresponding to the detail branch, boundary branch, and global branch respectively; B d 、 B b 、 B g Represent the bias items corresponding to the detail branch, boundary branch, and global branch respectively; R d (x,y) 、 R b (x,y) 、 R g (x,y) Respectively represent the corresponding structural complexity in the detail branch, boundary branch, and global branch S(x,y) , semantic confidence C(x,y) and texture intensity T(x,y) Combined into a regional structure information vector.

[0027] Based on the further optimization of the above solution, the step S2 of "using the hierarchical feature fusion module to realize multi-scale semantic information fusion" is specifically as follows: Introduce the hierarchical feature fusion module at the 1 / 8 and 1 / 16 resolution stages of the backbone network: Assume that the input features of the current backbone network, that is, the global branch features, are ,in, 、 Represent the edge semantic features of the boundary branch and the local structural features of the detail branch respectively, then:

[0028] Where: AvgPool represents average pooling; MaxPool represents maximum pooling; MLP represents a multilayer perceptron with shared weights; The channel attention weights are multiplied element-wise to regenerate the global features:

[0029] A spatial attention mechanism is used to highlight the structural salient areas of detail branches:

[0030]

[0031] Finally, the attention enhancement feature 、 Along the channel dimension and input features F i-1 Splice to form fused input features: .

[0032] Based on the further optimization of the above solution, the step S2 of "using residual reconstruction learning and self-repair mechanism to complete the self-identification of the model, predict errors, and perform self-repair" is specifically as follows: Set up two decoding paths, the main path SegHead , generate regular predictions, reconstruct paths ReconHead , output pseudo-supervisory prediction; after the output features of the three branches (ie, detail branch, boundary branch, global branch) are fused, the fused features F fused Send it to two decoders at the same time to obtain the corresponding prediction probabilities P seg (x,y) 、 P recon (x,y) :

[0033] Then, calculate the KL divergence between the two predicted probabilities and obtain the residual graph R KL (x,y) :

[0034] Afterwards, a local filter is used to smooth the residual image to avoid interference caused by isolated error points:

[0035] Then dynamically generate the mask:

[0036] Where: Indicates the 90th percentile of the smoothed residual of the entire image, that is, only the highest 10% of the residual area is retained for repair; ReLU represents a nonlinear activation function; Structural self-repair loss L repair , with the reconstruction path as the goal, the main path performs reverse repair in the high residual area:

[0037] Finally, build the optimization loss model:

[0038] Where: CE Cross-entropy loss function, used to improve pixel-level classification accuracy; IoU represents the intersection-over-union loss function, which is used to optimize the overlap between the predicted area and the true mask; Represents the weight coefficient of self-repair loss in the total loss; L repair It is used to guide the main prediction path to focus on unstable areas and correct prediction deviations.

[0039] Based on further optimization of the above scheme, step S3 is specifically as follows: training and validation are both performed on a single GPU, CUDNN Benchmark is enabled to improve performance, and 6 concurrent data loading threads are set; OHEM (Online Hard Example Mining) is enabled based on ImageNet initialization of pre-trained weights to improve the ability to learn difficult examples, where the OHEM threshold is set to 0.9 and the number of retained samples is 131,072; the input and benchmark sizes are 512×512, the number of GPU processing samples per batch is 6, the total number of training rounds is 300, the optimizer is SGD, the initial learning rate is 0.005, the weight decay is 0.0005, and the momentum is 0.9; data augmentation is enabled, including random flipping and multi-scale training.

[0040] The following are the effects of the technical solution of the present invention: This paper utilizes a hybrid feature framework composed of a convolutional neural network (CNN) and a SwiftFormer encoder, leveraging both the model's representational capabilities and structural robustness. This framework exhibits excellent engineering versatility and is adaptable to segmentation tasks involving diverse remote sensing data (urban, rural, and natural features). Furthermore, the paper designs a multi-scale feature extraction module based on the CNN and SwiftFormer encoder backbone. This module decouples semantic modeling from boundary perception through three parallel branches: boundary, global, and detail. Furthermore, an information mutualization mechanism is introduced within the detail branch, boundary branch, and global semantic branch structures. This dynamic information flow enables mutual guidance and feedback updates, breaking the static isolation between branches. This allows for more comprehensive fusion of branch features in the semantic space, enhancing the semantic consistency and expression diversity of feature fusion, and significantly improving boundary resolution and small object restoration capabilities. Furthermore, a region-aware gating mechanism, combining structural complexity, semantic confidence, and texture variation to construct regional vectors, enables local adaptive regulation of feature flow and enables adaptive feature retention strategies for different regions, thereby enhancing the model's responsiveness to varying segmentation difficulties. In addition, the present invention constructs a prediction stability residual map and dynamically generates a repair mask through a dual decoding path of the main prediction path and the pseudo reconstruction path, implements pseudo-supervision optimization on uncertain areas, and effectively improves the model's fault tolerance for complex or weakly supervised samples.

[0041] Unlike some enhancement methods that require complex auxiliary structures or multi-model reasoning, the self-repair module proposed in this invention is only activated during the training phase and can be completely removed during the reasoning phase. The model structure is simple and the deployment cost is low, making it suitable for large-scale and efficient applications in remote sensing scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 4 is a structural block diagram of the image segmentation method in an embodiment of the present invention.

[0043] Figure 2 4 is an overall flow chart of the image segmentation method in an embodiment of the present invention.

[0044] Figure 3 Flowchart of the convolutional neural network and SwiftFormer encoder in an embodiment of the present invention.

[0045] Figure 4 This is a flowchart of the hierarchical feature fusion module implementing multi-scale semantic information fusion in an embodiment of the present invention.

[0046] Figure 5 Flowchart of the residual reconstruction learning and self-repair mechanism in an embodiment of the present invention.

[0047] Figure 6 This is an improved flow chart of another embodiment of the present invention.

[0048] Figure 7 This is a comparison diagram of the image segmentation method of the present invention and the existing segmentation method. DETAILED DESCRIPTION

[0049] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0050] Example 1: An image segmentation method based on deep learning remote sensing images, the overall structure diagram is as follows Figure 1 Shown, including: Step S1: Collect remote sensing images, use them to construct a dataset, and complete data preprocessing. The remote sensing images use the LoveDA dataset, which contains large-scale, high-resolution remote sensing images. A total of 5,987 remote sensing images with a resolution of 1024×1024 are provided, along with their corresponding pixel-level semantic labels. This dataset covers multiple regions in China, encompassing urban, rural, and mixed areas, with significant urban-rural distribution differences, and can truly reflect the inter-domain offset problem commonly seen in actual remote sensing applications. The specific steps for constructing a dataset using remote sensing images are as follows: the dataset is cropped with a step size of 512 pixels to obtain images of 512×512 pixels each, and the images are preprocessed, ignoring the classes with pixel values ​​of 0. The remaining seven types of land features include background, building, road, water, barren soil, agriculture, and forest. The images are randomly rotated, translated, flipped, cropped, and other operations are performed to further increase the data volume to improve the generalization ability of the model; the dataset is then divided into training set, validation set, and test set (among which the training set, validation set, and test set can be randomly divided in a ratio of 7:2:1).

[0051] Step S2: Construct a remote sensing image semantic segmentation model (see Figure 2 As shown): First, build a convolutional neural network (CNN) and SwiftFormer encoder, specifically: The cascaded SwiftFormer encoder architecture uses an efficient attention mechanism to gradually extract multi-scale features at different network stages while maintaining the ability to model global context. This architecture enables the joint modeling of local detail features and long-range dependencies. For a given input feature , the SwiftFormer encoder first adopts a 3x3 depthwise separable convolution ( DWConv ) captures the spatial structure and is then passed through a 1x1 convolution ( Conv ) transforms the channel dimension, thereby compressing features while preserving local structural information (see Figure 3 (a)), specifically:

[0052] At the same time, the SwiftFormer encoder's efficient additive attention mechanism (EAA) significantly reduces computational costs while maintaining the ability to capture global context information. Figure 3 As shown in (b), the specific process is: First, for the input features , obtain its query matrix through two linear projections Q and bond matrix K :

[0053] Where: , represents the learning parameters; n represents the sequence length, d represents the embedding dimension; Then, each query matrix is ​​calculatedQ and learning vectors The scaled dot product between them generates a global attention score for each position :

[0054] Afterwards, the global attention score is used to aggregate the query matrix to obtain the global query vector q :

[0055] The global query vector is then fused with the key matrix through element-wise multiplication to encode the interaction between all spatial locations:

[0056] Where: Linear represents a linear transformation; Represents element-wise multiplication; Finally, the global context features are combined with the normalized query representation through a residual connection to generate the final output of the encoder:

[0057] Where: Norm () represents the normalization function.

[0058] Then, a three-branch channel is constructed and a three-branch mutual guidance fusion mechanism is used to realize the dynamic interaction modeling of local and global information. Specifically: First, a unified feature map is extracted from the backbone network of the convolutional neural network and the SwiftFormer encoder. F shared , and input it into three branches respectively: Detail branch ( Detail Branch ): , used to extract texture details; Boundary branch ( Boundary Branch ): , used to extract edge line information; Global branch ( Global Branch ): , used to extract context; Where: The feature mapping function representing the detail branch is composed of convolution, ReLU 、 BN etc. local extraction modules; Represents the boundary feature extraction module, including Sobe lConvolution, edge enhancement layer, etc.; express Transformer ,The global modeling module composed of attention mechanism, is biased towards semantic grouping and regional perception; Afterwards, the mutual guidance in the three branches is established through the guidance graph:

[0059] Where: 、 、 、 、 、 They represent the global branch guiding the detail branch graph, the boundary branch guiding the detail branch graph, the global branch guiding the boundary branch graph, the detail branch guiding the boundary branch graph, the detail branch guiding the global branch graph, and the boundary branch guiding the global branch graph respectively; Represents the Sigmoid activation function, whose range is [0,1]; Indicates channel dimension splicing; Finally, the outputs of the detail branch, boundary branch, and global branch are updated respectively through the guidance graph:

[0060] Where: 、 、 Represent the updated detail branch feature map, boundary branch feature map, and global branch feature map respectively; 、 、 、 、 、 They represent the learning parameters respectively.

[0061] Afterwards, a region-aware gating mechanism is used to achieve region-adaptive feature information flow control, specifically: First, set up the regional structure feature extraction module. Each pixel in the region has a set of values, and the structural complexity S(x,y) To reflect whether the current pixel position is at an edge, corner, or area with drastic texture changes:

[0062]

[0063] Where: G x Indicates a point I(x,y) exist x Gradient in the axial direction; G y Indicates a point I(x,y) exist y Gradient in the axial direction; By the semantic confidence of the current region C(x,y) To determine the accuracy of the prediction:

[0064] Where: Softmax() express Softmax function; F pred (x,y) Indicates that the model is at pixel point (x,y) The predicted feature vector at ; By local grayscale standard deviation T(x,y) Determine the texture intensity of the current area:

[0065] Where: I i Indicates (x,y) The pixel value of the small window centered at represents the pixel mean; N Indicates the number of pixels in the current area; The region-aware gating modules are deployed in detail branches, boundary branches, and global branches, and the structural complexity is adaptively selected. S(x,y) , semantic confidence C(x,y) and texture intensity T(x,y) , and dynamically adjust the information flow weight through the gating function, so as to complete the regional hierarchical difference modeling before feature fusion, improve the fusion effect and segmentation accuracy, specifically: Detail branch:

[0066]

[0067] Boundary branches:

[0068]

[0069] Global branch:

[0070]

[0071] Where: W d 、 W b 、 W g Represent the weight parameters corresponding to the detail branch, boundary branch, and global branch respectively; B d 、 B b 、 B gRepresent the bias items corresponding to the detail branch, boundary branch, and global branch respectively; R d (x,y) 、 R b (x,y) 、 R g (x,y) Respectively represent the corresponding structural complexity in the detail branch, boundary branch, and global branch S(x,y) , semantic confidence C(x,y) and texture intensity T(x,y) Combined into a regional structure information vector.

[0072] Then, a hierarchical feature fusion module is used to achieve multi-scale semantic information fusion, see Figure 4 As shown, specifically: Introduce the hierarchical feature fusion module at the 1 / 8 and 1 / 16 resolution stages of the backbone network: Assume that the input features of the current backbone network, that is, the global branch features, are ,in, 、 Represent the edge semantic features of the boundary branch and the local structural features of the detail branch respectively, then:

[0073] Where: AvgPool represents average pooling; MaxPool represents maximum pooling; MLP represents a multilayer perceptron with shared weights; The channel attention weights are multiplied element-wise to regenerate the global features:

[0074] A spatial attention mechanism is used to highlight the structural salient areas of detail branches:

[0075]

[0076] Finally, the attention enhancement feature 、 Along the channel dimension and input features F i-1 Splice to form fused input features: .

[0077] Finally, the residual reconstruction learning and self-repair mechanism is used to complete the self-identification of the model, predict errors, and perform self-repair. Figure 5 As shown, specifically: Set up two decoding paths, the main path SegHead , generate regular predictions, reconstruct paths ReconHead , output pseudo-supervisory prediction; after the output features of the three branches (ie, detail branch, boundary branch, global branch) are fused, the fused features F fused Send it to two decoders at the same time to obtain the corresponding prediction probabilities P seg (x,y) 、 P recon (x,y) :

[0078] Then, calculate the KL divergence between the two predicted probabilities and obtain the residual graph R KL (x,y) :

[0079] Afterwards, a local filter is used to smooth the residual image to avoid interference caused by isolated error points:

[0080] Then dynamically generate the mask:

[0081] Where: Indicates the 90th percentile of the smoothed residual of the entire image, that is, only the highest 10% of the residual area is retained for repair; ReLU represents a nonlinear activation function; Structural self-repair loss L repair , with the reconstruction path as the goal, the main path performs reverse repair in the high residual area:

[0082] Finally, build the optimization loss model:

[0083] Where: CE Cross-entropy loss function, used to improve pixel-level classification accuracy; IoU represents the intersection-over-union loss function, which is used to optimize the overlap between the predicted area and the true mask; Represents the weight coefficient of self-repair loss in the total loss; L repair It is used to guide the main prediction path to focus on unstable areas and correct prediction deviations.

[0084] Step S3: Use the dataset in step S1 to train, verify, and optimize the remote sensing image semantic segmentation model constructed in step S2 to complete model construction; wherein, training and verification are performed on a single GPU, CUDNN Benchmark is enabled to improve performance, and 6 concurrent data loading threads are set; pre-trained weights are initialized based on ImageNet and OHEM (Online Hard Example Mining) is enabled to improve the ability to learn difficult examples, wherein: the OHEM threshold is set to 0.9, and the number of retained samples is 131072; the input and benchmark sizes are 512×512, the number of GPU processing samples per batch is 6, the total number of training rounds is 300, the optimizer is SGD, the initial learning rate is 0.005, the weight decay is 0.0005, and the momentum is 0.9; data enhancement is enabled, including random flipping and multi-scale training (Multi-Scale).

[0085] Test size and benchmark size: 512×512, multi-scale and flipping test strategies are enabled to improve inference stability and accuracy, test batch size: 6.

[0086] Example 2: As a preferred embodiment of the present invention, based on Example 1, in order to further explore the up-down dependency relationship under different receptive fields, a multi-scale splitting and fusion mechanism is further introduced in the hierarchical feature fusion module in step S2, specifically: For the input feature map , which is divided into four channel groups, see Figure 6 (a) shows:

[0087] Each group performs convolution operations using convolution kernels of various sizes to capture different receptive field ranges, specifically:

[0088] Channel Group It is then concatenated with the responses from the four convolutional paths to form a multi-scale representation, specifically:

[0089] Finally, the enhanced sub-features are concatenated to generate the output of the fusion module: .

[0090] In order to evaluate the semantic segmentation method in this embodiment, the performance of the semantic segmentation method in this embodiment is compared with existing methods (including DDRNet, BiseNetV2 and HRNet). The comparison results are shown in the following table:

[0091] As can be clearly seen from the above table, the model in this embodiment achieved significant performance improvements in all categories, with an overall mIoU of 0.627, which is approximately 14.8% higher than the existing HRNet. At the same time, the model in this embodiment showed significant advantages in key heterogeneous surface cover types such as buildings, roads, and water bodies, with their intersection-over-union (IoU) scores reaching 0.647, 0.651, and 0.758, respectively. This shows that the model in this embodiment has obvious advantages in maintaining fine-grained details and global semantic consistency.

[0092] To more intuitively demonstrate the performance advantages of the method in this embodiment in semantic segmentation tasks, Figure 7 The method of the embodiment of the present invention and the prediction results of DDRNet and BiSeNetV2; Figure 7 As shown, DDRNet and BiSeNetV2 exhibit significant breakage and deformation in road areas, making it difficult to accurately restore their true shape (especially in complex structural areas such as curves and intersections), and boundaries are prone to blurring and merging. Furthermore, both methods suffer from varying degrees of outline distortion and misclassification around building edges, failing to effectively capture the detailed structural features of compact urban blocks. In contrast, the road predictions generated by the method in this embodiment are not only continuous but also effectively follow the curved layout, avoiding the local breakage and boundary offset issues seen in other methods. Furthermore, the predicted building areas exhibit clear and coherent contours, demonstrating the model's superior capabilities in multi-scale feature alignment and fine-grained boundary modeling.

Claims

1. An image segmentation method based on deep learning remote sensing images, characterized by: include: Step S1: collect remote sensing images, use remote sensing images to construct data sets, and complete data preprocessing; Step S2: Construct a remote sensing image semantic segmentation model: First, build a convolutional neural network and a SwiftFormer encoder; then, construct a three-branch channel and use a three-branch mutual guidance fusion mechanism to achieve dynamic interaction modeling of local and global information; then, use a region-aware gating mechanism to achieve region-adaptive feature information flow regulation; Then, a hierarchical feature fusion module is used to achieve multi-scale semantic information fusion. Finally, residual reconstruction learning and self-repair mechanisms are used to complete the model's self-identification, prediction errors, and self-repair. Step S3: Use the data set in step S1 to train, verify and optimize the remote sensing image semantic segmentation model constructed in step S2 to complete model construction.

2. The image segmentation method based on deep learning remote sensing images according to claim 1, characterized in that: In step S1, the remote sensing image adopts the LoveDA dataset, which contains large-scale, high-resolution remote sensing images, providing a total of 5987 remote sensing images with a resolution of 1024×1024 and their corresponding pixel-level semantic labels; the dataset is constructed using remote sensing images as follows: the dataset is cropped with a step size of 512 pixels to obtain images of 512×512 pixels each, and the images are preprocessed, ignoring the classes with pixel values ​​of 0. The remaining seven types of land features include background, buildings, roads, water bodies, barren soil, farmland and woodland, and the images are enhanced; then the dataset is divided into a training set, a validation set and a test set.

3. The image segmentation method based on deep learning remote sensing images according to claim 1 or 2, characterized in that: The "constructing convolutional neural network and SwiftFormer encoder" in step S2 is specifically as follows: Adopting a cascaded SwiftFormer encoder, this architecture gradually extracts multi-scale features at different network stages through an efficient attention mechanism while maintaining the ability to model global context; For a given input feature , the SwiftFormer encoder first uses 3x3 depth-wise separable convolution to capture the spatial structure, and then transforms the channel dimension through 1x1 convolution, thereby compressing features while retaining local structural information, specifically: At the same time, the SwiftFormer encoder's efficient additive attention mechanism significantly reduces computational costs while maintaining the ability to capture global context information. The specific process is as follows: First, for the input features , obtain its query matrix through two linear projections Q and bond matrix K : Where: , represents the learning parameters; n represents the sequence length, d represents the embedding dimension; Then, each query matrix is ​​calculated Q and learning vectors The scaled dot product between them generates a global attention score for each position : Afterwards, the global attention score is used to aggregate the query matrix to obtain the global query vector q : The global query vector is then fused with the key matrix through element-wise multiplication to encode the interaction between all spatial locations: Where: Linear represents a linear transformation; Represents element-wise multiplication; Finally, the global context features are combined with the normalized query representation through a residual connection to generate the final output of the encoder: Where: Norm () represents the normalization function.

4. The image segmentation method based on deep learning remote sensing images according to claim 3, characterized in that: The step S2 of "constructing a three-branch channel and using a three-branch mutual guidance fusion mechanism to achieve dynamic interactive modeling of local and global information" is specifically as follows: First, a unified feature map is extracted from the backbone network of the convolutional neural network and the SwiftFormer encoder. F shared , and input it into three branches respectively: Detail branch: , used to extract texture details; Boundary branches: , used to extract edge line information; Global branch: , used to extract context; Where: Feature mapping function representing detail branch; represents the boundary feature extraction module; express Transformer ,The global modeling module composed of attention mechanism, is biased towards semantic grouping and regional perception; Afterwards, the mutual guidance in the three branches is established through the guidance graph: Where: 、 、 、 、 、 They represent the global branch guiding the detail branch graph, the boundary branch guiding the detail branch graph, the global branch guiding the boundary branch graph, the detail branch guiding the boundary branch graph, the detail branch guiding the global branch graph, and the boundary branch guiding the global branch graph respectively; Represents the Sigmoid activation function, whose range is [0,1]; Indicates channel dimension splicing; Finally, the outputs of the detail branch, boundary branch, and global branch are updated respectively through the guidance graph: Where: 、 、 Represent the updated detail branch feature map, boundary branch feature map, and global branch feature map respectively; 、 、 、 、 、 They represent the learning parameters respectively.

5. The image segmentation method based on deep learning remote sensing images according to claim 4, characterized in that: In step S2, "using the region-aware gating mechanism to achieve region-adaptive feature information flow control" is specifically as follows: First, set up the regional structure feature extraction module. Each pixel in the region has a set of values, and the structural complexity S (x,y) To reflect whether the current pixel position is at an edge, corner, or area with drastic texture changes: Where: G x Indicates a point I(x,y) exist x Gradient in the axial direction; G y Indicates a point I(x,y) exist y Gradient in the axial direction; By the semantic confidence of the current region C(x,y) To determine the accuracy of the prediction: Where: Softmax() express Softmax function; F pred (x,y) Indicates that the model is at pixel point (x,y) The predicted feature vector at ; By local grayscale standard deviation T(x,y) Determine the texture intensity of the current area: Where: I i Indicates (x,y) The pixel value of the small window centered at represents the pixel mean; N Indicates the number of pixels in the current area; The region-aware gating modules are deployed in detail branches, boundary branches, and global branches, and the structural complexity is adaptively selected. S(x,y) , semantic confidence C(x,y) and texture intensity T(x,y) , and dynamically adjust the information flow weight through the gating function, so as to complete the regional hierarchical difference modeling before feature fusion, improve the fusion effect and segmentation accuracy, specifically: Detail branch: Boundary branches: Global branch: Where: W d 、 W b 、 W g Represent the weight parameters corresponding to the detail branch, boundary branch, and global branch respectively; B d 、 B b 、 B g Represent the bias items corresponding to the detail branch, boundary branch, and global branch respectively; R d (x,y) 、 R b (x,y) 、 R g (x,y) Respectively represent the corresponding structural complexity in the detail branch, boundary branch, and global branch S(x,y) , semantic confidence C(x,y) and texture intensity T (x,y) Combined into a regional structure information vector.

6. The image segmentation method based on deep learning remote sensing images according to claim 5, characterized in that: In step S2, "using a hierarchical feature fusion module to achieve multi-scale semantic information fusion" is specifically as follows: Introduce the hierarchical feature fusion module at the 1 / 8 and 1 / 16 resolution stages of the backbone network: Assume that the input features of the current backbone network, that is, the global branch features, are ,in, 、 Represent the edge semantic features of the boundary branch and the local structural features of the detail branch respectively, then: Where: AvgPool represents average pooling; MaxPool represents maximum pooling; MLP represents a multilayer perceptron with shared weights; The channel attention weights are multiplied element-wise to regenerate the global features: A spatial attention mechanism is used to highlight the structural salient areas of detail branches: Finally, the attention enhancement feature 、 Along the channel dimension and input features F i-1 Splice to form fused input features: 。 7. The image segmentation method based on deep learning remote sensing images according to claim 6, characterized in that: In step S2, "using residual reconstruction learning and self-repair mechanism to complete model self-identification, predict errors, and perform self-repair" is specifically as follows: Set up two decoding paths, the main path SegHead , generate regular predictions, reconstruct paths ReconHead , output pseudo-supervised predictions; After the three branches output features are fused, the fused features F fused Send it to two decoders at the same time to obtain the corresponding prediction probabilities P seg (x,y) 、 P recon (x,y) : Then, calculate the KL divergence between the two predicted probabilities and obtain the residual graph R KL (x,y) : Afterwards, a local filter is used to smooth the residual image to avoid interference caused by isolated error points: Then dynamically generate the mask: Where: Indicates the 90th percentile of the smoothed residual of the entire image, that is, only the highest 10% of the residual area is retained for repair; ReLU represents a nonlinear activation function; Structural self-repair loss L repair , with the reconstruction path as the goal, the main path performs reverse repair in the high residual area: Finally, build the optimization loss model: Where: CE Cross-entropy loss function, used to improve pixel-level classification accuracy; IoU represents the intersection-over-union loss function, which is used to optimize the overlap between the predicted area and the true mask; Represents the weight coefficient of self-repair loss in the total loss; L repair It is used to guide the main prediction path to focus on unstable areas and correct prediction deviations.

8. The image segmentation method based on deep learning remote sensing images according to claim 1, characterized in that: Specifically, step S3 is as follows: training and validation are performed on a single GPU, CUDNN Benchmark is enabled to improve performance, and 6 concurrent data loading threads are set; OHEM is enabled based on ImageNet initialization of pre-trained weights, wherein: the OHEM threshold is set to 0.9, and the number of retained samples is 131,072; the input and benchmark sizes are 512×512, the number of GPU processed samples per batch is 6, the total number of training rounds is 300, the optimizer is SGD, the initial learning rate is 0.005, the weight decay is 0.0005, and the momentum is 0.9; data augmentation is enabled, including random flipping and multi-scale training.

Citation Information

Patent Citations

  • Knee joint nuclear magnetic resonance image automatic segmentation method based on deep learning

    CN115953416A

  • Realization method of remote sensing image road extraction task based on Swift-SegEdgeNet

    CN118736434A

  • Method and system for constructing cross-scale large-kernel convolution corn leaf disease segmentation model based on attention coordination mechanism

    CN120014275A

Cited By

  • SPR image optimization processing method based on image segmentation and edge enhancement

    CN120876346A

  • Semantic association-based remote sensing fire burnout area segmentation method and system

    CN120912631A

  • A semantic correlation-based remote sensing fire burn area segmentation method and system

    CN120912631B

  • Historical building intelligent identification method and system based on multi-source spatio-temporal data

    CN121033679A

  • Shielding perception road intelligent extraction method

    CN121236618A