Remote sensing image segmentation method based on semantic space interaction and edge guidance

By using a semantic space interaction module and an edge-guided supervision mechanism, combined with a multi-level feature fusion framework, the problems of large-scale land cover classification confusion and rough boundaries in remote sensing image segmentation are solved, achieving high-precision remote sensing image segmentation and improving the robustness and boundary coherence of the model.

CN120953612AActive Publication Date: 2025-11-14耕宇牧星(北京)空间科技有限公司

Patent Information

Application Number
CN202511081373.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-11-14
Estimated Expiration
2045-08-04

AI Technical Summary

Technical Problem

Existing remote sensing image segmentation models struggle to effectively model the global semantic dependencies of large-scale features, leading to confusion in the classification of continuous areas such as farmland and forest. Furthermore, shallow spatial details are severely degraded during the deep feature abstraction process, resulting in the under-detection of targets such as small buildings and narrow roads. Boundary prediction is coarse, and there is a lack of explicit edge modeling and effective mechanisms for decoupling and enhancing the interaction between semantic information and spatial details.

Method used

By employing a semantic space interaction module and an edge-guided supervision mechanism, and through a multi-level feature fusion framework, we achieve collaborative optimization of deep semantic understanding and shallow details. Combined with a multi-scale fusion and loss collaborative optimization mechanism, we improve the model's robustness in segmenting multi-scale ground features, small targets, and fine-grained boundaries.

Benefits of technology

It significantly improves the segmentation accuracy and boundary coherence in complex remote sensing scenarios, solves the problems of boundary spikes, missed detection of small targets and discontinuous segmentation of fragmented features, and enhances the model's ability to represent multi-scale features and complex boundaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953612A_ABST
    Figure CN120953612A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image segmentation method based on semantic space interaction and edge guidance, which relates to the field of remote sensing image segmentation, and comprises the following steps: carrying out multi-level standard convolution block processing and semantic space interaction module processing on a preprocessed remote sensing image by utilizing a feature extraction network, extracting basic features, and carrying out semantic space information decoupling; and outputting multi-layer enhanced features, performing multi-level standard convolution block and up-sampling processing on the features through a feature fusion network, fusing the features through jump connection and combining a classification head to output a segmentation probability graph. According to the method, feature expression is enhanced through semantic space interaction decoupling, boundary recognition is enhanced through an edge supervision mechanism, and multi-scale fusion and joint loss optimization are carried out, so that the segmentation precision and boundary coherence of the remote sensing image in a complex scene are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing image segmentation, and more specifically to a remote sensing image segmentation method based on semantic space interaction and edge guidance. Background Technology

[0002] Remote sensing image segmentation is a core foundation for land use monitoring, disaster assessment, and smart city management, and its accuracy directly determines the reliability of geographic information extraction.

[0003] Current mainstream segmentation models, due to the locality limitation of convolutional operations, struggle to model the global semantic dependencies of large-scale features, leading to confusion in the classification of continuous areas such as farmland and woodland. Simultaneously, during the deep feature abstraction process, shallow spatial details are severely degraded, resulting in missed detections of small buildings, narrow roads, and other targets. Although attempts have been made to improve this through multi-scale fusion or attention mechanisms, semantic information and spatial details are still coupled in a simple superposition manner, lacking effective decoupling and interaction enhancement mechanisms, thus limiting their adaptability to complex scenes.

[0004] Meanwhile, existing methods typically rely on deep semantic features to directly generate segmentation boundaries, lacking explicit edge modeling. The disconnect between low-level feature-based edge auxiliary branches and high-level semantics leads to coarse predicted boundaries, making slender features such as roads and field ridges prone to breakage or jagged edges, and highlighting the problem of overlapping outlines in urban building complexes. Edge supervision mechanisms driven by high-level semantic features have not been fully explored, limiting the ability to express boundaries in a refined manner.

[0005] Furthermore, skip connections in the feature decoding stage often directly splice shallow and deep features, failing to effectively filter shallow noise and exacerbating the distortion of small target details; the loss function often uses a single cross-entropy, which is insufficient for elements with highly unbalanced pixel distribution, such as isolated vegetation and building edges, resulting in poor consistency in the model's segmentation of fragmented features.

[0006] Therefore, how to design a remote sensing image segmentation method based on semantic space interaction and edge guidance, establish an efficient decoupling interaction of semantic and spatial features and a multi-scale fusion and loss co-optimization mechanism, and improve the segmentation accuracy and boundary coherence in complex remote sensing scenarios are problems that urgently need to be solved by those skilled in the art. Summary of the Invention

[0007] In view of this, the present invention provides a remote sensing image segmentation method based on semantic space interaction and edge guidance. By constructing a semantic-spatial decoupling module, an edge guidance supervision mechanism, and a multi-level feature fusion framework, it achieves collaborative optimization of deep semantic understanding and shallow details, and ultimately improves the robustness of the model in segmenting multi-scale land features, small targets, and fine-grained boundaries, meeting the practical application needs of land use classification, building detection, and other applications.

[0008] To achieve the above objectives, the present invention adopts the following technical solution:

[0009] A remote sensing image segmentation method based on semantic space interaction and edge guidance includes the following steps:

[0010] S1. Using a feature extraction network, the preprocessed remote sensing image is processed through multi-level standard convolutional blocks and a semantic space interaction module to extract basic features and decouple semantic space information, outputting multi-layer enhanced features.

[0011] S2. Through a feature fusion network, the features are... Multi-level standard convolutional blocks and upsampling are performed, and features are fused through skip connections. It also combines the classification head to output a segmentation probability map.

[0012] Preferably, S1 includes:

[0013] S11. The preprocessed remote sensing image is sequentially processed through the first standard convolutional block and the first semantic space interaction module to generate features.

[0014] S12, Features Features are generated by sequentially processing the data through a second standard convolutional block and a second semantic space interaction module.

[0015] S12, Features Features are generated by sequentially processing the data through a third standard convolutional block and a third semantic space interaction module.

[0016] S12, Features Features are generated by sequentially processing the data through the fourth standard convolutional block and the fourth semantic space interaction module.

[0017] Preferably, in step S1, the semantic space interaction module processes the following:

[0018] The input feature F is segmented into semantic features along the channel dimension. and spatial features C represents the number of channels, H represents the height, W represents the width, and α represents the channel adjustment factor;

[0019] semantic feature F 1 Position embedding and 3×3 convolutional layer processing are performed to generate deep semantic features F. s For deep semantic features F s Global average pooling and activation function processing are performed to obtain the channel attention weights ω1.

[0020] For spatial features F 2 Perform pointwise convolutional layer processing to generate feature F c ; For feature F c A 1×1 convolutional layer, a normalization layer, and an activation function are applied to obtain the spatial attention weight ω2.

[0021] The attention weights ω1 and ω2 are respectively weighted onto feature F. c F s Obtain attention-enhancing feature F 1‘ F 2‘ And perform element-wise addition and fusion to obtain feature F * .

[0022] Preferably, step S1 further incorporates an edge guidance mechanism, including:

[0023] Features The image size is restored to the input image size through an upsampling operation, and features are generated. Upsample(·) represents the upsampling operation, and saze = (H, W) represents the size of the feature target;

[0024] Features Perform 1×1 convolution dimensionality reduction, and then activate with the Sigmoid function to generate a single-channel edge probability map. P edge ∈{0,1} H×W σ(·) represents the activation of the Sigmoid function;

[0025] Combined with the true edge label Y edge ∈{0,1} H×W Supervised learning is performed using binary cross-entropy loss.

[0026] Preferably, the binary cross-entropy loss is expressed as:

[0027]

[0028] Where i and j represent the row index and column index, respectively.

[0029] Preferably, S2 includes:

[0030] S21, Features The feature X5 is generated by sequentially passing it through the fifth standard convolutional block and upsampling.

[0031] S22, Regarding feature X5 and feature Feature X6 is generated by fusion through skip connections and sequentially through a sixth standard convolutional block and upsampling.

[0032] S23, Combine feature X6 with feature... The segmentation probability map is output by fusion through skip connections and sequentially processed through the seventh standard convolutional block and the segmentation head.

[0033] Preferably, step S2 further includes constructing a dice loss by combining the segmentation probability map P with the true label Y:

[0034]

[0035] Where i, j, and c represent row index, column index, and category index, respectively, and ∈ represents the smoothing factor.

[0036] Preferably, in step S2, the classification head includes a 1×1 convolutional layer and an activation function, and the Sigmoid activation function is used for binary segmentation, while the Softmax activation function is used for multi-class segmentation.

[0037] Preferably, the joint loss function L total Represented as:

[0038] L total =L dice +λ·L edge

[0039] Where λ is the balance coefficient.

[0040] Preferably, the standard convolutional block includes a convolutional layer, a normalized layer, and a SiLU activation function.

[0041] As can be seen from the above technical solution, compared with the prior art, the technical solution of the present invention has the following beneficial effects:

[0042] 1. By using the semantic space interaction module, the input features are segmented into semantic and spatial sub-features by channel, and channel attention and spatial attention mechanisms are applied respectively. This effectively solves the problem of insufficient coupling between semantic information and spatial details in traditional convolutional neural networks. Through dual attention weighting and feature fusion, the model's ability to express the features of multi-scale ground objects and irregularly structured targets in remote sensing images is significantly enhanced.

[0043] 2. An edge probability map generated by upsampling high-level features is introduced, and combined with real edge labels for binary cross-entropy loss supervision. This allows the model to explicitly focus on the boundary regions of ground objects during training, especially for fine edges, narrow roads, or small target outlines commonly found in remote sensing images. Through the synergistic optimization of the edge probability map and semantic segmentation, problems such as boundary spikes and target breaks in traditional methods are effectively alleviated, significantly improving the boundary coherence and pixel-level accuracy of the segmentation results.

[0044] 3. By combining the constructed multi-level skip connection fusion mechanism and joint loss function, the adaptability of the model to complex remote sensing scenes is optimized. Multi-scale fusion ensures that information of small targets and detailed structures is not lost, while Dice loss in the joint loss is sensitive to small targets and edge loss strengthens the boundary. The two are dynamically adjusted through the balance coefficient, which significantly improves the segmentation stability and generalization ability of the model on difficult samples such as vegetation patches and fragmented ground features. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0046] Figure 1 A flowchart of a remote sensing image segmentation method based on semantic space interaction and edge guidance provided in an embodiment of the present invention;

[0047] Figure 2 This is a schematic diagram of the feature extraction network structure provided in an embodiment of the present invention;

[0048] Figure 3 This is a schematic diagram of the feature fusion network structure provided in an embodiment of the present invention. Detailed Implementation

[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] like Figure 1 As shown, this embodiment provides a remote sensing image segmentation method based on semantic space interaction and edge guidance, including the following steps:

[0051] S1. Using a feature extraction network, the preprocessed remote sensing image is processed through multi-level standard convolutional blocks and a semantic space interaction module to extract basic features and decouple semantic space information, outputting multi-layer enhanced features.

[0052] S2. Through a feature fusion network, the features are... Multi-level standard convolutional blocks and upsampling are performed, and features are fused through skip connections. It also combines the classification head to output a segmentation probability map.

[0053] This method achieves feature decoupling and dual attention fusion through a semantic space interaction module, enhancing the model's ability to express multi-scale features and complex boundaries in remote sensing images. Combined with an edge guidance mechanism, it accurately strengthens the ability to identify feature boundaries. Furthermore, it fuses deep semantics and shallow details through multi-scale jump connections and dynamically optimizes the segmentation results using a joint loss function. This effectively solves the problems of boundary spikes, missed detection of small targets, and discontinuous segmentation of fragmented features, thus improving the segmentation robustness in complex remote sensing scenarios.

[0054] This method is based on the feature extraction network and feature fusion network in the remote sensing image segmentation model. It is divided into two core stages: feature extraction and feature fusion. The two stages work together to form a complete processing link from low-level feature capture to high-level semantic parsing, ensuring that the segmentation result contains both accurate category information and retains clear spatial details.

[0055] The following provides a further detailed explanation of each step and related technical features in the above method;

[0056] In this embodiment, remote sensing images can be obtained through satellite or airborne sensors for Earth observation. Specifically, multispectral, hyperspectral, or synthetic aperture radar imaging systems are used to collect raw data in a specific band range with preset spatial and radiometric resolutions, and simultaneously record metadata such as imaging time, sensor attitude, and solar altitude angle to form a raw digital quantization matrix. The data dimension depends on the number of channels. The raw data is stored in GeoTIFF or HDF5 format and includes geographic coordinate information.

[0057] Furthermore, radiometric correction, geometric correction, and feature enhancement are sequentially performed on the acquired remote sensing images: Radiometric correction converts DN values ​​into apparent reflectance or backscattering coefficients using sensor calibration coefficients, and combines this with an atmospheric correction model to eliminate the effects of aerosol scattering; Geometric correction, based on ground control points or sensor orbit parameters, eliminates terrain displacement and projection distortion through polynomial transformation or RPC models to achieve pixel-level georegistration; Feature enhancement employs adaptive filtering and spectral normalization, ultimately outputting a preprocessed image with uniform size, consistent radiometrics, and geometric accuracy, providing stable input for subsequent segmentation.

[0058] like Figure 2 As shown, a feature extraction network is used to process the preprocessed remote sensing image through multi-level standard convolutional blocks and a semantic space interaction module, extracting basic features and decoupling semantic space information to output multi-layer enhanced features. Specifically, it includes:

[0059] S11. The preprocessed remote sensing image is sequentially processed through the first standard convolutional block and the first semantic space interaction module to generate features.

[0060] S12, Features Features are generated by sequentially processing the data through a second standard convolutional block and a second semantic space interaction module.

[0061] S12, Features Features are generated by sequentially processing the data through a third standard convolutional block and a third semantic space interaction module.

[0062] S12, Features Features are generated by sequentially processing the data through the fourth standard convolutional block and the fourth semantic space interaction module.

[0063] Furthermore, the semantic space interaction module processes the following:

[0064] The input feature F is segmented into semantic features along the channel dimension. and spatial features C represents the number of channels, H represents the height, W represents the width, and α represents the channel adjustment factor, which ranges from 0.6 to 0.8 and is used to control the ratio of semantic channels to spatial channels.

[0065] semantic feature F 1 Position embedding and 3×3 convolutional layer processing are performed to generate deep semantic features F. s For deep semantic features F s Global average pooling and activation function processing are performed to obtain channel attention weights ω1; among them, the position embedding is implemented using sinusoidal position encoding or a learnable position encoding matrix to enhance the spatial structure perception capability of semantic features.

[0066] For spatial features F 2 Perform pointwise convolutional layer processing to generate feature F c ; For feature F c A 1×1 convolutional layer, a normalization layer, and an activation function are applied to obtain the spatial attention weight ω2.

[0067] The attention weights ω1 and ω2 are respectively weighted onto feature F. c F s Obtain attention-enhancing feature F 1‘ F 2‘ And perform element-wise addition and fusion to obtain feature F * ;

[0068] Introducing this module during the feature extraction stage significantly improves the model's adaptability to multi-scale features through channel decoupling and dual attention mechanisms. Semantic paths enhance the class discrimination of large-scale features, while spatial paths preserve the geometric details of small targets. Dual attention weighted fusion effectively suppresses background interference, solving the bottlenecks of semantic confusion and detail loss in traditional methods.

[0069] Furthermore, it also incorporates edge-guided mechanisms, including:

[0070] Features The image size is restored to the input image size through an upsampling operation, and features are generated. Hpsample(·) represents the upsampling operation, and saze = (H, W) represents the size of the feature target;

[0071] The upsampling operation here is implemented using bilinear interpolation or transposed convolution. Bilinear interpolation performs linear weighted calculations based on local neighborhood pixels of the input feature map, restoring resolution with geometric rules and without parameters, ensuring spatial consistency of edge prediction. Transposed convolution, on the other hand, inserts zero values ​​between input features using a learnable convolution kernel before performing standard convolution operations, controlling the output size with stride and padding, and adaptively reconstructing semantic features through gradient optimization. Both methods use interpolation factors to scale or dynamically crop and pad to strictly align with the target size. Bilinear interpolation is suitable for edge generation that requires structural fidelity, while transposed convolution is suitable for feature decoding that requires semantic learning.

[0072] Features Perform 1×1 convolution dimensionality reduction, and then activate with the Sigmoid function to generate a single-channel edge probability map. P edge ∈{0,1} H×W σ(·) represents the activation of the Sigmoid function;

[0073] Combined with the true edge label Y edge ∈{0,1} H×W Supervised learning is performed using binary cross-entropy loss;

[0074] In the process of generating real edge labels, edge points are marked first by comparing the differences in neighborhood categories pixel by pixel based on real segmentation labels, and skeletonization is used to ensure the width of a single pixel; or the original image is processed by Canny and Sobel operators to generate initial edges, and then morphological closing operations are used to repair breaks and opening operations are used to eliminate noise. Finally, hole filling and size alignment are performed uniformly to output a strictly binarized matrix. All operations are constrained to have a single morphological iteration to maintain the authenticity of edge topology and meet the remote sensing annotation specifications.

[0075] The edge guidance mechanism here generates an edge probability map based on high-level features, making full use of the boundary abstraction capabilities of deep semantics, avoiding noise interference introduced by low-level features, and combining the binary cross-entropy supervision of real edge labels to force the model to focus on complex boundary areas such as road intersections and field ridge fracture zones, significantly improving the topological coherence of the segmentation results and effectively overcoming the problems of burrs and fractures.

[0076] Furthermore, the binary cross-entropy loss is expressed as:

[0077]

[0078] Where i and j represent the row index and column index, respectively.

[0079] like Figure 3 As shown, through a feature fusion network, features are processed... Multi-level standard convolutional blocks and upsampling are performed, and features are fused through skip connections. It also combines the classification head to output a segmentation probability map; specifically including:

[0080] S21, Features The feature X5 is generated by sequentially passing it through the fifth standard convolutional block and upsampling.

[0081] S22, Regarding feature X5 and feature Feature X6 is generated by fusion through skip connections and sequentially through a sixth standard convolutional block and upsampling.

[0082] S23, Combine feature X6 with feature... The segmentation probability map is output by fusing the data through skip connections and processing it sequentially through the seventh standard convolutional block and the segmentation head.

[0083] The decoding stage integrates shallow high-resolution features and deep strong semantic features layer by layer to achieve complementary enhancement of details and semantics. Specifically, shallow features can provide fine structures such as building edges and river directions, while deep features can contribute global context such as land parcel categories and vegetation cover, which can synergistically improve the segmentation integrity of fragmented land features.

[0084] Furthermore, it also includes constructing the Dice loss by combining the segmentation probability map P with the true label Y:

[0085]

[0086] Where i, j, and c represent row index, column index, and category index, respectively, and ∈ represents a smoothing factor used to avoid numerical instability when the denominator is zero during the segmentation of small targets.

[0087] Furthermore, the classification head includes a 1×1 convolutional layer and an activation function. For binary segmentation, the Sigmoid activation function is used, while for multi-class segmentation, the Softmax activation function is used. First, the input features are adjusted in channel dimension using a 1×1 convolutional layer, mapping the features to a dimension matching the number of target classes. Then, the corresponding activation function is selected based on the segmentation task type. In binary segmentation, the Sigmoid activation function maps the output value to the [0,1] interval to represent the probability of belonging to the target class. In multi-class segmentation, the Softmax activation function normalizes the output value. Finally, the sum of the probabilities of each class at each pixel location is equal to 1, resulting in a segmentation probability map that can be directly used to determine the pixel's class affiliation.

[0088] Based on its specific application scenarios, in binary segmentation tasks such as building extraction and water body detection, the classification head inputs the fused features into a 1×1 convolutional layer for channel compression to generate a single-channel feature map; then, the Sigmoid activation function is used to calculate the independent probability value of the target category pixel by pixel, and outputs a binary probability map, such as the probability of building areas approaching 1 and the probability of background approaching 0, which directly supports thresholding to generate binary masks, meeting the needs of efficient segmentation of single targets in disaster assessment and urban planning;

[0089] For various scenarios such as land use classification and crop identification, the classification head uses a 1×1 convolutional layer to expand the number of feature channels to the total number of categories, and then performs normalization along the channel dimension through the Softmax function: the mutually exclusive probability distribution of all categories is output at each pixel position, such as the probability of farmland, forest land and road of a certain pixel is 1, to ensure the exclusivity of pixel-level classification results.

[0090] Furthermore, the joint loss function L total Represented as:

[0091] L total =L dice +λ·L edge

[0092] Wherein, λ is the balance coefficient, which is dynamically adjusted according to the training stage. The initial value is 1.0, and it is reduced to 0.5 in the later stage of training to enhance the semantic segmentation accuracy. Dice loss optimizes the internal consistency of the target, and edge loss enhances the boundary accuracy. The two are trained together through dynamic balance coefficient, so that the model maintains high robustness in small target-dense scenarios such as crop monitoring.

[0093] Furthermore, a standard convolutional block includes a convolutional layer, a normalized layer, and a SiLU activation function;

[0094] Its data processing includes: the standard convolutional block first aggregates spatial neighborhood information of the input feature map through convolutional layers, extracts local patterns using a sliding window mechanism, and generates a new feature representation with translation invariance, providing basic feature support for subsequent semantic interaction; the convolutional output is normalized by a normalization layer to adjust the data distribution and suppress covariate shift during training; then it is fed into the SiLU activation function, which balances the linear transfer of features and nonlinear expression through an adaptive gating mechanism, outputting a smooth and discriminative feature map, effectively improving the network's generalization ability in complex terrain modeling.

[0095] This embodiment presents a remote sensing image segmentation method based on semantic space interaction and edge guidance. Through a semantic space interaction module, it achieves feature decoupling and adaptive fusion, overcoming the semantic generalization bottleneck of traditional models in complex scenarios. Combined with an edge guidance mechanism to drive refined boundary learning, it significantly improves the topological coherence of linear features such as road networks and field ridges. Furthermore, by combining multi-scale skip connections and joint loss optimization, it enhances detail restoration and category discrimination capabilities. In scenarios such as refined agricultural classification, urban building cluster segmentation, and disaster damage assessment, it effectively solves problems such as missed detection of small targets, boundary burrs, and fragmented feature segmentation, providing reliable technical support for high-precision intelligent remote sensing interpretation.

[0096] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.

[0097] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A remote sensing image segmentation method based on semantic space interaction and edge guidance, characterized in that, Includes the following steps: S1. Using a feature extraction network, the preprocessed remote sensing image is processed through multi-level standard convolutional blocks and a semantic space interaction module to extract basic features and decouple semantic space information, outputting multi-layer enhanced features. S2. Through a feature fusion network, the features are... Multi-level standard convolutional blocks and upsampling are performed, and features are fused through skip connections. It also combines the classification head to output a segmentation probability map.

2. The remote sensing image segmentation method based on semantic space interaction and edge guidance according to claim 1, characterized in that, S1 includes: S11. The preprocessed remote sensing image is sequentially processed through the first standard convolutional block and the first semantic space interaction module to generate features. S12, Features Features are generated by sequentially processing the data through a second standard convolutional block and a second semantic space interaction module. S12, Features Features are generated by sequentially processing the data through a third standard convolutional block and a third semantic space interaction module. S12, Features Features are generated by sequentially processing the data through the fourth standard convolutional block and the fourth semantic space interaction module.

3. The remote sensing image segmentation method based on semantic space interaction and edge guidance according to claim 1, characterized in that, In step S1, the semantic space interaction module processes the following: The input feature F is segmented into semantic features along the channel dimension. and spatial features C represents the number of channels, H represents the height, W represents the width, and α represents the channel adjustment factor; semantic feature F 1 Position embedding and 3×3 convolutional layer processing are performed to generate deep semantic features F. s For deep semantic features F s Global average pooling and activation function processing are performed to obtain the channel attention weights ω1. For spatial features F 2 Perform pointwise convolutional layer processing to generate feature F c ; for feature F c A 1×1 convolutional layer, a normalization layer, and an activation function are applied to obtain the spatial attention weight ω2. The attention weights ω1 and ω2 are respectively weighted onto feature F. c F s Obtain attention-enhancing feature F 1‘ F 2‘ And perform element-wise addition and fusion to obtain feature F * .

4. The remote sensing image segmentation method based on semantic space interaction and edge guidance according to claim 1, characterized in that, The S1 also incorporates an edge-guided mechanism, including: Features The image size is restored to the input image size through an upsampling operation, and features are generated. Upsample(·) represents the upsampling operation, and saze = (H, W) represents the size of the feature target; Features Perform 1×1 convolution dimensionality reduction, and then activate with the Sigmoid function to generate a single-channel edge probability map. P edge ∈{0,1} H×W σ(·) represents the activation of the Sigmoid function; Combined with the true edge label Y edge ∈{0,1} H×W Supervised learning is performed using binary cross-entropy loss.

5. The remote sensing image segmentation method based on semantic space interaction and edge guidance according to claim 1, characterized in that, The binary cross-entropy loss is expressed as: Where i and j represent the row index and column index, respectively.

6. The remote sensing image segmentation method based on semantic space interaction and edge guidance according to claim 1, characterized in that, S2 includes: S21, Features The feature X5 is generated by sequentially passing it through the fifth standard convolutional block and upsampling. S22, Regarding feature X5 and feature Feature X6 is generated by fusion through skip connections and sequentially through a sixth standard convolutional block and upsampling. S23, Combine feature X6 with feature... The segmentation probability map is output by fusion through skip connections and sequentially processed through the seventh standard convolutional block and the segmentation head.

7. The remote sensing image segmentation method based on semantic space interaction and edge guidance according to claim 1, characterized in that, S2 further includes constructing the dice loss by combining the segmentation probability map P with the true label Y: Where i, j, and c represent row index, column index, and category index, respectively, and ∈ represents the smoothing factor.

8. The remote sensing image segmentation method based on semantic space interaction and edge guidance according to claim 1, characterized in that, In S2, the classification head includes a 1×1 convolutional layer and an activation function. The Sigmoid activation function is used for binary segmentation, and the Softmax activation function is used for multi-class segmentation.

9. The remote sensing image segmentation method based on semantic space interaction and edge guidance according to claim 1, characterized in that, Joint loss function L total Represented as: L total =L dice +λ·L edge Where λ is the balance coefficient.

10. The remote sensing image segmentation method based on semantic space interaction and edge guidance according to claim 1, characterized in that, The standard convolutional block includes a convolutional layer, a normalized layer, and a SiLU activation function.

Citation Information

Patent Citations

  • High-resolution remote sensing image land coverage classification method based on local detail enhancement and edge constraint

    CN113343789A

  • Multi-scale fusion remote sensing image semantic segmentation method and system

    CN115512103A

  • Cross-scale graph similarity guide aggregation system, method and application

    CN115880552A

  • Remote sensing image semantic segmentation method for guiding multi-information fusion based on boundary information

    CN116797792A

  • Edge auxiliary feature calibration method for real-time semantic segmentation

    CN119180959A

Cited By

  • Remote sensing image segmentation method based on multi-scale gating bottleneck convolution scanning

    CN121640277A

  • Land coverage classification method based on edge-guided cross-modal interactive fusion

    CN121884012A