Remote sensing image segmentation method combined with non-additional parameter similarity attention
Through the attention mechanism without additional parameter similarity attention mechanism combined with hollow convolution and depth separation convolution for multi-scale feature extraction, and the diffusion model is used to restore details, solving the problems of boundary blurring and detail loss in complex scenes in remote sensing image segmentation, achieving efficient and robust segmentation effect.
Patent Information
- Application Number
- CN202510341872.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The existing remote sensing image segmentation method is difficult to effectively solve the problems of detail loss and boundary fracture in complex scenarios, especially in high noise and low contrast areas, and the traditional method has high computational complexity and poor real-time performance.
The attention mechanism without additional parameter similarity is adopted, combined with cavity convolution and depth separation convolution for multi-scale feature extraction, details are restored through diffusion model, and segmentation results are optimized using the composite loss function.
It significantly reduces the computational complexity, improves segmentation accuracy and robustness, enhances adaptability to complex scenes, solves the problems of boundary blur and details loss, and improves the clarity and consistency of segmentation results.
Smart Images

Figure CN120259661A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing image segmentation, and more specifically, to a remote sensing image segmentation method combining similarity attention without additional parameters. Background Art
[0002] Remote sensing image segmentation is widely used in the fields of land cover classification, change detection, environmental monitoring, urban planning, etc. in the field of remote sensing image processing. With the rapid development of remote sensing technology, the image resolution continues to improve, and the scene complexity increases significantly, posing higher requirements for the accuracy, detail recovery ability, and computational efficiency of segmentation algorithms.
[0003] However, traditional segmentation methods have obvious limitations when dealing with complex remote sensing scenes. High-complexity backgrounds, small targets, and fuzzy boundary features make it difficult for traditional algorithms to accurately distinguish land cover categories, especially in detail areas such as farmland edges and building contours, where segmentation breaks or misjudgments are likely to occur. Although convolutional neural networks based on deep learning have improved segmentation performance through multi-layer feature extraction, their ability to capture global context information is still insufficient, making it difficult to balance the expression of local details and overall structure.
[0004] Existing methods based on the attention mechanism can improve the segmentation effect by adaptive feature enhancement, but generally rely on additional trainable parameters to adjust the attention weights. This not only increases the number of model parameters and computational complexity but also easily causes overfitting problems due to parameter redundancy, limiting their application in remote sensing tasks with high real-time requirements. In addition, the existing attention mechanism lacks robustness in low-contrast regions and under noise interference, resulting in the loss of key boundary information.
[0005] On the other hand, existing methods still face challenges in detail recovery and noise suppression. The high noise, uneven illumination, and blurred texture problems commonly present in high-resolution remote sensing images make it difficult for traditional segmentation networks to effectively recover detail information, resulting in blurred or broken edges in the segmentation results. Although diffusion models have shown strong detail recovery capabilities in the field of image generation, their integrated application in remote sensing segmentation tasks is not yet mature, and existing methods lack systematic optimization strategies to balance detail accuracy and global consistency.
[0006] Therefore, how to design a remote sensing image segmentation method combining similarity attention without additional parameters to effectively solve the problems of detail loss and boundary breakage in complex scenarios and improve the adaptability and segmentation quality for low-contrast regions and high-noise remote sensing images is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0007] In view of this, the present invention provides a remote sensing image segmentation method combining similarity attention without additional parameters, which dynamically strengthens boundary and texture features by constructing an adaptive attention unit without additional parameters; combines dilated convolution, depthwise separable convolution and diffusion model to achieve efficient extraction and detail restoration of multi-scale features, significantly improving the segmentation accuracy and robustness in complex scenes while reducing the computational complexity.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] A remote sensing image segmentation method combining similarity attention without additional parameters, comprising the following steps:
[0010] S1. Use a feature extraction and fusion optimization network to extract multi-scale features of the remote sensing image and perform fusion optimization to generate initial features;
[0011] S2. Based on the initial features, perform sensitivity enhancement through a similarity attention unit without additional parameters in the feature enhancement network, and perform boundary and texture feature extraction through a boundary and texture feature extraction unit, and fuse them to generate enhanced features;
[0012] S3. Combine a segmentation network to optimize the enhanced features and perform class prediction, and output the segmentation result of the remote sensing image.
[0013] Further, the S1 includes:
[0014] S11. Use a Mamba pre-trained model to extract features from the remote sensing image to obtain high-level semantic features φ;
[0015] S12. Based on the high-level semantic features φ, perform multi-scale feature extraction through a multi-scale convolutional layer to obtain multi-level features φ multi ;
[0016] S13. Perform weighted summation or splicing fusion on the multi-level features φ multi to generate a comprehensive feature φ fused ;
[0017] S14. Perform dimensionality reduction optimization on the comprehensive feature φ fused through a convolutional layer and an activation function to obtain initial features φ 0 .
[0018] Further, in the S12, the multi-scale convolutional layer includes at least three different sizes of convolutional kernels, and the multi-level features φ multi are calculated by the following formula:
[0019] φ multi = [Conv k1 , Convk2 , Conv k2
[0020] Among them, Conv k1 , Conv k2 , Conv k2 respectively represent convolution operations with convolution kernel sizes of k1, k2, and k3, and [·] represents the concatenation operation.
[0021] Furthermore, in the S2, a sensitivity enhancement is performed by a similarity attention unit without additional parameters, including:
[0022] S21. Normalize the initial feature φ 0 to obtain a normalized feature
[0023] S22. Calculate the mean and variance corresponding to each neuron based on the normalized feature to evaluate the relative difference from other neurons; among them, the neuron represents the feature value corresponding to any position or channel;
[0024] S23. Quantify the neuron difference degree through an energy function to generate an energy matrix E;
[0025] S24. Calculate the importance weight of the neuron according to the energy matrix E and dynamically adjust the feature response through the Sigmoid function;
[0026] S25. Multiply the adjusted feature by the original feature to generate an attention-enhanced feature
[0027]
[0028] Among them, sigmoid(·) represents the sigmoid activation function.
[0029] Furthermore, in the S23, the energy function is expressed as:
[0030]
[0031] Among them, t represents the current neuron, represents the mean of the remaining neurons after removing the current neuron, represents the variance of the remaining neurons after removing the current neuron, and λ represents the regularization parameter.
[0032] Furthermore, in the S2, a boundary and texture feature extraction unit performs boundary and texture feature extraction, including:
[0033] Take the initial feature φ 0 The input dilated convolutional layer expands the receptive field, and through the LeakyReLU activation function, convolutional layer, and ReLU activation function, the boundary feature φ is generated. D ;
[0034] φ D = ReLU(Conv(LeakyReLU(DConv(φ 0 ))))
[0035] Among them, Conv(·) represents the convolutional layer operation, and DConv(·) represents the dilated convolutional layer operation;
[0036] And, the initial feature φ 0 is input into the depthwise separable convolutional layer and processed through the Mish activation function and fully connected layer to generate the texture feature φ DW ;
[0037] φ DW = FC(Mish(DWConv(φ 0 )))
[0038] Among them, FC(·) represents the fully connected layer operation, and DWConv(·) represents the depthwise separable convolutional layer operation.
[0039] Furthermore, it also includes:
[0040] The boundary feature φ D and the texture feature φ D are concatenated to obtain the comprehensive feature φ DDW ;
[0041] The comprehensive feature φ DDW is successively subjected to convolution, batch normalization, and ReLU activation processing, and fused with the boundary feature φ D and the texture feature φ DW to obtain the enhanced feature
[0042] Furthermore, in S2, it also includes:
[0043] The enhanced feature is processed through the ReLU activation function and fused with the attention enhanced feature after convolution and batch normalization processing to obtain the fused feature;
[0044] The fused feature is pixel rearranged to obtain the enhanced feature φ + ;
[0045]
[0046] Among them, PShuffle(·) represents the pixel rearrangement operation, and BN(·) represents the batch normalization operation.
[0047] Further, the S3 includes:
[0048] S31. Initially optimize the enhanced feature φ + through a convolutional layer, a batch normalization layer, and an activation function to improve the feature expression ability;
[0049] S32. Based on the initially optimized enhanced feature, perform multi-step denoising processing through a diffusion model, gradually recover the detailed information, and obtain the optimized enhanced feature;
[0050] S33. Combine the segmentation head to convert the optimized enhanced feature into a pixel-level classification result.
[0051] Further, the segmentation network adopts a composite loss function L total ;
[0052] L total = λ1L CE + λ2L Dice + λ3L Boundary
[0053] where λ1, λ2, and λ3 are weight coefficients, and L CE represents the cross-entropy loss function, L Dice represents the Dice loss function, and L Boundary represents the boundary loss function.
[0054] It can be seen from the above technical solutions that compared with the prior art, the technical solutions of the present invention have the following beneficial effects:
[0055] 1. By constructing a similarity attention unit without additional parameters, this solution dynamically evaluates the feature differences using normalization, mean square error calculation, and energy function, without introducing additional trainable parameters. Compared with the traditional attention mechanism, it significantly reduces the number of model parameters and computational complexity, while maintaining the ability to adaptively focus on the key regions of the image. In complex scenes of remote sensing images, it can not only efficiently enhance the boundary and detail information, but also avoid the overfitting risk caused by parameter redundancy, improving the real-time performance and robustness of the segmentation task.
[0056] 2. Further expand the receptive field through dilated convolutions to capture global boundary information, combine depthwise separable convolutions to extract local texture details, and fuse these two types of features to form a comprehensive feature representation. Dilated convolutions cover a larger context while maintaining the resolution, and depthwise separable convolutions reduce the computational cost through channel separation. The two complement each other to optimize the feature extraction ability for complex backgrounds and small targets. In addition, the features are fused with attention enhancement through pixel rearrangement operations, further improving the model's adaptability to low-contrast regions and high-noise scenes in remote sensing images, ensuring the detail integrity and boundary accuracy of the segmentation results.
[0057] 3. Introduce a diffusion model at the backend of the segmentation network, and gradually refine the feature map through a multi-step reverse denoising process to recover the detailed information lost due to noise or blur. This process combines multi-objective optimization of cross-entropy, Dice, and boundary loss, not only strengthening the classification accuracy of the foreground and background, but also enhancing the sharpness of the segmentation edges through boundary constraints. Especially when dealing with high-resolution remote sensing images, the diffusion model can effectively solve the problems of detail blurring and boundary breakage in traditional methods, significantly improving the segmentation clarity and overall consistency in complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0059] Figure 1 It is a flowchart of a remote sensing image segmentation method combining similarity attention without additional parameters provided by an embodiment of the present invention;
[0060] Figure 2 It is a schematic diagram of the process of generating enhanced features based on a feature enhancement network provided by an embodiment of the present invention;
[0061] Figure 3 It is a schematic diagram of the process of outputting a segmentation result based on a segmentation network provided by an embodiment of the present invention;
[0062] Figure 4 It is a schematic diagram of the process of segmenting urban remote sensing images combining similarity attention without additional parameters provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0064] Embodiment 1;
[0065] As Figure 1 shown, this embodiment provides a remote sensing image segmentation method combining similarity attention without additional parameters, including the following steps:
[0066] S1. Use a feature extraction and fusion optimization network to extract multi-scale features of the remote sensing image and perform fusion optimization to generate initial features;
[0067] S2. Based on the initial features, perform sensitivity enhancement through a similarity attention unit without additional parameters in the feature enhancement network, and extract boundary and texture features through a boundary and texture feature extraction unit, and fuse them to generate enhanced features;
[0068] S3. Combine with a segmentation network to optimize and perform class prediction on the enhanced features, and output the segmentation result of the remote sensing image.
[0069] This method significantly improves the model efficiency through an efficient attention mechanism without additional parameters. At the same time, it combines various technologies such as dilated convolution and depthwise separable convolution to enhance the feature expression ability, and introduces a diffusion model in the segmentation network to achieve detail restoration and optimization. It not only improves the accuracy and robustness of remote sensing image segmentation, but also greatly reduces the computational complexity, effectively solving the problems of blurred boundaries and lost details in complex scenes.
[0070] The following further elaborates on each of the above steps in detail;
[0071] In S1 of this embodiment, a feature extraction and fusion optimization network is used to extract multi-scale features of the remote sensing image and perform fusion optimization to generate initial features; specifically including:
[0072] S11. Use the Mamba pre-trained model to extract features from the remote sensing image to obtain high-level semantic features φ; the Mamba model is trained based on a large-scale remote sensing dataset, and its deep network structure is good at capturing complex patterns in the remote sensing image and can efficiently extract high-quality features from the remote sensing image, providing a solid foundation for subsequent feature enhancement and segmentation tasks;
[0073] S12. Based on the high-level semantic features φ, perform multi-scale feature extraction through a multi-scale convolutional layer to obtain multi-level features φ containing global informationmulti ;
[0074] The multi-scale convolutional layer includes at least three different sizes of convolutional kernels, and the multi-level feature φ multi is calculated by the following formula:
[0075] φ multi = [Conv k1 , Conv k2 , Conv k2
[0076] where Conv k1 , Conv k2 , Conv k2 represent the convolutional operations with convolutional kernel sizes of k1, k2, and k3 respectively, and [·] represents the concatenation operation.
[0077] Specifically, it can use 3×3, 5×5, and 7×7 convolutional kernels to extract local, mid-range, and global features respectively, and integrate them into multi-level features through the concatenation operation; the multi-scale design covers ground object targets at different scales in remote sensing images, avoids feature omission caused by a single-scale convolutional kernel, and enhances the generalization ability of the model to complex scenes;
[0078] S13. Perform weighted summation or concatenation fusion on the multi-level feature φ multi to generate a comprehensive feature φ fused ; it can assign adaptive weights to different-scale features to avoid redundancy;
[0079] S14. Perform dimensionality reduction optimization on the comprehensive feature φ fused through a convolutional layer and an activation function to reduce the subsequent computational complexity and obtain the initial feature φ 0 .
[0080] In this step, the application of the Mamba pre-trained model significantly improves the efficiency and accuracy of feature extraction. At the same time, the design of the multi-scale convolutional layer enables the model to capture multi-scale information in the image, enhancing the adaptability of the model to complex scenes. Further, the comprehensive feature generated by weighted summation or concatenation fusion retains both global information and rich detail information, providing comprehensive feature support for subsequent feature enhancement and segmentation tasks.
[0081] As Figure 2 shown, in this embodiment S2, based on the initial feature, sensitivity enhancement is performed through the parameter-free similarity attention unit in the feature enhancement network, and boundary and texture feature extraction is performed through the boundary and texture feature extraction unit, and the enhanced feature is generated by fusion;
[0082] Traditional attention mechanisms rely on trainable fully connected layers or convolutional layers to generate weights. In this embodiment, the feature differences are directly quantified through an energy function, avoiding parameter redundancy and reducing the risk of overfitting. The sensitivity enhancement is performed by a similarity attention unit without additional parameters, specifically including:
[0083] S21. Normalize the initial feature φ 0 to obtain the normalized feature This operation ensures that the mean and variance of subsequent calculations are not affected by the feature dimension, improving the stability of the attention weights.
[0084] S22. Calculate the mean and variance corresponding to each neuron based on the normalized feature to evaluate the relative difference from other neurons. Here, a neuron represents the feature value corresponding to any position or channel.
[0085] Specifically, for neuron t, calculate the mean and variance of the remaining neurons after removing neuron t. By excluding the current neuron t, calculating the mean and variance of the remaining features can reflect the relative relationship between this neuron and the overall features. If the feature value of a certain neuron is much higher than that of other neurons, the mean of the remaining features will decrease significantly and the variance will also decrease after removing this value, thus reflecting its particularity. By excluding itself, it can accurately evaluate the relative difference between the current neuron and the overall features, avoiding the interference of self-correlation.
[0086] S23. Quantify the neuron difference degree through the energy function to generate the energy matrix E. The energy function is expressed as:
[0087]
[0088] where t represents the current neuron, represents the mean of the remaining neurons after removing the current neuron, represents the variance of the remaining neurons after removing the current neuron, and λ represents the regularization parameter;
[0089] In the specific energy function , the coefficient 4 is used to adjust the overall range of the energy value so that it is distributed in a more reasonable interval; and is designed to normalize the energy value to ensure that the function is insensitive to scale. Combining is used to measure the deviation degree of the current neuron from the mean of the remaining neurons. The greater the deviation, the higher the energy value; and is designed to balance the local difference and the overall stability by introducing the variance of the remaining features and the regularization term. This function objectively evaluates the feature importance through statistical methods, avoiding subjective parameter settings. The regularization parameter λ enhances the numerical stability and is applicable to high-noise remote sensing scenarios.
[0090] S24. Calculate the importance weights of neurons according to the energy matrix E, and dynamically adjust the feature responses through the Sigmoid function;
[0091] S25. Multiply the adjusted features by the original features to generate attention-enhanced features
[0092]
[0093] Among them, sigmoid(·) represents the sigmoid activation function; the sigmoid activation function is used to smoothly adjust the importance of each neuron, so that the activation value of each neuron is dynamically adjusted according to its relative importance, improving the attention to detail and boundary information.
[0094] Furthermore, the boundary and texture feature extraction unit performs boundary and texture feature extraction, including:
[0095] Input the initial feature φ 0 into the dilated convolutional layer to expand the receptive field, and generate the boundary feature φ D through the LeakyReLU activation function, convolutional layer and ReLU activation function; specifically, it uses a 3×3 convolutional kernel with a dilation rate r = 2 to expand the receptive field to (2r + 1)×(2r + 1), capturing long-range boundary information without reducing the resolution;
[0096] φ D = ReLU(Conv(LeakyReLU(DConv(φ 0 ))))
[0097] Among them, Conv(·) represents the convolutional layer operation, and DConv(·) represents the dilated convolutional layer operation;
[0098] And, input the initial feature φ 0 into the depthwise separable convolutional layer, and process it through the Mish activation function and fully connected layer to generate the texture feature φ DW ; specifically, decompose the standard convolution into a channel-wise convolution with a 3×3 kernel and a point-wise convolution with a 1×1 kernel, reducing the number of parameters to 1 / N, reducing the computational amount while retaining texture details;
[0099] φ DW = FC(Missh(DWConv(φ 0 )))
[0100] Among them, FC(·) represents the fully connected layer operation, and DWConv(·) represents the depthwise separable convolutional layer operation.
[0101] The atrous convolution here focuses on the global boundary, and the depthwise separable convolution extracts local textures. The two complement each other in features through splicing and optimization, solving the problem that it is difficult for a single convolution kernel to balance long-range and short-range dependencies.
[0102] Furthermore, it also includes:
[0103] Splice the boundary feature φ D and the texture feature φ DW to obtain the comprehensive feature φ DDW ;
[0104] Perform convolution, batch normalization, and ReLU activation processing on the comprehensive feature φ DDW in sequence, and fuse it with the boundary feature φ D and the texture feature φ DW to obtain the enhanced feature The splicing operation integrates boundary and texture information, and the convolutional layer further optimizes the feature representation, enhancing the model's segmentation ability for complex ground objects.
[0105] Furthermore, the above steps also include:
[0106] Process the enhanced feature through the ReLU activation function, and fuse it with the attention-enhanced feature after convolution and batch normalization to obtain the fused feature;
[0107] Perform pixel rearrangement on the fused feature to obtain the enhanced feature φ + ;
[0108]
[0109] Among them, PShuffle(·) represents the pixel rearrangement operation, and BN(·) represents the batch normalization operation; the pixel rearrangement operation restores high-resolution details, and the feature addition fuses global attention and local details, significantly improving the clarity of the edges of the segmentation result.
[0110] In this step, the design of the similarity attention unit without additional parameters significantly reduces the computational complexity of the model and avoids the overfitting problem, while maintaining the ability to adaptively focus on key regions of the image. The combination of the boundary and texture feature extraction units enables the model to capture both boundary and texture information in the image, improving the richness and expressiveness of the features. Through the splicing, fusion, and optimization processing of the features, the robustness and segmentation accuracy of the features are further enhanced.
[0111] As Figure 3 shown, in step S3 of this embodiment, in combination with the segmentation network, the enhanced feature is optimized and class prediction is performed to output the segmentation result of the remote sensing image; specifically, it includes:
[0112] S31. Initially optimize the enhanced feature φ + through a convolutional layer, a batch normalization layer, and an activation function to improve the feature expression ability;
[0113] S32. Based on the initially optimized enhanced feature, perform multi-step denoising processing through a diffusion model to gradually restore the detailed information and obtain the optimized enhanced feature; specifically, in this embodiment, the diffusion model adopts a U-Net structure, which includes an encoder-decoder module. The encoder extracts multi-scale features through downsampling, and the decoder gradually restores the resolution in combination with skip connections; the diffusion model gradually refines the features through multi-step denoising. Compared with single-step post-processing, it can more accurately restore small targets and blurred boundaries in complex scenes, and will be deeply integrated with traditional segmentation networks to perform denoising in the feature space rather than the pixel space, which not only avoids the interference of the original image noise on segmentation but also significantly improves the detailed restoration ability through iterative optimization.
[0114] S33. Combine with a segmentation head to convert the optimized enhanced feature into a pixel-level classification result.
[0115] Further, the above segmentation network adopts a composite loss function L total ;
[0116] L total = λ1L CE + λ2L Dice + λ3L Boundary
[0117] where λ1, λ2, and λ3 are weight coefficients, L CE represents the cross-entropy loss function, L Dice represents the Dice loss function, and L Boundary represents the boundary loss function; specifically, the cross-entropy loss optimizes the classification accuracy, the Dice loss alleviates the class imbalance, and the boundary loss strengthens the edge sharpness. During the training process, the three cooperate to improve the comprehensive performance of the model in complex scenes and can more comprehensively consider various requirements of the segmentation task.
[0118] In this step, the introduction of the diffusion model enables the model to gradually denoise and restore the detailed information in the image, especially significantly improving the clarity and accuracy of the segmentation result when processing complex remote sensing images. At the same time, the design of the segmentation head enables the model to accurately convert the optimized enhanced feature into a pixel-level classification result, realizing the accurate segmentation of remote sensing images.
[0119] This embodiment provides a remote sensing image segmentation method combined with a similarity attention mechanism without additional parameters. Initial features are obtained through a feature extraction and fusion optimization network, and a similarity attention unit without additional parameters is used to enhance the sensitivity of key regions. At the same time, a boundary and texture feature extraction unit is combined to capture richer detail information. Further, a diffusion model is used for multi-step denoising processing to gradually restore and optimize the detail information, and finally, a high-precision segmentation result is output through a segmentation network.
[0120] This method not only significantly reduces the computational complexity and overfitting risk, but also improves the ability to capture boundaries and details in complex scenes through the integration of multiple technologies, achieving efficient and accurate remote sensing image segmentation. The use of a composite loss function further ensures the classification accuracy and edge sharpness in the case of class imbalance, enhancing the robustness and adaptability of the model in practical applications.
[0121] Embodiment 2;
[0122] As Figure 4 shown, this embodiment takes the building segmentation in high-resolution urban remote sensing images as a specific scenario, elaborates the implementation process of the technical solution in detail, and focuses on optimizing challenges such as dense buildings, shadow occlusion, and complex boundaries in the urban environment.
[0123] The overall process is as follows:
[0124] 1) Data preprocessing and feature extraction;
[0125] Input a city remote sensing image with a resolution of 0.5m / pixel into the feature extraction and fusion optimization network. This image contains dense building clusters, roads, vegetation, and shadow areas. First, perform image preprocessing, adopt a radiation correction method based on a physical model to reduce the interference of building shadows on feature extraction; further apply the CLAHE algorithm to enhance the visibility of low-contrast areas; and normalize the pixel values to the [0,1] interval to eliminate illumination differences.
[0126] Subsequently, load the Mamba model pre-trained on the Urban3D dataset to extract high-level semantic features from the preprocessed image; use convolutional kernels of different sizes of 3×3, 5×7, and 7×7 to extract local window frame details, building outlines, and block global layouts respectively to generate multi-level features; generate a comprehensive feature through channel attention weighted summation to retain the global structure and local details of the building; finally, use a 1×1 convolution to reduce the number of channels from 512 to 256 to generate initial features.
[0127] 2) Feature enhancement and detail restoration;
[0128] In the feature enhancement stage, first, parameter-free similarity attention enhancement is performed; the features are layer-normalized, the mean and variance excluding itself are calculated for each feature point to quantify the difference from surrounding features; a weight matrix is generated through an energy function, adjusted by Sigmoid, and multiplied by the normalized features to obtain attention-enhanced features, with a focus on strengthening the building edge area.
[0129] Next, a dual-path feature extraction and fusion method is adopted; a 3×3 dilated convolution with a dilation rate r = 2 is used to expand the receptive field to 7×7 to extract the building contour; at the same time, depthwise separable convolution (3×3 pointwise convolution + 1×1 pointwise convolution) is used to extract the building surface texture; after splicing the two, through convolution + BN + ReLU operations and pixel rearrangement to restore high-resolution features, further sharpening the building edge.
[0130] 3) Diffusion model optimization and segmentation prediction;
[0131] In the diffusion model optimization stage, based on the diffusion model, the initial noise step t = 100, and a 50-step fast denoising process is carried out. A lightweight U-Net structure is adopted, with 4 layers of downsampling in the encoder and 4 layers of upsampling in the decoder, and multi-scale features are fused through skip connections; during the denoising process, the building contour breakage caused by shadow occlusion is effectively repaired, and the window frame texture is enhanced.
[0132] Furthermore, 2 layers of 3×3 convolution + 1×1 convolution are used to map the optimized features to pixel-level classification results. In loss calculation, a composite loss function is adopted, including cross-entropy loss (weight = 0.5), Dice loss (weight = 0.3), and boundary loss (weight = 0.2), to jointly optimize the segmentation accuracy and edge consistency.
[0133] 4) Post-processing;
[0134] Finally, post-processing steps are carried out, including filling small holes, smoothing edge serrations, converting the segmentation result into a vector polygon, and removing abnormal fragments, to obtain the final building segmentation result.
[0135] This embodiment aims at building segmentation in high-resolution urban remote sensing images. Through data preprocessing, feature extraction, feature enhancement and detail restoration, diffusion model optimization and segmentation prediction, and post-processing steps, it effectively addresses challenges such as dense buildings, shadow occlusion, and complex boundaries, and finally obtains accurate building segmentation results.
[0136] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0137] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A remote sensing image segmentation method combining similarity attention without additional parameters, characterized in that, It includes the following steps: S1. Use the feature extraction and fusion optimization network to extract multi-scale features of the remote sensing image and perform fusion optimization to generate initial features; S2. Based on the initial features, perform sensitivity enhancement through the parameter-free similarity attention unit in the feature enhancement network, and extract boundary and texture features through the boundary and texture feature extraction unit, and fuse them to generate enhanced features; S3. Combine with the segmentation network to optimize the enhanced features and perform class prediction, and output the segmentation result of the remote sensing image.
2. The remote sensing image segmentation method combining similarity attention without additional parameters according to claim 1, wherein, The S1 includes: S11. Use the Mamba pre-trained model to extract features from the remote sensing image to obtain high-level semantic features φ; S12. Based on the high-level semantic feature φ, multi-scale feature extraction is performed through a multi-scale convolutional layer to obtain a multi-level feature φ containing global information multi ; S13. Perform weighted summation or splicing fusion on the multi-level feature φ multi to generate a comprehensive feature φ fused ; S14. Perform dimensionality reduction optimization on the comprehensive feature φ through a convolutional layer and an activation function to obtain the initial feature φ fused 0 . 3. A remote sensing image segmentation method combining similarity attention without additional parameters according to claim 2, characterized in that, In the S12, the multi-scale convolutional layer includes at least three different-sized convolutional kernels, and the multi-level feature φ multi is calculated by the following formula: φ multi = [Conv k1 , Conv k2 , Conv k2 Among them, Conv k1 , Conv k2 , Conv k2 respectively represent convolution operations with convolution kernel sizes of k1, k2, and k3, and [·] represents the concatenation operation.
4. A remote sensing image segmentation method combining similarity attention without additional parameters according to claim 1, characterized in that In the S2, the parameter-free similarity attention unit performs sensitivity enhancement, including: S21. Normalize the initial feature φ 0 to obtain a normalized feature S22. Based on the normalized features Calculate the mean and variance corresponding to each neuron, and evaluate the relative difference from other neurons; wherein, the neuron represents the eigenvalue corresponding to any position or channel. S23. Through the energy function Quantify the neuron difference degree to generate an energy matrix E; S24. Calculate the neuron importance weights according to the energy matrix E, and dynamically adjust the feature response through the Sigmoid function; S25. Multiply the adjusted feature by the original feature to generate an attention-enhanced feature Among them, sigmoid(·) represents the sigmoid activation function.
5. A remote sensing image segmentation method combining similarity attention without additional parameters according to claim 4, characterized in that In S23, the energy function is expressed as: Among them, t represents the current neuron, represents the mean of the remaining neurons after removing the current neuron, represents the variance of the remaining neurons after removing the current neuron, and λ represents the regularization parameter.
6. A remote sensing image segmentation method combining similarity attention without additional parameters according to claim 4, characterized in that, In the S2, the boundary and texture feature extraction unit extracts boundary and texture features, including: Input the initial feature φ 0 into the dilated convolutional layer to expand the receptive field, and generate the boundary feature φ through the LeakyReLU activation function, convolutional layer and ReLU activation function D ; φ D = ReLU(Conv(LeakyReLU(DConv(φ 0 )))) Among them, Conv(·) represents the convolutional layer operation, and DConv(·) represents the dilated convolutional layer operation; And, the initial feature φ 0 is input into the depthwise separable convolutional layer and processed through the Mish activation function and the fully connected layer to generate the texture feature φ DW ; φ DW = FC(Mish(DWConv(φ 0 ))) Among them, FC(·) represents the fully connected layer operation, and DWConv(·) represents the depthwise separable convolutional layer operation.
7. A remote sensing image segmentation method combining similarity attention without additional parameters according to claim 6, characterized in that, It also includes: Concatenate the boundary feature φ D and the texture feature φ DW to obtain the comprehensive feature φ DDW ; For the comprehensive feature φ DDW perform convolution, batch normalization, and ReLU activation processing in sequence, and fuse with the boundary feature φ D and the texture feature φ DW to obtain an enhanced feature 8. A remote sensing image segmentation method combining similarity attention without additional parameters according to claim 7, characterized in that, In the S2, it also includes: The enhanced features are processed by the ReLU activation function and fused with the attention-enhanced features processed by convolution and batch normalization to obtain the fused features; Perform pixel rearrangement on the fused feature to obtain the enhanced feature φ + ; Among them, PShuffle(·) represents the pixel shuffle operation, and BN(·) represents the batch normalization operation.
9. A remote sensing image segmentation method combining similarity attention without additional parameters according to claim 1, characterized in that, The S3 includes: S31. Initially optimize the enhanced feature φ + through a convolutional layer, a batch normalization layer, and an activation function to improve the feature expression ability; S32. Based on the initially optimized enhanced features, perform multi-step denoising processing through the diffusion model, gradually restore the detailed information, and obtain the optimized enhanced features; S33. Combine with the segmentation head to convert the optimized enhanced features into pixel-level classification results.
10. A remote sensing image segmentation method combining similarity attention without additional parameters according to claim 9, characterized in that, The segmentation network adopts a composite loss function L total ; L total = λ1L CE + λ2L Dice + λ3L Boundary Among them, λ1, λ2, and λ3 are weight coefficients, and L CE represents the cross-entropy loss function, and L Dice represents the Dice loss function, and L Boundary represents the boundary loss function.
Citation Information
Patent Citations
Method for extracting building change area in double-time-phase remote sensing image based on twinborn mixed attention mechanism and multi-scale feature fusion
CN118212532A
Remote sensing image semantic segmentation method based on semantic adaptive edge enhancement network
CN118781596A
Complex scene remote sensing image segmentation method based on multi-context U-Net network
CN118898712A
Urban streetscape image real-time semantic segmentation method based on attention boundary enhancement and aggregation pyramid
CN119478401A
Method and equipment for classifying hepatocellular carcinoma images by combining computer vision features and radiomics features
US20210200988A1
Cited By
Ground surface coverage remote sensing dynamic monitoring method and system based on self-adaptive irregular unit grid
CN120495904A
Remote sensing image segmentation method based on multilayer feature fusion and prior guidance
CN120852454A
Remote sensing image segmentation method based on global dependence and local texture fusion attention
CN121725246A