Remote sensing image semantic segmentation method based on feature interaction enhancement

By adopting feature interaction enhancement module in semantic segmentation of remote sensing images, combined with standard convolution, depth separation convolution and expansion convolution, the problems of low computational efficiency and insufficient processing of feature balance for different scales in the existing technology are solved, and higher semantic segmentation accuracy and robustness are achieved.

CN119992099AActive Publication Date: 2025-05-13耕宇牧星(北京)空间科技有限公司

Patent Information

Application Number
CN202510174992.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-13
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The existing deep learning technology has problems such as low computational efficiency, insufficient balance processing of features of different scales, and limited performance improvement space in the semantic segmentation of remote sensing images.

Method used

Using a method based on feature interaction enhancement, a feature interaction enhancement module is constructed, combined with standard convolution, deep separable convolution and expanded convolution, etc., to enhance the capture ability of multi-scale features in remote sensing images, and achieve coordinated extraction of global and local features.

Benefits of technology

It significantly improves the accuracy and robustness of semantic segmentation, makes the model more adaptable and accurate when dealing with complex scenarios and different scale goals, and reduces the resource requirements for model training and deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992099A_ABST
    Figure CN119992099A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image semantic segmentation method based on feature interaction enhancement, and the method comprises the steps: obtaining a remote sensing image, and carrying out the feature extraction, and obtaining a preliminary feature; performing feature interaction enhancement based on the preliminary features to obtain enhanced features; and performing pixel-level segmentation according to the enhanced features. According to the method, the preliminary features are interactively enhanced in combination with operations such as standard convolution, depth separable convolution and expansion convolution, so that the capability of capturing multi-scale features in the remote sensing image is improved, collaborative extraction of global and local features is realized, and the precision and robustness of semantic segmentation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and more particularly to a remote sensing image semantic segmentation method based on feature interaction enhancement. Background Art

[0002] With the rapid development of remote sensing technology, high-resolution remote sensing images have been widely used in environmental monitoring, disaster assessment, crop yield assessment and other fields. As one of the key tasks in computer vision, semantic segmentation can classify each pixel in an image.

[0003] However, high-resolution remote sensing images have the characteristics of complex background and large differences in target scales. In traditional semantic segmentation, feature extraction methods mainly rely on spectral features and spatial features, but these methods often do not work well when dealing with complex scenes.

[0004] In recent years, deep learning technology has made significant progress in the field of semantic segmentation, especially the application of convolutional neural networks (CNN) and Transformer methods, such as attention mechanism, multi-scale feature fusion, etc. In addition, Transformer-based semantic segmentation methods have also made significant progress, such as Swin-Transformer, which performs well in object detection and segmentation tasks.

[0005] Although the above methods significantly improve segmentation accuracy and robustness, existing deep learning technologies (such as attention mechanisms and multi-scale feature fusion) have some shortcomings and drawbacks in global and local feature extraction. For example, the attention mechanism can effectively enhance the feature expression of important areas, but its computational complexity is high, especially when processing large-scale remote sensing images, which may lead to slower reasoning speed. In addition, excessive attention to local features may lead to neglect of global background information, thereby affecting the understanding of the overall semantics of the target.

[0006] In terms of segmentation accuracy, although multi-scale feature fusion can effectively extract image information of different scales, in complex scenes, objects with large scale differences may still be affected by the inconsistent information between features of different scales, resulting in reduced accuracy of segmentation results. In addition, these methods often rely on a large amount of training data and computing resources, and there is limited room for performance improvement when data is scarce or computing power is limited.

[0007] In general, although deep learning technology has shown excellent performance in the semantic segmentation of remote sensing images, there is still room for improvement when processing images with complex backgrounds and targets of different scales, such as computational efficiency and balanced processing of features of different scales. Summary of the invention

[0008] In view of this, in order to at least partially solve the above technical problems, the present invention provides a remote sensing image semantic segmentation method based on feature interaction enhancement, which aims to combine operations such as standard convolution, depthwise separable convolution and dilated convolution to enhance the ability to capture multi-scale features in remote sensing images and realize the collaborative extraction of global and local features, thereby improving the accuracy and robustness of semantic segmentation.

[0009] In order to achieve the above object, the present invention adopts the following technical solution:

[0010] A remote sensing image semantic segmentation method based on feature interaction enhancement, comprising:

[0011] Acquire remote sensing images for feature extraction to obtain preliminary features;

[0012] Perform feature interaction enhancement based on preliminary features to obtain enhanced features;

[0013] and pixel-level segmentation based on enhanced features;

[0014] Feature interaction enhancement, the steps include:

[0015] The preliminary features are input into the upper branch and middle branch networks for local feature extraction. The obtained results are multiplied element by element and then activated to obtain the attention weight map.

[0016] Inputting the preliminary features into the lower branch network to extract global features, and multiplying the global features with the attention weight map and performing a jump connection with the preliminary features to obtain preliminary enhanced features;

[0017] The preliminary enhanced features are batch normalized and then input into the spatial channel reconstruction convolution branch and the depth-wise separable convolution branch respectively. The results are multiplied and jump-connected with the preliminary enhanced features to obtain the enhanced features.

[0018] Preferably, the upper branch network includes a first convolution and a full-dimensional dynamic convolution;

[0019] The middle branch network includes a second convolution and a first depth-wise separable convolution;

[0020] The lower branch network includes dilated convolution and a second depthwise separable convolution.

[0021] Preferably, the spatial channel reconstruction convolution branch includes a third convolution and a spatial channel reconstruction convolution;

[0022] The depthwise separable convolution branch includes the fourth convolution, the third depthwise separable convolution, and an activation function.

[0023] Preferably, the step of extracting features to obtain preliminary features includes:

[0024] Extract local features of remote sensing images and reduce their dimensionality;

[0025] Use the ResNet backbone network to gradually extract high-level semantic features from the dimensionality reduction features;

[0026] Capture target details of different scales in high-level semantic features based on attention mechanism;

[0027] Preferably, the dimensionality reduction includes performing batch normalization, ReLU activation function and maximum pooling layer on the local features in sequence.

[0028] As a preference, the attention mechanism is used to capture target details of different scales in high-level semantic features, including:

[0029] Generate multi-scale features based on high-level semantic features;

[0030] Determine the attention weights of features at each scale through the attention network;

[0031] An interactive weighted approach is used to fuse multi-scale features.

[0032] Preferably, pixel-level segmentation is performed according to the enhanced features, including:

[0033] Through multiple atrous convolutional layers, features of different scales are generated according to the enhanced features to improve the perception field of the enhanced features;

[0034] Fusion of features at different scales;

[0035] Upsample the fused features.

[0036] The upsampled features are classified through the convolutional layer to obtain the category label of each pixel.

[0037] Preferably, after passing through the atrous convolution layer, activation function and batch normalization are added in sequence.

[0038] Preferably, each pixel is normalized using a Softmax function to generate a probability distribution for each category to obtain a segmentation result.

[0039] Compared with the existing technology,

[0040] 1) By constructing a feature interaction enhancement module, the present invention can effectively capture multi-scale features in remote sensing images and realize the collaborative extraction of global and local features. Compared with traditional methods, the accuracy and robustness of semantic segmentation are significantly improved, making the model more adaptable and accurate when dealing with complex scenes and targets of different scales.

[0041] 2) The invention adopts batch normalization, skip links and other technical means in the feature extraction process, which not only stabilizes the network training and accelerates convergence, but also retains the detailed information in the original features, and at the same time integrates high-level features into the output features to form enhanced features containing rich contextual information and detailed features, providing better input for subsequent segmentation networks, further improving the segmentation effect, reducing the gradient vanishing or gradient exploding problems caused by inconsistent input feature distribution, and optimizing the model training process and performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0043] Figure 1 It is a flow chart of a remote sensing image semantic segmentation method based on feature interaction enhancement of the present invention;

[0044] Figure 2 Interactive enhancement flow chart for the features of the present invention. DETAILED DESCRIPTION

[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] The present invention discloses a remote sensing image semantic segmentation method based on feature interaction enhancement. By constructing a hybrid feature interaction module and combining operations such as standard convolution, depthwise separable convolution and dilated convolution, the ability to capture multi-scale features in remote sensing images is enhanced, and the coordinated extraction of global and local features is realized, thereby improving the accuracy and robustness of semantic segmentation, providing richer and more expressive input data for subsequent semantic segmentation networks, so as to better serve the task of accurate segmentation of ship targets and the like in remote sensing images.

[0047] In one embodiment, a remote sensing image semantic segmentation method based on feature interaction enhancement includes the following three steps:

[0048] Step 1: Obtain remote sensing images for feature extraction to obtain preliminary features;

[0049] Step 2: Perform feature interaction enhancement based on preliminary features to obtain enhanced features;

[0050] Step 3: Perform pixel-level segmentation based on enhanced features.

[0051] The present invention not only improves the ability to capture multi-scale features in remote sensing images and realizes the coordinated extraction of global and local features, but also improves the accuracy and robustness of semantic segmentation.

[0052] In one embodiment, referring to Figure 1 , the execution process of each step is as follows:

[0053] Step 1: Extract features from remote sensing images;

[0054] 1.1 Extract local features of remote sensing images and reduce dimensionality; dimensionality reduction includes batch normalization, ReLU activation function and maximum pooling layer on local features in sequence.

[0055] In this embodiment, the input remote sensing image is a 3-channel RGB image with a size of H×W×3, where H and W represent the height and width of the image respectively, and 3 represents the color channel of the image. The remote sensing image carries the original visual information of the target, including the texture, color and other features of the target and background elements.

[0056] The present application extracts local features of remote sensing images through a convolution layer. In one embodiment, the convolution kernel of the convolution layer is 7×7, the stride is set to 2, and the padding is set to 3; the size of the convolution kernel determines the range of the local area of ​​the image perceived by the model. In this embodiment, the use of a 7x7 convolution kernel can capture a larger range of local features, such as edges and textures in the image, and the stride 2 and padding 3 ensure that the convolution operation does not lose important spatial information.

[0057] At the same time, the image is downsampled to reduce the resolution of the image, making subsequent feature extraction more efficient.

[0058] Through the initial convolution operation, the model can capture the low-level features in the image (such as edges, corners, textures, etc.), while reducing the resolution of the image and reducing the amount of calculation. The size of the image is reduced from H×W×3 to about Where C1 is the number of channels after the convolution operation (usually 64 or 128).

[0059] Furthermore, the dimension of local features is reduced;

[0060] In this embodiment, after the convolution operation, a batch normalization layer is first applied. The function of this layer is to normalize the output of each layer to the same scale, thereby reducing the gradient vanishing or gradient exploding problem, helping to accelerate training and improve stability.

[0061] Secondly, the ReLU activation function (Rectified Linear Unit) is applied, which converts all negative values ​​to zero, introduces nonlinearity, and enables the model to learn more complex feature expressions. The purpose of this step is to enhance the model's expressive power so that it can extract more meaningful low-level features;

[0062] Again, a max pooling operation is applied. This example uses a 3×3 pooling window with a stride of 2. The purpose of this step is to further reduce the spatial resolution of the image, retain the most important feature information, and enhance the spatial invariance of the model (i.e., robustness to image displacement). The pooling operation reduces the image size by selecting the maximum value in the local area while maintaining the salient features of the image. The output feature map size of the pooling layer will become The amount of image information is further compressed.

[0063] 1.2 Use the ResNet backbone network to gradually extract high-level semantic features from the dimensionality reduction features;

[0064] This application introduces ResNet (Residual Network) as the backbone network. ResNet gradually extracts high-level semantic features of images through multiple stages. Each stage contains several convolutional layers, and residual connections are used to avoid the gradient vanishing problem in deep networks, so that the network can be trained deeper and extract more complex and abstract features.

[0065] Each residual module ensures that information can flow in the network by adding the input directly to the output, thereby enhancing the network's expressiveness and effectively preventing performance degradation during deep network training.

[0066] As the number of network layers increases, ResNet gradually extracts high-level semantic information of the image. For example, the initial convolutional layer mainly captures local edge and texture features, and as the network goes deeper, ResNet begins to understand more complex patterns, such as the shape and texture of the target and the distinction between the background. The output feature map of the ResNet network has a size of 4H×4W×C, where C is the final number of channels, usually 2048 or more. At this point, the feature map already contains rich semantic information, including the shape of the target, the texture of the background, etc.

[0067] 1.3 Capture target details of different scales in high-level semantic features based on the attention mechanism;

[0068] The high-level semantic features extracted by the ResNet backbone network have been generated, and the size of the feature map is 4H×4W×C, where H and W are the height and width of the image, and C is the final number of channels (usually 2048 or more). These feature maps contain the semantic information of the ship targets in the remote sensing images, but there may be target details of different scales that are not fully captured.

[0069] Therefore, this application introduces the MAMBA (Multi-scale Attention Module with BilinearAttention) module, the core of which is the multi-scale attention mechanism, which improves the expressiveness of the image by fusing features of different scales.

[0070] The MAMBA module mainly consists of the following two parts:

[0071] 1) Multi-scale attention mechanism: assign different weights to feature maps of different scales to enhance the fusion ability of multi-scale information;

[0072] 2) Bilinear attention mechanism: Enhance the interaction between each scale, finely capture local and global information at different scales, and avoid information loss.

[0073] In one embodiment, the capturing step includes:

[0074] 1.3.1 Generate multiple scale feature maps based on high-level semantic features;

[0075] This embodiment uses multiple convolutional layers or pooling layers to generate feature maps of multiple scales from the original feature map. For example, feature maps of multiple scales can be generated by convolution kernels of different sizes (such as 3x3, 5x5, etc.) or different pooling operations (such as maximum pooling, average pooling, etc.). Feature maps of these scales can capture the details of objects of different sizes.

[0076] 1.3.2 Determine the attention weight of each scale feature through the attention network;

[0077] The attention network is a convolutional layer or a fully connected layer, and the weights obtained reflect the importance of features at different scales. In remote sensing images, the scales of objects vary greatly, and MAMBA dynamically adjusts the weight of each scale based on the content of the feature map.

[0078] 1.3.3 Use interactive weighted method to fuse multi-scale features.

[0079] In this embodiment, a bilinear attention mechanism is used to enable effective fusion of feature maps of different scales through an interactive weighting method, thereby enhancing the expression ability of the model at different scales.

[0080] Furthermore, the final fused feature map is further processed by an activation function (such as ReLU) and outputs feature maps for subsequent networks. These feature maps are richer and can capture the details and background information of ship targets at different scales.

[0081] The size of the output feature map after the MAMBA module is still 4H×4W×C, where C is the number of channels after the MAMBA module (usually remains unchanged, but can be adjusted as needed). These output feature maps are more refined than those without the MAMBA module and contain multi-scale information, which can better support subsequent segmentation tasks.

[0082] The feature map output by the MAMBA module will be used as the input for subsequent feature interaction enhancement and then passed to the decoder for further upsampling and segmentation operations.

[0083] Through the multi-scale feature enhancement of the MAMBA module, the network can better restore the detailed information of the target and avoid the loss of important target information during the upsampling process.

[0084] This application gradually extracts low-level to high-level semantic features by applying convolutional layers, batch normalization, ReLU activation function, maximum pooling, and using the ResNet backbone network, which can efficiently capture the morphology and background information of the ship target.

[0085] Step 2: Perform feature interaction enhancement based on the preliminary feature map;

[0086] In one embodiment, the present application constructs a feature interaction enhancement module, the structure of which is as follows: Figure 2 As shown in the figure; this module enhances the ability to capture multi-scale features in remote sensing images by combining operations such as standard convolution, depthwise separable convolution and dilated convolution, realizes the coordinated extraction of global and local features, and improves the accuracy and robustness of semantic segmentation.

[0087] Reference Figure 2 , the interactive enhancement steps are:

[0088] 2.1 Input the preliminary feature map into the upper branch and middle branch network for local feature extraction, and obtain the attention weight map through the activation function after element-by-element multiplication.

[0089] In a preferred embodiment, the preliminary feature map is first Perform batch normalization and get To stabilize network training and accelerate convergence.

[0090] Further, They are input into the upper branch and middle branch networks respectively to extract local features and reduce computational complexity. The upper branch network includes the first convolution and the full-dimensional dynamic convolution; the middle branch network includes the second convolution and the first depth-separable convolution;

[0091] Then, the features extracted by the two branches are multiplied element by element, and the Softmax activation function is applied to obtain the attention weight map. The expression is:

[0092]

[0093] Among them, DWConv represents depthwise separable convolution, which extracts local features, Conv represents convolution, and OCONV represents full-dimensional dynamic convolution operation. represents the multiplication operation, and σ represents the Softmax activation function.

[0094] 2.2 In order to make up for the lack of attention paid to local features by the upper and middle branches, the feature interaction enhancement module of this application designs a lower branch to capture global context information; that is, the preliminary feature map is input into the lower branch network to extract global features, and the lower branch network includes dilated convolution and a second depth-separable convolution.

[0095] In one embodiment, A 7×7 dilated convolution operation is applied, which can expand the receptive field of the convolution kernel without increasing the amount of calculation, thereby effectively capturing a wider range of global features. Then, local features are extracted through depthwise separable convolution to obtain the output of the lower branch.

[0096]

[0097] Among them, DConv represents dilated convolution. By combining dilated convolution and depthwise separable convolution, the long-distance dependencies in remote sensing images can be effectively captured, and the recognition ability of complex scenes can be improved.

[0098] Further, the global feature is multiplied by the attention weight map and then multiplied by the preliminary feature map Perform skip connection to obtain the initial enhanced feature map Preferably, before performing the skip connection, the multiplication result is applied through a convolutional layer, and the expression is:

[0099]

[0100] in, Represents an addition operation. By skipping links, the detailed information in the original features can be retained, and the high-level features extracted by the feature enhancement module can be integrated into the output features, thus forming an enhanced feature map containing rich contextual information and detailed features.

[0101] 2.3 Preliminary Enhanced Feature Map Batch normalization is performed to accelerate the convergence of the model and reduce the gradient vanishing or gradient exploding problems caused by inconsistent input feature distribution during training. The feature map after batch normalization is denoted as It adjusts the mean and variance to make the numerical range of each feature channel more consistent, which helps the network learn more expressive features.

[0102] Then they are respectively input into the spatial channel reconstruction convolution branch and the depthwise separable convolution branch. The spatial channel reconstruction convolution branch includes the third convolution and the spatial channel reconstruction convolution. The spatial channel reconstruction convolution reduces redundant information by reconstructing the spatial and channel dimensions; the depthwise separable convolution branch includes the fourth convolution, the third depthwise separable convolution and the activation function.

[0103] Further multiply the results to get

[0104]

[0105] Among them, SCConv stands for spatial channel reconstruction convolution, which means point-by-point multiplication.

[0106] Then, Perform convolutional layer operations to extract higher-level features and combine them with the initial enhanced features Add and get enhanced features

[0107] This addition operation not only retains the information of the original features, but also integrates the advanced features processed by multi-layer convolution and depth-separable convolution, thereby enhancing the model's ability to understand complex scenes. Through such operations, the expressive power of the feature map has been significantly enhanced, which can better serve the semantic segmentation task of remote sensing images.

[0108] Finally, the enhanced feature map obtained It combines the original feature map with the new features obtained through complex convolution operations. The feature map of size H×W×C contains rich semantic information of the target in the image. The size H×W represents the spatial resolution of the image, while C represents the number of channels, which is usually 2048 or more, depending on the configuration of the ResNet backbone network. It provides richer input data for the subsequent semantic segmentation network.

[0109] Step 3: Perform pixel-level segmentation based on enhanced features.

[0110] In one embodiment, the enhanced feature Input to the subsequent convolutional network to generate the results of the segmentation task. Feature map High-level semantic information has been extracted through the previous steps and can therefore be used as input to the segmentation network for further processing and to produce a segmentation map.

[0111] In this embodiment, dilated convolution is applied to capture a wider range of contextual information and further enhance the resolution and details of the model. Dilated convolution expands the receptive field of the convolution kernel by inserting intervals between the elements of the convolution kernel without increasing the amount of computation or pooling operations. Dilated convolution can effectively retain the details in the feature map while capturing greater contextual information in the image, enabling the model to understand the structure in the image more comprehensively. By applying dilated convolutions of different scales, the model can simultaneously obtain local details and global semantic information, which is crucial for accurate segmentation of the target.

[0112] In an exemplary embodiment, the segmentation step includes:

[0113] 3.1 Through multiple atrous convolutional layers, feature maps of different scales are generated according to the enhanced features to gradually improve the receptive field of view of the enhanced features and capture a wider range of contextual information; preferably, after each convolution operation, activation functions (such as ReLU) and batch normalization are added to ensure stable network training.

[0114] 3.2 Fuse feature maps of different scales to make full use of contextual information of different scales.

[0115] In this application, feature fusion methods include feature concatenation and feature summation. Through multi-scale feature fusion, the model can extract information from different scales and improve the accuracy of segmentation results, especially when facing small targets or complex backgrounds, which can effectively improve the performance of detection and segmentation.

[0116] Furthermore, the present application processes the fused feature map through a convolutional layer to further compress the size of the feature map while retaining important multi-scale information.

[0117] 3.3 Upsample the fusion features,

[0118] After the dilated convolution and feature fusion, the resolution of the feature map may be low. In order to generate the final segmentation result, it is necessary to upsample the feature map through the decoding process to restore the resolution of the feature map to the same size as the input image, that is, H×W×C. out , where Cout is the number of channels of the segmentation result, which is usually equal to the number of categories.

[0119] In this application, upsampling can be achieved by deconvolution, upsampling layers, or bilinear interpolation. Upsampling restores the spatial resolution of the feature map for pixel-level segmentation. In this process, the model gradually restores the details of the image and generates a segmentation map that is consistent with the size of the input image.

[0120] 3.4 Classify the upsampled features through a 1x1 convolutional layer to obtain the category label of each pixel.

[0121] This step maps the features of each pixel to the final category space to obtain the category prediction for each pixel. 1x1 convolution can classify each pixel independently and output the probability or label of the corresponding category. Since the size of the convolution kernel is 1x1, it does not affect the spatial resolution of the image, but can assign a classification label to each pixel based on the previously extracted features.

[0122] Through the 1x1 convolution layer, the category label of each pixel is generated to obtain the final segmentation map; then each pixel is normalized through the Softmax function to generate the probability distribution of each category, and finally a classification result for each pixel is obtained.

[0123] Compared with the prior art, the present invention has the following beneficial effects:

[0124] The present invention not only improves the accuracy of segmentation, but also reduces the resource requirements for model training and deployment through efficient parameter fine-tuning and the use of adapters, making it suitable for remote sensing image processing tasks in various computing environments. This method has stronger adaptability and accuracy when processing complex scenes and objects of different scales, and significantly improves the performance of semantic segmentation of remote sensing images.

[0125] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0126] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A remote sensing image semantic segmentation method based on feature interaction enhancement, characterized in that: Acquire remote sensing images for feature extraction to obtain preliminary features; Perform feature interaction enhancement based on preliminary features to obtain enhanced features; and pixel-level segmentation based on enhanced features; Feature interaction enhancement, the steps include: The preliminary features are input into the upper branch and middle branch networks for local feature extraction. The obtained results are multiplied element by element and then activated to obtain the attention weight map. Inputting the preliminary features into the lower branch network to extract global features, and multiplying the global features with the attention weight map and performing a jump connection with the preliminary features to obtain preliminary enhanced features; The preliminary enhanced features are batch normalized and then input into the spatial channel reconstruction convolution branch and the depth-wise separable convolution branch respectively. The results are multiplied and jump-connected with the preliminary enhanced features to obtain the enhanced features.

2. The remote sensing image semantic segmentation method according to claim 1, characterized in that: The upper branch network includes the first convolution and the full-dimensional dynamic convolution; The middle branch network includes a second convolution and a first depth-wise separable convolution; The lower branch network includes dilated convolution and a second depthwise separable convolution.

3. The remote sensing image semantic segmentation method according to claim 1, characterized in that: The spatial channel reconstruction convolution branch includes the third convolution and the spatial channel reconstruction convolution; The depthwise separable convolution branch includes the fourth convolution, the third depthwise separable convolution, and an activation function.

4. The remote sensing image semantic segmentation method according to claim 1, characterized in that: The steps for feature extraction and obtaining preliminary features include: Extract local features of remote sensing images and reduce their dimensionality; Use the ResNet backbone network to gradually extract high-level semantic features from the dimensionality reduction features; Capture object details of different scales in high-level semantic features based on attention mechanism.

5. The remote sensing image semantic segmentation method according to claim 4, characterized in that: Dimensionality reduction includes batch normalization, ReLU activation function and maximum pooling layer on local features.

6. The remote sensing image semantic segmentation method according to claim 4, characterized in that: The attention mechanism is used to capture target details of different scales in high-level semantic features, including: Generate multi-scale features based on high-level semantic features; Determine the attention weights of features at each scale through the attention network; An interactive weighted approach is used to fuse multi-scale features.

7. The remote sensing image semantic segmentation method according to claim 1, characterized in that: Pixel-level segmentation based on enhanced features, including: Through multiple atrous convolutional layers, features of different scales are generated according to the enhanced features to improve the perception field of the enhanced features; Fusion of features at different scales; Upsample the fused features. The upsampled features are classified through the convolutional layer to obtain the category label of each pixel.

8. The remote sensing image semantic segmentation method according to claim 7, characterized in that: After passing through the atrous convolution layer, activation function and batch normalization are added in sequence.

9. The remote sensing image semantic segmentation method according to claim 7, characterized in that: The Softmax function is used to normalize each pixel to generate the probability distribution of each category and obtain the segmentation result.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method based on attention multi-scale feature fusion

    CN111127493A

  • Remote sensing image semantic segmentation method and device

    CN111582104A

  • Real-time semantic segmentation method based on multi-scale feature interaction and enhancement

    CN116385719A

  • Remote sensing image semantic segmentation method, device and equipment based on mixed feature extraction

    CN116778169A

  • Road extraction method and system based on dynamic and deformation crossover Transformer

    CN117437550A

Cited By

  • Remote sensing image segmentation method and system based on cross-normal-form feature fusion and alignment

    CN120807934A

  • A remote sensing image segmentation method and system based on cross-paradigm feature fusion and alignment

    CN120807934B

  • Remote sensing image semantic segmentation method based on CNN-Transform-SAM dynamic collaboration and scene adaptation

    CN121147523A