Remote sensing image road semantic segmentation method based on polarization self-attention feature enhancement

By introducing polarization self-attention feature enhancement mechanism in remote sensing image processing, generating channel self-attention and spatial self-attention, dynamic weighting and adjustment features, the problem of poor feature capture in remote sensing image path segmentation is solved, and the accuracy and robustness of segmentation are significantly improved.

CN119992101AActive Publication Date: 2025-05-13耕宇牧星(北京)空间科技有限公司
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510175926.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-13
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

When traditional CNN models deal with road segmentation tasks in remote sensing images, it is difficult to accurately capture small and unevenly distributed road features, and under complex backgrounds and high noise conditions, missegment is prone to occur.

Method used

Using a method based on polarization self-attention feature enhancement, channel self-attention and spatial self-attention are generated through convolution and reshaping operations. Combined with Softmax and Sigmoid activation functions, the extracted features are dynamically weighted and adjusted to enhance feature representation ability and adapt to output distribution.

Benefits of technology

It significantly improves the accuracy of road semantic segmentation in remote sensing images, enhances the network's ability to capture key features of road segmentation, and improves the accuracy and robustness of segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992101A_ABST
    Figure CN119992101A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image road semantic segmentation method based on polarization self-attention feature enhancement. The method comprises the following steps: carrying out preprocessing and feature extraction on an original road remote sensing image; constructing a polarization self-attention feature enhancement module, respectively generating channel self-attention and space self-attention through convolution and remodeling operations, and dynamically weighting and adjusting the extracted features in combination with an activation function to obtain an enhanced feature map; and recovering the spatial resolution of the enhanced feature map through up-sampling and convolution operations, and outputting a segmentation result with the same size as the original road remote sensing image. According to the method, a polarization self-attention mechanism is introduced, so that the feature representation capability and the adaptive capacity to output distribution are enhanced, the network can more accurately capture the key features for road segmentation in the remote sensing image, and the method has remarkable advantages when the remote sensing image with complex background and more noise is processed; the precision and robustness of road semantic segmentation can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of remote sensing image processing, and in particular to a remote sensing image road semantic segmentation method based on polarization self-attention feature enhancement. Background Art

[0002] Remote sensing images are widely used in geographic information systems, environmental monitoring, urban planning and other fields. With the development of remote sensing technology, it is becoming easier to obtain high-resolution remote sensing images, but how to extract useful information from these images, especially for road semantic segmentation, is still a challenging task. Traditional image processing methods often perform poorly when processing remote sensing images with complex backgrounds and high noise, and it is difficult to accurately segment road areas.

[0003] In recent years, deep learning technology has made significant progress in the field of image segmentation. Convolutional neural networks (CNNs) gradually extract feature information from images through multi-layer convolution operations, and can effectively capture complex structures and details in images. However, in the task of road segmentation in remote sensing images, traditional CNN models still have some limitations. First, the road features in remote sensing images are often small and unevenly distributed, and traditional CNN models tend to ignore key information when processing these features. Second, the background in remote sensing images is complex and noisy, and traditional CNN models tend to produce mis-segmentation when processing these complex backgrounds.

[0004] In order to solve the above problems, researchers have proposed various improvement methods. For example, the U-Net model can effectively capture the detailed information in the image by gradually extracting and restoring image features through the encoder-decoder structure. However, when processing remote sensing images, the U-Net model still has the problem of insufficient feature representation ability, and it is difficult to accurately segment the road area. In addition, the attention mechanism has also been widely used in image segmentation tasks. The attention mechanism can enhance the representation ability of key features by weighting and adjusting the features, thereby improving the accuracy of segmentation. However, when processing remote sensing images, the existing attention mechanism often ignores the spatial relationship and channel relationship of the features, making it difficult to fully capture the road features. Summary of the invention

[0005] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides a remote sensing image road semantic segmentation method based on polarization self-attention feature enhancement, which aims to improve the accuracy of road semantic segmentation in remote sensing images by enhancing feature representation capability and adapting output distribution.

[0006] To achieve the above object, the technical solution adopted by the present invention is:

[0007] In a first aspect, an embodiment of the present invention provides a remote sensing image road semantic segmentation method based on polarization self-attention feature enhancement, the method comprising the following steps:

[0008] Step 1: Preprocess and extract features of the original road remote sensing image;

[0009] Step 2: Construct a polarized self-attention feature enhancement module, generate channel self-attention and spatial self-attention respectively through convolution and reshaping operations, combine the activation function, dynamically weight and adjust the extracted features, and obtain an enhanced feature map;

[0010] Step 3: Perform a decoding operation to restore the spatial resolution of the enhanced feature map through upsampling and convolution operations, and output a segmentation result with the same size as the original road remote sensing image.

[0011] Furthermore, in step 1, the original road remote sensing image is preprocessed and feature extracted, and the specific process includes:

[0012] Step 1.1: First, preprocess the input image, including normalization and denoising; then extract the basic feature map through a convolutional layer;

[0013] Step 1.2: Apply an activation function to the basic feature map to introduce nonlinearity into the network and obtain a second feature map;

[0014] Step 1.3: Through the downsampling operation, the spatial resolution of the second feature map is reduced to aggregate context information and reduce the amount of calculation to obtain the third feature map;

[0015] Step 1.4: Repeat the operations of steps 1.2 and 1.3 to extract deeper abstract features.

[0016] Furthermore, in step 2, the specific process of the polarization self-attention feature enhancement module performing feature enhancement includes:

[0017] Assume the input feature map is Where C represents the number of channels of the input image, H and W represent the height and width of the image respectively;

[0018] ①Construct channel self-attention:

[0019] The input feature map Perform convolution and reshape operations to obtain

[0020]

[0021] in, Reshape means reshaping operation, that is, changing the shape of tensor; Conv 1×1Represents a 1×1 convolutional layer;

[0022] right Using convolutional layers, reshape operations, and activation functions, we get

[0023]

[0024] in, Softmax represents the softmax activation function;

[0025] Will and After matrix multiplication, the convolution layer, batch normalization layer and activation function are used to obtain channel self-attention.

[0026]

[0027] in, Sigmoid represents the Sigmoid activation function; LayerNorm represents the normalization layer;

[0028] ②Construct spatial self-attention:

[0029] right Perform convolution, global average pooling and reshaping operations to obtain

[0030]

[0031] in, GAP represents the global average pooling operation;

[0032] right Using convolutional layers, reshape operations, and activation functions, we get

[0033]

[0034] in, Softmax represents the softmax activation function;

[0035] Will and By performing matrix multiplication, reshaping and activation function, we can get spatial self-attention

[0036]

[0037] in,

[0038] ③ Channel self-attention Spatial Self-Attention Respectively with the input feature map Multiply to obtain enhanced feature map

[0039]

[0040] in, ⊙ represents element-wise multiplication; Represents matrix addition.

[0041] Furthermore, in step 3, a decoding operation is performed to restore the spatial resolution of the enhanced feature map through upsampling and convolution operations, and a segmentation result having the same size as the original road remote sensing image is output, which specifically includes:

[0042] Feature map upsampling: The spatial resolution of the feature map is gradually restored through upsampling operations, making the feature map gradually approach the size of the original image:

[0043] Convolution and activation: Apply multiple convolutional layers and activation functions to the upsampled feature map to extract finer-grained features;

[0044] Fusion of multi-scale features: The upsampled feature map obtained in the decoding process is fused with the low-level feature map in the encoding stage to form a multi-scale feature map;

[0045] Output segmentation results: The multi-scale feature maps are passed through several convolutional layers and activation functions, and finally the segmentation results are output.

[0046] Furthermore, in step 3, a composite loss function is constructed by combining cross entropy loss, Dice coefficient loss and boundary perception loss to perform remote sensing image road semantic segmentation training.

[0047] Furthermore, the constructed composite loss function is:

[0048]

[0049] in, represents compound loss; represents the cross entropy loss; Represents Dice coefficient loss; represents the boundary perception loss; α, β, and γ represent the weight coefficients of each loss, respectively.

[0050] In a second aspect, the present invention further provides a remote sensing image road semantic segmentation network based on polarization self-attention feature enhancement, which applies the above-mentioned remote sensing image road semantic segmentation method based on polarization self-attention feature enhancement to realize remote sensing image road semantic segmentation, and the network includes:

[0051] Image preprocessing and feature extraction module, used for preprocessing and feature extraction of original road remote sensing images;

[0052] The polarized self-attention feature enhancement module is used to generate channel self-attention and spatial self-attention through convolution and reshaping operations, respectively, and dynamically weight and adjust the extracted features in combination with Softmax and Sigmoid activation functions to obtain enhanced feature maps;

[0053] The road semantic segmentation module is used to perform decoding operations, restore the spatial resolution of the enhanced feature map through upsampling and convolution operations, and output a segmentation result with the same size as the original road remote sensing image.

[0054] In a third aspect, the present invention also provides an electronic device comprising a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the above-mentioned remote sensing image road semantic segmentation method based on polarization self-attention feature enhancement.

[0055] Compared with the prior art, the present invention has at least the following beneficial effects:

[0056] 1. The present invention provides a remote sensing image road semantic segmentation method based on polarization self-attention feature enhancement. By introducing the polarization self-attention mechanism, the feature representation ability and the adaptability to the output distribution are enhanced, so that the network can more accurately capture the key features for road segmentation in remote sensing images, thereby significantly improving the accuracy of semantic segmentation.

[0057] 2. The present invention uses a polarized self-attention feature enhancement module to generate channel self-attention and spatial self-attention respectively, and combines Softmax and Sigmoid activation functions to achieve dynamic weighting and adjustment of features. This process enables the network to "focus" on features that are more critical to road segmentation in remote sensing images, while reducing the impact of irrelevant areas, thereby more accurately capturing road features in the image. This method has significant advantages when processing remote sensing images with complex backgrounds and more noise, and can effectively improve the accuracy and robustness of road semantic segmentation.

[0058] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.

[0059] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0061] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0062] Figure 1 A schematic flow chart of a method for remote sensing image road semantic segmentation based on polarization self-attention feature enhancement provided in an embodiment of the present invention.

[0063] Figure 2 A schematic diagram of a remote sensing image road semantic segmentation network architecture provided in an embodiment of the present invention.

[0064] Figure 3 A schematic diagram of the structure and principle of the polarization self-attention feature enhancement module provided in an embodiment of the present invention.

[0065] Figure 4 A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0066] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.

[0067] In the description of the present invention, it should be noted that: in some processes described in the specification and drawings of this application, multiple operations appearing in a specific order are included, but it should be clearly understood that these operations may not be performed in the order in which they appear in this document or may be performed in parallel. In addition, various serial numbers are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0068] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0069] See also Figure 1As shown, the present invention provides a remote sensing image road semantic segmentation method based on polarization self-attention feature enhancement, which aims to improve the accuracy of road segmentation in remote sensing images by enhancing feature representation capabilities and adapting output distribution. The method mainly includes the following steps:

[0070] Step 1: This step describes the process of extracting features from the original remote sensing image, including preprocessing, convolution, activation function and downsampling operations, gradually extracting higher-level abstract features to prepare for the subsequent polarization self-attention module.

[0071] Step 2: Enhance feature representation and adapt output distribution through polarized self-attention. Through convolution and reshaping operations, channel self-attention and spatial self-attention are generated respectively, and combined with Softmax and Sigmoid activation functions, dynamic weighting and adjustment of features are achieved. Finally, self-attention is multiplied with the original features, so that the network can accurately capture the key features for road segmentation in remote sensing images and improve the accuracy of semantic segmentation.

[0072] Step 3: Decoding and composite loss function construction. This step gradually restores the spatial resolution of the enhanced feature map through upsampling and convolution operations, outputs a segmentation result of the same size as the original image, and combines cross entropy loss, Dice coefficient loss, and boundary perception loss to construct a composite loss function to optimize the network training process, thereby improving the accuracy and robustness of road semantic segmentation.

[0073] The remote sensing image road semantic segmentation method based on polarization self-attention feature enhancement of the present invention not only improves the accuracy of segmentation, but also reduces the resource requirements for model training and deployment through efficient parameter fine-tuning, making it suitable for remote sensing image processing tasks in various computing environments.

[0074] Combine the following Figure 2 and Figure 3 As shown, the specific implementation mode and working principle of the method of the present invention are described in detail:

[0075] Step 1 described: original image feature extraction and processing, such as Figure 2 As shown on the left side of the figure. In this step, the present invention starts with the original remote sensing image and gradually extracts the feature information in the image through multi-layer convolution operations. The feature map of each layer not only retains the spatial information of the image, but also learns more advanced abstract features through the gradually increasing convolution layers. Finally, the output feature map will be used in the subsequent polarization self-attention module. The specific process is as follows:

[0076] Step 1.1: Preprocessing and preliminary convolution of input image: Assume that the original input image is Where C0 represents the number of channels of the input image, H and W are the height and width of the image respectively. First, the input image is preprocessed, such as normalization and denoising, and then a 3×3 convolutional layer is used to extract basic features. C1 is the number of output channels after the convolution operation. This operation extracts low-level features (such as edges, textures, etc.) from the original image through preliminary convolution, laying the foundation for the processing of subsequent layers. This step is similar to the first layer of convolution operation in U-Net, and its purpose is to convert the input image into a feature representation that is easier to learn.

[0077] Step 1.2: Convolution and activation function: For the feature map in step 1.1 Applying activation functions (such as ReLU) enables the network to introduce nonlinearity and thus better handle complex image features. (i.e. the second feature map, the same applies below). The ReLU activation function introduces nonlinearity, which enables the network to better fit complex image features. It effectively avoids the gradient vanishing problem and accelerates convergence during training.

[0078] Step 1.3: Downsampling operation (pooling): Through downsampling operations (such as 2×2 maximum pooling), the spatial resolution of the feature map is reduced to aggregate more contextual information and reduce the amount of calculation, resulting in The pooling operation reduces the spatial resolution of the feature map, allowing the model to focus on a larger receptive field (i.e., a larger range in the image). In U-Net, this downsampling layer helps extract deeper features while reducing the complexity of subsequent calculations.

[0079] Step 1.4: Repeat the convolution and activation operations of steps 1.2 and 1.3 to further abstract and extract deeper features. The output feature map of the i-th layer is

[0080]

[0081] Among them, Conv represents the convolution layer; ReLU represents the ReLU activation function; Maxpool represents the maximum pooling operation. Through multiple convolutions and nonlinear activations, the network can gradually extract higher-level abstract features. This operation is similar to the deep encoder in U-Net, which gradually learns more complex structures and details in the image.

[0082] Step 2 described: Construct polarized self-attention to enhance the representation ability of features and the adaptability to output distribution. Process each layer of feature maps obtained in step 1, and the size of the output feature map is consistent with the size of the input feature map, such as Figure 3 The specific process is as follows:

[0083] Step 2.1: Assume that the input feature map is Construct channel self-attention: First, perform convolution and reshape operations to obtain

[0084]

[0085] in, Reshape means reshaping operation, that is, changing the shape of tensor; Conv 1×1 Represents a 1×1 convolutional layer. This step can mix the channels of the feature map in the spatial dimension (height and width), thereby reducing the number of channels of the feature map and enhancing the expression of important channels.

[0086] Step 2.2: Applying convolutional layers, reshape operations, and activation functions, we get

[0087] in, Softmax represents the softmax activation function.

[0088] Step 2.3: Get the value from step 2.1 and step 2.2 After matrix multiplication, convolution layer, batch normalization layer and activation function, channel self-attention can be obtained.

[0089]

[0090] in, Sigmoid represents the Sigmoid activation function; LayerNorm represents the normalization layer. In the process of generating self-attention, the nonlinear combination of Softmax and Sigmoid can better adapt to the output distribution in typical fine-grained regression tasks.

[0091] Step 2.4: Construct spatial self-attention: First, perform convolution, global average pooling and reshaping operations to obtain

[0092]

[0093] in, GAP stands for global average pooling. The core idea of ​​spatial self-attention is to model the global context of the input image in the spatial dimension. In remote sensing image road segmentation, the features at different locations in the image may be significantly different. Therefore, by performing convolution and global average pooling operations on the feature map, the overall spatial information of the image can be extracted, which helps to capture long-distance dependencies. The reshaping operation makes the feature dimension suitable for subsequent attention calculations, thereby enhancing the spatial relationship between features.

[0094] Step 2.5: Applying convolutional layers, reshape operations, and activation functions, we get

[0095] in, Softmax represents the softmax activation function.

[0096] Step 2.6: Add the value obtained in step 2.4 to and step 2.5 By performing matrix multiplication, reshaping and activation function, we can get spatial self-attention

[0097]

[0098] in, This series of operations enables the network to accurately capture the spatial relationships of objects such as roads in remote sensing images, thereby effectively performing semantic segmentation.

[0099] Step 2.7: Channel self-attention Spatial Self-Attention Respectively with the original features Multiplication:

[0100]

[0101] in, ⊙ represents element-wise multiplication; Represents matrix addition. This process enables the network to "focus" on features that are more critical to road segmentation in remote sensing images, while reducing the impact of irrelevant areas. Self-attention helps the network capture road features in images more accurately, especially in remote sensing images with complex backgrounds and more noise. Therefore, for each layer of feature maps The corresponding features will be enhanced.

[0102] In the polarized self-attention of the present invention, "polarization" is mainly reflected in the process of selectively weighting and adjusting features through the self-attention mechanism. Through this mechanism, the model can adaptively adjust the "importance" of input features, so as to better capture the features in remote sensing images that are critical to road segmentation. The term polarization is used here more as a metaphor, reflecting the dynamics and directionality of the feature weighting, adjustment and optimization process.

[0103] Step 3 described: decoding and composite loss function construction, such as Figure 2 As shown on the right. The purpose of step 3 is to restore the enhanced feature map obtained in step 2 to a segmentation result of the same size as the original image through the decoding process, and at the same time design a composite loss function for training the network. This step is similar to the decoding process in U-Net. The spatial resolution is gradually restored through upsampling and convolution operations, and finally a prediction map of the same size as the input image is output. In order to optimize the training effect, a composite loss function is also used to combine different loss terms to guide the model to learn more accurate segmentation results. The specific process is as follows:

[0104] Step 3.1: Feature map upsampling. In this process, the spatial resolution of the feature map is gradually restored through upsampling. First, the feature map processed by the polarized self-attention feature enhancement module is upsampled (e.g., bilinear interpolation or deconvolution) to obtain the upsampled feature map. This step gradually restores the spatial resolution so that the feature map gradually approaches the size of the original image.

[0105] Step 3.2: Convolution and Activation. On the upsampled feature map, multiple convolution layers and activation functions (such as ReLU or Leaky ReLU) are applied for further processing to extract finer-grained features. In this step, the convolution operation is not only used to extract high-level features in the image, but also to make the information in the feature map more compact. Next, the feature map is further processed by the activation function to introduce nonlinearity and enhance the network's expressiveness.

[0106] Step 3.3: Fusion of multi-scale features. In the decoding process, multi-scale feature fusion technology is used to combine feature information at different levels to improve the segmentation effect. Specifically, the upsampled feature map obtained in the decoding process is fused with the low-level feature map in the encoder stage to form a multi-scale feature map. In this way, the network can integrate low-level and high-level information to better capture the detailed features in the image.

[0107] That is, the formulas of steps 3.1, 3.2 and 3.3 are as follows:

[0108]

[0109] Among them, Upsample represents the upsampling operation, Conv represents the convolution operation, and ReLU represents the ReLU activation function. Represents the feature map addition operation.

[0110] Step 3.4: Output layer and prediction results. After several convolutional layers and activation functions, the segmentation result is finally output. The result is a class probability map, which indicates the probability of each pixel belonging to the road or the background. In this process, the final output feature map is activated by softmax or sigmoid function, and the output is the class probability of each pixel. Assume that the output map is Where C is the number of categories (2 in the road semantic segmentation task, corresponding to the road and background categories respectively), and H and W are the spatial dimensions of the image.

[0111] Step 3.5: Composite loss function. In order to effectively train the network, a composite loss function is used to combine multiple loss terms. The composite loss function includes cross entropy loss, Dice coefficient loss, and boundary-aware loss to optimize the model training process. Cross entropy loss Cross entropy loss is used to measure the difference between the predicted class probability and the true label. Dice coefficient loss The Dice coefficient loss is used to evaluate the overlap between the segmentation result and the true label, which is especially suitable for cases with imbalanced categories. In semantic segmentation tasks, especially road segmentation, there is often an imbalance in the ratio of background to road. Boundary perception loss Boundary-aware loss is used to enhance the network's performance when processing image boundaries (such as road edges). This loss can help the network better learn the boundary information in the image and improve segmentation accuracy.

[0112] The final training loss function is the weighted sum of the above losses, and the formula is as follows:

[0113]

[0114] Among them, α, β and γ are weight coefficients, which are adjusted through methods such as cross-validation to balance the impact of each loss term during the training process. By optimizing the total loss function, the model can learn more accurate road segmentation results, especially in complex backgrounds and road edge areas. Through this step, the network's decoding process and the composite loss function work together, allowing the network to accurately segment roads in remote sensing images and improve the accuracy and robustness of road semantic segmentation.

[0115] From the description of the above embodiments, those skilled in the art can know that: the present invention provides a method for semantic segmentation of roads in remote sensing images based on polarization self-attention feature enhancement, which generates channel self-attention and spatial self-attention respectively through convolution and reshaping operations, and combines Softmax and Sigmoid activation functions to achieve dynamic weighting and adjustment of features. This process enables the network to "focus" on features that are more critical to road segmentation in remote sensing images, while reducing the impact of irrelevant areas, thereby more accurately capturing road features in images, especially in remote sensing images with complex backgrounds and more noise.

[0116] Specifically, channel self-attention enhances the expression of important channels by mixing channels of feature maps in the spatial dimension, thereby improving the representation ability of features. Spatial self-attention extracts the overall spatial information of the image by performing convolution and global average pooling operations on feature maps, captures long-distance dependencies, and enhances the spatial relationship between features. By multiplying channel self-attention and spatial self-attention with the original features respectively, the network can better adapt to the output distribution in typical fine-grained regression tasks, thereby improving the accuracy of semantic segmentation.

[0117] In summary, the present invention enhances the feature representation capability and adaptability to output distribution by introducing the polarized self-attention mechanism, enabling the network to more accurately capture the key features for road segmentation in remote sensing images, thereby significantly improving the accuracy of semantic segmentation. This method has significant advantages when processing remote sensing images with complex backgrounds and more noise, and can effectively improve the accuracy and robustness of road semantic segmentation.

[0118] Further, refer to Figure 2 As shown, the present invention also provides a remote sensing image road semantic segmentation network based on polarization self-attention feature enhancement, and applies a remote sensing image road semantic segmentation method based on polarization self-attention feature enhancement in the above embodiment to realize remote sensing image road semantic segmentation. The network includes:

[0119] Image preprocessing and feature extraction module, used for preprocessing and feature extraction of original road remote sensing images;

[0120] The polarized self-attention feature enhancement module is used to generate channel self-attention and spatial self-attention through convolution and reshaping operations, respectively, and dynamically weight and adjust the extracted features in combination with Softmax and Sigmoid activation functions to obtain enhanced feature maps;

[0121] The road semantic segmentation module is used to perform decoding operations, restore the spatial resolution of the enhanced feature map through upsampling and convolution operations, and output a segmentation result with the same size as the original road remote sensing image.

[0122] The network provided in the embodiment of the present invention has the same implementation principle and technical effects as those in the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the system embodiment, reference can be made to the corresponding contents in the aforementioned method embodiment, which will not be repeated here.

[0123] In addition, refer to Figure 4 As shown, an embodiment of the present invention further provides an electronic device, which may include a processor 10, a memory 11, a communication bus 12 and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, wherein the processor executes the computer program to implement a remote sensing image road semantic segmentation method based on polarization self-attention feature enhancement in the above method embodiment.

[0124] The processor 10 may be composed of an integrated circuit in some embodiments, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips. The processor 10 is the control core of the electronic device, and uses various interfaces and lines to connect various components of the entire electronic device, and executes various functions of the electronic device and processes data by running or executing programs or modules stored in the memory 11, and calling data stored in the memory 11.

[0125] The memory 11 may be, for example, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples of storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (RAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, and any suitable combination thereof.

[0126] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, network systems, electronic devices or computer program products, etc. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0127] It should be noted that the word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several distinct components, and by means of a suitably programmed computer.

[0128] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0129] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for road semantic segmentation in remote sensing images based on polarization self-attention feature enhancement, characterized in that: The method comprises the following steps: Step 1: Preprocess and extract features of the original road remote sensing image; Step 2: Construct a polarized self-attention feature enhancement module, generate channel self-attention and spatial self-attention respectively through convolution and reshaping operations, combine the activation function, dynamically weight and adjust the extracted features, and obtain an enhanced feature map; Step 3: Perform a decoding operation to restore the spatial resolution of the enhanced feature map through upsampling and convolution operations, and output a segmentation result with the same size as the original road remote sensing image.

2. The method for remote sensing image road semantic segmentation based on polarization self-attention feature enhancement according to claim 1, characterized in that: In step 1, the original road remote sensing image is preprocessed and feature extracted, and the specific process includes: Step 1.1: First, preprocess the input image, including normalization and denoising; then extract the basic feature map through a convolutional layer; Step 1.2: Apply an activation function to the basic feature map to introduce nonlinearity into the network and obtain a second feature map; Step 1.3: Through the downsampling operation, the spatial resolution of the second feature map is reduced to aggregate context information and reduce the amount of calculation to obtain the third feature map; Step 1.4: Repeat the operations of steps 1.2 and 1.3 to extract deeper abstract features.

3. The method for remote sensing image road semantic segmentation based on polarization self-attention feature enhancement according to claim 1, characterized in that: In step 2, the specific process of the polarization self-attention feature enhancement module performing feature enhancement includes: Assume the input feature map is Where C represents the number of channels of the input image, H and W represent the height and width of the image respectively; ①Construct channel self-attention: The input feature map Perform convolution and reshape operations to obtain in, Reshape means reshaping operation, that is, changing the shape of tensor; Conv 1×1 Represents a 1×1 convolutional layer; right Using convolutional layers, reshape operations, and activation functions, we get in, Softmax represents the softmax activation function; Will and After matrix multiplication, the convolution layer, batch normalization layer and activation function are used to obtain channel self-attention. in, Sigmoid represents the Sigmoid activation function; LayerNorm represents the normalization layer; ②Construct spatial self-attention: right Perform convolution, global average pooling and reshaping operations to obtain in, GAP represents the global average pooling operation; right Using convolutional layers, reshape operations, and activation functions, we get in, Softmax represents the softmax activation function; Will and By performing matrix multiplication, reshaping and activation function, we can get spatial self-attention in, ③ Channel self-attention Spatial Self-Attention Respectively with the input feature map Multiply to obtain enhanced feature map in, ⊙ represents element-wise multiplication; Represents matrix addition.

4. The method for remote sensing image road semantic segmentation based on polarization self-attention feature enhancement according to claim 1, characterized in that: In step 3, a decoding operation is performed to restore the spatial resolution of the enhanced feature map through upsampling and convolution operations, and a segmentation result having the same size as the original road remote sensing image is output, which specifically includes: Feature map upsampling: The spatial resolution of the feature map is gradually restored through upsampling operations, making the feature map gradually approach the size of the original image: Convolution and activation: Apply multiple convolutional layers and activation functions to the upsampled feature map to extract finer-grained features; Fusion of multi-scale features: The upsampled feature map obtained in the decoding process is fused with the low-level feature map in the encoding stage to form a multi-scale feature map; Output segmentation results: The multi-scale feature maps are passed through several convolutional layers and activation functions, and finally the segmentation results are output.

5. The method for remote sensing image road semantic segmentation based on polarization self-attention feature enhancement according to claim 1, characterized in that: In step 3, a composite loss function is constructed by combining cross entropy loss, Dice coefficient loss and boundary perception loss to perform remote sensing image road semantic segmentation training.

6. The method for remote sensing image road semantic segmentation based on polarization self-attention feature enhancement according to claim 5, characterized in that: The constructed composite loss function is: in, represents compound loss; represents the cross entropy loss; Represents Dice coefficient loss; represents the boundary perception loss; α, β, and γ represent the weight coefficients of each loss, respectively.

7. A remote sensing image road semantic segmentation network based on polarization self-attention feature enhancement, characterized in that: A remote sensing image road semantic segmentation method based on polarization self-attention feature enhancement as described in any one of claims 1 to 6 is applied to realize remote sensing image road semantic segmentation, and the network includes: Image preprocessing and feature extraction module, used for preprocessing and feature extraction of original road remote sensing images; The polarized self-attention feature enhancement module is used to generate channel self-attention and spatial self-attention through convolution and reshaping operations, respectively, and dynamically weight and adjust the extracted features in combination with Softmax and Sigmoid activation functions to obtain enhanced feature maps; The road semantic segmentation module is used to perform decoding operations, restore the spatial resolution of the enhanced feature map through upsampling and convolution operations, and output a segmentation result with the same size as the original road remote sensing image.

8. An electronic device, characterized in that: It includes a processor and a memory, the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement a remote sensing image road semantic segmentation method based on polarization self-attention feature enhancement as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Remote sensing image road segmentation method combining super-resolution and attention mechanism

    CN113888550A

  • Remote sensing image road segmentation method fusing multi-scale features and double attention mechanism

    CN117078943A

  • Remote sensing image semantic segmentation method based on spatial detail perception and attention guidance

    CN117274608A

  • Remote sensing image urban water area segmentation method based on attention regulation and related device

    CN118691809A

  • Remote sensing image semantic segmentation method based on Mama and Transform architecture fusion

    CN119206227A