Ground feature semantic segmentation method and system based on local-global remote sensing image
By introducing spatial attention and channel attention mechanisms into semantic segmentation of aircraft remote sensing images, capturing local details and global features, and performing feature fusion, the problem of poor semantic segmentation performance in complex environments is solved, and higher accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510114469.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The semantic segmentation task of aircraft remote sensing images faces problems such as complex surface environment, diverse landform types and noise interference, resulting in poor semantic segmentation performance.
The local-global semantic segmentation method of remote sensing images is adopted to introduce spatial attention and channel attention mechanisms. Through the channel attention mechanism and spatial attention mechanism, local details and global features in the image are captured and fused to improve the accuracy and robustness of semantic segmentation.
It improves the accuracy and robustness of semantic segmentation of buildings, roads and water bodies in aircraft remote sensing images, maintains high semantic segmentation performance in complex environments, and enhances the generalization ability of the model.
Smart Images

Figure CN120047685A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and more specifically, to a method and system for land object semantic segmentation based on local-global remote sensing images. Background Art
[0002] With the continuous development of remote sensing technology, aerial remote sensing images are increasingly widely used in fields such as geographic information systems, environmental monitoring, urban planning, and agricultural management. Aerial remote sensing images are obtained by remote sensors carried by aircraft, and have the advantages of high resolution, wide coverage, and fast acquisition speed, and can provide rich surface information. However, the land object semantic segmentation task of aerial remote sensing images faces many challenges, such as complex surface environments, diverse land object types, noise interference, etc., and these factors will affect the accuracy and efficiency of semantic segmentation.
[0003] Traditional methods for land object semantic segmentation of aerial remote sensing images mainly rely on manually designed feature extraction algorithms when performing semantic segmentation of buildings, roads, or water bodies, such as texture features, color features, shape features, etc. Although these methods perform well in certain specific scenarios, they are often difficult to capture all the information in the image in complex environments, resulting in poor semantic segmentation performance. In addition, the feature extraction process for buildings, roads, or water bodies by traditional methods usually requires a large amount of manual participation, with low efficiency and difficulty in meeting the requirements of large-scale data processing.
[0004] In recent years, with the rapid development of deep learning technology, convolutional neural networks (CNNs) have demonstrated powerful capabilities in the task of land object semantic segmentation of images. CNNs automatically extract features in images of buildings, roads, or water bodies through multiple convolutional operations, and can capture high-level information in the image, thereby significantly improving the accuracy of semantic segmentation. However, there are still some limitations in using only the global feature extraction method of CNNs when processing aerial remote sensing images. The land objects in aerial remote sensing images usually have multi-scale features, and it is difficult to capture these detailed information only relying on global features, resulting in limited performance of semantic segmentation of buildings, roads, or water bodies.
[0005] To solve the above problems, researchers have proposed various improvement methods. For example, by introducing an attention mechanism, the model can pay more attention to important regions in the image, thereby improving the effect of feature extraction for images of buildings, roads, or water bodies. The attention mechanism mainly includes two forms: spatial attention and channel attention. The spatial attention mechanism can capture local details in the image, such as edges and corners, thus improving the accuracy of semantic segmentation; the channel attention mechanism can highlight important features of different channels, such as colors and textures, thereby enhancing the discrimination ability of features. However, existing attention mechanisms usually only focus on local or global information and are difficult to comprehensively utilize local and global information, resulting in incomplete feature extraction.
[0006] In addition, the task of ground object semantic segmentation for aircraft remote sensing images also faces the problem of robustness in complex environments. For example, interference factors such as shadows and cloud cover will affect the quality of the image, thereby reducing the accuracy of ground object semantic segmentation. Traditional methods usually have difficulty effectively dealing with these interferences, resulting in a decline in the performance of the model in complex environments.
[0007] Therefore, how to extract multi-scale features and maintain high ground object semantic segmentation performance in complex environments is an urgent problem to be solved by those skilled in the art in the task of aircraft remote sensing image semantic segmentation. Summary of the Invention
[0008] In view of this, the present invention provides a method and system for ground object semantic segmentation based on local-global remote sensing images, introducing spatial attention and channel attention mechanisms to more accurately capture local details and important features in the image, thereby improving the accuracy and robustness of remote sensing image semantic segmentation.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] In a first aspect, the present invention proposes a method for ground object semantic segmentation based on local-global remote sensing images, including the following steps:
[0011] S1. Obtain the remote sensing image to be segmented in the target area for preprocessing, and perform feature extraction through a convolutional neural network to obtain a feature map F;
[0012] S2. According to the global information of the feature map F, adopt a channel attention mechanism and a spatial attention mechanism to obtain a globally enhanced feature;
[0013] S3. Divide the feature map F into multiple feature vectors, adopt the channel attention mechanism and the spatial attention mechanism for each feature vector and re-integrate them to obtain a complete locally enhanced feature;
[0014] S4. Fuse the global enhanced feature and the complete local enhanced feature to obtain an enhanced feature F'.
[0015] S5. Input the enhanced feature F' into the ground object semantic segmentation model to obtain a ground object semantic segmentation result P; the ground object semantic segmentation result P includes buildings, roads, and water bodies.
[0016] Further, in step S1, obtain a remote sensing image to be segmented in the target area, perform preprocessing, and extract features through a convolutional neural network to obtain a feature map F; specifically including:
[0017] Obtain a remote sensing image to be segmented in the target area, perform image denoising, perform image enhancement on the denoised image, and then perform image normalization.
[0018] Segment the normalized image, remove the background part in the segmented image, and use the remaining segmented image as the original image.
[0019] Input the original image into an initial convolutional neural network to obtain low-level features.
[0020] Perform non-linear transformation, pooling layer, and batch normalization layer processing on the low-level features to obtain a feature F0.
[0021] Perform multi-level feature extraction on the feature F 0 to obtain high-level features as the feature map F; where the shape of the feature map F is C×H×W, C is the number of channels, H is the height, and W is the width.
[0022] Further, in step S2, according to the global information of the feature map F, adopt a channel attention mechanism and a spatial attention mechanism to obtain a global enhanced feature; specifically including:
[0023] Adopt a channel attention mechanism, perform average pooling and max pooling on each channel of the feature map F, splice them to obtain a global channel attention map; multiply it with the global matrix of the feature map F to obtain an enhanced feature of the global channel.
[0024] Adopt a spatial attention mechanism, perform average pooling and max pooling on the enhanced feature of the global channel in the channel dimension, splice them to obtain a global spatial attention map; then multiply it with the enhanced feature of the global channel to obtain a global enhanced feature.
[0025] Further, the adopting of the channel attention mechanism to perform average pooling and max pooling on each channel of the feature map F specifically includes:
[0026] Using adaptive average pooling, compress the feature map of each channel into a single average value to obtain a channel average pooling tensor with a shape of C×1×1;
[0027] Using adaptive max pooling, compress the feature map of each channel into a single maximum value to obtain a channel max pooling tensor with a shape of C×1×1;
[0028] Perform convolution, activation, and convolution operations on the channel average pooling tensor and the channel max pooling tensor in sequence to splice and compress the number of channels.
[0029] Furthermore, the spatial attention mechanism is adopted to perform average pooling and max pooling on the enhanced features of the global channels in the channel dimension, specifically including:
[0030] Perform average pooling on the enhanced features of the global channels in the channel dimension to obtain a spatial average pooling tensor with a shape of 1×H×W;
[0031] At the same time, perform max pooling in the channel dimension to obtain a spatial max pooling tensor with a shape of 1×H×W;
[0032] Perform convolution, activation, and convolution operations on the spatial average pooling tensor and the spatial max pooling tensor in sequence to splice and compress the number of channels.
[0033] Furthermore, in step S3, divide the feature map F into multiple feature vectors, and adopt the channel attention mechanism and the spatial attention mechanism for each feature vector and re-integrate them to obtain complete local enhanced features; specifically including:
[0034] Divide the feature map F into multiple feature vectors f along the height and width directions i , where i is a positive integer greater than 1;
[0035] Adopt the channel attention mechanism to perform average pooling and max pooling on each channel of the feature vector f i , and obtain a channel weight vector w after splicing ci ; multiply with the matrix of the feature vector f i to obtain enhanced features of different channels;
[0036] Adopt the spatial attention mechanism to perform average pooling and max pooling on the enhanced features of different channels in the channel dimension, and obtain a spatial weight vector after splicing; then multiply with the enhanced features of different channels to obtain local enhanced features;
[0037] Based on the spatial position encoding information during the division of the feature map F, re-integrate the local enhanced features to obtain complete local enhanced features.
[0038] Further, in step S5, input the enhanced feature F' into the ground object semantic segmentation model to obtain the ground object semantic segmentation result P, which specifically includes:
[0039] Input the enhanced feature F' into the ground object semantic segmentation model, perform global average pooling operation on each feature map of the enhanced feature F' to obtain a feature map compressed into a single average value;
[0040] Perform a flattening operation on the compressed feature map to obtain the corresponding one-dimensional vector;
[0041] Input the one-dimensional vector into the fully connected layer, perform linear transformation and non-linear activation, and then apply the Softmax activation function to obtain the final ground object semantic segmentation result P.
[0042] Further, step S5 further includes:
[0043] Adopt cross-entropy loss to evaluate the difference between the semantic segmentation result and the true label as the first loss function;
[0044] Apply L2 regularization loss to the global channel attention map, global spatial attention map, channel weight vector, and spatial weight vector as the second loss function;
[0045] Calculate the overlap between the predicted region and the true label region through the Dice coefficient, threshold the probability, and obtain the Dice loss as the third loss function;
[0046] Combine the first loss function, the second loss function, and the third loss function to obtain a composite loss function and optimize the ground object semantic segmentation model.
[0047] In a second aspect, the present invention proposes a ground object semantic segmentation system based on local-global remote sensing images, including the following modules:
[0048] Acquisition module: used to acquire the remote sensing image to be segmented in the target area for preprocessing, and extract features through a convolutional neural network to obtain a feature map F;
[0049] Global information-guided feature enhancement module: used to obtain global enhanced features according to the global information of the feature map F by adopting a channel attention mechanism and a spatial attention mechanism;
[0050] Local information-guided feature enhancement module: used to divide the feature map F into multiple feature vectors, adopt the channel attention mechanism and the spatial attention mechanism for each feature vector and re-integrate them to obtain a complete local enhanced feature;
[0051] Enhanced Feature Fusion Module: It is used to fuse the global enhanced feature and the complete local enhanced feature to obtain the enhanced feature F'.
[0052] Geometric Object Semantic Segmentation Module: It is used to input the enhanced feature F' into the geometric object semantic segmentation model to obtain the geometric object semantic segmentation result P; the geometric object semantic segmentation result P includes buildings, roads, and water bodies.
[0053] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a method and system for geometric object semantic segmentation based on local-global remote sensing images, having the following beneficial effects:
[0054] 1) By introducing the spatial attention and channel attention mechanisms, the present invention can more accurately capture the local details and important features of buildings, roads, and water bodies in the image, thereby improving the accuracy and robustness of semantic segmentation of buildings, roads, and water bodies in aerial remote sensing images.
[0055] 2) By fusing local and global information, the module can maintain a high semantic segmentation performance in complex environments and improve the generalization ability of the model.
[0056] 3) Finally, through feature fusion, comprehensively using feature information of different scales and levels, the overall performance of the semantic segmentation model for buildings, roads, and water bodies is further improved. Description of the Drawings
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0058] Figure 1 It is a flowchart of a method for geometric object semantic segmentation based on local-global remote sensing images provided by an embodiment of the present invention.
[0059] Figure 2 It is a schematic diagram of a module for local-global information-guided feature enhancement provided by an embodiment of the present invention.
[0060] Figure 3 It is an architecture diagram of a neural network model provided by an embodiment of the present invention.
[0061] Figure 4 It is a framework diagram of a system for geometric object semantic segmentation based on local-global remote sensing images provided by an embodiment of the present invention. Detailed Embodiments
[0062] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0063] Embodiment 1
[0064] An embodiment of the present invention discloses a method for semantic segmentation of ground objects based on local-global remote sensing images. Referring to Figure 1 as shown, it includes the following steps S1 to S5:
[0065] S1. Obtain the remotely sensed image to be segmented in the target area for preprocessing, and perform feature extraction through a convolutional neural network to obtain a feature map F;
[0066] S2. According to the global information of the feature map F, adopt a channel attention mechanism and a spatial attention mechanism to obtain a globally enhanced feature;
[0067] S3. Divide the feature map F into multiple feature vectors, adopt the channel attention mechanism and the spatial attention mechanism for each feature vector and re-integrate them to obtain a complete locally enhanced feature;
[0068] S4. Fuse the globally enhanced feature and the complete locally enhanced feature to obtain an enhanced feature F';
[0069] S5. Input the enhanced feature F' into the ground object semantic segmentation model to obtain a ground object semantic segmentation result P; the ground object semantic segmentation result P includes buildings, roads, and water bodies.
[0070] In the embodiment of the present invention, according to the above steps, first obtain a remotely sensed image over the city. For example, the image includes clouds, buildings, roads, water bodies, vegetation, vehicles, pedestrians, etc., perform image processing and feature extraction to obtain a feature map F.
[0071] Secondly, referring to Figure 2As shown in the figure, by introducing the spatial attention (SA) and channel attention (CA) mechanisms, the local details and important features in the feature map F are captured more accurately. This enables the semantic segmentation model of ground objects to more accurately identify and semantically segment various targets in aircraft remote sensing images, such as buildings, roads, water bodies, etc. By fusing local and global information, the module can maintain high semantic segmentation performance in complex environments. For example, for interference factors such as shadows and cloud occlusions that are common in remote sensing images, the module can reduce the impact of these interferences through multi-scale feature learning and attention mechanisms, thereby improving the robustness of the model. The local attention mechanism can accurately capture the boundary information of the target, which is particularly important for the recognition of small-scale targets (such as small buildings, vehicles, etc.) in remote sensing images. By refining the target boundaries, the module can improve the recognition ability of the semantic segmentation model for small-scale targets. By comprehensively utilizing local and global information, the module can improve the generalization ability of the model, so that it can maintain high semantic segmentation performance on different scenes and data sets. This is of great significance for the diversity and complexity of the semantic segmentation task of remote sensing images.
[0072] After the semantic segmentation of the land object is performed using the method of the present invention, each pixel can be assigned to the land object category to which it belongs. Through the semantic segmentation of land objects, an image labeled with various land object categories can be obtained, in which each pixel is assigned a category label. This fine-level classification can provide basic data for subsequent quantitative analysis, such as calculating the proportion of various types of land use, estimating the building area, and evaluating urban expansion trends. In addition, high-quality land object semantic segmentation results are also the key support for realizing intelligent decision-making in multiple fields such as precision agriculture, environmental protection, and disaster response.
[0073] The following are detailed descriptions of each of the above steps:
[0074] Step S1 specifically includes:
[0075] Acquire remote sensing images over the city, such as images including clouds, buildings, roads, water bodies, vegetation, vehicles and pedestrians, and perform denoising, enhancement and normalization on the remote sensing images to improve image quality and reduce noise interference.
[0076] Regarding the denoising operation, the input aircraft remote sensing image is denoised using methods such as median filtering to reduce the impact of noise on subsequent processing. Denoising can improve the clarity of the image, reduce noise interference, and thus improve the accuracy of feature extraction.
[0077] Image enhancement is performed on the denoised image. The contrast and brightness of the image are enhanced through methods such as histogram equalization and contrast adjustment. Image enhancement can highlight important details in the image and make feature extraction more effective. Finally, image normalization is performed to normalize the pixel values of the image to a fixed range.
[0078] Next, the image is separated into multiple regions of interest through image segmentation, and background noise is removed.
[0079] The normalized image is segmented using methods such as threshold segmentation, edge detection, or region growing to divide the image into multiple regions of interest (ROIs). Image segmentation can separate different target regions in the image, facilitating subsequent feature extraction and semantic segmentation.
[0080] The background part of the image is removed through methods such as background modeling or background subtraction, and the regions of interest are retained. Background removal can reduce the impact of background noise on feature extraction and improve the accuracy of feature extraction.
[0081] Furthermore, as shown in Figure 3 a convolutional neural network (CNN) is used for preliminary feature extraction. Through convolutional, activation, pooling, and batch normalization operations, high-level features in the image are gradually extracted.
[0082] First, the original image I is input into the initial convolutional layer of the convolutional neural network (CNN). In this embodiment, a relatively small convolutional kernel (3x3) is used in the initial convolutional layer to capture low-level features in the image, such as edges and textures.
[0083] Next, after the convolutional operation, an activation function (such as ReLU) is applied to perform a non-linear transformation on the convolutional result to introduce non-linearity. Then, after the activation function, a pooling layer (such as max pooling or average pooling) is applied to downsample the feature map to reduce the size of the feature map while retaining important features. After the pooling layer, a batch normalization layer can be introduced to normalize the feature map to accelerate the training process and improve the stability of the model, obtaining the feature F 0 :
[0084]
[0085] where Conv represents the convolutional operation, ReLU represents the rectified linear unit activation function, Pool represents the pooling operation, and BatchNormalization represents the batch normalization layer.
[0086] Finally, multi-layer feature extraction is performed. The above-mentioned convolutional, activation, pooling, and batch normalization operations are repeated to construct a multi-layer convolutional neural network, and high-level features F in the image are gradually extracted.
[0087] Step S2 specifically includes:
[0088] First, the channel attention mechanism and the spatial attention mechanism are described based on the input feature map F.
[0089] (1)Regarding the channel attention mechanism, for the input feature map F, which is a feature map with the shape of C×H×W; where C represents the number of channels, H is the height, and W is the width.
[0090] First, use adaptive average pooling to compress the feature map of each channel into a single average value, obtaining a tensor with the shape of C×1×1. Meanwhile, use adaptive max pooling to also compress the feature map of each channel into a single maximum value, obtaining a tensor with the shape of C×1×1.
[0091] Next, combine the above two tensors. Use convolution (1x1 convolution) to compress the number of channels. This operation reduces the number of channels to reduce the computational load. After convolution, apply the ReLU activation function to increase non-linearity. Subsequently, use another 1x1 convolution to restore the number of channels to the original scale, obtaining the final channel attention feature map.
[0092] Finally, add and in the channel dimension to obtain which fuses the feature information from average pooling and max pooling. Then apply the Sigmoid activation function to map the output to the range [0,1], obtaining a channel attention map with the shape of C×1×1. This attention map represents the importance weights of each channel.
[0093] (2)Regarding the spatial attention mechanism, for the input feature map F, which is a feature map with the shape of C×H×W; where C represents the number of channels, H is the height, and W is the width.
[0094] First, perform average pooling on the input feature map F in the channel dimension (dim = 1), obtaining a tensor with the shape of 1×H×W. It contains the channel mean of each position; meanwhile, perform max pooling on the input feature map F in the channel dimension (dim = 1), obtaining a tensor with the shape of 1×H×W. It contains the maximum value of each position.
[0095] Next, concatenate the results of average pooling and max pooling, concatenate and in the channel dimension (dim = 1), obtaining a tensor with the shape of 2×H×W. It contains the average pooling and max pooling information of each spatial position. Then, a convolution operation is performed, and the concatenated feature map is processed using a convolutional layer with a kernel size of 7. This convolutional layer compresses the number of input channels from 2 to 1 and outputs a tensor with a shape of 1×H×W.
[0096] Finally, the Sigmoid activation function is applied to map the output to the range [0,1], representing the attention weights of each spatial position. The output is an attention map with a shape of 1×H×W for weighting the input feature map.
[0097] In this embodiment, a channel attention mechanism is adopted. Average pooling and max pooling are performed on each channel of the feature map F, and after concatenation, a global channel attention map is obtained which is multiplied by the global matrix of the feature map F to obtain the enhanced feature of the global channel
[0098] The global information of the feature map F The feature undergoes average pooling and max pooling on each channel through the channel attention mechanism to obtain The channel attention mechanism can highlight the important features of different channels, such as color, texture, etc., thereby enhancing the discrimination ability of the features. This is particularly important for semantic segmentation of complex scenes in remote sensing images. Then, the channel attention is applied to the features of the global information guidance branch to obtain the features after channel enhancement
[0099]
[0100] where, represents the matrix multiplication operation. By applying the channel attention weights, the important channel information in the global features can be highlighted, thereby improving the expression ability of the features.
[0101] Next, through the spatial attention mechanism, average pooling and max pooling are performed in the channel dimension, and after concatenation, a global spatial attention map is obtained Then the spatial attention is applied to the feature to obtain the enhanced feature after enhancing the global information guidance feature as
[0102]
[0103] where, Represents a matrix multiplication operation. By comprehensively utilizing channel attention and spatial attention, the module can improve the generalization ability of the model, enabling it to maintain high semantic segmentation performance on different scenarios and datasets. This is of great significance for the diversity and complexity in the semantic segmentation task of remote sensing images.
[0104] Step S3 specifically includes:
[0105] The feature map F is divided into n equally sized small block feature vectors f along the height H and width W directions i , where i = 1, 2, 3...n; By dividing the feature map, local features can be processed more meticulously, thereby improving the recognition ability for small-scale targets.
[0106] Next, for obtaining the channel weight vector, each feature vector f is learned through the channel attention mechanism i , and average pooling and max pooling are performed on each channel of the feature vector f i , and after concatenation, the channel weight vector w is obtained ci .
[0107] The channel attention of each small block feature is applied to the small block feature to obtain
[0108]
[0109] where, Represents a matrix multiplication operation. Through the channel weight vector, the important features of different channels can be highlighted, thereby enhancing the discrimination ability of the features.
[0110] Then, by applying the spatial attention mechanism to , average pooling and max pooling are performed in the channel dimension, and after concatenation, the spatial weight vector is obtained Then, feature enhancement is performed on to obtain
[0111]
[0112] where, Represents a matrix multiplication operation. Such an operation can capture local details in the image, thereby improving the accuracy of semantic segmentation.
[0113] Finally, based on the spatial position encoding information during feature map segmentation, all the small block features are re-integrated to obtain the feature with complete local information-guided feature enhancement
[0114] Step S4, the global enhanced feature and the complete local enhanced feature Perform fusion to obtain the enhanced feature F'.
[0115]
[0116] Among them, F' is the result after the final fusion.
[0117] By fusing global and local features, it is possible to comprehensively utilize feature information at different scales and levels, thereby improving the overall performance of the semantic segmentation model. The fused features can more comprehensively express various feature information in the image, including global structure and local details. This comprehensive feature expression ability can significantly improve the adaptability of the semantic segmentation model to complex scenes, thereby improving the accuracy and robustness of semantic segmentation.
[0118] Step S5 specifically includes:
[0119] Input the enhanced feature F' into the ground object semantic segmentation model. In this embodiment, refer to Figure 3 As shown, first, enter the global average pooling layer, perform global average pooling operation on each feature map, and compress each feature map into a single average value. Then, perform a flattening operation on the feature map after global average pooling to convert it into a one-dimensional vector. Input the flattened feature vector into the fully connected layer for linear transformation and non-linear activation to obtain the ground object semantic segmentation result.
[0120] After the fully connected layer, apply the Softmax activation function to convert the output into a probability distribution to obtain the final semantic segmentation result P:
[0121]
[0122] Among them, GlobalAvgPool represents the global average pooling operation, Flatten represents the flattening operation, FC represents the fully connected layer, Softmax represents the Softmax activation function, and P is the final semantic segmentation probability distribution.
[0123] Step S5 also includes constructing a composite loss function to optimize the ground object semantic segmentation model.
[0124] In this embodiment, first, use the cross-entropy loss as the semantic segmentation loss. The cross-entropy loss can measure the difference between the semantic segmentation result and the true label, guiding the training process of the model.
[0125] The semantic segmentation loss is used as the first loss function; as follows:
[0126]
[0127] Among them, yj is the true label input to the model, is the predicted probability of the model, and N is the number of samples.
[0128] Next, apply the L2 regularization loss to the weights generated by channel attention and spatial attention. The L2 regularization loss can prevent the model from overfitting and improve the generalization ability of the model;
[0129] The L2 regularization loss is used as the second loss function; as follows:
[0130]
[0131] where λ is the regularization coefficient and Γ is all the weights of the model.
[0132] In this embodiment, the output probability represents the probability that the model predicts each pixel belongs to a certain class (usually the foreground class), and the range is between [0, 1]. The true label represents the actual label, usually 0 or 1, indicating the background or foreground respectively.
[0133] The foreground class refers to the object or region of interest in the image. In this embodiment, the foreground may include buildings, vegetation, vehicles, and pedestrians; simply put, the foreground is the main object to be recognized and extracted.
[0134] The background class refers to all other parts except the foreground. In this embodiment, it may include clouds, roads, water bodies, etc.; it represents the area that is not directly involved in the main analysis or processing.
[0135] This embodiment also obtains the third loss function through probability thresholding; since the Dice coefficient calculates the overlap between the predicted region and the true label region, it is necessary to convert the output probability value into a binary prediction result when calculating If the output probability p is greater than a certain threshold (e.g., 0.5), then it is considered that the pixel belongs to the foreground (1); otherwise, it belongs to the background (0).
[0136] The Dice loss is used as the third loss function; as follows:
[0137]
[0138] Combine the above three loss functions to construct a composite loss function:
[0139]
[0140] where α is the weight coefficient used to balance the contributions of different loss functions. The composite loss function can comprehensively consider the contributions of multiple loss functions and improve the semantic segmentation performance and robustness of the model.
[0141] Embodiment 2
[0142] An embodiment of the present invention discloses a ground object semantic segmentation system based on local-global remote sensing images. Referring to Figure 4 as shown, it includes the following modules:
[0143] Acquisition module: It is used to acquire the remote sensing image to be segmented in the target area for preprocessing, and perform feature extraction through a convolutional neural network to obtain a feature map F;
[0144] Global information-guided feature enhancement module: It is used to obtain global enhanced features by adopting a channel attention mechanism and a spatial attention mechanism according to the global information of the feature map F
[0145] Local information-guided feature enhancement module: It is used to divide the feature map F into multiple feature vectors, and adopt a channel attention mechanism and a spatial attention mechanism for each feature vector and re-integrate them to obtain a complete local enhanced feature
[0146] Enhanced feature fusion module: It is used to fuse the global enhanced feature and the complete local enhanced feature to obtain an enhanced feature F';
[0147] Ground object semantic segmentation module: It is used to input the enhanced feature F' into a ground object semantic segmentation model to obtain a ground object semantic segmentation result P; the ground object semantic segmentation result P includes buildings, roads and water bodies.
[0148] In this embodiment, the acquisition module acquires a remote sensing image for preprocessing according to the method of Embodiment 1, and performs feature extraction through a convolutional neural network to obtain a feature map F with a shape of C×H×W.
[0149] The feature input into the global information-guided feature enhancement module is According to the global information of the feature map F, a channel attention mechanism and a spatial attention mechanism are adopted to obtain global enhanced features
[0150] The feature map F is divided into multiple feature vectors f i , where i = 1, 2, 3...n, and input into the local information-guided feature enhancement module. A channel attention mechanism and a spatial attention mechanism are adopted for each feature vector and re-integrated to obtain a complete local enhanced feature
[0151] Through the enhanced feature fusion module, the global enhanced feature and the complete local enhanced feature Fuse them to obtain the enhanced feature F'; and input it into the ground object semantic segmentation module for prediction to obtain the ground object semantic segmentation result P.
[0152] By introducing the spatial attention and channel attention mechanisms, the embodiments of the present invention can capture local details and important features in the image more precisely. This enables the semantic segmentation model to more accurately identify and semantically segment various targets in the aircraft remote sensing image, such as buildings, roads, water bodies, etc.
[0153] By fusing local and global information, the module can maintain a high semantic segmentation performance in complex environments. For example, in the presence of interference factors such as shadows and cloud occlusions commonly found in remote sensing images, the module can reduce the impact of these interferences through multi-scale feature learning and attention mechanisms, thereby improving the robustness of the model.
[0154] By comprehensively utilizing local and global information, the present invention can improve the generalization ability of the model, enabling it to maintain a high semantic segmentation performance in different scenarios and datasets. This is of great significance for the diversity and complexity in the remote sensing image semantic segmentation task. In addition, through feature fusion, by comprehensively using feature information of different scales and levels, the overall performance of the semantic segmentation model can be further enhanced. The method of the present invention not only improves the accuracy of semantic segmentation, but also reduces the resource requirements for model training and deployment through parameter-efficient fine-tuning and the use of adapters, making it applicable to remote sensing image processing tasks in various computing environments.
[0155] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method section.
[0156] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for semantic segmentation of ground objects based on local-global remote sensing images, characterized in that: The following steps are involved: S1, obtain the remote sensing image to be segmented in the target area for preprocessing, and extract features through a convolutional neural network to obtain a feature map F; S2. According to the global information of the feature map F, a channel attention mechanism and a spatial attention mechanism are used to obtain a global enhanced feature; S3, dividing the feature map F into multiple feature vectors, applying the channel attention mechanism and the spatial attention mechanism to each feature vector and reintegrating them to obtain a complete local enhanced feature; S4, fusing the global enhancement feature and the complete local enhancement feature to obtain an enhancement feature F'; S5. Input the enhanced feature F' into a ground object semantic segmentation model to obtain a ground object semantic segmentation result P; the ground object semantic segmentation result P includes buildings, roads and water bodies.
2. The method for semantic segmentation of ground objects based on local-global remote sensing images according to claim 1, characterized in that: Step S1, obtaining a remote sensing image to be segmented of the target area for preprocessing, and extracting features through a convolutional neural network to obtain a feature map F; specifically comprising: Obtain the remote sensing image to be segmented in the target area, perform image denoising, perform image enhancement processing on the denoised image, and then perform image normalization processing; Segment the normalized image, remove the background in the segmented image, and use the remaining segmented image as the original image; Inputting the original image into an initial convolutional neural network to obtain low-level features; The low-level features are processed by nonlinear transformation, pooling layer and batch normalization layer to obtain feature F 0 ; For the feature F 0 Perform multi-layer feature extraction to obtain high-level features as a feature map F; wherein the shape of the feature map F is C×H×W, where C is the number of channels, H is the height, and W is the width.
3. The method for semantic segmentation of land objects based on local-global remote sensing images according to claim 2, characterized in that: Step S2: According to the global information of the feature map F, a channel attention mechanism and a spatial attention mechanism are used to obtain a global enhanced feature; Specifically include: Using the channel attention mechanism, average pooling and maximum pooling are performed on each channel of the feature map F, and the global channel attention map is obtained after splicing; and the enhanced features of the global channel are obtained by multiplying the global matrix of the feature map F; The spatial attention mechanism is adopted to perform average pooling and maximum pooling on the enhanced features of the global channel in the channel dimension, and the global spatial attention map is obtained after splicing; and then the global enhanced features are multiplied with the enhanced features of the global channel to obtain the global enhanced features.
4. The method for semantic segmentation of ground objects based on local-global remote sensing images according to claim 3, characterized in that: The channel attention mechanism is adopted to perform average pooling and maximum pooling on each channel of the feature map F, specifically including: Use adaptive average pooling to compress the feature map of each channel into a single average value, obtaining a channel average pooling tensor of shape C×1×1; Use adaptive max pooling to compress the feature map of each channel into a single maximum value, and obtain a channel max pooling tensor with a shape of C×1×1; For the channel average pooling tensor and the channel maximum pooling tensor, convolution, activation, and re-convolution operations are performed in sequence to splice and compress the number of channels.
5. The method for semantic segmentation of ground objects based on local-global remote sensing images according to claim 4, characterized in that: The spatial attention mechanism is used to perform average pooling and maximum pooling on the enhanced features of the global channel in the channel dimension, specifically including: The enhanced features of the global channel are average pooled in the channel dimension to obtain a spatial average pooling tensor with a shape of 1×H×W; At the same time, maximum pooling is performed in the channel dimension to obtain a spatial maximum pooling tensor with a shape of 1×H×W; For the spatial average pooling tensor and the spatial maximum pooling tensor, convolution, activation, and re-convolution operations are performed in sequence to splice and compress the number of channels.
6. The method for semantic segmentation of ground objects based on local-global remote sensing images according to claim 5, characterized in that: Step S3, dividing the feature map F into multiple feature vectors, applying the channel attention mechanism and the spatial attention mechanism to each feature vector and reintegrating them to obtain a complete local enhanced feature; Specifically include: The feature map F is divided into multiple feature vectors f along the height and width directions i , where i is a positive integer greater than 1; Using the channel attention mechanism, the feature vector f i Average pooling and maximum pooling are performed on each channel of , and the channel weight vector w is obtained after splicing. ci ; and the feature vector f i The matrix multiplication of obtains the enhanced features of different channels; The spatial attention mechanism is used to perform average pooling and maximum pooling on the enhanced features of different channels in the channel dimension, and the spatial weight vector is obtained after splicing; and then the spatial weight vector is multiplied with the enhanced features of different channels to obtain the local enhanced features; Based on the spatial position encoding information when the feature map F is segmented, the local enhanced features are reintegrated to obtain complete local enhanced features.
7. The method for semantic segmentation of ground objects based on local-global remote sensing images according to claim 6, characterized in that: Step S5, inputting the enhanced feature F' into the ground object semantic segmentation model to obtain the ground object semantic segmentation result P; specifically comprising: The enhanced feature F' is input into the ground feature semantic segmentation model, and each feature map of the enhanced feature F' is subjected to a global average pooling operation to obtain a feature map compressed into a single average value; Flatten the compressed feature map to obtain the corresponding one-dimensional vector; The one-dimensional vector is input into the fully connected layer, linearly transformed and nonlinearly activated, and then the Softmax activation function is applied to obtain the final semantic segmentation result P of the object.
8. The method for semantic segmentation of ground objects based on local-global remote sensing images according to claim 7, characterized in that: Step S5 further includes: Cross entropy loss is used to evaluate the difference between the semantic segmentation result and the true label as the first loss function; Applying L2 regularization loss as a second loss function to the global channel attention map, the global spatial attention map, the channel weight vector, and the spatial weight vector; The overlap between the predicted area and the true label area is calculated by the Dice coefficient, and the probability is thresholded to obtain the Dice loss as the third loss function; The first loss function, the second loss function and the third loss function are combined to obtain a composite loss function to optimize the semantic segmentation model of the land object.
9. A ground object semantic segmentation system based on local-global remote sensing images, characterized in that: Includes the following modules: Acquisition module: used to obtain the remote sensing image to be segmented in the target area for preprocessing, and extract features through convolutional neural network to obtain feature map F; A global information guided feature enhancement module is used to obtain global enhanced features according to the global information of the feature map F by using a channel attention mechanism and a spatial attention mechanism; Local information guided feature enhancement module: used to divide the feature map F into multiple feature vectors, use channel attention mechanism and spatial attention mechanism for each feature vector and reintegrate them to obtain complete local enhanced features; Enhanced feature fusion module: used to fuse the global enhanced feature and the complete local enhanced feature to obtain the enhanced feature F'; Land object semantic segmentation module: used for inputting the enhanced feature F' into the land object semantic segmentation model to obtain the land object semantic segmentation result P; the land object semantic segmentation result P includes buildings, roads and water bodies.
Citation Information
Patent Citations
Remote sensing image road segmentation method combining intensive attention and parallel upsampling
CN114092824A
Marine remote sensing image semantic segmentation method and network based on statistical texture
CN117058374A
Method for segmenting water surface feasible region and target by using MCSSNet network
CN117409197A
Remote sensing image double-branch feature fusion solid waste identification method and system and electronic equipment
CN117765410A
Remote sensing image semantic segmentation method, device and system, and storage medium
CN118470327A