Remote sensing image segmentation method based on multi-scale feature fusion and index selection mechanism
Through the combination of multi-scale convolutional network and sparse attention mechanism, the problems of accuracy and efficiency in remote sensing image segmentation are solved, and efficient remote sensing image segmentation is achieved.
Patent Information
- Application Number
- CN202510333470.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-20
AI Technical Summary
The existing remote sensing image segmentation method has low segmentation accuracy and low computational efficiency in processing complex scenarios, making it difficult to effectively extract multi-scale information and focus on important areas.
Multi-scale convolutional network is used to extract feature information of different scales of remote sensing images, and feature aggregation is performed through index selection enhancement modules in combination with sparse attention mechanism, and decoded using a decoder to improve segmentation accuracy.
It significantly improves the accuracy of remote sensing image segmentation, reduces the demand for computing resources, and is suitable for image processing tasks in various computing environments.
Smart Images

Figure CN120279266A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and more specifically, to a remote sensing image segmentation method based on multi-scale feature fusion and index selection mechanism. Background Technique
[0002] Remote sensing image segmentation technology plays a crucial role in remote sensing image processing and analysis, and is widely used in fields such as environmental monitoring, urban planning, land use investigation, and disaster assessment. Traditional remote sensing image segmentation methods mainly rely on pixel-level feature extraction and classification techniques, including classic methods such as threshold segmentation, edge detection, and region growing. Although these methods have solved the remote sensing image segmentation problem to a certain extent, they face challenges such as low segmentation accuracy and low computational efficiency in complex scenarios.
[0003] With the rapid development of deep learning technology, remote sensing image segmentation methods based on convolutional neural networks (CNNs) have gradually become a research hotspot. Especially with the continuous optimization of network structures, many deep learning models have achieved remarkable results in remote sensing image segmentation. However, when dealing with remote sensing images, these methods still face some main problems: First, remote sensing images usually have complex backgrounds and diverse ground object types, and traditional convolutional neural networks have certain limitations in extracting local features of images. Second, although the existing attention mechanisms can capture key information in local regions, due to limited computing resources, it is difficult to efficiently process a large amount of feature information in large-scale remote sensing images. In order to improve the segmentation performance under limited computing resources, how to effectively extract multi-scale information and focus on important regions has become an important research direction in remote sensing image segmentation methods. Summary of the Invention
[0004] In view of this, the present invention provides a remote sensing image segmentation method based on multi-scale feature fusion and index selection mechanism, which can not only efficiently extract multi-scale information in remote sensing images, but also effectively focus on key regions in the images through a sparse attention mechanism, thereby significantly improving the segmentation accuracy.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] A remote sensing image segmentation method based on multi-scale feature fusion and index selection mechanism, including the following steps:
[0007] S1. Use a multi-scale convolutional network to extract different scale feature information in the remote sensing image, and obtain a fused feature map of the remote sensing image;
[0008] S2. Use an index selection enhancement module based on a sparse attention mechanism to perform feature aggregation on the fused feature map of the remote sensing image to obtain an enhanced feature map;
[0009] S3. Use a decoder to decode the enhanced feature map to obtain the remote sensing image segmentation result.
[0010] Furthermore, before step S1 uses a multi-scale convolutional network to extract different-scale features from the remote sensing image, it also includes preprocessing the collected remote sensing image, specifically including:
[0011] Perform geometric correction, radiometric correction, image enhancement and denoising, size normalization and cropping operations on the collected remote sensing image in sequence.
[0012] Furthermore, step S1 specifically includes:
[0013] S11. Input the remote sensing image into the initial convolutional layer for preliminary feature extraction to obtain the initial feature map
[0014] S12. Input the initial feature map into the spatial global average pooling layer to extract the global information of the remote sensing image and generate the global context feature map
[0015] At the same time, input the initial feature map into the windmill-shaped convolutional layer to extract the detailed information of the local area of the remote sensing image and obtain the local feature map
[0016] S13. Multiply the global context feature map and the local feature map element-wise, and add the multiplication result to the local feature map element-wise to generate the initial fusion feature map
[0017] S14. Input the initial fusion feature map into the convolutional layer, ReLu activation function layer, normalization layer, convolutional layer, and Sigmoid activation function layer in sequence to obtain the activation feature map a;
[0018] S15. Multiply the activation feature map a with the initial feature map element-wise to obtain the final fusion feature map
[0019] Furthermore, step S2 specifically includes:
[0020] S21. Perform normalization processing on the fusion feature map obtained in step S1, and input the normalized feature map into different convolutional branches respectively to extract the query Q, key K, and value Va of the multi-scale attention mechanism;
[0021] Meanwhile, the normalized feature maps are sequentially input into a convolutional layer and a full-dimensional dynamic convolutional layer to obtain a normalized feature map V;
[0022] S22. Multiply the query Q and the key K to calculate the attention matrix A, where each element in the attention matrix A represents the similarity between the query and the key;
[0023] S23. Use a masking operation to set the values at irrelevant positions in the attention matrix A to zero, generating an initial sparse attention matrix A';
[0024] S24. From the initial sparse attention matrix A', select the K largest attention scores and their corresponding index positions in each row;
[0025] S25. Through a hashing operation, redistribute the attention scores of the retained K maximum values to the corresponding positions in the attention matrix, and set the scores at the remaining positions to zero, obtaining a secondary sparse attention matrix A Sparse ;
[0026] S26. Multiply A Sparse element-wise with the normalized feature map V to obtain a weighted feature map X:
[0027] S27. Input the weighted feature map X into a convolutional layer for convolution operation, and perform a residual connection between the result of the convolution operation and the fused feature map obtained in step S1 to obtain an enhanced feature map
[0028] Furthermore, step S3 specifically includes:
[0029] S31. Gradually restore the resolution of the enhanced feature map to the resolution of the original image through an upsampling operation, obtaining an enhanced feature map with the original resolution;
[0030] S32. Use a standard convolutional layer to further process the enhanced feature map with the original resolution output in step S31 to obtain a processed feature map;
[0031] S33. Use a convolution operation to convert the feature map processed by the convolution operation in step S32 into an output map with the same size as the input remote sensing image, and use a 1x1 convolutional kernel to map the converted feature map to the required number of classes, and perform a softmax or sigmoid activation function operation on the mapped output result to obtain the class prediction probability for each pixel.
[0032] Furthermore, the loss function in the process of obtaining the class prediction probability specifically includes the following loss functions:
[0033] Loss total = λ1·LossCE + λ2·Loss Dice ;
[0034] Where Loss CE represents the cross-entropy loss function; Loss Dice represents the Dice loss function, and λ1 and λ2 are weight coefficients.
[0035] On the other hand, the present invention also discloses a remote sensing image segmentation system based on a multi-scale feature fusion and index selection mechanism, including,
[0036] Feature extraction module: extracting different-scale feature information in the input remote sensing image by using a multi-scale convolutional network, and obtaining a fused feature map of the remote sensing image;
[0037] Index selection enhancement module based on a sparse attention mechanism: used for feature aggregation of the fused feature map of the remote sensing image to obtain an enhanced feature map;
[0038] Feature decoding module: decoding the enhanced feature map by using a decoder to obtain the remote sensing image segmentation result.
[0039] Preferably, the above remote sensing image segmentation system further includes a remote sensing image preprocessing module, which is used to perform geometric correction, radiometric correction, image enhancement and denoising, size normalization and cropping operations on the remote sensing image before feature extraction in sequence.
[0040] Preferably, the feature extraction module in the above remote sensing image segmentation system specifically performs the following steps:
[0041] S11. Input the remote sensing image into the initial convolutional layer for preliminary feature extraction to obtain an initial feature map
[0042] S12. Input the initial feature map into the spatial global average pooling layer to extract the global information of the remote sensing image and generate a global context feature map
[0043] At the same time, input the initial feature map into the windmill-shaped convolutional layer to extract the detailed information of the local area of the remote sensing image to obtain a local feature map
[0044] S13. Multiply the global context feature map and the local feature map element-wise, and add the multiplication result to the local feature map element-wise to generate an initial fused feature map
[0045] S14. Input the initial fused feature map Input the convolutional layer, ReLu activation function layer, normalization layer, convolutional layer, and Sigmoid activation function layer in sequence to obtain the activation feature map a;
[0046] S15. Multiply the activation feature map a element-wise with the initial feature map to obtain the final fused feature map
[0047] Preferably, in the above remote sensing image segmentation system, the index selection enhancement module based on the sparse attention mechanism obtains the enhanced feature map, and specifically performs the following steps:
[0048] S21. Normalize the fused feature map obtained by the feature extraction module, and input the normalized feature map into different convolutional branches respectively to extract the query Q, key K, and value Va of the multi-scale attention mechanism;
[0049] At the same time, input the normalized feature map into the convolutional layer and the full-dimensional dynamic convolutional layer in sequence to obtain the normalized feature map V;
[0050] S22. Multiply the query Q and the key K to calculate the attention matrix A, and each element in the attention matrix A represents the similarity between the query and the key;
[0051] S23. Use the masking operation to set the values at the irrelevant positions in the attention matrix A to zero to generate the initial sparse attention matrix A';
[0052] S24. From the initial sparse attention matrix A', select the K largest attention scores and their corresponding index positions in each row;
[0053] S25. Redistribute the attention scores of the remaining K maximum values to the corresponding positions in the initial sparse attention matrix A' through the hashing operation, and the scores at the remaining positions are set to zero to obtain the secondary sparse attention matrix A Sparse ;
[0054] S26. Multiply the secondary sparse attention matrix A Sparse element-wise with the normalized feature map V to obtain the weighted feature map X:
[0055] S27. Input the weighted feature map X into a convolutional layer for convolution operation, and perform residual connection on the convolution operation result and the fused feature map obtained by the feature extraction module to obtain the enhanced feature map
[0056] Preferably, in the above remote sensing image segmentation system, the feature decoding module decodes the enhanced feature map using the decoder, which specifically includes the following steps:
[0057] S31. Restore the resolution of the enhanced feature map step by step to the resolution of the original remote sensing image through upsampling operations, obtaining an enhanced feature map with the original resolution;
[0058] S32. Further process the enhanced feature map with the original resolution output in step S31 using a standard convolutional layer to obtain a processed feature map;
[0059] S33. Use convolutional operations to convert the feature map processed in step S32 into an output map with the same size as the input remote sensing image, and use a 1x1 convolutional kernel to map the converted feature map to the required number of classes, and perform softmax or sigmoid activation function operations on the mapped output results to obtain the class prediction probability for each pixel.
[0060] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a remote sensing image segmentation method and related system based on multi-scale feature fusion and index selection mechanism, having the following beneficial effects:
[0061] 1) Improve the accuracy of remote sensing image segmentation: By introducing multi-scale feature extraction and sparse attention mechanism, it can better focus on important regions in remote sensing images, thereby improving the segmentation accuracy.
[0062] 2) Reduce the demand for computing resources: Using the sparse attention mechanism, it reduces the interference of unimportant information, significantly improves the computing efficiency, reduces the demand for computing resources, and is applicable to image processing tasks in various computing environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for description in the embodiments or the prior art. Obviously, the following drawings are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0064] Figure 1 It is a schematic diagram of the overall process of the remote sensing image segmentation method provided by the present invention.
[0065] Figure 2 It is a schematic diagram of the feature extraction module provided by the present invention.
[0066] Figure 3 It is a schematic diagram of the index selection enhancement module provided by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0067] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0068] The present invention discloses a remote sensing image segmentation method and system based on multi-scale feature fusion and index selection mechanism, aiming to effectively improve the segmentation accuracy by combining multi-scale feature extraction with sparse attention mechanism and making full use of information at different levels and scales in remote sensing images. In terms of multi-scale feature extraction, local details and global information are captured through a multi-branch convolution structure, and at the same time, technologies such as dilated convolution and depthwise separable convolution are introduced to enhance the perception ability of features at different scales. The sparse attention mechanism calculates important regions in the multi-scale feature maps, adaptively selects key features for weighted aggregation, suppresses irrelevant information, and improves the computational efficiency. The organic combination of this multi-scale and attention mechanism can not only help the model capture richer detailed features in complex remote sensing images, but also effectively improve the segmentation accuracy of the model in images with different resolutions and different target sizes.
[0069] First, an embodiment of the present invention discloses a remote sensing image segmentation method based on multi-scale feature fusion and index selection mechanism, as Figure 1 shown, including:
[0070] S1. Extract different-scale feature information in the remote sensing image using a multi-scale convolutional network, and obtain a fused feature map of the remote sensing image;
[0071] S2. Aggregate features of the fused feature map of the remote sensing image using an index selection enhancement module based on a sparse attention mechanism to obtain an enhanced feature map;
[0072] S3. Decode the enhanced feature map using a decoder to obtain a pixel-level segmentation result.
[0073] Each step in the embodiment will be further described below.
[0074] Step S1 can specifically include preprocessing and feature extraction of the original remote sensing image, and the specific implementation steps are as follows.
[0075] Step S10: After a series of pre - processing on the collected remote - sensing image, the input image I to be fed into the neural network is obtained. First, the image undergoes geometric correction to eliminate geometric distortions caused by sensor or platform movement. Then, radiometric correction techniques are applied to remove sensor noise and atmospheric interference, ensuring the accuracy of the image data. Next, image enhancement and denoising processes are carried out to improve the image quality so that the neural network can better identify and process features. Finally, the image undergoes size normalization and cropping operations, making the size of the input image meet the requirements of the neural network while maintaining the key information of the image without loss. After these pre - processing steps, the obtained image I is high - quality data suitable for inputting into the neural network for analysis and inference.
[0076] Step S11: Initial feature extraction. As Figure 2 shown, the input image I obtained after pre - processing of the remote - sensing image extracts preliminary features through the initial convolutional layer, obtaining the initial feature map This step is mainly used for low - level feature extraction of the input image, laying a foundation for subsequent high - level feature operations.
[0077] Step S12: Perform spatial global average pooling and windmill - shaped convolution operations on the initial feature map respectively.
[0078] The feature map passes through the spatial global average pooling layer to extract the global information of the image, generating the global context feature map
[0079]
[0080] Among them, Conv represents the convolutional layer, and GAP represents the spatial global average pooling layer.
[0081] Meanwhile, the feature map passes through the windmill - shaped convolutional layer (PSConv) to extract the detailed information of the local area, obtaining the feature map
[0082]
[0083] Among them, PSConv represents the windmill - shaped convolutional layer. This windmill - shaped convolutional layer aims to capture multi - scale and local pattern information through a specific convolutional pattern and can effectively extract complex detailed features from remote - sensing images.
[0084] Step S13: Feature fusion to generate the final fused feature map Multiply and element - by - element to obtain the fused feature map. By element - by - element multiplication, it can capture and The relationship characteristics between. Then, the result of the multiplication is added to to generate the final fused feature map
[0085]
[0086] where, ⊙ represents element-wise multiplication, represents the matrix addition operation.
[0087] Step S14: Nonlinear transformation and feature enhancement. Perform a convolution operation on the fused feature map and further enhance the features through the activation function (ReLU) and the normalization layer. Subsequently, perform a second convolution operation, and finally process it through the Sigmoid activation function to obtain the activation feature map a:
[0088]
[0089] where, Sigmoid represents the Sigmoid activation function, ReLU represents the ReLU activation function, and BN represents the normalization layer. This operation further enhances the expressive power of the features through the combination of multiple layers of convolution, nonlinear activation, and the normalization layer, enabling the network to better adapt to the complex patterns in remote sensing images.
[0090] Step S15: Final output feature map. Multiply the activation feature map a element-wise with the initial feature map to obtain the final fused feature map as the output of the network. This process is equivalent to strengthening the original feature map with the enhanced feature information to obtain a more accurate representation.
[0091] The feature extraction process can be implemented through the Figure 2 feature extraction module shown. This feature extraction module effectively extracts the high-level features of remote sensing images through a series of convolutional layers, pooling layers, windmill-shaped convolutions, and feature fusion operations. In this process, global and local information is extracted through spatial global average pooling and windmill-shaped convolutions, and the correlation between the two is strengthened through element-wise multiplication and addition fusion, finally generating feature maps with high expressive power. These feature maps will provide richer information for subsequent classification or segmentation tasks.
[0092] Step S2: Use the index selection enhancement module based on the sparse attention mechanism to perform feature aggregation on the fused feature map of the remote sensing image to obtain the enhanced feature map. The overall framework diagram of the index selection enhancement module based on the sparse attention mechanism is as shown in Figure 3 and the specific steps of this module are as follows.
[0093] S21. Normalize the fused feature map obtained by the feature extraction module, and input the normalized feature maps into different convolutional branches respectively to extract the query Q, key K, and value Va of the multi-scale attention mechanism;
[0094] Meanwhile, input the normalized feature maps into the convolutional layer and the full-dimensional dynamic convolutional layer in sequence to obtain the normalized feature map V;
[0095] In the above steps, the normalization operation helps to standardize the input features, reduce the scale differences between different features, and thus improve the training stability of the network. Then, the normalized feature maps are input into different convolutional branches respectively to extract the query (Q), key (K), and value (Va). Specifically, the feature map passes through different convolutional layers and depthwise separable convolutional layers to obtain Q, passes through the convolutional layer and the dilated convolutional layer to obtain K, and passes through the convolutional layer and the full-dimensional dynamic convolutional layer to obtain the feature map V: where DilatedConv represents the dilated convolution and DynamicConv represents the full-dimensional dynamic convolution operation.
[0096] Then calculate the attention and mask operations to generate the initial sparse attention matrix A′, calculate the attention scores, multiply the query Q and the key K, and calculate the attention matrix A, where each element represents the similarity between the query and the key. Next, in order to mask out the unimportant attention scores, apply the mask operation. The mask operation sets the unimportant attention scores to zero and retains the most critical scores, thus avoiding the interference of these irrelevant information on the subsequent feature aggregation. Here, the mask operation generates a sparse attention matrix A′ by setting the values at the irrelevant positions to zero. The specific implementation steps are as follows:
[0097] S22. Multiply the query Q and the key K to calculate the attention matrix A, where each element in the attention matrix A represents the similarity between the query and the key;
[0098] S23. Use the mask operation to set the values at the irrelevant positions in the attention matrix A to zero, generating the initial sparse attention matrix A′.
[0099] Next, index selection and hashing operations are performed. Select the top K maximum values and their indices. From the attention matrix A′, select the top K maximum attention scores and their corresponding index positions in each row. These strongest attention scores can better represent the relationships between image features. By selecting the top K maximum values, the model can accurately retain the key features relevant to the current task. According to the selected top K maximum values and their indices, retain these important attention scores and ignore the other unimportant parts. This step is the core of implementing the sparse attention mechanism. Through index selection, the model can precisely locate the most valuable information. The hashing operation redistributes the attention scores of the retained K maximum values to their corresponding positions in the attention matrix, and the scores at the remaining positions are set to zero, obtaining A Sparse This operation makes the attention matrix sparser, which helps improve computational efficiency and the model's attention ability. This process sparsifies the attention matrix through the hashing operation, making the computational process more efficient and focusing on important feature information. The specific implementation steps are as follows:
[0100] S24. From the initial sparse attention matrix A′, select the top K maximum attention scores and their corresponding index positions in each row;
[0101] S25. Through the hashing operation, redistribute the attention scores of the retained K maximum values to their corresponding positions in the initial sparse attention matrix A′, and set the scores at the remaining positions to zero, obtaining the secondarily sparse attention matrix A Sparse ;
[0102] Then, weighted feature aggregation and residual connection operations are performed. This operation includes steps S26 and S27, and the specific steps are as follows:
[0103] S26. Element-wise multiply the sparsified attention matrix A Sparse and the feature map V to obtain the weighted feature map
[0104]
[0105] where ⊙ represents element-wise multiplication.
[0106] S27. After obtaining the weighted feature map, to avoid information loss, pass the weighted feature map through a convolutional layer and perform a residual connection with the original fused feature map to finally obtain the enhanced feature map
[0107] The residual connection ensures that the input features and the enhanced features can be effectively fused in the output, thus helping the network better perform feature extraction and learning.
[0108] Through the above steps, the method of network application index selection successfully implements the sparse attention mechanism, which can effectively focus on important image regions in the remote sensing image segmentation task. The mask operation, index selection, and hashing operation work together to ensure the efficiency and accuracy of feature aggregation. On this basis, through weighted feature aggregation and residual connection, the network's ability to express key features is enhanced, thereby improving the performance of remote sensing image segmentation.
[0109] Step S3: The enhanced feature map obtained in step S2 is decoded by the decoder to convert it into a pixel-level segmentation result.
[0110] Step S31: Upsampling of the feature map. The feature map in is gradually restored to the resolution of the original image. The feature map is gradually restored from a lower resolution to a higher resolution through a series of upsampling operations. Deconvolution or bilinear interpolation is used for the upsampling operation. The deconvolution operation is usually carried out jointly with the learning of the convolution kernel to extract more meaningful features while restoring the spatial resolution.
[0111] Step S32: Feature processing by convolution operation. The decoded features are further processed in the decoder part to accurately restore the structure and details of the image. The standard convolution layer is used to further process the feature map output in step S31. Several convolution layers are used to extract more advanced feature information and gradually reduce the number of channels. A non-linear activation function, such as the ReLU or LeakyReLU activation function, can be added after each convolution layer.
[0112] Step S33: Final pixel-level classification. The feature map is converted into a class prediction for each pixel. In the last layer of the decoder, a convolution operation is used to convert the feature map into an output map with the same size as the input image. A 1x1 convolution kernel is used to map the final feature map to the required number of classes (i.e., the number of output channels is equal to the number of classes). The softmax or sigmoid activation function is applied to the output to obtain the class prediction probability for each pixel.
[0113] During the class prediction process, the loss function is obtained through the following steps. First, the cross-entropy loss is used to measure the difference between the predicted class of each pixel and the true label. Assume that the output of the network is the class prediction probability distribution for each pixel, denoted as The true label given to this neural network by the present invention is y(i), where i represents the pixel position. The cross-entropy loss is used for each pixel:
[0114]
[0115] Next, the Dice loss is used to optimize the overlapping region between the predicted segmentation result and the ground truth segmentation result, which is particularly suitable for dealing with the problem of class imbalance. Calculate the Dice coefficient between the predicted segmentation map and the ground truth label map, which measures the similarity between two sets:
[0116]
[0117] Then calculate the Dice loss:
[0118]
[0119] Finally, the total loss function combines multiple loss functions to obtain a final loss value for optimization. The final loss function is a weighted sum of multiple losses:
[0120] Loss total =λ1·Loss CE +λ2·Loss Dice
[0121] where λ1 and λ2 are weight coefficients that can be adjusted according to needs.
[0122] For the remote sensing image segmentation method of the present invention, first, different scale information in the image is extracted through a multi-scale convolutional network in the feature extraction stage, which can capture both the global features of the image and finely process local detail features. Specifically, the global information of the image is extracted through a spatial global average pooling (GAP) layer, and the detail information of the local area is extracted through a pinwheel-shaped convolutional layer (PSConv). The fusion of these two features provides rich feature expressions for subsequent image segmentation tasks.
[0123] Based on the above method, the present invention also discloses a remote sensing image segmentation system based on multi-scale feature fusion and index selection mechanism, including,
[0124] Feature extraction module: Use a multi-scale convolutional network to extract different scale feature information from the input remote sensing image and obtain a fused feature map of the remote sensing image;
[0125] Index selection enhancement module based on sparse attention mechanism: Used to perform feature aggregation on the fused feature map of the remote sensing image to obtain an enhanced feature map;
[0126] Feature decoding module: Use a decoder to decode the enhanced feature map to obtain a pixel-level segmentation result.
[0127] The specific execution process of each module can refer to the description in the method of this application and will not be repeated here.
[0128] To solve the efficiency problem of traditional attention mechanisms in large-scale remote sensing image processing, the present invention adopts a sparse attention mechanism. After normalizing the input features, convolutional branches are used to generate query, key, and value feature maps, and the attention matrix is calculated. Unimportant attention scores are masked out through a masking operation, thereby reducing the computational amount. Then, the strongest K attention scores are retained through an index selection operation to further focus on important regions in the image. Finally, weighted feature aggregation and residual connection methods are used to enhance the network's ability to express key features and ensure more accurate segmentation results.
[0129] Compared with traditional methods, the present invention not only improves the accuracy of remote sensing image segmentation, but also, in the case of limited computing resources, through the combination of the sparse attention mechanism and multi-scale feature extraction, enables the method to perform large-scale image segmentation tasks more efficiently. Especially in complex environments, it can effectively distinguish ground object features and avoid segmentation errors caused by the neglect of local features in traditional methods.
[0130] The innovation of the present invention lies in its combination of multi-scale feature extraction and the sparse attention mechanism. It not only overcomes the limitations of traditional methods, but also greatly improves the performance of remote sensing image segmentation through an efficient computing method, and has broad application prospects.
[0131] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0132] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A remote sensing image segmentation method based on multi-scale feature fusion and index selection mechanism, characterized in that It includes the following steps: S1. Extract different-scale feature information from the remote sensing image using a multi-scale convolutional network, and obtain a fused feature map of the remote sensing image; S2. Aggregate features of the fused feature map of the remote sensing image using an index selection enhancement module based on a sparse attention mechanism to obtain an enhanced feature map; S3. Decode the enhanced feature map using a decoder to obtain the remote sensing image segmentation result.
2. The remote sensing image segmentation method based on multi-scale feature fusion and index selection mechanism according to claim 1, characterized in that, Before step S1 extracts different-scale features from the remote sensing image using a multi-scale convolutional network, it also includes preprocessing the collected remote sensing image, specifically including: Performing geometric correction, radiometric correction, image enhancement and denoising, size normalization and cropping operations on the collected remote sensing image in sequence.
3. The remote sensing image segmentation method based on multi-scale feature fusion and index selection mechanism according to claim 1, wherein Step S1 specifically includes: S11. Input the remote sensing image into an initial convolutional layer for preliminary feature extraction to obtain an initial feature map; S12. Input the initial feature map into a spatial global average pooling layer to extract the global information of the remote sensing image and generate a global context feature map; At the same time, input the initial feature map into a windmill-shaped convolutional layer to extract the detailed information of the local area of the remote sensing image to obtain a local feature map; S13. Multiply the global context feature map and the local feature map element by element, and add the multiplication result to the local feature map element by element to generate an initial fused feature map; S14. Input the initial fused feature map into a convolutional layer, a ReLu activation function layer, a normalization layer, a convolutional layer, and a Sigmoid activation function layer in sequence to obtain an activated feature map; S15. Multiply the activated feature map obtained in step S14 by the initial feature map element by element to obtain the final fused feature map.
4. The remote sensing image segmentation method based on multi-scale feature fusion and index selection mechanism according to claim 1, characterized in that Step S2 specifically includes: S21. Normalize the fused feature map obtained in step S1, and input the normalized feature map into different convolutional branches respectively to extract the query Q, key K, and value Va of the multi-scale attention mechanism; At the same time, input the normalized feature map into a convolutional layer and a full-dimensional dynamic convolutional layer in sequence to obtain a normalized feature map V; S22. Multiply the query Q and the key K to calculate the attention matrix A, and each element in the attention matrix A represents the similarity between the query and the key; S23. Use a masking operation to set the values at irrelevant positions in the attention matrix A to zero to generate an initial sparse attention matrix A'; S24. Select the largest K attention scores and their corresponding index positions in each row from the initial sparse attention matrix A'; S25. Redistribute the attention scores of the retained K maximum values to the corresponding positions in the initial sparse attention matrix A' through a hashing operation, and set the scores at the remaining positions to zero to obtain the quadratic sparse attention matrix A Sparse ; S26. Multiply the quadratic sparse attention matrix A Sparse and the normalized feature map V element-wise to obtain the weighted feature map X: S27. Input the weighted feature map X into a convolutional layer for convolution operation, and perform residual connection on the result of the convolution operation and the fused feature map obtained in step S1 to obtain an enhanced feature map 5. The remote sensing image segmentation method based on multi-scale feature fusion and index selection mechanism according to claim 1, wherein Step S3 specifically includes: S31. Gradually restore the resolution of the enhanced feature map to the resolution of the original remote sensing image through upsampling operations to obtain an enhanced feature map with the original resolution; S32. Further process the enhanced feature map with the original resolution output in step S31 using a standard convolutional layer to obtain a processed feature map; S33. Use a convolutional operation to convert the feature map processed by the convolutional operation in step S32 into an output map with the same size as the input remote sensing image, and use a 1x1 convolutional kernel to map the converted feature map to the required number of classes, and perform a softmax or sigmoid activation function operation on the mapped output result to obtain the class prediction probability of each pixel.
6. A remote sensing image segmentation system based on multi-scale feature fusion and index selection mechanism, characterized in that including, Feature extraction module: It uses a multi-scale convolutional network to extract different-scale feature information from the input remote sensing image and obtains a fused feature map of the remote sensing image; Index selection enhancement module based on sparse attention mechanism: It is used to aggregate features of the fused feature map of the remote sensing image to obtain an enhanced feature map; Feature decoding module: It uses a decoder to decode the enhanced feature map to obtain the remote sensing image segmentation result.
7. The remote sensing image segmentation method based on multi-scale feature fusion and index selection mechanism according to claim 6, wherein It also includes a remote sensing image preprocessing module, which is used to perform geometric correction, radiometric correction, image enhancement and denoising, size normalization and cropping operations on the remote sensing image before feature extraction in sequence.
8. The remote sensing image segmentation method based on multi-scale feature fusion and index selection mechanism according to claim 6, characterized in that The feature extraction module specifically performs the following steps: S11. Input the remote sensing image into the initial convolutional layer for preliminary feature extraction to obtain an initial feature map S12. Input the initial feature map into the spatial global average pooling layer to extract the global information of the remote sensing image and generate a global context feature map Meanwhile, the initial feature map is input into the windmill-shaped convolutional layer to extract the detailed information of the local area of the remote sensing image, and a local feature map S13. Multiply the global context feature map and the local feature map element - by - element, and add the multiplication result to the local feature map element - by - element to generate the initial fusion feature map S14. Input the initial fused feature map into the convolutional layer, ReLu activation function layer, normalization layer, convolutional layer, and Sigmoid activation function layer in sequence to obtain the activation feature map a; S15. Multiply the activated feature map a element-wise with the initial feature map to obtain the final fused feature map 9. The remote sensing image segmentation method based on multi-scale feature fusion and index selection mechanism according to claim 6, wherein The index selection enhancement module based on the sparse attention mechanism aggregates the features of the fused feature map of the remote sensing image to obtain an enhanced feature map, and specifically performs the following steps: S21. Normalize the fused feature map obtained by the feature extraction module, and input the normalized feature map into different convolutional branches respectively to extract the query Q, key K, and value Va of the multi-scale attention mechanism; At the same time, input the normalized feature map into a convolutional layer and a full-dimensional dynamic convolutional layer in sequence to obtain the normalized feature map V; S22. Multiply the query Q and the key K to calculate the attention matrix A, and each element in the attention matrix A represents the similarity between the query and the key; S23. Use a masking operation to set the values at irrelevant positions in the attention matrix A to zero to generate an initial sparse attention matrix A'; S24. Select the largest K attention scores and their corresponding index positions in each row from the initial sparse attention matrix A'; S25. Redistribute the attention scores of the retained K maximum values to the corresponding positions in the initial sparse attention matrix A' through a hashing operation, and set the scores at the remaining positions to zero to obtain the quadratic sparse attention matrix A Sparse ; S26. Multiply the secondary sparse attention matrix A Sparse and the normalized feature map V element by element to obtain the weighted feature map X: S27. Input the weighted feature map X into a convolutional layer for convolution operation, and perform a residual connection between the result of the convolution operation and the fused feature map obtained by the feature extraction module to obtain an enhanced feature map 10. The remote sensing image segmentation method based on multi-scale feature fusion and index selection mechanism according to claim 6, characterized in that, The feature decoding module uses a decoder to decode the enhanced feature map, which specifically includes the following steps: S31. Gradually restore the resolution of the enhanced feature map to the resolution of the original remote sensing image through upsampling operations to obtain an enhanced feature map with the original resolution; S32. Use a standard convolutional layer to further process the enhanced feature map with the original resolution output in step S31 to obtain a processed feature map; S33. Use a convolutional operation to convert the feature map processed by the convolutional operation in step S32 into an output map with the same size as the input remote sensing image, and use a 1x1 convolutional kernel to map the converted feature map to the required number of classes, and perform a softmax or sigmoid activation function operation on the mapped output result to obtain the class prediction probability of each pixel.
Citation Information
Patent Citations
Remote sensing image classification method based on multi-scale sparse cross fusion and semantic enhancement
CN118154986A
Semantic segmentation model and segmentation method for high-resolution remote sensing image
CN119206229A
Remote sensing image marine and non-marine area segmentation method based on pyramid mechanism
WO2023039959A1
Cited By
Image diffusion enhancement method and device based on multi-scale feature extraction and fusion
CN120976039A
Remote sensing image segmentation method and system based on fine screening double-domain attention mechanism
CN121582565A