Remote sensing image segmentation method and system of fusion stage perception and multi-dimensional orientation mechanism

By integrating staged perception and multi-dimensional orientation mechanisms, this remote sensing image segmentation method utilizes a multi-stage perception enhancer and a multi-dimensional orientation cyclic key-value module to address the issues of insufficient multi-scale target adaptability and orientation modeling in remote sensing image segmentation. This achieves high-precision segmentation of remote sensing images, with a significant improvement in segmentation performance, particularly in complex terrain areas.

CN121053541BActive Publication Date: 2026-04-17耕宇牧星(北京)空间科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
耕宇牧星(北京)空间科技有限公司
Filing Date
2025-08-29
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing remote sensing image segmentation methods have weak adaptability to multi-scale targets and insufficient ability to model directional structures, resulting in low boundary recognition accuracy. Furthermore, they lack effective multi-level semantic fusion and directional perception linkage mechanisms, leading to problems such as broken target contours and category confusion.

Method used

A remote sensing image segmentation method that integrates stage perception and multi-dimensional orientation mechanisms is adopted. Through a multi-stage perception enhancer and a multi-dimensional orientation cyclic key-value module, it can achieve accurate segmentation of multi-scale, multi-directional, and multi-class targets in remote sensing images. The combination of stage perception enhancer and multi-dimensional orientation cyclic key-value module improves the model's segmentation consistency and robustness of remote sensing images.

Benefits of technology

It significantly improves the recognition accuracy and segmentation consistency of multi-scale ground targets in remote sensing images, especially in transitional areas such as buildings and roads, where it exhibits better segmentation consistency and robustness. It also enhances the global modeling capability for linear structures and directional targets, and improves edge fit and class differentiation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053541B_ABST
    Figure CN121053541B_ABST
Patent Text Reader

Abstract

This invention discloses a remote sensing image segmentation method and system that integrates stage perception and multidimensional orientation mechanisms, belonging to the field of remote sensing image processing technology. The method includes the following steps: S1: preprocessing the remote sensing image to be segmented; S2: inputting the preprocessed remote sensing image to be segmented into a segmentation model integrating stage perception and multidimensional orientation mechanisms to obtain a segmentation result image of the remote sensing image to be segmented; wherein, the segmentation model integrating stage perception and multidimensional orientation mechanisms includes several stage perception enhancers and a multidimensional orientation cyclic key-value module. This invention achieves accurate segmentation of multi-scale, multi-directional, and multi-category targets in remote sensing images while maintaining a lightweight model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing image processing technology, and more specifically to a remote sensing image segmentation method and system based on a fusion-stage perception and multi-dimensional orientation mechanism. Background Technology

[0002] Currently, with the continuous advancement of remote sensing imaging technology and the widespread acquisition of high-resolution remote sensing data, remote sensing image segmentation is increasingly being applied in numerous fields such as urban planning, land use, disaster monitoring, and agricultural resource analysis. The task of remote sensing image segmentation requires assigning semantic labels to each pixel in a high-dimensional image to identify different land cover types such as buildings, roads, vegetation, and water bodies. This task is characterized by high resolution, large differences in target scale, and complex spatial structure. Most existing remote sensing image segmentation methods are based on convolutional neural network (CNN) structures, such as FCN, U-Net, and the DeepLab series. These methods rely on deep convolutional layers to extract semantic features and perform segmentation and recovery through upsampling or skip connections. While these methods perform well in conventional natural images, they suffer from significant problems in remote sensing image applications, including weak multi-scale target adaptability, insufficient directional structure modeling capabilities, and low boundary recognition accuracy. Furthermore, although segmentation methods incorporating Transformer structures have achieved some breakthroughs in semantic modeling in recent years, their high computational complexity and lack of spatial local modeling capabilities limit their deployment and application in high-resolution remote sensing image scenarios.

[0003] Furthermore, existing improvement strategies for addressing the aforementioned issues still have many shortcomings. For example, some studies have attempted to improve the situation through attention mechanisms and multi-scale feature fusion. While channel or spatial attention modules such as SE and CBAM can enhance feature representation capabilities to some extent, they lack inter-stage adaptive modeling capabilities and struggle to accurately capture the contributions of features at different levels to segmentation. In terms of orientation modeling, most methods still rely on fixed-direction convolution or unstructured fully connected modeling, failing to balance orientation consistency with efficient computation. Traditional methods lack an effective linkage mechanism between multi-level semantic fusion and orientation-aware modeling, often resulting in issues such as broken target contours, boundary jagged edges, or category confusion in the final segmentation image.

[0004] Therefore, how to provide a remote sensing image segmentation method and system that can achieve accurate segmentation of multi-scale, multi-directional, and multi-category targets in remote sensing images while maintaining a lightweight model is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a remote sensing image segmentation method and system that integrates stage perception and multidimensional orientation mechanisms.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] Firstly, a remote sensing image segmentation method integrating stage perception and multi-dimensional orientation mechanisms is provided, comprising the following steps:

[0008] S1: Preprocess the remote sensing image to be segmented;

[0009] S2: Input the preprocessed remote sensing image to be segmented into the segmentation model of the fusion stage perception and multidimensional orientation mechanism to obtain the segmentation result map of the remote sensing image to be segmented; wherein, the segmentation model of the fusion stage perception and multidimensional orientation mechanism includes several stage perception enhancers and a multidimensional orientation cyclic key value module.

[0010] Preferably, S2 specifically includes the following steps:

[0011] The preprocessed remote sensing image to be segmented is input into the initial convolutional layer to obtain the first feature map f1;

[0012] The first feature map f1 is input into the first stage perception enhancer to obtain the first enhanced feature map f1. * ;

[0013] For the first enhanced feature map f1 * Perform downsampling to obtain the second feature map f2;

[0014] The second feature map f2 is input into the second stage perceptual enhancer to obtain the second enhanced feature map f2. * ;

[0015] For the second enhanced feature map f2 * Perform downsampling to obtain the third feature map f3;

[0016] The third feature map f3 is input into the third stage perception enhancer to obtain the third enhanced feature map f3. * ;

[0017] For the third enhanced feature map f3 * Perform downsampling to obtain the fourth feature map f4;

[0018] The fourth feature map f4 is input into the fourth stage perception enhancer to obtain the fourth enhanced feature map f4. * ;

[0019] For the fourth enhanced feature map f4 * Perform downsampling to obtain the fifth feature map f5;

[0020] The fifth feature map f5 is input into the fifth stage perception enhancer to obtain the fifth enhanced feature map f5. *;

[0021] The fifth enhanced feature map f5 * The input is fed into the multidimensional orientation cyclic key-value module to obtain the global enhanced feature map f. + ;

[0022] For the global enhanced feature map f + After upsampling, it is then compared with the fourth enhanced feature map f4. * Perform element-wise addition to obtain the sixth feature map f6;

[0023] The sixth feature map f6 is input into the sixth stage perception enhancer to obtain the sixth enhanced feature map f6. * ;

[0024] For the sixth enhanced feature map f6 * After upsampling, it is then compared with the third enhanced feature map f3 * Perform element-wise addition to obtain the seventh feature map f7;

[0025] The seventh feature map f7 is input into the seventh stage perception enhancer to obtain the seventh enhanced feature map f7. * ;

[0026] For the seventh enhanced feature map f7 * After upsampling, it is then compared with the second enhanced feature map f2 * Perform element-wise addition to obtain the eighth feature map f8;

[0027] The eighth feature map f8 is input into the eighth stage perception enhancer to obtain the eighth enhanced feature map f8. * ;

[0028] For the eighth enhanced feature map f8 * After upsampling, it is then compared with the first enhanced feature map f1 * Perform element-wise addition to obtain the ninth feature map f9;

[0029] The ninth feature map f9 is input into the ninth stage perception enhancer to obtain the ninth enhanced feature map f9. * ;

[0030] The ninth enhanced feature map f9 * The data is input into the segmentation head to obtain a segmentation prediction map of the remote sensing image to be segmented.

[0031] Preferably, S2 further includes the following steps:

[0032] After normalizing the segmentation prediction map of the remote sensing image to be segmented using the Softmax activation function, the category with the highest probability value at each pixel position is selected as the final segmentation label to obtain the segmentation result map of the remote sensing image to be segmented.

[0033] Preferably, S2 further includes the following steps:

[0034] feature map f i The first intermediate feature map f is obtained by sequentially processing the data through depthwise separable convolutional layers and the GELU activation function. i 1 ;

[0035] Where i = 1, 2, 3, 4, 5, 6, 7, 8, 9;

[0036] When i = 1, the feature map f i This represents the first feature map f1;

[0037] When i = 2, the feature map f i This represents the second feature map f2;

[0038] When i = 3, the feature map f i This represents the third feature map f3;

[0039] When i = 4, the feature map f i This represents the fourth feature map f4;

[0040] When i = 5, the feature map f i This represents the fifth feature map f5;

[0041] When i = 6, the feature map f i This represents the sixth feature map f6;

[0042] When i = 7, the feature map f i This represents the seventh feature map f7;

[0043] When i = 8, the feature map f i This represents the eighth feature map f8;

[0044] When i = 9, the feature map f i This represents the ninth feature map f9;

[0045] The first intermediate feature map f i 1 The second intermediate feature map f is obtained by sequentially processing the convolutional layer, batch normalization layer, and GELU activation function of the first standard convolutional module. i 2 ;

[0046] The first intermediate feature map f i 1 and the second intermediate feature map f i 2 Perform element-wise addition to obtain the fused feature map f. i 3 ;

[0047] The fused feature map f i 3 The enhanced feature map f is obtained by sequentially processing the convolutional layer, batch normalization layer, and GELU activation function of the second standard convolutional module, and the convolutional layer, batch normalization layer, and GELU activation function of the third standard convolutional module. i * ;

[0048] Wherein, when i = 1, the enhanced feature map f i * This represents the first enhanced feature map f1 * ;

[0049] When i = 2, the enhanced feature map f i * This represents the second enhanced feature map f2. * ;

[0050] When i = 3, the enhanced feature map f i * This represents the third enhanced feature map f3. * ;

[0051] When i = 4, the enhanced feature map f i * This represents the fourth enhanced feature map f4. * ;

[0052] When i = 5, the enhanced feature map f i * This represents the fifth enhanced feature map f5. * ;

[0053] When i = 6, the enhanced feature map f i * This represents the sixth enhanced feature map f6. * ;

[0054] When i = 7, the enhanced feature map f i * This represents the seventh enhanced feature map f7. * ;

[0055] When i = 8, the enhanced feature map f i *This represents the eighth enhanced feature map f8. * ;

[0056] When i = 9, the enhanced feature map f i * This represents the ninth enhanced feature map f9. * .

[0057] Preferably, the fifth enhanced feature map f5 * The input is fed into the multidimensional orientation cyclic key-value module to obtain the global enhanced feature map f. + Specifically, it includes the following steps:

[0058] The fifth enhanced feature map f5 * Expand from left to right to obtain the first token sequence d1;

[0059] The fifth enhanced feature map f5 * Expand from right to left to obtain the second token sequence d2;

[0060] The fifth enhanced feature map f5 * Expand from top to bottom to obtain the third token sequence d3;

[0061] The fifth enhanced feature map f5 * Expanding from bottom to top, we obtain the fourth token sequence d4;

[0062] The first token sequence d1 is weighted and aggregated in the spatial dimension to obtain the first aggregated feature d1'; the first token sequence d1 is weighted and aggregated in the reverse spatial dimension to obtain the first reverse aggregated feature d1".

[0063] The second token sequence d2 is weighted and aggregated in the spatial dimension to obtain the second aggregated feature d'2; the second token sequence d2 is weighted and aggregated in the reverse spatial dimension to obtain the second reverse aggregated feature d2".

[0064] The third token sequence d3 is weighted and aggregated in the spatial dimension to obtain the third aggregated feature d3'; the third token sequence d3 is weighted and aggregated in the reverse spatial dimension to obtain the third reverse aggregated feature d3".

[0065] The fourth token sequence d4 is weighted and aggregated in the spatial dimension to obtain the fourth aggregated feature d'4; the fourth token sequence d4 is weighted and aggregated in the reverse spatial dimension to obtain the fourth reverse aggregated feature d4".

[0066] After the first aggregated feature d1' and the first reverse aggregated feature d1" are calculated by bidirectional weighted key value, they are inversely mapped together to obtain the first two-dimensional feature map F1.

[0067] After the second aggregated feature d'2 and the second reverse aggregated feature d2" are respectively calculated by bidirectional weighted key value, they are then inversely mapped together to obtain the second two-dimensional feature map F2.

[0068] After the third aggregation feature d3' and the third reverse aggregation feature d3" are calculated by bidirectional weighted key value, they are then inversely mapped together to obtain the third two-dimensional feature map F3.

[0069] After the fourth aggregation feature d'4 and the fourth reverse aggregation feature d4" are respectively calculated by bidirectional weighted key value, they are then inversely mapped together to obtain the fourth two-dimensional feature map F4.

[0070] Pixel-level averaging is performed on the first two-dimensional feature map F1, the second two-dimensional feature map F2, the third two-dimensional feature map F3, and the fourth two-dimensional feature map F4 to obtain the average two-dimensional feature map F. * ;

[0071] The average two-dimensional feature map F * The global enhanced feature map f is obtained by flattening, weighted aggregation and reshaping of the channel dimensions. + .

[0072] Preferably, the preprocessing in S1 includes size normalization, histogram equalization, or normalization.

[0073] Preferably, the total loss function used when training the segmentation model for the fusion stage perception and multi-dimensional orientation mechanism is:

[0074]

[0075] Among them, L total The total loss function is represented by λ1 and λ2, which are adjustable weight parameters. Indicates the need to obtain the segmentation prediction graph The cross-entropy loss between the actual segmentation map Y labeled by humans; Indicates the need to obtain the segmentation prediction graph The Dice loss of the real segmentation map Y with manual annotation.

[0076] In a second aspect, a remote sensing image segmentation system that integrates stage perception and multidimensional orientation mechanism is provided, characterized in that it is used to implement the remote sensing image segmentation method described in the first aspect, including a preprocessing unit and a segmentation model that integrates stage perception and multidimensional orientation mechanism.

[0077] The preprocessing unit is used to preprocess the remote sensing image to be segmented;

[0078] The segmentation model based on the fusion stage perception and multidimensional orientation mechanism is used to segment the preprocessed remote sensing image to obtain the segmentation result image of the remote sensing image to be segmented.

[0079] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a remote sensing image segmentation method and system, which achieves the following beneficial effects:

[0080] This invention introduces a stage-aware enhancer into remote sensing image segmentation for the first time. By adopting a structure-adaptive feature recalibration method for different coding stages, it effectively preserves shallow texture edge information and enhances deep semantic expression capabilities, significantly improving the recognition accuracy of multi-scale ground objects. In particular, it has better segmentation consistency and robustness in transitional areas such as buildings, roads, and river networks.

[0081] The multidimensional orientation cyclic key-value module constructed in this invention can expand and bidirectionally aggregate image features from four main directions (left→right, right→left, up→down, down→up), effectively compensating for the shortcomings of convolutional networks in perceiving long-range dependencies and directional information. It improves the model's global modeling ability for heterogeneous spatial patterns such as linear structures (e.g., roads, canals) and highly directional targets (e.g., farmland, building arrangements) in remote sensing images, and improves the edge fit and category discrimination of the segmentation results while ensuring inference efficiency.

[0082] In summary, this invention achieves accurate segmentation of multi-scale, multi-directional, and multi-category targets in remote sensing images while maintaining a lightweight model. Attached Figure Description

[0083] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0084] Figure 1 A schematic diagram illustrating the processing of the segmentation model for the fusion-stage perception and multi-dimensional orientation mechanism provided by this invention;

[0085] Figure 2 A schematic diagram of the process of the stage perception enhancer provided by the present invention;

[0086] Figure 3 This is a schematic diagram of the processing of the multidimensional orientation cyclic key value module provided by the present invention. Detailed Implementation

[0087] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0088] In a first aspect, embodiments of the present invention disclose a remote sensing image segmentation method that integrates stage-based perception and multi-dimensional orientation mechanisms, comprising the following steps:

[0089] S1: Preprocess the remote sensing image to be segmented;

[0090] In one or more embodiments, the preprocessing in S1 includes size normalization, histogram equalization, or normalization.

[0091] It is understandable that preprocessing can improve image quality and the stability of feature extraction.

[0092] S2: Input the preprocessed remote sensing image to be segmented into the segmentation model of the fusion stage perception and multidimensional orientation mechanism to obtain the segmentation result map of the remote sensing image to be segmented; wherein, the segmentation model of the fusion stage perception and multidimensional orientation mechanism includes several stage perception enhancers and a multidimensional orientation cyclic key value module.

[0093] In one or more embodiments, such as Figure 1 As shown, S2 specifically includes the following steps:

[0094] 1) Input Image and Feature Extraction

[0095] The preprocessed remote sensing image to be segmented is input into the initial convolutional layer to obtain the first feature map f1;

[0096] Understandably, the preprocessed remote sensing image to be segmented is input into the initial convolutional layer of the neural network. Standard convolution operations are used to extract low-level edge and texture features, outputting the first feature map f1. The first feature map f1 preserves the original spatial structure information of the remote sensing image to be segmented, serving as the foundation for subsequent multi-scale semantic modeling.

[0097] 2) Multi-stage feature encoding and enhancement processing

[0098] The first feature map f1 is input into the first stage perception enhancer to obtain the first enhanced feature map f1. * ;

[0099] For the first enhanced feature map f1 * Perform downsampling to obtain the second feature map f2;

[0100] The second feature map f2 is input into the second stage perceptual enhancer to obtain the second enhanced feature map f2. * ;

[0101] For the second enhanced feature map f2 * Perform downsampling to obtain the third feature map f3;

[0102] The third feature map f3 is input into the third stage perception enhancer to obtain the third enhanced feature map f3. * ;

[0103] For the third enhanced feature map f3 * Perform downsampling to obtain the fourth feature map f4;

[0104] The fourth feature map f4 is input into the fourth stage perception enhancer to obtain the fourth enhanced feature map f4. * ;

[0105] For the fourth enhanced feature map f4 * Perform downsampling to obtain the fifth feature map f5;

[0106] The fifth feature map f5 is input into the fifth stage perception enhancer to obtain the fifth enhanced feature map f5. * ;

[0107] It is understandable that:

[0108] The enhanced feature map output by the stage-aware augmenter has the ability to enhance attention in stages.

[0109] The downsampling is a single spatial downsampling, used to reduce resolution and expand the sensing area to capture a larger scale of remotely sensed ground features.

[0110] The aforementioned top-down multi-level feature extraction and enhancement structure can effectively adapt to the characteristics of remote sensing images, such as large differences in object scale, blurred category boundaries, and complex terrain, ensuring the stable acquisition of deep semantic features.

[0111] 3) Direction perception and semantic fusion

[0112] The fifth enhanced feature map f5 * The input is fed into the multidimensional orientation cyclic key-value module to obtain the global enhanced feature map f. + ;

[0113] It is understandable that the multidimensional orientation cyclic key value module can effectively extract spatial structural relationships with directional consistency (such as the arrangement of roads, rivers, and buildings) in remote sensing images through a four-way scanning mechanism and a two-way weight aggregation strategy.

[0114] 4) Feature fusion and enhancement during the decoding process

[0115] In the decoding stage, a combination of feature fusion and stage-aware enhancer is used to achieve effective reconstruction of multi-scale features. The specific steps are as follows:

[0116] For the global enhanced feature map f + After upsampling, it is then compared with the fourth enhanced feature map f4. * Perform element-wise addition to obtain the sixth feature map f6;

[0117] The sixth feature map f6 is input into the sixth stage perception enhancer to obtain the sixth enhanced feature map f6. * ;

[0118] For the sixth enhanced feature map f6 * After upsampling, it is then compared with the third enhanced feature map f3 * Perform element-wise addition to obtain the seventh feature map f7;

[0119] The seventh feature map f7 is input into the seventh stage perception enhancer to obtain the seventh enhanced feature map f7. * ;

[0120] For the seventh enhanced feature map f7 * After upsampling, it is then compared with the second enhanced feature map f2 * Perform element-wise addition to obtain the eighth feature map f8;

[0121] The eighth feature map f8 is input into the eighth stage perception enhancer to obtain the eighth enhanced feature map f8. * ;

[0122] For the eighth enhanced feature map f8 * After upsampling, it is then compared with the first enhanced feature map f1 * Perform element-wise addition to obtain the ninth feature map f9;

[0123] The ninth feature map f9 is input into the ninth stage perception enhancer to obtain the ninth enhanced feature map f9. * ;

[0124] It is understandable that upsampling is used to gradually restore spatial resolution, and spatial expansion is performed during the upsampling process using methods such as transposed convolution or bilinear interpolation.

[0125] It is understandable that the above decoding process fully combines shallow high-resolution structural information with deep semantic features, and uses a staged perceptual enhancer to semantically recalibrate the fusion results at each level, thereby improving edge continuity, target consistency, and class segmentation accuracy in remote sensing images. It exhibits significant segmentation stability advantages, especially for low-contrast target boundaries or regions with small inter-class differences.

[0126] It is understandable that the ninth enhanced feature map f9 * It has the same or similar spatial resolution as the remote sensing image to be segmented, and integrates shallow texture information and deep semantic information, thus possessing good spatial consistency and category discrimination ability.

[0127] 5) Feature mapping and segmentation head generation of segmentation prediction map

[0128] The ninth enhanced feature map f9 * The data is input into the segmentation head to obtain a segmentation prediction map of the remote sensing image to be segmented.

[0129] It is understandable that: the ninth enhanced feature map f9 * The dimensions are C*H*W (C, H, and W represent the ninth enhanced feature map f9, respectively). * (Number of channels, height, and width); the segmentation head consists of a set of 1×1 convolutional layers, its function being to process the ninth enhanced feature map f9. * The number of channels is mapped from C to the number of land cover categories N. The segmentation prediction map obtained through this linear mapping operation... ( The size is N*H*W, where N, H, and W represent the segmentation prediction map, respectively. The number of channels, height, and width of the segmentation head represent the probability that each pixel in the remote sensing image to be segmented belongs to each land cover category. It is suitable for the fine segmentation of multiple land cover categories (such as buildings, roads, vegetation, water bodies, etc.) in remote sensing images. The segmentation head adopts a lightweight structure, has good inference efficiency, and is suitable for large-scale remote sensing scene deployment.

[0130] In one or more embodiments, S2 further includes the following steps:

[0131] After normalizing the segmentation prediction map of the remote sensing image to be segmented using the Softmax activation function, the category with the highest probability value at each pixel location is selected as the final segmentation label, thus obtaining the segmentation result map of the remote sensing image to be segmented.

[0132] Understandably, to improve the geometric consistency and boundary connectivity of the segmentation results, optional post-processing modules, such as Conditional Random Fields (CRF) or morphological operations, can be added to smooth coarse areas and correct edge details. The final output segmentation results can be directly used for downstream remote sensing tasks, such as land cover classification, change detection, and urban planning analysis.

[0133] In one or more embodiments, such as Figure 2 As shown, S2 further includes the following steps:

[0134] S21: Transfer feature map f i The first intermediate feature map f is obtained by sequentially processing the data through depthwise separable convolutional layers and the GELU activation function. i 1 ;

[0135] The corresponding expression is: f i 1 =GELU(DSConv(f i ));

[0136] Where GELU represents the GELU activation function; DSConv represents a depthwise separable convolutional layer; and feature map f i The dimensions are C*H*W, where C, H, and W represent the feature map f, respectively. i The number of channels, height, and width;

[0137] Depthwise Separable Convolution (DSConv) can reduce computational complexity and improve spatial feature extraction capabilities.

[0138] Where i = 1, 2, 3, 4, 5, 6, 7, 8, 9;

[0139] When i = 1, the feature map f i This represents the first feature map f1;

[0140] When i = 2, the feature map f i This represents the second feature map f2;

[0141] When i = 3, the feature map f i This represents the third feature map f3;

[0142] When i = 4, the feature map f i This represents the fourth feature map f4;

[0143] When i = 5, the feature map f i This represents the fifth feature map f5;

[0144] When i = 6, the feature map f i This represents the sixth feature map f6;

[0145] When i = 7, the feature map f i This represents the seventh feature map f7;

[0146] When i = 8, the feature map f i This represents the eighth feature map f8;

[0147] When i = 9, the feature map f i This represents the ninth feature map f9;

[0148] The first intermediate feature map f i 1 The second intermediate feature map f is obtained by sequentially processing the convolutional layer, batch normalization layer, and GELU activation function of the first standard convolutional module. i 2 ;

[0149] The corresponding expression is: f i 2 =GELU(BN(Conv(f i 1 )));

[0150] Where GELU represents the GELU activation function; BN represents batch normalization; and Conv represents convolution.

[0151] This process can enhance nonlinear expressive power and stabilize gradient propagation.

[0152] S21 can effectively extract local details (such as road boundaries and water textures) and their important semantic structures at high spatial resolution in remote sensing images.

[0153] S22: Transfer the first intermediate feature map f i 1 and the second intermediate feature map f i 2 Perform element-wise addition to obtain the fused feature map f. i 3 ;

[0154] The corresponding expression is: f i 3 =f i 1 +f i 2 ;

[0155] This process can enhance the robustness of features and the transfer of information between layers;

[0156] The fused feature map fi 3 The enhanced feature map f is obtained by sequentially processing the convolutional layer, batch normalization layer, and GELU activation function of the second standard convolutional module, and the convolutional layer, batch normalization layer, and GELU activation function of the third standard convolutional module. i * ;

[0157] The corresponding expression is: f i * =GELU(BN(Conv(GELU(BN(Conv(f i 3 ))))));

[0158] This process can further refine spatial structure and semantic relationships;

[0159] Enhanced feature map f i * It has good scale adaptability and category sensitivity, and is especially suitable for processing spatial transition areas between complex features in remote sensing images.

[0160] Wherein, when i = 1, the enhanced feature map f i * This represents the first enhanced feature map f1 * ;

[0161] When i = 2, the enhanced feature map f i * This represents the second enhanced feature map f2. * ;

[0162] When i = 3, the enhanced feature map f i * This represents the third enhanced feature map f3. * ;

[0163] When i = 4, the enhanced feature map f i * This represents the fourth enhanced feature map f4. * ;

[0164] When i = 5, the enhanced feature map f i * This represents the fifth enhanced feature map f5. * ;

[0165] When i = 6, the enhanced feature map f i * This represents the sixth enhanced feature map f6. * ;

[0166] When i = 7, the enhanced feature map fi * This represents the seventh enhanced feature map f7. * ;

[0167] When i = 8, the enhanced feature map f i * This represents the eighth enhanced feature map f8. * ;

[0168] When i = 9, the enhanced feature map f i * This represents the ninth enhanced feature map f9. * .

[0169] In one or more embodiments, such as Figure 3 As shown:

[0170] The fifth enhanced feature map f5 * The input is fed into the multidimensional orientation cyclic key-value module to obtain the global enhanced feature map f. + Specifically, it includes the following steps:

[0171] The fifth enhanced feature map f5 * Expand from left to right to obtain the first token sequence d1;

[0172] The fifth enhanced feature map f5 * Expand from right to left to obtain the second token sequence d2;

[0173] The fifth enhanced feature map f5 * Expand from top to bottom to obtain the third token sequence d3;

[0174] The fifth enhanced feature map f5 * Expanding from bottom to top, we obtain the fourth token sequence d4;

[0175] It is understandable that the fifth enhanced feature map f5 of 2-D is generated along the four main directions (left→right, right→left, up→down, down→up). * Flatten into four 1-D token sequences (i.e., first token sequence d1, second token sequence d2, token sequence d3, and fourth token sequence d4).

[0176] The corresponding expression is: d3=θ↓(f5 * ); d4=θ↑(f5) * );

[0177] in, This indicates an expansion from left to right; θ indicates expansion from right to left; θ↓ indicates expansion from top to bottom; θ↑ indicates expansion from bottom to top.

[0178] The first token sequence d1 is weighted and aggregated in the spatial dimension to obtain the first aggregated feature d1'; the first token sequence d1 is weighted and aggregated in the reverse spatial dimension to obtain the first reverse aggregated feature d1".

[0179] The second token sequence d2 is weighted and aggregated in the spatial dimension to obtain the second aggregated feature d'2; the second token sequence d2 is weighted and aggregated in the reverse spatial dimension to obtain the second reverse aggregated feature d2".

[0180] The third token sequence d3 is weighted and aggregated in the spatial dimension to obtain the third aggregated feature d3'; the third token sequence d3 is weighted and aggregated in the reverse spatial dimension to obtain the third reverse aggregated feature d3".

[0181] The fourth token sequence d4 is weighted and aggregated in the spatial dimension to obtain the fourth aggregated feature d'4; the fourth token sequence d4 is weighted and aggregated in the reverse spatial dimension to obtain the fourth reverse aggregated feature d4".

[0182] The corresponding expression is:

[0183] d' j =spa(d j );

[0184] d j =spa(reverse(d j ));

[0185] Where j = 1, 2, 3, 4; spa represents the weighted aggregation of spatial dimensions; reverse represents the reverse direction.

[0186] After the first aggregated feature d1' and the first reverse aggregated feature d1" are calculated by bidirectional weighted key value, they are inversely mapped together to obtain the first two-dimensional feature map F1.

[0187] After the second aggregated feature d'2 and the second reverse aggregated feature d2" are respectively calculated by bidirectional weighted key value, they are then inversely mapped together to obtain the second two-dimensional feature map F2.

[0188] After the third aggregation feature d3' and the third reverse aggregation feature d3" are calculated by bidirectional weighted key value, they are then inversely mapped together to obtain the third two-dimensional feature map F3.

[0189] After the fourth aggregation feature d'4 and the fourth reverse aggregation feature d4" are respectively calculated by bidirectional weighted key value, they are then inversely mapped together to obtain the fourth two-dimensional feature map F4.

[0190] The corresponding expression is:

[0191]

[0192] Where j = 1, 2, 3, 4; RWKV represents bidirectional weighted key value calculation; This represents the inverse mapping.

[0193] The above process fully integrates remote dependence and boundary information from all directions, and can effectively address the problem of anisotropic layout of objects in remote sensing images.

[0194] Pixel-level averaging is performed on the first two-dimensional feature map F1, the second two-dimensional feature map F2, the third two-dimensional feature map F3, and the fourth two-dimensional feature map F4 to obtain the average two-dimensional feature map F. * ;

[0195] The corresponding expression is: F * =F1+F2+F3+F4.

[0196] The average two-dimensional feature map F * The global enhanced feature map f is obtained by flattening, weighted aggregation and reshaping of the channel dimensions. + .

[0197] The corresponding expression is: f + =cha(Flatten(F * );

[0198] Here, Flatten means flattening; cha means weighted aggregation and reshaping of the channel dimension.

[0199] The aforementioned feature fusion strategy fully utilizes the information redundancy and complementarity in different directions, thereby improving the segmentation consistency and boundary accuracy of multi-category ground objects in remote sensing images.

[0200] The multidimensional orientation cyclic key-value module can improve the model's ability to perceive complex directional structures (such as intersecting roads, river flow direction, and building layout) in remote sensing images.

[0201] In one or more embodiments, the total loss function used when training the segmentation model of the fusion stage perception and multidimensional orientation mechanism is:

[0202]

[0203] Among them, L totalThe total loss function is represented by λ1 and λ2, which are adjustable weight parameters. Indicates the need to obtain the segmentation prediction graph The cross-entropy loss between the actual segmentation map Y labeled by humans; Indicates the need to obtain the segmentation prediction graph The Dice loss of the real segmentation map Y with manual annotation.

[0204] λ1 and λ2 are dynamically adjusted based on the degree of class imbalance in the remote sensing dataset and the sensitivity to boundary errors. This total loss function design can effectively suppress the gradient shift phenomenon dominated by the background class and enhance the perception capability of small features (such as bridges, narrow roads, and canals).

[0205] In a second aspect, embodiments of the present invention also disclose a remote sensing image segmentation system that integrates stage perception and multidimensional orientation mechanisms, characterized in that it is used to implement the remote sensing image segmentation method described in the first aspect, including a preprocessing unit and a segmentation model that integrates stage perception and multidimensional orientation mechanisms.

[0206] The preprocessing unit is used to preprocess the remote sensing image to be segmented;

[0207] The segmentation model based on the fusion stage perception and multidimensional orientation mechanism is used to segment the preprocessed remote sensing image to obtain the segmentation result image of the remote sensing image to be segmented.

[0208] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0209] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A remote sensing image segmentation method integrating stage perception and multi-dimensional orientation mechanisms, characterized in that, Includes the following steps: S1: Preprocess the remote sensing image to be segmented; S2: Input the preprocessed remote sensing image to be segmented into the segmentation model of the fusion stage perception and multidimensional orientation mechanism to obtain the segmentation result image of the remote sensing image to be segmented; wherein, the segmentation model of the fusion stage perception and multidimensional orientation mechanism includes several stage perception enhancers and a multidimensional orientation cyclic key value module, specifically including: The preprocessed remote sensing image to be segmented is input into the initial convolutional layer to obtain the first feature map f1; The first feature map f1 is input into the first stage perception enhancer to obtain the first enhanced feature map f1. ; For the first enhanced feature map f1 Perform downsampling to obtain the second feature map f2; The second feature map f2 is input into the second stage perceptual enhancer to obtain the second enhanced feature map f2. ; For the second enhanced feature map f2 Perform downsampling to obtain the third feature map f3; The third feature map f3 is input into the third stage perception enhancer to obtain the third enhanced feature map f3. ; For the third enhanced feature map f3 Perform downsampling to obtain the fourth feature map f4; The fourth feature map f4 is input into the fourth stage perception enhancer to obtain the fourth enhanced feature map f4. ; For the fourth enhanced feature map f4 Perform downsampling to obtain the fifth feature map f5; The fifth feature map f5 is input into the fifth stage perception enhancer to obtain the fifth enhanced feature map f5. ; The fifth enhanced feature map f5 The input is fed into the multidimensional orientation cyclic key-value module to obtain the global enhanced feature map f. + Specifically, it includes the following steps: The fifth enhanced feature map f5 Expand from left to right to obtain the first token sequence. ; The fifth enhanced feature map f5 Expand from right to left to obtain the second token sequence. ; The fifth enhanced feature map f5 Expand from top to bottom to obtain the third token sequence. ; The fifth enhanced feature map f5 Expand from bottom to top to obtain the fourth token sequence. ; The first token sequence After weighted aggregation along the spatial dimension, the first aggregated feature is obtained. ; the first token sequence After weighted aggregation in the reverse spatial dimension, the first reverse aggregation feature is obtained. ; The second token sequence After weighted aggregation along the spatial dimension, the second aggregated feature is obtained. ; the second token sequence After weighted aggregation in the reverse spatial dimension, the second reverse aggregation feature is obtained. ; The third token sequence After weighted aggregation along spatial dimensions, the third aggregation feature is obtained. ; the third token sequence After weighted aggregation in the reverse spatial dimension, the third reverse aggregation feature is obtained. ; The fourth token sequence After weighted aggregation along the spatial dimension, the fourth aggregation feature is obtained. The fourth token sequence After weighted aggregation in the reverse spatial dimension, the fourth reverse aggregation feature is obtained. ; The first aggregation feature and the first reverse aggregation feature After being calculated using bidirectional weighted key values, they are combined and inversely mapped to obtain the first two-dimensional feature map F1. The second aggregation feature and the second reverse aggregation feature After each key value is calculated using a bidirectional weighted method, they are then inversely mapped together to obtain the second two-dimensional feature map F2. The third aggregation feature and the third reverse aggregation feature After each key value is calculated using a bidirectional weighted method, they are then inversely mapped together to obtain the third two-dimensional feature map F3. The fourth aggregation feature and the fourth reverse aggregation feature After each key value is calculated using a bidirectional weighted method, they are then inversely mapped together to obtain the fourth two-dimensional feature map F4. Pixel-level averaging is performed on the first two-dimensional feature map F1, the second two-dimensional feature map F2, the third two-dimensional feature map F3, and the fourth two-dimensional feature map F4 to obtain the average two-dimensional feature map F. ; The average two-dimensional feature map F The global enhanced feature map f is obtained by flattening, weighted aggregation and reshaping of the channel dimensions. + ; For the global enhanced feature map f + After upsampling, it is then compared with the fourth enhanced feature map f4. Perform element-wise addition to obtain the sixth feature map f6; The sixth feature map f6 is input into the sixth stage perception enhancer to obtain the sixth enhanced feature map f6. ; For the sixth enhanced feature map f6 After upsampling, it is then compared with the third enhanced feature map f3 Perform element-wise addition to obtain the seventh feature map f7; The seventh feature map f7 is input into the seventh stage perception enhancer to obtain the seventh enhanced feature map f7. ; For the seventh enhanced feature map f7 After upsampling, it is then compared with the second enhanced feature map f2 Perform element-wise addition to obtain the eighth feature map f8; The eighth feature map f8 is input into the eighth stage perception enhancer to obtain the eighth enhanced feature map f8. ; For the eighth enhanced feature map f8 After upsampling, it is then compared with the first enhanced feature map f1 Perform element-wise addition to obtain the ninth feature map f9; The ninth feature map f9 is input into the ninth stage perception enhancer to obtain the ninth enhanced feature map f9. ; The ninth enhanced feature map f9 The data is input into the segmentation head to obtain a segmentation prediction map of the remote sensing image to be segmented.

2. The remote sensing image segmentation method based on the fusion of stage perception and multidimensional orientation mechanism according to claim 1, characterized in that, S2 further includes the following steps: After normalizing the segmentation prediction map of the remote sensing image to be segmented using the Softmax activation function, the category with the highest probability value at each pixel position is selected as the final segmentation label to obtain the segmentation result map of the remote sensing image to be segmented.

3. The remote sensing image segmentation method based on the fusion of stage perception and multi-dimensional orientation mechanism according to claim 1, characterized in that, S2 also includes the following steps: The feature map f i The first intermediate feature map f i 1 ; Where i = 1, 2, 3, 4, 5, 6, 7, 8, 9; When i=1, the feature map f i This represents the first feature map f1; When i=2, the feature map f i This represents the second feature map f2; When i=3, the feature map f i This represents the third feature map f3; When i=4, the feature map f i This represents the fourth feature map f4; When i=5, the feature map f i This represents the fifth feature map f5; When i=6, the feature map f i This represents the sixth feature map f6; When i=7, the feature map f i This refers to the seventh feature map f7; When i=8, the feature map f i This represents the eighth feature map f8; When i=9, the feature map f i This represents the ninth feature map f9; The first intermediate feature map f i 1 The second intermediate feature map f is obtained by sequentially processing the convolutional layer, batch normalization layer, and GELU activation function of the first standard convolutional module. i 2 ; The first intermediate feature map f i 1 and the second intermediate feature map f i 2 Perform element-wise addition to obtain the fused feature map f. i 3 ; The fused feature map f i 3 The enhanced feature map f is obtained by sequentially processing the convolutional layer, batch normalization layer, and GELU activation function of the second standard convolutional module, and the convolutional layer, batch normalization layer, and GELU activation function of the third standard convolutional module. i ; Wherein, when i=1, the enhanced feature map f i This represents the first enhanced feature map f1 ; When i=2, the enhanced feature map f i This represents the second enhanced feature map f2. ; When i=3, the enhanced feature map f i This represents the third enhanced feature map f3. ; When i=4, the enhanced feature map f i This represents the fourth enhanced feature map f4. ; When i=5, the enhanced feature map f i This represents the fifth enhanced feature map f5. ; When i=6, the enhanced feature map f i This represents the sixth enhanced feature map f6. ; When i=7, the enhanced feature map f i This represents the seventh enhanced feature map f7. ; When i=8, the enhanced feature map f i This represents the eighth enhanced feature map f8. ; When i=9, the enhanced feature map f i This represents the ninth enhanced feature map f9. .

4. The remote sensing image segmentation method based on the fusion of stage perception and multi-dimensional orientation mechanism according to claim 1, characterized in that, Preprocessing in S1 includes size normalization, histogram equalization, or normalization.

5. The remote sensing image segmentation method based on the fusion of stage perception and multi-dimensional orientation mechanism according to claim 1, characterized in that, The total loss function used when training the segmentation model for the fusion stage perception and multi-dimensional orientation mechanism is: ; in, Represents the total loss function; Indicates an adjustable weight parameter; Indicates the need to obtain the segmentation prediction graph and manually labeled real segmentation map Cross-entropy loss; Indicates the need to obtain the segmentation prediction graph and manually labeled real segmentation map The loss of Dice.

6. A remote sensing image segmentation system integrating stage perception and multi-dimensional orientation mechanisms, characterized in that, The remote sensing image segmentation method according to any one of claims 1-5 includes a preprocessing unit and a segmentation model with a fusion stage perception and multidimensional orientation mechanism. The preprocessing unit is used to preprocess the remote sensing image to be segmented; The segmentation model based on the fusion stage perception and multidimensional orientation mechanism is used to segment the preprocessed remote sensing image to obtain the segmentation result image of the remote sensing image to be segmented.

Citation Information

Patent Citations

  • High-resolution remote sensing image semantic segmentation method based on class-level context aggregation

    CN116563535A

  • Remote sensing image target segmentation method based on multistage feature aggregation and segmentation

    CN119919658A