Remote sensing image segmentation method based on heterogeneous double-flow feature extraction and adaptive feature lifter

Through heterogeneous dual-stream feature extraction and adaptive feature lifter, the problems of insufficient local information capture and insufficient feature fusion in remote sensing image segmentation are solved, and high-precision and efficient remote sensing image segmentation are achieved.

CN120279267AActive Publication Date: 2025-07-08耕宇牧星(北京)空间科技有限公司
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510349485.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

Traditional remote sensing image segmentation methods are difficult to effectively process multi-scale and multi-morphological landform information, and lack local detail features capture and robustness, especially in complex backgrounds and noise environments.

Method used

Heterogeneous dual-stream feature extraction and adaptive feature lifter are used to optimize the segmentation effect by using upper and lower branch feature extraction networks, combining dynamic convolution and full-dimensional dynamic convolution, and using composite loss function.

Benefits of technology

It significantly improves the accuracy and robustness of remote sensing image segmentation, can improve segmentation accuracy under variable land types and complex backgrounds, reduce computing resource consumption, and adapt to different computing platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279267A_ABST
    Figure CN120279267A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image segmentation method based on heterogeneous double-flow feature extraction and an adaptive feature lifter. The method comprises the following steps: inputting a target remote sensing image into a trained remote sensing image segmentation model; the remote sensing image segmentation model is composed of a heterogeneous double-flow feature extraction unit, a self-adaptive feature lifting unit and a mapping unit; performing feature extraction on the target remote sensing image through a heterogeneous double-flow feature extraction unit to obtain an upper branch output feature and a lower branch output feature; performing fusion and enhancement processing on the upper branch output features and the lower branch output features through a self-adaptive feature lifting unit to obtain enhanced features; and performing mapping processing on the enhanced features through a mapping unit to obtain a segmentation result corresponding to the target remote sensing image. The remote sensing image segmentation method can effectively solve the problems of insufficient local information capture, insufficient feature fusion, unbalanced categories and the like in a traditional remote sensing image segmentation method, thereby remarkably improving the precision and robustness of remote sensing image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and more specifically, to a remote sensing image segmentation method based on heterogeneous two-stream feature extraction and an adaptive feature booster. Background Art

[0002] Remote sensing image segmentation aims to extract significant ground object information from remote sensing images, such as buildings, roads, water bodies, farmlands, etc. This ground object information is of great significance for applications such as environmental monitoring, urban planning, and disaster assessment. However, the segmentation task of remote sensing images faces many challenges, mainly reflected in the diversity of ground objects in the images, background complexity, boundary ambiguity between ground objects, and scale changes. In order to effectively solve these problems, in recent years, image segmentation methods based on deep learning have been widely applied, especially the advantages of convolutional neural networks (CNNs) in remote sensing image segmentation, which have attracted the attention of a large number of researchers.

[0003] However, traditional CNN-based remote sensing image segmentation methods still have several limitations. First, existing convolutional networks mostly rely on fixed feature extraction patterns and are difficult to process multi-scale and multi-form ground object information in remote sensing images. Second, there are large differences in the spatial structures of different regions in remote sensing images, and traditional convolutional networks are difficult to effectively capture local detail features in the images, resulting in limited segmentation accuracy. In addition, background noise, complex lighting conditions, and boundary ambiguity of ground objects in remote sensing images often lead to insufficient robustness of existing segmentation methods in practical applications.

[0004] Therefore, how to solve the problems of insufficient capture of local information, insufficient feature fusion, and class imbalance in traditional remote sensing image segmentation methods, thereby significantly improving the accuracy and robustness of remote sensing image segmentation, is an urgent problem for those skilled in the art to solve. Summary of the Invention

[0005] In view of the above problems, the present invention provides a remote sensing image segmentation method based on heterogeneous two-stream feature extraction and an adaptive feature booster to at least solve some of the technical problems mentioned in the above background art.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] The present invention provides a remote sensing image segmentation method based on heterogeneous two-stream feature extraction and an adaptive feature booster, including the following steps:

[0008] Input the target remote sensing image into a trained remote sensing image segmentation model; the remote sensing image segmentation model is composed of a heterogeneous two-stream feature extraction unit, an adaptive feature boosting unit, and a mapping unit;

[0009] Through the heterogeneous two-stream feature extraction unit, feature extraction is performed on the target remote sensing image to obtain upper-branch output features and lower-branch output features;

[0010] Through the adaptive feature enhancement unit, the upper-branch output features and the lower-branch output features are fused and enhanced to obtain enhanced features;

[0011] Through the mapping unit, mapping processing is performed on the enhanced features to obtain the segmentation result corresponding to the target remote sensing image.

[0012] Further, the step of performing feature extraction on the target remote sensing image through the heterogeneous two-stream feature extraction unit to obtain upper-branch output features and lower-branch output features specifically includes:

[0013] Use an initial convolutional layer to perform sliding convolution on the target remote sensing image to generate a preliminary feature map;

[0014] Through the upper-branch feature extraction module, local feature extraction is performed on the preliminary feature map to obtain upper-branch output features;

[0015] Through the lower-branch feature extraction module, semantic feature extraction is performed on the preliminary feature map to obtain lower-branch output features.

[0016] Further, the step of performing local feature extraction on the preliminary feature map through the upper-branch feature extraction module to obtain upper-branch output features specifically includes:

[0017] Pass the preliminary feature map sequentially through a first normalization layer, a first convolutional layer, a first depthwise separable convolutional layer, a dropout layer, and a windmill-shaped convolutional layer to obtain a direction-aware feature map;

[0018] Through a first skip connection layer, the direction-aware feature map and the preliminary feature map are fused to obtain upper-branch output features.

[0019] Further, the step of performing semantic feature extraction on the preliminary feature map through the lower-branch feature extraction module to obtain lower-branch output features specifically includes:

[0020] Perform tokenization processing on the preliminary feature map to obtain label features;

[0021] Pass the label features sequentially through a Kolmogorov-Arnold layer, a second depthwise separable convolutional layer, and a second normalization layer to obtain normalized features;

[0022] Through a second skip connection layer, the normalized features and the preliminary feature map are fused to obtain lower-branch output features.

[0023] Further, through the adaptive feature enhancement unit, the output features of the upper branch and the output features of the lower branch are fused and enhanced to obtain enhanced features; specifically including:

[0024] The output features of the upper branch and the output features of the lower branch are concatenated through a concatenation layer to obtain concatenated features;

[0025] The concatenated features are sequentially passed through a global average pooling layer, a second convolutional layer, and a first activation function to generate global context features;

[0026] The global context features are multiplied by the concatenated features and then passed through a fourth convolutional layer to obtain adaptive features;

[0027] The output features of the upper branch are processed by a dynamic deformable convolutional layer to generate an upper branch feature map;

[0028] Full-dimensional dynamic convolution is used to extract features from the output features of the lower branch to obtain a lower branch feature map;

[0029] Through a third skip connection layer, the upper branch feature map and the lower branch feature map are fused to obtain a fused feature;

[0030] The fused features are sequentially passed through a third convolutional layer, a spiking neuron layer, and a second activation function to generate a weight matrix;

[0031] The weight matrix is multiplied by the adaptively adjusted features to obtain enhanced features.

[0032] Further, through the mapping unit, the enhanced features are mapped to obtain the segmentation result corresponding to the target remote sensing image; specifically including:

[0033] After spatially convolving and expanding the enhanced features through a fifth convolutional layer, spatial upsampling is performed through a deconvolution layer to obtain a reconstructed feature map;

[0034] The reconstructed feature map is subjected to feature diffusion and consistency enhancement processing to obtain consistency-enhanced features;

[0035] The consistency-enhanced features are sequentially passed through a sixth convolutional layer and a third activation function to obtain the predicted segmentation result corresponding to the target remote sensing image.

[0036] Further, the reconstructed feature map is subjected to feature diffusion and consistency enhancement processing to obtain consistency-enhanced features; specifically including:

[0037] The reconstructed feature map is convolved with a Sobel operator to obtain four-neighborhood gradients;

[0038] Dynamically adjust the diffusion coefficient according to the current feature edge intensity;

[0039] Based on the four-neighborhood gradient and the diffusion coefficient, update the reconstructed feature map to obtain an updated feature map;

[0040] Weightedly fuse the updated feature map and the reconstructed feature map through a gating mechanism to generate a consistency-enhanced feature.

[0041] Furthermore, the composite loss function of the remote sensing image segmentation model is expressed as:

[0042]

[0043] Wherein, represents the composite loss function; is the cross-entropy loss; represents the Dice coefficient loss; λ CE represents the weighting coefficient of the cross-entropy loss; λ Dice represents the weighting coefficient of the Dice coefficient loss; Y represents the true label; represents the predicted value; Y i represents the probability value of the i-th class in the true label; represents the probability value of the i-th class in the predicted value; C represents a total of C classes; Y j represents the j-th pixel value of the true label; represents the j-th pixel value of the predicted value; N represents the total number of pixels in the target remote sensing image.

[0044] Through the above technical solutions, compared with the prior art, the present invention discloses a remote sensing image segmentation method based on heterogeneous two-stream feature extraction and an adaptive feature booster, which has the following beneficial effects:

[0045] By combining the heterogeneous two-stream feature extraction unit and the adaptive feature boosting unit, the present invention can effectively capture the multi-scale features of the image and enhance the features of key regions, improving the accuracy and detail expressiveness of remote sensing image segmentation.

[0046] By combining the diffusion mechanism with adaptive weighting, the model of the present invention can effectively suppress background noise and enhance the consistency of features, thereby optimizing the spatial continuity of remote sensing image segmentation. Moreover, this method can reduce the consumption of computing resources through efficient feature processing, improve the computing efficiency of the model, and meet the requirements of different computing platforms.

[0047] The technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0049] Figure 1 Schematic diagram of the remote sensing image segmentation method based on heterogeneous dual-stream feature extraction and adaptive feature booster provided by the embodiment of the present invention.

[0050] Figure 2 Schematic diagram of the working process of the heterogeneous dual-stream feature extraction unit provided by the embodiment of the present invention.

[0051] Figure 3 Schematic diagram of the working process of the adaptive feature booster unit provided by the embodiment of the present invention. Detailed implementation manners

[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0053] The embodiment of the present invention discloses a remote sensing image segmentation method based on heterogeneous dual-stream feature extraction and adaptive feature booster. Refer to Figure 1 As shown, it includes the following steps:

[0054] S1. Input the target remote sensing image into the trained remote sensing image segmentation model; the remote sensing image segmentation model is composed of a heterogeneous dual-stream feature extraction unit, an adaptive feature booster unit, and a mapping unit;

[0055] S2. Through the heterogeneous dual-stream feature extraction unit, extract features from the target remote sensing image to obtain the upper branch output features and the lower branch output features;

[0056] S3. Through the adaptive feature booster unit, fuse and enhance the upper branch output features and the lower branch output features to obtain enhanced features;

[0057] S4. Through the mapping unit, perform mapping processing on the enhanced features to obtain the segmentation result corresponding to the target remote sensing image.

[0058] The remote sensing image segmentation method based on heterogeneous dual-stream feature extraction and adaptive feature booster provided by the embodiments of the present invention aims to improve the accuracy and robustness of remote sensing image segmentation through innovative feature extraction and enhancement mechanisms. Especially in the face of diverse land cover types and complex backgrounds, it can significantly improve the segmentation accuracy while reducing the resource requirements for model training and deployment, enabling it to adapt to remote sensing image processing tasks in various computing environments. Next, each of the above steps will be described in detail.

[0059] In the above step S1, the target remote sensing image is input into the trained remote sensing image segmentation model; wherein, the remote sensing image segmentation model is composed of a heterogeneous dual-stream feature extraction unit, an adaptive feature booster unit, and a mapping unit; the input target remote sensing image usually contains various land cover types and different geographical information.

[0060] In the above step S2, through the heterogeneous dual-stream feature extraction unit, feature extraction is performed on the target remote sensing image to obtain the upper-branch output feature and the lower-branch output feature; see Figure 2 shown, specifically including:

[0061] (1) Use the initial convolutional layer to perform sliding convolution on the target remote sensing image to generate the preliminary feature map φ o ; specifically:

[0062] The initial convolutional layer performs sliding convolution on the target remote sensing image through multiple convolutional kernels to extract basic image features such as edges, textures, color distributions, etc., and generates the preliminary feature map φ o ; this process lays the foundation for subsequent feature processing and remote sensing image segmentation tasks, helping the network capture the global and local information of the image.

[0063] (2) Perform local feature extraction on the preliminary feature map φ o through the upper-branch feature extraction module to obtain the upper-branch output feature φ u ; specifically including:

[0064] The preliminary feature map φ o is sequentially passed through the first normalization layer, the first convolutional layer, the first depthwise separable convolutional layer, the dropout layer, and the windmill-shaped convolutional layer to obtain the direction-aware feature map; wherein, the first convolutional layer includes multiple convolutional layers for extracting the preliminary feature map φ oThe spatial features enhance the network's ability to capture local details of remote sensing images; the first depthwise separable convolutional layer decomposes the traditional convolutional operation into channel-wise convolution and point-wise convolution, thus effectively reducing the computational load while maintaining a high feature extraction ability; the Dropout layer can effectively prevent the network from overfitting during training. This layer randomly discards some neuron connections, forcing the network to learn more generalizable features; the windmill-shaped convolutional layer is introduced to capture features in specific directions or patterns in remote sensing images, especially local directional structures that may exist in the images, such as linear features like roads and buildings.

[0065] Through the first skip connection layer, the direction-aware feature map and the preliminary feature map φ o are fused to obtain the upper-branch output feature φ u to ensure the flow and fusion of information.

[0066] (3) The lower-branch feature extraction module performs semantic feature extraction on the preliminary feature map φ o to obtain the lower-branch output feature φ d . Specifically, it includes:

[0067] The preliminary feature map φ o is tokenized to obtain label features; this tokenization process converts different regions in the remote sensing image into label representations in a certain way. Usually, the pixels or features of the image are mapped to a discrete label space to facilitate subsequent classification or segmentation tasks of the network. For example, using a region segmentation-based method, different land cover types or geographical features are labeled with different class labels, so that each pixel of the image can be associated with a specific label. In this process, pixel-level annotation or region-level annotation methods can be used to ensure that the tokenized image features can be effectively used for classification or segmentation.

[0068] The label features are successively passed through the Kolmogorov - Arnold layer, the second depth - separable convolutional layer, and the second normalization layer to obtain normalized features. Among them, the Kolmogorov - Arnold layer is based on the Kolmogorov - Arnold representation theorem. The introduction of this layer is to enhance the network's ability in complex pattern recognition, especially in the recognition of highly non - linear and complex ground object boundaries that may exist in remote sensing images. The Kolmogorov - Arnold layer can better represent the high - order relationships and detailed information in the image through complex mathematical mappings, thereby enhancing the network's expressive ability. The second depth - separable convolutional layer further extracts and compresses the extracted features, not only reducing the computational amount but also effectively extracting highly abstract features, which is convenient for subsequent image segmentation tasks. To ensure the stability of the network, the second normalization layer is used to standardize the features, helping to adjust the distribution of the features so that features at different levels are on the same scale, avoiding the problems of gradient disappearance or explosion that may occur during the training process, and improving the training speed and performance of the model.

[0069] Through the second skip connection layer, the normalized features and the preliminary feature map φ o are fused to obtain the lower - branch output feature φ d . This design of skip connection helps to retain the initially extracted detailed information in the network, while fusing the high - level features of the lower branch, enhancing the accuracy of the segmentation task.

[0070] In the above step S3, through the adaptive feature enhancement unit, the upper - branch output feature φ u and the lower - branch output feature φ d are fused and enhanced to obtain enhanced features; as shown in Figure 3 :

[0071] (1) Feature concatenation and adaptive weighting:

[0072] The upper - branch output feature φ u and the lower - branch output feature φ d are concatenated through the concatenation layer to obtain the concatenated feature φ c ; this process fuses local features from different scales, providing more comprehensive information for subsequent processing.

[0073] The concatenated feature φ c is successively passed through the global average pooling layer, the second convolutional layer, and the first activation function to generate the global context feature A; expressed as:

[0074] A = σ1(Conv2(GlobalAvgPool(φ c )))

[0075] Among them, σ1 represents the first activation function, specifically the Sigmoid activation function, and Conv2 represents the second convolutional layer; GlobalAvgPool represents the global average pooling layer.

[0076] Multiply the global context feature A with the concatenated feature φ c and then pass it through the fourth convolutional layer to obtain the adaptive feature φ ca ; By introducing global context information, it can flexibly strengthen the key region features in the remote sensing image, suppress background noise, and thus effectively improve the accuracy of remote sensing image segmentation.

[0077] (2) Dynamic deformable and full-dimensional dynamic convolution feature fusion:

[0078] Use the dynamic deformable convolutional layer (DCNv4) to process the output feature φ of the upper branch u to extract adaptable features and generate the upper branch feature map; At the same time, use the omnidirectional dynamic convolution (ODConv) to perform feature extraction on the output feature φ of the lower branch d to obtain the lower branch feature map; The combination of these two operations can adaptively adjust the shape and size of the convolutional kernel under multi-scale and complex backgrounds to better capture the important spatial information in the remote sensing image.

[0079] Through the third skip connection layer, perform weighted fusion processing on the upper branch feature map and the lower branch feature map to obtain the fusion feature φ ud+ ; It is expressed as:

[0080]

[0081] Among them, DCNv4 represents the dynamic deformable convolutional layer; ODConv represents the omnidirectional dynamic convolutional layer; represents the addition operation. This fusion method can not only enhance the expression ability of features but also reduce the interference of redundant information, thereby further optimizing the effect of image segmentation. By dynamically adjusting the shape and scale of the convolutional kernel, the module can more accurately capture the complex spatial changes in the remote sensing image, especially when processing images with multiple land cover types or irregular shapes, showing excellent feature extraction capabilities.

[0082] (3) Feature enhancement and output generation:

[0083] Pass the fusion feature φ ud+ sequentially through the third convolutional layer, the spiking neuron layer, and the second activation function to generate the weight matrix It is expressed as:

[0084]

[0085] Among them, SNN represents the spiking neuron layer; σ2 represents the second activation function; Conv3 represents the third convolutional layer.

[0086] Multiply the weight matrix by the adaptively adjusted feature φ ca to obtain the enhanced feature φ + ; which is expressed as:

[0087]

[0088] Among them, represents matrix multiplication. This process further refines the feature representation of the remote sensing image, enabling the model to more accurately segment different regions in the image. Especially in the case of diverse land cover types and complex backgrounds, it can effectively improve the accuracy and robustness of remote sensing image segmentation. By adding the spiking neuron layer, the module not only enhances the feature expression but also improves the flexibility and adaptability of the model. In the remote sensing image segmentation task, especially when facing images with complex textures and shape variations, it can significantly improve the segmentation accuracy.

[0089] In the above step S4, through the mapping unit, the enhanced feature is mapped to obtain the segmentation result corresponding to the target remote sensing image; specifically including:

[0090] (1) After spatially convolving and expanding the enhanced feature φ + through the fifth convolutional layer, perform spatial upsampling through the deconvolution layer to restore the image resolution and obtain the reconstructed feature map F;

[0091] Specifically, the enhanced feature φ + obtained through the adaptive feature enhancement unit already contains rich context information and multi-scale features. In the embodiments of the present invention, the fifth convolutional layer and the deconvolution layer are used to perform spatial reconstruction on the enhanced feature φ + to ensure that the features can adapt to remote sensing images of different resolutions and restore spatial details.

[0092] (2) Perform feature diffusion and consistency enhancement processing on the reconstructed feature map F to obtain the consistency-enhanced feature; specifically, in order to further strengthen the spatial context information, a diffusion-based feature fusion mechanism is introduced.

[0093] This feature diffusion and consistency enhancement mechanism constructs a feature diffusion field based on the heat conduction equation, including:

[0094] a) Diffusion kernel construction: Define the anisotropic diffusion kernel K d , whose weight distribution follows the edge-aware function where represents the spatial gradient of the reconstructed feature map F; z is the adaptive smoothing coefficient;

[0095] b) Iterative diffusion process: Establish partial differential equations Use the explicit Euler method for multiple iterations (e.g., 3 - 5 times), and each iteration includes:

[0096] Neighborhood feature difference calculation: Perform Sobel operator convolution on the reconstructed feature map to obtain the four - neighborhood gradient;

[0097] Diffusion coefficient update: Dynamically adjust the diffusion coefficient according to the current feature edge intensity;

[0098] Eigenvalue update: Based on the four - neighborhood gradient and the diffusion coefficient, update the reconstructed feature map to obtain the updated feature map; expressed as:

[0099]

[0100] where, Δt represents the time step, which represents the time interval from the current time t to the next time t + 1. It is a discretized time step used to gradually update the state of the field F in numerical simulation; F t+1 represents the updated eigenvalue at time t + 1; F t represents the eigenvalue before update at time t;

[0101] c) Weightedly fuse the updated feature map and the reconstructed feature map through a gating mechanism to generate the consistency - enhanced feature F';

[0102] The diffusion process aims to smooth the spatial information in the feature map and make the feature information of adjacent pixels more consistent, thereby improving the spatial continuity of the segmentation result. This process is similar to noise removal and texture smoothing in images, making the image segmentation smoother and more accurate in the edge region. After the diffusion process, the feature information will become more consistent, reducing noise and unnecessary local details.

[0103] (3) Pass the consistency - enhanced feature F' through the sixth convolutional layer and the third activation function in sequence to obtain the predicted segmentation result corresponding to the target remote - sensing image

[0104] To effectively train the model, a composite loss function combining cross - entropy loss and Dice coefficient loss is introduced in the embodiments of the present invention, aiming to optimize the segmentation accuracy and handle the problem of class imbalance. Specifically, the composite loss function is expressed as:

[0105]

[0106]

[0107] where, represents the composite loss function; is the cross-entropy loss, which is used to calculate the difference between the prediction result and the true label; represents the Dice coefficient loss, which is used to measure the overlap degree between the prediction result and the true label; λ CE represents the weighting coefficient of the cross-entropy loss; λ Dice represents the weighting coefficient of the Dice coefficient loss; Y represents the true label; represents the predicted value; Y i represents the probability value of the i-th class in the true label; represents the probability value of the i-th class in the predicted value; C represents a total of C classes; Y j represents the j-th pixel value of the true label; represents the j-th pixel value of the predicted value; N represents the total number of pixels in the target remote sensing image.

[0108] In the embodiments of the present invention, by introducing the weighting coefficients λ CE and λ Dice , the weights of the cross-entropy loss and the Dice coefficient loss can be flexibly adjusted, so as to adjust the optimization intensity of the loss function for different targets according to the characteristics of the dataset. During the actual training process, the values of λ CE and λ Dice can be adjusted according to the experimental results, and usually the most suitable weight coefficients are selected through cross-validation.

[0109] To further improve the performance of the model, the present invention adopts the Adam optimizer for training. The Adam optimizer can adaptively adjust the learning rate to meet the training requirements at different stages. During the training process, the parameters are optimized through multiple iterations to minimize the total loss function. During the training process, the model will continuously adjust the feature fusion method to ensure that the feature expression ability is sufficient to handle complex remote sensing image segmentation tasks, and at the same time gradually improve the segmentation accuracy through the optimization of the loss function.

[0110] In summary, a remote sensing image segmentation method based on heterogeneous two-stream feature extraction and adaptive feature booster provided by the present invention aims to improve the accuracy and robustness of remote sensing image segmentation. This method combines the advantages of the heterogeneous two-stream feature extraction network and the adaptive feature booster, and effectively captures local and global information in the image through multi-scale and multi-dimensional feature learning, thereby improving the segmentation accuracy. Specifically, the present invention solves the challenges in remote sensing image segmentation through the following key technical solutions:

[0111] First, the present invention employs a heterogeneous dual-stream feature extraction unit, which deeply mines the spatial features and semantic features in remote sensing images through the feature extraction networks of the upper branch and the lower branch respectively. The upper branch network focuses on capturing local detailed information in the image through modules such as depthwise separable convolution and windmill-shaped convolution; the lower branch network further extracts high-level semantic information in the image through the Kolmogorov-Arnold layer and depthwise separable convolution. Through the feature fusion of the upper and lower branches, the network can obtain rich features of remote sensing images at multiple scales and angles, thereby improving the segmentation performance.

[0112] Secondly, in order to further optimize the feature representation, the present invention proposes an adaptive feature enhancer. Through the combination of dynamic convolution and full-dimensional dynamic convolution, the network can adaptively adjust the shape and size of the convolution kernel to better capture the spatial information in remote sensing images. The operations of feature splicing and adaptive weighting enable features at different scales to be effectively fused under the guidance of global context information, reducing the interference of background noise and simultaneously enhancing the feature representation ability of the target region.

[0113] Generally speaking, through the innovative design of heterogeneous dual-stream feature extraction and the adaptive feature enhancer, the present invention solves the problems existing in traditional remote sensing image segmentation methods, such as insufficient capture of local information, insufficient feature fusion, and class imbalance, thereby significantly improving the accuracy and robustness of remote sensing image segmentation and having important application value.

[0114] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.

[0115] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A remote sensing image segmentation method based on heterogeneous dual-stream feature extraction and an adaptive feature enhancer, characterized in that It includes the following steps: Input the target remote sensing image into the trained remote sensing image segmentation model; the remote sensing image segmentation model is composed of a heterogeneous dual-stream feature extraction unit, an adaptive feature enhancement unit, and a mapping unit; Through the heterogeneous dual-stream feature extraction unit, extract features from the target remote sensing image to obtain upper-branch output features and lower-branch output features; Through the adaptive feature enhancement unit, fuse and enhance the upper-branch output features and the lower-branch output features to obtain enhanced features; Through the mapping unit, perform mapping processing on the enhanced features to obtain the segmentation result corresponding to the target remote sensing image.

2. The remote sensing image segmentation system based on heterogeneous two-stream feature extraction and adaptive feature booster according to claim 1, wherein The step of extracting features from the target remote sensing image through the heterogeneous dual-stream feature extraction unit to obtain upper-branch output features and lower-branch output features specifically includes: Use an initial convolutional layer to perform sliding convolution on the target remote sensing image to generate a preliminary feature map; Through the upper-branch feature extraction module, perform local feature extraction on the preliminary feature map to obtain upper-branch output features; Through the lower-branch feature extraction module, perform semantic feature extraction on the preliminary feature map to obtain lower-branch output features.

3. The remote sensing image segmentation system based on heterogeneous two-stream feature extraction and adaptive feature booster according to claim 2, wherein The step of performing local feature extraction on the preliminary feature map through the upper-branch feature extraction module to obtain upper-branch output features specifically includes: Pass the preliminary feature map through a first normalization layer, a first convolutional layer, a first depthwise separable convolutional layer, a dropout layer, and a windmill-shaped convolutional layer in sequence to obtain a direction-aware feature map; Through a first skip connection layer, fuse the direction-aware feature map and the preliminary feature map to obtain upper-branch output features.

4. The remote sensing image segmentation system based on heterogeneous two-stream feature extraction and adaptive feature booster according to claim 2, characterized in that, The step of performing semantic feature extraction on the preliminary feature map through the lower-branch feature extraction module to obtain lower-branch output features specifically includes: Perform tokenization processing on the preliminary feature map to obtain label features; Pass the label features through a Kolmogorov-Arnold layer, a second depthwise separable convolutional layer, and a second normalization layer in sequence to obtain normalized features; Through a second skip connection layer, fuse the normalized features and the preliminary feature map to obtain lower-branch output features.

5. The remote sensing image segmentation system based on heterogeneous two-stream feature extraction and adaptive feature booster according to claim 1, wherein The step of fusing and enhancing the upper-branch output features and the lower-branch output features through the adaptive feature enhancement unit to obtain enhanced features; Specifically includes: Through a concatenation layer, concatenate the upper-branch output features and the lower-branch output features to obtain concatenated features; Pass the concatenated features through a global average pooling layer, a second convolutional layer, and a first activation function in sequence to generate global context features; Multiply the global context features by the concatenated features and then pass through a fourth convolutional layer to obtain adaptive features; Use a dynamic deformable convolutional layer to process the upper-branch output features to generate an upper-branch feature map; Use a full-dimensional dynamic convolution to extract features from the lower-branch output features to obtain a lower-branch feature map; Through a third skip connection layer, fuse the upper-branch feature map and the lower-branch feature map to obtain fused features; Pass the fusion feature through a third convolutional layer, a spiking neuron layer, and a second activation function in sequence to generate a weight matrix; Multiply the weight matrix by the adaptive adjustment feature to obtain an enhanced feature.

6. The remote sensing image segmentation system based on heterogeneous two-stream feature extraction and adaptive feature booster according to claim 1, wherein Through the mapping unit, perform a mapping process on the enhanced feature to obtain the segmentation result corresponding to the target remote sensing image; specifically including: After spatially convolving and expanding the enhanced feature through a fifth convolutional layer, perform spatial upsampling through a deconvolution layer to obtain a reconstructed feature map; Perform feature diffusion and consistency enhancement processing on the reconstructed feature map to obtain a consistency-enhanced feature; Pass the consistency-enhanced feature through a sixth convolutional layer and a third activation function in sequence to obtain the predicted segmentation result corresponding to the target remote sensing image.

7. The remote sensing image segmentation system based on heterogeneous dual-stream feature extraction and adaptive feature booster according to claim 6, wherein The performing feature diffusion and consistency enhancement processing on the reconstructed feature map to obtain a consistency-enhanced feature; specifically including: Perform Sobel operator convolution processing on the reconstructed feature map to obtain a four-neighborhood gradient; Dynamically adjust the diffusion coefficient according to the current feature edge intensity; Based on the four-neighborhood gradient and the diffusion coefficient, update the reconstructed feature map to obtain an updated feature map; Perform weighted fusion on the updated feature map and the reconstructed feature map through a gating mechanism to generate a consistency-enhanced feature.

8. The remote sensing image segmentation system based on heterogeneous two-stream feature extraction and adaptive feature booster according to claim 1, wherein The composite loss function of the remote sensing image segmentation model is expressed as: Among them, represents the composite loss function; is the cross-entropy loss; represents the Dice coefficient loss; λ CE represents the weighting coefficient of the cross-entropy loss; λ Dice represents the weighting coefficient of the Dice coefficient loss; Y represents the true label; represents the predicted value; Y i represents the probability value of the i-th class in the true label; represents the probability value of the i-th class in the predicted value; C represents a total of C classes; Y j represents the j-th pixel value of the true label; represents the j-th pixel value of the predicted value; N represents the total number of pixels in the target remote sensing image.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method based on double-branch feature fusion

    CN115797931A

  • Multi-fusion comparison network based on full-dimensional dynamic convolution

    CN116757955A

  • Remote sensing image road segmentation method fusing multi-scale features and double attention mechanism

    CN117078943A

  • Remote sensing image semantic segmentation method based on double-branch dynamic attention

    CN117422878A

  • Remote sensing image semantic segmentation method and device based on dual-path encoder

    CN118314342A