Unmanned aerial vehicle remote sensing image segmentation method based on multi-modal fine-grained feature mining
By collaboratively mining fine-grained heterogeneous information in DSM through multiple modules, the problem of unbalanced information processing in multimodal remote sensing image segmentation is solved, and the accuracy and detail processing capabilities of UAV remote sensing image segmentation are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG UNIV
- Filing Date
- 2026-01-13
- Publication Date
- 2026-05-19
AI Technical Summary
Existing multimodal remote sensing image segmentation methods fail to distinguish between bimodal information processing and lack contextual modeling, resulting in auxiliary modes failing to effectively complement the main mode or interfering with it, thus limiting the improvement of UAV remote sensing image segmentation performance.
We employ a multi-module collaborative approach to mine fine-grained heterogeneous information in the DSM. By using a multi-scale local feature mining module and a correlation-driven local feature mining module, we enhance the model's ability to distinguish details. Furthermore, by combining a global feature mining module and a multi-source, multi-scale feature fusion module, we strengthen contextual modeling and feature fusion.
It significantly improves the accuracy and detail processing capabilities of UAV remote sensing image segmentation. Through multi-module collaborative mining and feature fusion, it enhances the model's global feature discrimination and detail processing capabilities.
Smart Images

Figure CN122066940A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a UAV remote sensing image segmentation method based on multimodal fine-grained feature mining, belonging to the field of UAV application technology. Background Technology
[0002] UAV remote sensing image segmentation is a fundamental and highly challenging task in the field of computer vision. Its core is to semantically group image pixels into meaningful regions, supporting numerous upstream applications such as land monitoring, agricultural management, road and building extraction, change detection, and disaster assessment. Compared to natural and medical images, remote sensing images present unique challenges, including significant changes in target scale, small inter-class differences, and large intra-class differences. Directly using image segmentation methods from other fields often fails to achieve optimal results.
[0003] In recent years, significant progress has been made in related research. Inspired by the success of fully convolutional networks (FCNs) in natural image segmentation, researchers have developed various end-to-end frameworks for remote sensing image segmentation, demonstrating superior performance. However, methods based on convolutional neural networks (CNNs) are limited by the receptive domain and have inherent limitations. With the emergence of visual transducers (ViTs), their powerful global context modeling capabilities have been successfully integrated into remote sensing image segmentation, effectively capturing long-range dependencies and significantly outperforming traditional CNN methods in complex scenes, achieving state-of-the-art performance.
[0004] Single-modal remote sensing images often fail to provide sufficient discriminative features, especially when distinguishing visually similar categories (such as buildings and roads). Therefore, multimodal methods combining remote sensing images with depth-sensing models (DSMs) offer a solution. The two provide surface detail information (texture, boundaries, etc.) and key depth information, respectively, enriching the data dimensions and enhancing category discrimination capabilities. However, existing multimodal methods process the two modalities indiscriminately and lack sufficient contextual modeling when mining complementary multimodal features. This results in the auxiliary modality DSM either failing to complement the primary modality or negatively impacting the main modality remote sensing image, thus limiting further improvements in UAV remote sensing image segmentation performance. Summary of the Invention
[0005] To address the problems of existing multimodal remote sensing image segmentation methods failing to distinguish between dual-modal information processing and lacking sufficient contextual modeling, which leads to auxiliary modalities failing to effectively complement the main modality or interfering with it, this invention proposes a UAV remote sensing image segmentation method based on multimodal fine-grained feature mining.
[0006] This invention emphasizes the dominant role of Relative Sensing Index (RSI), collaboratively mining fine-grained heterogeneous information in Distributed Sensitive Models (DSMs) through multiple modules, while strengthening contextual modeling and feature fusion to improve the accuracy and detail processing capabilities of UAV remote sensing image segmentation. Its core is multimodal fine-grained feature mining using only RSI features as query features. Specifically, this invention employs a multi-scale local feature mining module and a correlation-driven local feature mining module to fully mine local complementary features and enhance the model's detail discrimination capabilities. The multi-scale local feature mining module is used to construct robust multi-scale query features to mine multi-scale local complementary features from deep features. Under the supervision of the proposed correlation separation loss, the correlation-driven local feature mining module can extract deep local detail features with non-linear correlations from the DSM and provide rich representations for complementary features. Simultaneously, this method uses a global feature mining module to comprehensively extract deep and shallow global complementary features between modalities, thereby improving the model's global feature discrimination capability. Furthermore, this method proposes a multi-source, multi-scale feature fusion module to retain richer detail information during contextual modeling and improve the model's detail processing capabilities.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: A UAV remote sensing image segmentation method based on multimodal fine-grained feature mining includes the following steps: An image segmentation model is constructed, which includes a multi-scale feature extractor for RSI and DSM, a multi-scale fusion module and shallow global feature encoder for RSI, a multi-scale fusion module and shallow global feature encoder for DSM, a multi-scale local feature mining module, a correlation-driven local feature mining module, a global feature mining module, and a multi-source multi-scale feature fusion module. Train an image segmentation model; UAV remote sensing image segmentation is performed using a trained image segmentation model; including: Step S1: Extract multi-scale features of RSI and DSM using a multi-scale feature extractor for RSI and DSM; Step S2: Through the multi-scale fusion module of RSI and the shallow global feature encoder, the multi-scale fusion module generates two sizes of fused features, which are respectively fed into the shallow global feature encoder for shallow global feature encoding and the multi-source multi-scale feature fusion encoder and decoder of step S7 for context modeling. This provides a foundation for deep mining of multi-scale fine-grained features in auxiliary modal DSM and participates in global context modeling to enhance the expressive power of features. Step S3: Multi-scale feature fusion is performed through the multi-scale fusion module of DSM to generate a size feature, which is only fed into the shallow global feature encoder for encoding, providing a basis for the local and global complementary feature mining in steps S5 and S6. Step S4: Through the multi-scale local feature mining module, the scale-level fusion method is used to enhance the query features of the main modality RSI features, thereby deeply mining the multi-scale fine-grained features in the auxiliary modality DSM. Step S5: Through the correlation-driven local feature mining module, use a CNN-based deep detail encoder to encode the RSI and DSM detailed features, use the encoded RSI features as the main modality, mine the DSM encoded features through multimodal cross attention, and mine non-linearly correlated complementary detailed features under the supervision of correlation separation loss. Step S6: Through the global feature mining module, firstly, the RSI and DSM are encoded at a deep level using a Transformer-based deep feature encoder to obtain high-level visual semantic information. Then, multi-head cross attention is used to obtain deep global complementary features from the DSM. Secondly, multi-head cross attention is used to obtain shallow global complementary features from the DSM on the shallow global features output in step S2. Step S7: Improve the U-net structure based on the Swin-Transformer framework through the multi-source multi-scale feature fusion module to form a multi-source multi-scale feature fusion encoder and decoder. Perform contextual global modeling on the multi-scale features of the original RSI and DSM, as well as the mined local and global DSM heterogeneous information that is complementary to the RSI features, to obtain robust visual representation features.
[0008] According to a preferred embodiment of the present invention, the specific implementation process of step S1 includes: A multimodal remote sensing dataset is defined as a set of remote sensing data pairs, including RSI and DSM, where RSI is the primary mode and DSM is the secondary mode; this definition is expressed as: Each remote sensing data pair includes a multimodal image pair. and the corresponding ground truth semantic segmentation map ,in, RSI data refers to remote sensing images. This indicates DSM data, or Digital Surface Model data, where H and W represent height and width, respectively. Representing the set of real numbers; a bi-branch residual backbone network pre-trained using a semi-supervised method is used to extract multi-scale features from RSI and DSM, expressed as: ; in, and These represent the multi-scale feature extractors used for RSI and DSM, namely the bi-branch residual backbone network; This represents different residual blocks in a two-branch residual backbone network. ; and Representing the RSI and DSM respectively Individual scale features.
[0009] According to a preferred embodiment of the present invention, the specific implementation process of step S2 includes: The multi-scale fusion module FA-RSI of RSI is used to fuse multi-scale features of RSI and output features of two sizes.
[0010] First, the multi-scale features are upsampled or downsampled to scale to the same feature dimension, with sizes of [sizes to be filled in]. and Then, feature concatenation is performed, concatenating the features along the channel dimension. Subsequently, a feature mapping operation consisting of 3*3 convolution and batch normalization is used to map the number of feature channels to 256. The shallow global feature encoder performs shallow encoding and is then used for feature selection between the primary and secondary modalities. The expression is as follows: ; in, FA-RSI, representing the multi-scale fusion module of RSI; It is a feature mapping operation consisting of 3*3 convolution and batch normalization, which maps the number of feature channels to 256. This indicates a bilinear interpolation upsampling or downsampling operation. Indicates a connection operation; , That is, the characteristics of the two sizes; Shallow global feature encoder pairs We perform global context modeling and extract shallow global features of RSI, as shown in the following expression: ; in, This represents the shallow global features of the extracted RSI. This represents a shallow global feature encoder.
[0011] According to a preferred embodiment of the present invention, the specific implementation process of step S3 includes: The multi-scale fusion module FA-DSM of DSM fuses the multi-scale features of DSM. First, it upsamples or downsamples the multi-scale features of DSM to scale them to the same feature dimension. Then, feature concatenation is performed, stitching the features together along the channel dimension. Subsequently, feature mapping operations, including 3x3 convolution and batch normalization, are used to map the number of feature channels to 256, resulting in an output size of [missing information]. The fused features are then processed by a shallow global feature encoder for shallow feature encoding; the expression is as follows: ; in, This refers to the multi-scale fusion module FA-DSM. The representative size is The fusion characteristics It is a feature mapping operation consisting of 3*3 convolution and batch normalization, which maps the number of feature channels to 256. This indicates a bilinear interpolation upsampling or downsampling operation. Indicates a connection operation; A shallow global feature encoder is constructed to perform global context modeling on the fused features, and shallow global features of the DSM are extracted, as shown in the following expression: ; in, This represents the shallow global features extracted from the DSM. This represents a shallow global feature encoder.
[0012] According to a preferred embodiment of the present invention, the specific implementation process of step S4 includes: A multi-scale local feature mining module is constructed, and a multi-scale fusion query method is adopted to sequentially process the adjacent multi-scale local features of RSI from front to back. To enhance the main modality query features, the following expression is used: ; in, This indicates a multi-scale fusion query. Indicates a downsampled convolutional block. and These represent batch normalization and linear mapping, respectively. Indicates a connection operation. These represent enhanced main modality query features at different scales of mining. ; Before performing feature mining, first... and Perform flattening and transformation operations to generate a 2D block sequence. and ,in, , For dimensions; Next, we will use multi-head cross-attention to leverage As a query, it extracts refined complementary features from the DSM; It is each head for each scale feature, and for each head for each scale feature, it will and Mapped to a new space to obtain queries from RSI. and keys from DSM Sum : ; in, , and All dimensions are ; The number of attention heads is set to 8; , and These represent the learnable weights for different queries, keys, and values in each attention head; concatenating the queries, keys, and values of all heads yields the following new representation: ; in, , and These represent queries related to multimodal shallow local features. ,key Sum Corresponding learnable weights; Next, the modal relation matrix is calculated using the queries associated with RSI and the keys associated with DSM. Then, using the relation matrix The values associated with the DSM are used to obtain multi-scale local complementary features; the expression is as follows: ; in, This represents multi-scale local complementary features. express operate, Indicates matrix transpose. This represents the sum of the embedding dimensions of all heads; Finally, use a linear mapping to... Convert to Using the same dimensions, complementary features are ultimately extracted.
[0013] According to a preferred embodiment of the present invention, the specific implementation process of step S5 includes: A correlation-driven local feature mining module is constructed to obtain deep, detailed features with nonlinear correlations and extract complementary features with rich representation forms from the depth information. Using a detail convolutional neural network encoder as a deep detail encoder, it utilizes invertible neural networks to preserve input features and has near-lossless feature extraction capabilities: ; in, This indicates a depth detail encoder. and These represent the depth detail features of RSI and DSM, respectively.
[0014] Under the supervision of correlation separation loss, deep detail features of nonlinear correlation are obtained. and The expressions for query, key, value, and feature mining are as follows: ; ; in, , It is the discovery of complementary details with rich representational forms.
[0015] According to a preferred embodiment of the present invention, the specific implementation process of step S6 includes: A global feature mining module is constructed to perform complementary feature mining on deep and shallow global features. Shallow global features are obtained through a shallow global feature encoder, using RSI features as queries, and global complementary features are mined from the DSM using MCA. The expressions for query, key, value, and feature mining are as follows: ; ; in, , This represents the shallow global complementary features discovered. For mining deep global features, a lightweight ViT-based deep global encoder (DGE) is used; this design achieves a balance between the ability to extract long-range dependent features and computational complexity, as shown in the following expression:
[0016] in, This represents the Depth Global Encoder (DGE). and These represent the deep global features extracted from RSI and DSM, respectively. MCA is used to extract deep global complementary features from deep information. The formulas for query, key, value, and feature mining are as follows: ; ; in, , This represents the deep, globally complementary features that have been discovered.
[0017] According to a preferred embodiment of the present invention, step S7 is specifically implemented as follows: All extracted local and global complementary features are fused using a multi-scale feature fusion module, outputting fused features. The multi-scale feature fusion module upsamples all extracted local and global complementary features to a size of [size missing]. Then, the components are concatenated along the channel dimension; next, dimensionality reduction is performed, as shown in the following expression:
[0018] ; in, This represents the fusion result of local and global complementary features from the DSM. This represents the fusion result of features output from the multi-scale local feature mining module. This indicates a connection operation; Up indicates an upsampling operation. Finally, RSI features The data is then fused, and a tokenization operation is performed to obtain the input tokens for the feature encoder: ; in, Indicates the input marker, It is a flattening operation. Indicates a convolutional layer. Indicates a cascading operation. This indicates a tokenization operation; During the encoding phase, MMF integrates multimodal local detail features and global deep encoding features obtained from the Swing Transformer at different scales; the expression is: ; in, This represents the Multi-Source Multi-Scale Feature Fusion Module (MMF). Indicates the first in Swing-T The encoded result of each block output, and Indicates the first The fusion information of each MMF, This represents the i-th complementary feature mined by the multi-scale local feature mining module. This represents the (i+1)th local feature of the remote sensing image; The entire encoding and decoding process can be represented in the following form: ; in, This represents a multi-source, multi-scale feature fusion encoder / decoder. This indicates the final segmentation result. Indicates the number of channels. It equals 1.
[0019] According to a preferred embodiment of the present invention, the decoder of the Swin-Unet is used for the corresponding RSI and DSM feature extractors. and Construct auxiliary decoding branches, denoted as follows: and The segmentation result based on the auxiliary decoding branch is represented as follows: ; ; in, and These represent the segmentation results of the auxiliary decoding branches of RSI and DSM, respectively.
[0020] According to a preferred embodiment of the present invention, the loss function of the image segmentation model includes relevance separation loss, semantic segmentation loss, and loss of auxiliary training strategies, thereby supervising the fine-grained feature extraction of the image segmentation model from multiple perspectives; specifically including: loss function It consists of two parts: semantic segmentation loss. and correlation separation loss Semantic segmentation uses the cross-entropy loss function. Cross-entropy loss function This includes the main branch decoding loss and the auxiliary branch decoding loss for RSI and DSM; the expression is as follows: ; ; in, , and Balance coefficients representing different losses; and These represent the number of pixels and the number of categories, respectively. This represents the semantic segmentation loss of the main branch. This represents the RSI auxiliary branch splitting loss. This represents the DSM auxiliary branch splitting loss. and Representing pixels Category The true confidence level and the confidence level predicted by the image segmentation model; Correlation separation loss The expression is as follows: ; in, This indicates the correlation coefficient calculation. It is a scaling factor. and These represent the depth detail features of RSI and DSM, respectively; The overall loss function of the image segmentation model Represented as: ; in, and These represent the balance coefficients.
[0021] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the above-described UAV remote sensing image segmentation method based on multimodal fine-grained feature mining.
[0022] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described UAV remote sensing image segmentation method based on multimodal fine-grained feature mining.
[0023] Compared with the prior art, the present invention has the following advantages: 0. Compared with existing technologies, this invention proposes a UAV remote sensing image segmentation method based on multimodal fine-grained feature mining. This method effectively solves the problems of existing multimodal remote sensing image segmentation methods failing to distinguish between dual-modal information processing and lacking sufficient contextual modeling, thus leading to auxiliary modalities failing to play a complementary role or interfering with the primary modality. This method consistently emphasizes the dominant role of RSI (Real-Time Sequence Index), focusing solely on using RSI features as query features for multimodal fine-grained feature mining. Through multi-module collaborative mining of fine-grained heterogeneous information in the DSM (Discrete Synthetic Model), and strengthening contextual modeling and feature fusion, it significantly improves the accuracy and detail processing capabilities of UAV remote sensing image segmentation.
[0024] 2. Compared with existing technologies, this invention proposes a UAV remote sensing image segmentation method based on multimodal fine-grained feature mining. It innovatively employs a multi-scale local feature mining module and a correlation-driven local feature mining module to fully mine local complementary features and enhance the model's detail discrimination capability. The correlation-driven local feature mining module can construct robust multi-scale query features and, under the supervision of correlation separation loss, extract deep local detail features with non-linear correlations from the DSM. Simultaneously, it is combined with a global feature mining module to comprehensively extract deep and shallow global complementary features between modalities. Furthermore, a multi-source, multi-scale feature fusion module preserves detailed information from contextual modeling, comprehensively improving the model's global feature discrimination and detail processing capabilities. Attached Figure Description
[0025] Figure 1 This is the overall network structure diagram of the image segmentation model of the present invention; Figure 2 This is a qualitative performance comparison graph on the Potsdam dataset. Detailed Implementation
[0026] The present invention will be further defined below with reference to the accompanying drawings and embodiments, but is not limited thereto.
[0027] Terminology Explanation: 1. RSI, remote sensing image; 2. DSM, Digital Surface Model; 3. Dual-branch Residual Neural Network (ResNet) backbone network; Example 1 A UAV remote sensing image segmentation method based on multimodal fine-grained feature mining includes the following steps: An image segmentation model is constructed, which includes a multi-scale feature extractor for RSI and DSM, a multi-scale fusion module and shallow global feature encoder for RSI, a multi-scale fusion module and shallow global feature encoder for DSM, a multi-scale local feature mining module, a correlation-driven local feature mining module, a global feature mining module, and a multi-source multi-scale feature fusion module. Train an image segmentation model; UAV remote sensing image segmentation is performed using a trained image segmentation model; including: Step S1: Extract multi-scale features of RSI and DSM using a multi-scale feature extractor for RSI and DSM; Step S2: Through the multi-scale fusion module of RSI and the shallow global feature encoder, the multi-scale fusion module generates two sizes of fused features, which are respectively fed into the shallow global feature encoder for shallow global feature encoding and the multi-source multi-scale feature fusion encoder and decoder of step S7 for context modeling. This provides a foundation for deep mining of multi-scale fine-grained features in auxiliary modal DSM and participates in global context modeling to enhance the expressive power of features. Step S3: The multi-scale fusion module of DSM and the shallow global feature encoder provide the foundation for deep mining of multi-scale fine-grained features in the auxiliary modality DSM in the main modality; multi-scale feature fusion is performed by the multi-scale fusion module of DSM to generate a size feature, which is only fed into the shallow global feature encoder for encoding, providing the foundation for the mining of complementary local and global features in steps S5 and S6; the subsequent correlation-driven local feature mining module uses RSI and shallow global features of DSM to perform deep detail encoding respectively, and then uses cross attention on the supervision line of the correlation separation loss to extract non-linear correlation complementary details.
[0028] The multi-scale fusion modules of RSI and DSM only perform multi-scale feature fusion. Specifically, the multi-scale feature fusion of RSI produces two sizes of features, which are respectively fed into the multi-source multi-scale feature fusion encoder / decoder for context modeling and the shallow global feature encoder for shallow global feature encoding.
[0029] Step S4: Through the multi-scale local feature mining module, the scale-level fusion method is used to enhance the query features of the main modality RSI features, thereby deeply mining the multi-scale fine-grained features in the auxiliary modality DSM. Step S5: Through the correlation-driven local feature mining module, use a CNN-based deep detail encoder to encode the RSI and DSM detailed features, use the encoded RSI features as the main modality, mine the DSM encoded features through multimodal cross attention, and mine non-linearly correlated complementary detailed features under the supervision of correlation separation loss. Step S6: Through the global feature mining module, firstly, the RSI and DSM are encoded at a deep level using a Transformer-based deep feature encoder to obtain high-level visual semantic information. Then, multi-head cross attention is used to obtain deep global complementary features from the DSM. Secondly, multi-head cross attention is used to obtain shallow global complementary features from the DSM on the shallow global features output in step S2. Step S7: Improve the U-net structure based on the Swin-Transformer framework through the multi-source multi-scale feature fusion module to form a multi-source multi-scale feature fusion encoder and decoder. Perform contextual global modeling on the multi-scale features of the original RSI and DSM, as well as the mined local and global DSM heterogeneous information that is complementary to the RSI features, to obtain robust visual representation features.
[0030] Example 2 The difference between the UAV remote sensing image segmentation method based on multimodal fine-grained feature mining described in Example 1 and the method described in Example 1 is as follows: The specific implementation process of step S1 includes: A multimodal remote sensing dataset is defined as a set of remote sensing data pairs, including RSI and DSM, where RSI is the primary mode and DSM is the secondary mode; this definition is expressed as: Each remote sensing data pair includes a multimodal image pair. and the corresponding ground truth semantic segmentation map ,in, RSI data refers to remote sensing images. This indicates DSM data, or Digital Surface Model data, where H and W represent height and width, respectively. Represents the set of real numbers; MFE is a multi-scale feature extractor, including a two-branch residual backbone network, a special RSI feature fusion module and a DSM feature fusion module, as well as corresponding shallow global feature encoders for RSI and DSM; a two-branch residual backbone network (ResNet) pre-trained based on a semi-supervised method is used to extract multi-scale features from RSI and DSM, expressed as: ; in, and These represent the multi-scale feature extractors used for RSI and DSM, namely the bi-branch residual backbone network; This represents different residual blocks in a two-branch residual backbone network. ; and Representing the RSI and DSM respectively Individual scale features.
[0031] , All of these are publicly available CNN-based residual networks used for feature extraction.
[0032] The specific implementation process of step S2 includes: The multi-scale fusion module FA-RSI of RSI is used to fuse multi-scale features of RSI and output features of two sizes.
[0033] First, the multi-scale features are upsampled or downsampled to scale to the same feature dimension, with sizes of [sizes to be filled in]. and Then, feature concatenation is performed, concatenating the features along the channel dimension. Subsequently, a feature mapping operation consisting of 3*3 convolution and batch normalization is used to map the number of feature channels to 256. A shallow global feature encoder (existing, see Cddfuse: Correlation-driven dual-branch feature decomposition for multi-modality image fusion, which uses the shared feature encoder from that paper as a shallow global feature encoder) is used for shallow encoding and then used for feature selection between the primary and secondary modalities. The expression is as follows: ; in, FA-RSI, representing the multi-scale fusion module of RSI; It is a feature mapping operation consisting of 3*3 convolution and batch normalization, which maps the number of feature channels to 256. This indicates a bilinear interpolation upsampling or downsampling operation. Indicates a connection operation; , That is, the characteristics of the two sizes; Shallow global feature encoder pairs We perform global context modeling and extract shallow global features of RSI, as shown in the following expression: ; in, This represents the shallow global features of the extracted RSI. This represents a shallow global feature encoder.
[0034] The specific implementation process of step S3 includes: The multi-scale fusion module FA-DSM of DSM fuses the multi-scale features of DSM. First, it upsamples or downsamples the multi-scale features of DSM to scale them to the same feature dimension. Then, feature concatenation is performed, stitching the features together along the channel dimension. Subsequently, feature mapping operations, including 3x3 convolution and batch normalization, are used to map the number of feature channels to 256, resulting in an output size of [missing information]. The fused features are then processed by a shallow global feature encoder for shallow feature encoding; this provides the foundation for subsequent deep mining of multi-scale fine-grained features in the main modality DSM, as shown in the following expression: ; in, This refers to the multi-scale fusion module FA-DSM. The representative size is The fusion characteristics It is a feature mapping operation consisting of 3*3 convolution and batch normalization, which maps the number of feature channels to 256. This indicates a bilinear interpolation upsampling or downsampling operation. Indicates a connection operation; A shallow global feature encoder is constructed to perform global context modeling on the fused features, and shallow global features of the DSM are extracted, as shown in the following expression: ; in, This represents the shallow global features extracted from the DSM. This represents a shallow global feature encoder.
[0035] The specific implementation process of step S4 includes: A multi-scale local feature mining module is constructed, and a multi-scale fusion query method is adopted to sequentially process the adjacent multi-scale local features of RSI from front to back. To enhance the main modality query features, the following expression is used: ; in, This indicates a multi-scale fusion query. Indicates a downsampled convolutional block. and These represent batch normalization and linear mapping, respectively. Indicates a connection operation. These represent enhanced main modality query features at different scales of mining. ; It is worth noting that, in order to reduce computational costs while maintaining the richness of the mined information, feature mining is performed only on multi-scale features, excluding features from the first stage. Before performing feature mining, the features are first... and Perform flattening and transformation operations to generate a 2D block sequence. and ,in, , For dimensions;
[0036] Next, we will use Multi-Head Cross Attention (MCA) to leverage As a query, it extracts refined complementary features from the DSM; It is each head for each scale feature, and for each head for each scale feature, it will and Mapped to a new space to obtain queries from RSI. and keys from DSM Sum : ; in, , and All dimensions are ; The number of attention heads is set to 8; , and These represent the learnable weights for different queries, keys, and values in each attention head; since the operations of all heads are the same, subscripts will be omitted for ease of discussion and description. By concatenating all the header queries, keys, and values, we obtain the following new representation:
[0037] ; in, , and These represent queries related to multimodal shallow local features. ,key Sum Corresponding learnable weights; Next, the modal relation matrix is calculated using the queries associated with RSI and the keys associated with DSM. Then, using the relation matrix The values associated with the DSM are used to obtain multi-scale local complementary features; the expression is as follows: ; in, This represents multi-scale local complementary features. express operate, Indicates matrix transpose. This represents the sum of the embedding dimensions of all heads; Finally, use a linear mapping to... Convert to Using the same dimensions, complementary features are ultimately extracted.
[0038] The specific implementation process of step S5 includes: A correlation-driven local feature mining module is constructed to obtain deep, detailed features with nonlinear correlations and extract complementary features with rich representation forms from the depth information; the correlation-driven local feature mining module is supervised by correlation separation loss.
[0039] Using a detail convolutional neural network encoder (see Cddfuse: Correlation-driven dual-branch feature decomposition for multi-modality image fusion) as a deep detail encoder (DDE) leverages invertible neural networks to better preserve input features and possesses near-lossless feature extraction capabilities. ; in, This indicates a depth detail encoder. and These represent the depth detail features of RSI and DSM, respectively.
[0040] Under the supervision of correlation separation loss, deep detail features of nonlinear correlation are obtained. and Similar to the feature mining process in the multi-scale local feature mining section, MCA is still used for feature mining. The core idea is to use... Mining as a query Advanced complementary features with nonlinear correlations. The expressions for query, key, value, and feature mining are as follows:
[0041] ; ; in, , It is the discovery of complementary details with rich representational forms.
[0042] The specific implementation process of step S6 includes: A global feature mining module is constructed to perform complementary feature mining on deep and shallow global features. This enables a more comprehensive capture of the global structural features of multimodal data and enhances the model's global discriminative ability. Shallow global features are obtained through a shallow global feature encoder, using RSI features as queries, and global complementary features are mined from the DSM using MCA. The expressions for query, key, value, and feature mining are as follows:
[0043] ; ; in, , This represents the shallow global complementary features discovered. For mining deep global features, a lightweight ViT-based deep global encoder, DGE (using the base Transformer encoder from Cddfuse: Correlation-driven dual-branch feature decomposition for multi-modality image fusion), is used. This design achieves a balance between the ability to extract long-range dependent features and computational complexity, as shown in the following expression:
[0044] in, This represents the Depth Global Encoder (DGE). and These represent the deep global features extracted from RSI and DSM, respectively. MCA is used to extract deep global complementary features from deep information. The formulas for query, key, value, and feature mining are as follows: ; ; in, , This represents the deep, globally complementary features that have been discovered.
[0045] The specific implementation process of step S7 includes: A multi-source, multi-scale feature fusion codec is constructed. This codec integrates the Multi-Source, Multi-Scale Feature Fusion (MMF) module into the Swin-Unet framework. Its core framework utilizes the advanced Swin-Unet framework, upon which the MMF module is built to enhance the model's detail processing capabilities. Figure 1In the multi-source, multi-scale feature fusion encoder-decoder section, the encoder is located in the middle and on the left, while the decoder is on the far right. Before encoding, the feature fusion module fuses all mined complementary features with features from RSI. During the encoding stage, four Swin Transformers (Swin-T) are used to encode the features. First, a Swin-T block is used for encoding, then downsampled by the block merging module. The output downsampled features are then fused with the fine-grained local features mined by the multi-scale local feature mining module and the features extracted by the multi-scale feature extractor through the multi-source, multi-scale feature fusion module. The fused features are then fed back into the Swin-T for encoding. This process is repeated until encoding is complete. It is worth noting that the fine-grained local features mined by the multi-scale local feature mining module and the features extracted by the multi-scale feature extractor in each multi-source, multi-scale feature fusion module are features from different feature stages and are therefore different. For details, please refer to the specific implementation formula of MMF. In the decoding stage, four Swing-T layers are used to decode the model. The Swing-T results from the encoder stage at the corresponding level are concatenated to the Swing-T layers in the decoding stage. Each Swing-T layer is then followed by a block expansion module for upsampling. Finally, the segmentation result is output through a linear mapping layer. For detailed implementation of Swing-Unet, please refer to the paper (Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation).
[0046] Specifically, all extracted local and global complementary features are fused using a multi-scale feature fusion module to output the fused features. The multi-scale feature fusion module upsamples all extracted local and global complementary features to a size of [size missing]. Then, the components are concatenated along the channel dimension; next, dimensionality reduction is performed to improve computational efficiency; the expression is as follows:
[0047] ; in, This represents the fusion result of local and global complementary features from the DSM. This represents the fusion result of features output from the multi-scale local feature mining module. This indicates a connection operation; Up indicates an upsampling operation. Finally, RSI features The input data is fused and tokenization is performed (the input data is parsed into a normalized token sequence using a tokenizer) to obtain the input tokens for the feature encoder: ; in, Indicates the input marker, It is a flattening operation. Indicates a convolutional layer. Indicates a cascading operation. This indicates a tokenization operation; During the encoding phase, MMF integrates multimodal local detail features and global deep encoding features obtained at different scales from the Swing Transformer (Swin-T); this effectively enhances the preservation of local details, thereby significantly improving the model's detail processing capabilities, as expressed in the following expression: ; in, This represents the Multi-Source Multi-Scale Feature Fusion Module (MMF). Indicates the first in Swing-T The encoded result of each block output, and Indicates the first The fusion information of each MMF, This represents the i-th complementary feature mined by the multi-scale local feature mining module. This represents the (i+1)th local feature of the remote sensing image; The entire encoding and decoding process can be represented in the following form: ; in, This represents the Multi-Source Multi-Scale Feature Fusion Encoder-Decoder (MMF-SUnet). This indicates the final segmentation result. Indicates the number of channels. It equals 1.
[0048] To extract richer local detail information, the Swin-Unet decoder is used to power the corresponding RSI and DSM feature extractors. and Construct auxiliary decoding branches, denoted as follows: and The segmentation result based on the auxiliary decoding branch is represented as follows: ; ; in, and These represent the segmentation results of the auxiliary decoding branches of RSI and DSM, respectively.
[0049] It's important to note that the auxiliary decoders are only used during training. They are removed during inference, so they do not add any additional computational cost during inference.
[0050] The loss function of the image segmentation model includes relevance separation loss, semantic segmentation loss, and loss from auxiliary training strategies, supervising the fine-grained feature extraction of the image segmentation model from multiple perspectives; specifically including: loss function It consists of two parts: semantic segmentation loss. and correlation separation loss Semantic segmentation uses the cross-entropy loss function. Cross-entropy loss function This includes the main branch decoding loss and the auxiliary branch decoding loss for RSI and DSM; the expression is as follows: ; ; in, , and The balance coefficients representing different losses are 2, 0.5, and 0.5, respectively. and These represent the number of pixels and the number of categories, respectively. This represents the semantic segmentation loss of the main branch. This represents the RSI auxiliary branch splitting loss. This represents the DSM auxiliary branch splitting loss. and Representing pixels Category The true confidence level and the confidence level predicted by the image segmentation model; Correlation separation loss The expression is as follows: ; in, This indicates the correlation coefficient calculation. It is a scaling factor. and These represent the depth detail features of RSI and DSM, respectively; The overall loss function of the image segmentation model Represented as: ; in, and These represent balance coefficients, 1 and 0.1 respectively.
[0051] Table 1 presents the quantitative results of the model; Table 1. Comparison of the proposed method with other advanced methods on the Potsdam dataset;
[0052] On public datasets, different methods train their models using the same training set and test their models using the same test set. Following the domain's setup on public datasets, comparisons are made across five foreground categories: impervious surface, building, low vegetation, tree, and car (or, for brevity, background). The IoU and F1 scores for each category are compared, as well as the mF1 and mIoU scores across multiple categories. Specifically, the F1 score is calculated as follows:
[0053] ; ; ; in, , and These represent true positive, false positive, and false negative, respectively.
[0054] The formula for calculating IoU is as follows: ; mF1 and mIoU can be represented in the following forms: ; ; in, It represents the number of semantic segmentation categories.
[0055] MFFMNet is the model of this invention, and it is compared with several other models. As shown in Table 1, on the Potsdam dataset, the method of this invention achieves the best IoU and F1 score for each class, and the best mF1 and mIoU for multiple classes. This demonstrates the advancement of the method of this invention, which can improve the performance of the model through multimodal complementary feature mining.
[0056] A direct visual comparison of different methods, such as... Figure 2 As shown. Figure 2 In the middle (a), the image is a remote sensing image. The area within the red rectangle contains the impervious surface and low vegetation categories, which are visually similar, making accurate segmentation challenging. Figure 2 In the middle (b), the true value of the segmentation result is shown. Figure 2 (c) shows the performance graph obtained using the existing ABCNet; Figure 2(d) shows the performance graph obtained using the existing Unetformer; Figure 2 (e) shows the performance graph obtained using the existing DC-Swin. Figure 2 (f) shows the performance graph obtained using the existing RS3Mamba; Figure 2 (g) shows the performance graph obtained using the existing ASMFNet; Figure 2 The middle (h) graph shows the performance obtained using the existing CMFNet. Figure 2 In the middle (i), the performance graph obtained using the existing FTransUNet is shown. Figure 2 (j) is the performance graph obtained using the MFFMNet of this invention; from Figure 2 As can be seen from (c)-(i), none of the methods can achieve accurate semantic segmentation. Figure 2 As shown in Figure (j), MFFMNet can accurately distinguish easily confused categories and effectively handle edges. This further demonstrates that the proposed MFFMNet can extract and fully utilize multimodal discriminative features and fine-grained local details, thereby improving the performance of dense pixel classification tasks.
[0057] Example 3 A computer device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the UAV remote sensing image segmentation method based on multimodal fine-grained feature mining as described in Embodiment 1 or 2.
[0058] Example 4 A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the UAV remote sensing image segmentation method based on multimodal fine-grained feature mining as described in Embodiment 1 or 2.
Claims
1. A UAV remote sensing image segmentation method based on multimodal fine-grained feature mining, characterized in that, Includes the following steps: An image segmentation model is constructed, which includes a multi-scale feature extractor for RSI and DSM, a multi-scale fusion module and shallow global feature encoder for RSI, a multi-scale fusion module and shallow global feature encoder for DSM, a multi-scale local feature mining module, a correlation-driven local feature mining module, a global feature mining module, and a multi-source multi-scale feature fusion module. Train an image segmentation model; UAV remote sensing image segmentation is performed using a trained image segmentation model; including: Step S1: Extract multi-scale features of RSI and DSM using a multi-scale feature extractor for RSI and DSM; Step S2: Through the multi-scale fusion module of RSI and the shallow global feature encoder, the multi-scale fusion module generates two sizes of fused features, which are respectively fed into the shallow global feature encoder for shallow global feature encoding and the multi-source multi-scale feature fusion encoder and decoder of step S7 for context modeling. This provides a foundation for deep mining of multi-scale fine-grained features in auxiliary modal DSM and participates in global context modeling to enhance the expressive power of features. Step S3: Multi-scale feature fusion is performed through the multi-scale fusion module of DSM to generate a size feature, which is only fed into the shallow global feature encoder for encoding, providing a foundation for the local and global complementary feature mining in steps S5 and S6; Step S4: The query features of the main modality RSI feature are enhanced by the scale-level fusion method through the multi-scale local feature mining module, thereby deeply mining the multi-scale fine-grained features in the auxiliary modality DSM; Step S5: Through the correlation-driven local feature mining module, use a CNN-based deep detail encoder to encode the RSI and DSM detailed features, use the encoded RSI features as the main modality, mine the DSM encoded features through multimodal cross attention, and mine non-linearly correlated complementary detailed features under the supervision of correlation separation loss. Step S6: Through the global feature mining module, firstly, the RSI and DSM are encoded at a deep level using a Transformer-based deep feature encoder to obtain high-level visual semantic information. Then, multi-head cross attention is used to obtain deep global complementary features from the DSM. Secondly, multi-head cross attention is used to obtain shallow global complementary features from the DSM on the shallow global features output in step S2. Step S7: Improve the U-net structure based on the Swin-Transformer framework through the multi-source multi-scale feature fusion module to form a multi-source multi-scale feature fusion encoder and decoder. Perform contextual global modeling on the multi-scale features of the original RSI and DSM, as well as the mined local and global DSM heterogeneous information that is complementary to the RSI features, to obtain robust visual representation features.
2. The UAV remote sensing image segmentation method based on multimodal fine-grained feature mining according to claim 1, characterized in that, The specific implementation process of step S1 includes: A multimodal remote sensing dataset is defined as a set of remote sensing data pairs, including RSI and DSM, where RSI is the primary mode and DSM is the secondary mode; this definition is expressed as: Each remote sensing data pair includes a multimodal image pair. and the corresponding ground truth semantic segmentation map ,in, RSI data refers to remote sensing images. This indicates DSM data, or Digital Surface Model data, where H and W represent height and width, respectively. Representing the set of real numbers; a bi-branch residual backbone network pre-trained using a semi-supervised method is used to extract multi-scale features from RSI and DSM, expressed as: ; in, and These represent the multi-scale feature extractors used for RSI and DSM, namely the bi-branch residual backbone network; This represents different residual blocks in a two-branch residual backbone network. ; and Representing the RSI and DSM respectively Individual scale features.
3. The UAV remote sensing image segmentation method based on multimodal fine-grained feature mining according to claim 1, characterized in that, The specific implementation process of step S2 includes: The multi-scale fusion module FA-RSI of RSI is used to fuse multi-scale features of RSI and output features of two sizes. First, the multi-scale features are upsampled or downsampled to scale to the same feature dimension, with sizes of [sizes to be filled in]. and Then, feature concatenation is performed, concatenating the features along the channel dimension. Subsequently, a feature mapping operation consisting of 3*3 convolution and batch normalization is used to map the number of feature channels to 256. The shallow global feature encoder performs shallow encoding and is then used for feature selection between the primary and secondary modalities. The expression is as follows: ; in, FA-RSI, representing the multi-scale fusion module of RSI; It is a feature mapping operation consisting of 3*3 convolution and batch normalization, which maps the number of feature channels to 256. This indicates a bilinear interpolation upsampling or downsampling operation. Indicates a connection operation; , That is, the characteristics of the two sizes; Shallow global feature encoder pairs We perform global context modeling and extract shallow global features of RSI, as shown in the following expression: ; in, This represents the shallow global features of the extracted RSI. This represents a shallow global feature encoder.
4. The UAV remote sensing image segmentation method based on multimodal fine-grained feature mining according to claim 1, characterized in that, The specific implementation process of step S3 includes: The multi-scale fusion module FA-DSM of DSM fuses the multi-scale features of DSM. First, it upsamples or downsamples the multi-scale features of DSM to scale them to the same feature dimension. Then, feature concatenation is performed, stitching the features together along the channel dimension. Subsequently, feature mapping operations, including 3x3 convolution and batch normalization, are used to map the number of feature channels to 256, resulting in an output size of [missing information]. The fused features are then processed by a shallow global feature encoder for shallow feature encoding; the expression is as follows: ; in, This refers to the multi-scale fusion module FA-DSM. The representative size is The fusion characteristics It is a feature mapping operation consisting of 3*3 convolution and batch normalization, which maps the number of feature channels to 256. This indicates a bilinear interpolation upsampling or downsampling operation. Indicates a connection operation; A shallow global feature encoder is constructed to perform global context modeling on the fused features, and shallow global features of the DSM are extracted, as shown in the following expression: ; in, This represents the shallow global features extracted from the DSM. This represents a shallow global feature encoder.
5. The UAV remote sensing image segmentation method based on multimodal fine-grained feature mining according to claim 1, characterized in that, The specific implementation process of step S4 includes: A multi-scale local feature mining module is constructed, and a multi-scale fusion query method is adopted to sequentially process the adjacent multi-scale local features of RSI from front to back. To enhance the main modality query features, the following expression is used: ; in, This indicates a multi-scale fusion query. Indicates a downsampled convolutional block. and These represent batch normalization and linear mapping, respectively. Indicates a connection operation. These represent enhanced main modality query features at different scales of mining. ; Before performing feature mining, first... and Perform flattening and transformation operations to generate a 2D block sequence. and ,in, , For dimensions; Next, we will use multi-head cross-attention to leverage As a query, it extracts refined complementary features from the DSM; It is each head for each scale feature, and for each head for each scale feature, it will and Mapped to a new space to obtain queries from RSI. and keys from DSM Sum : ; in, , and All dimensions are ; The number of attention heads is set to 8; , and These represent the learnable weights for different queries, keys, and values in each attention head; concatenating the queries, keys, and values of all heads yields the following new representation: ; in, , and These represent queries related to multimodal shallow local features. ,key Sum Corresponding learnable weights; Next, the modal relation matrix is calculated using the queries associated with RSI and the keys associated with DSM. Then, using the relation matrix The values associated with the DSM are used to obtain multi-scale local complementary features; the expression is as follows: ; in, This represents multi-scale local complementary features. express operate, Indicates matrix transpose. This represents the sum of the embedding dimensions of all heads; Finally, use a linear mapping to... Convert to Using the same dimensions, complementary features are ultimately extracted.
6. The UAV remote sensing image segmentation method based on multimodal fine-grained feature mining according to claim 1, characterized in that, The specific implementation process of step S5 includes: A correlation-driven local feature mining module is constructed to obtain deep, detailed features with nonlinear correlations and extract complementary features with rich representation forms from the depth information. Using a detail convolutional neural network encoder as a deep detail encoder, it utilizes invertible neural networks to preserve input features and has near-lossless feature extraction capabilities: ; in, This indicates a depth detail encoder. and These represent the depth detail features of RSI and DSM, respectively; Under the supervision of correlation separation loss, deep detail features of nonlinear correlation are obtained. and The expressions for query, key, value, and feature mining are as follows: ; ; in, , It is the discovery of complementary details with rich representational forms; The specific implementation process of step S6 includes: A global feature mining module is constructed to perform complementary feature mining on deep and shallow global features. Shallow global features are obtained through a shallow global feature encoder, using RSI features as queries, and global complementary features are mined from the DSM using MCA. The expressions for query, key, value, and feature mining are as follows: ; ; in, , This represents the shallow global complementary features discovered. For mining deep global features, a lightweight ViT-based deep global encoder (DGE) is used; this design achieves a balance between the ability to extract long-range dependent features and computational complexity, as shown in the following expression: ; in, This represents the Depth Global Encoder (DGE). and These represent the deep global features extracted from RSI and DSM, respectively. MCA is used to extract deep global complementary features from deep information. The formulas for query, key, value, and feature mining are as follows: ; ; in, , This represents the deep, globally complementary features that have been discovered.
7. The UAV remote sensing image segmentation method based on multimodal fine-grained feature mining according to claim 1, characterized in that, The specific implementation process of step S7 includes: All extracted local and global complementary features are fused using a multi-scale feature fusion module to output fused features; the multi-scale feature fusion module upsamples all extracted local and global complementary features to a size of Then, the components are concatenated along the channel dimension; next, dimensionality reduction is performed, as shown in the following expression: ; in, This represents the fusion result of local and global complementary features from the DSM. This represents the fusion result of features output from the multi-scale local feature mining module. This indicates a connection operation; Up indicates an upsampling operation. Finally, RSI features The data is then fused, and a tokenization operation is performed to obtain the input tokens for the feature encoder: ; in, Indicates the input marker, It is a flattening operation. Indicates a convolutional layer. Indicates a cascading operation. This indicates a tokenization operation; During the encoding phase, MMF integrates multimodal local detail features and global deep encoding features obtained from the Swing Transformer at different scales; the expression is: ; in, This represents the Multi-Source Multi-Scale Feature Fusion Module (MMF). Indicates the first in Swing-T The encoded result of each block output, and Indicates the first The fusion information of each MMF, This represents the i-th complementary feature mined by the multi-scale local feature mining module. This represents the (i+1)th local feature of the remote sensing image; The entire encoding and decoding process can be represented in the following form: ; in, This represents a multi-source, multi-scale feature fusion encoder / decoder. This indicates the final segmentation result. Indicates the number of channels. It equals 1; Utilizing the Swin-Unet decoder for the corresponding RSI and DSM feature extractors and Construct auxiliary decoding branches, denoted as follows: and The segmentation result based on the auxiliary decoding branch is represented as follows: ; ; in, and These represent the segmentation results of the auxiliary decoding branches of RSI and DSM, respectively.
8. A UAV remote sensing image segmentation method based on multimodal fine-grained feature mining according to claims 1-7, characterized in that, The loss function of the image segmentation model includes relevance separation loss, semantic segmentation loss, and loss of auxiliary training strategy, which supervises the fine-grained feature extraction of the image segmentation model from multiple perspectives. Specifically, it includes: loss function It consists of two parts: semantic segmentation loss and correlation separation loss Semantic segmentation uses the cross-entropy loss function. Cross-entropy loss function This includes the main branch decoding loss and the auxiliary branch decoding loss for RSI and DSM; the expression is as follows: ; ; in, , and Balance coefficients representing different losses; and These represent the number of pixels and the number of categories, respectively. This represents the semantic segmentation loss of the main branch. Represents the RSI auxiliary branch splitting loss. This represents the DSM auxiliary branch splitting loss. and Representing pixels Category The true confidence level and the confidence level predicted by the image segmentation model; Correlation separation loss The expression is as follows: ; in, This indicates the correlation coefficient calculation. It is a scaling factor. and These represent the depth detail features of RSI and DSM, respectively; The overall loss function of the image segmentation model Represented as: ; in, and These represent the balance coefficients.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the UAV remote sensing image segmentation method based on multimodal fine-grained feature mining as described in any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements any of the steps of the UAV remote sensing image segmentation method based on multimodal fine-grained feature mining described in 1-8.