A U-shaped multimodal fusion segmentation method combining a graph neural network and a Mamba model

By combining the U-shaped multimodal fusion segmentation method of graph neural network and Mamba model, the problems of insufficient global modeling capabilities and high computational complexity in multimodal fusion segmentation of remote sensing images are solved, and efficient remote sensing image segmentation is realized, which is suitable for large-scale remote sensing image processing.

CN119006813BActive Publication Date: 2025-07-18SHEN ZHEN WAN ZHI DA XIN XI ZI XUN YOU XIAN GONG SI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411067820.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-06
Publication Date
2025-07-18
Estimated Expiration
2044-08-06

AI Technical Summary

Technical Problem

The existing multimodal fusion segmentation method for remote sensing images has problems such as insufficient global modeling capabilities, high computational complexity, large dependence on marked data, and manual parameter optimization and adjustment, resulting in degradation of segmentation performance.

Method used

The U-shaped multimodal fusion segmentation method combined with graph neural network and Mamba model is adopted to capture global information through the Mamba-Graph encoder, the Graph encoder captures local details, and uses parallel fusion module to establish cross-modal long-distance dependence, and deeply fusion is performed through graph neural network technology, so that the decoder recovers hidden features.

Benefits of technology

Reliance on hardware configuration and computing resources is reduced, efficient application of large-scale remote sensing image segmentation tasks is realized, and segmentation performance and computing efficiency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119006813B_ABST
    Figure CN119006813B_ABST
Patent Text Reader

Abstract

The present invention proposes a U-shaped multi-modal fusion segmentation method combining a graph neural network and a Mamba model, belonging to the field of multi-modal fusion segmentation of remote sensing images. First, obtain and label a remote sensing image set and divide them into a training set and a validation set. Then, construct a multi-modal fusion segmentation network, whose main architecture includes an encoder, a fusion layer, a decoder, and a segmentation head. Furthermore, perform class balancing and enhancement, and then train in the multi-modal fusion segmentation network. Finally, segment the entire set of remote sensing images and input them into the network to obtain semantic segmentation results. The detailed process includes the construction of various feature representations and fusion modules. The encoding and decoding parts of the constructed network use the Mamba-Graph combination method. The present invention takes into account the characteristics of different modalities, makes full use of the complementarity between modalities, and establishes long-range dependencies across modalities. It alleviates the poor feature representation caused by the incompatibility between modalities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of remote sensing image multi-modal fusion segmentation, and more specifically relates to a U-shaped multi-modal fusion segmentation method combining a graph neural network and a Mamba model. Background Art

[0002] Remote sensing, as an important means of earth observation, collects a vast amount of ground object data. Accurate semantic segmentation of remote sensing images is the key to utilizing this data. Multi-modal fusion segmentation methods have attracted wide attention because they can make full use of existing data and achieve better performance than traditional single-modal technologies. However, most existing methods are based on the CNN and Transformer architectures, which results in limited global-local modeling capabilities and excessively high computational complexity, making it impossible to support the input of large images.

[0003] In addition, there are some other problems with existing multi-modal fusion segmentation methods. For example, such methods often rely on a large amount of labeled data, and in actual remote sensing image processing tasks, it is very difficult to obtain a large amount of high-quality labeled data. On the other hand, most existing multi-modal fusion segmentation methods use a manually adjusted approach for parameter optimization, which not only involves a large amount of work but may also lead to a decrease in segmentation performance. Summary of the Invention

[0004] The present invention proposes a U-shaped multi-modal fusion segmentation method combining a graph neural network and a Mamba model. It can solve the following problems: insufficient global modeling ability of pure CNN methods; the computational complexity of CNN+Transformer methods being exponential with the input length; difficulties in fusion due to large modal differences, or a decrease in accuracy due to insufficient utilization of each modal information after fusion.

[0005] To achieve the above object, the present invention is implemented by the following technical solutions: The method includes:

[0006] Obtain a remote sensing image set, label it to obtain a class label set, and after cutting the remote sensing image set and the class label set into a fixed size, divide the cut data set into a training set and a validation set;

[0007] Construct a multi-modal fusion segmentation network, the main architecture of which consists of an encoder, a fusion layer, a decoder, and a segmentation head;

[0008] Balance and enhance the classes of the training set and the validation set, input them into the multi-modal fusion segmentation network for training, and take the parameter weights at the highest accuracy; divide the entire remote sensing image set and input it into the multi-modal fusion segmentation network to obtain the semantic segmentation result of the remote sensing image ground object interpretation.

[0009] In one solution, the encoding part of the constructed multi-modal fusion segmentation network adopts the Mamba-Graph combination method;

[0010] The decoder restores the hidden fusion features through the use of multiple upsampling modules for the final segmentation process; after the output result of the decoder is processed by a two-layer fully connected MLP in the segmentation head, the final output is obtained.

[0011] In one solution, the construction of the multi-modal fusion segmentation network includes the following steps:

[0012] The training set including RGB and DSM remote sensing images is sequentially input into the encoder, and the current feature representation is retained before each downsampling, and 4 different-resolution feature representations of RGB and DSM are obtained in sequence;

[0013] The feature representations RGB4 and DSM4 output by the last layer of the encoder are used as the inputs of the parallel fusion module to obtain Fusion4, and so on to obtain Fusion1, Fusion2, and Fusion3;

[0014] The first layer of the decoder takes Fusion4 ∈ R C*H / 16*W / 16 as the input, where H*W is the size of the original input image and C is the number of channels. The number of channels in the first layer of the decoder corresponds to the number of channels in the last layer of the encoder and so on The first layer of the decoder consists of residual blocks, specifically Conv 3x3 +BN+ReLU)+X. Since the second layer of the decoder, after the decoder incorporates the Fusion i features, it is then input into the residual block.

[0015] In one solution, the first three layers of the encoder are composed of Mamba, and the last layer of the encoder is composed of Graph; the fusion layer is composed of parallel fusion modules.

[0016] In one solution, the training set includes RGB and DSM remote sensing images.

[0017] In one solution, the ratio of the training set to the validation set is 7:3.

[0018] Advantages of the present invention:

[0019] In a U-shaped multimodal fusion segmentation method that combines a graph neural network and a Mamba model in the present invention, the Mamba (the first three layers) and Graph (the last three layers) encoders in the multimodal fusion segmentation network take into account the global-local idea. The first three-layer encoder is composed of Mamba to capture global information, and the last-layer encoder is composed of Graph to capture irregular local details.

[0020] In a U-shaped multimodal fusion segmentation method that combines a graph neural network and a Mamba model in the present invention, the parallel fusion module in the multimodal fusion segmentation network fully considers the characteristics between different modalities, uses parallel cross and concatMamba blocks to establish long-range dependencies between cross-modalities and limits the parameters within an acceptable range, and on this basis, uses graph neural network technology to enable deep fusion of modalities.

[0021] In a U-shaped multimodal fusion segmentation method that combines a graph neural network and a Mamba model in the present invention, in the decoder module of the multimodal fusion segmentation network, starting from the second layer of the decoder, it uses adaptive weight fusion to fuse the features output by the decoder and the fusion feature Fusion, and uses residual blocks to alleviate the overfitting caused by the excessive depth of the network.

[0022] The method of the present invention reduces the dependence on hardware configuration and computing resources compared with the prior art, realizes the application in actual large-scale remote sensing image segmentation tasks, and has high practicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a flowchart of the method of the present invention;

[0024] Figure 2 It is a schematic diagram of the overall network structure of the present invention;

[0025] Figure 3 It is a schematic diagram of the Mamba-graph Fusion (MGF) structure proposed by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0026] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Typical embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.

[0027] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as understood by those skilled in the technical field to which this invention belongs. The terms used in the description of this invention in this specification are only for the purpose of describing specific embodiments and are not intended to limit this invention.

[0028] The overall concept of this invention is as follows: In this patent, a U-shaped multi-modal fusion segmentation method combining a graph neural network and a Mamba model is proposed. First, this paper introduces the Mamba modality to capture global information. Compared with Transformer, the complexity of the Mamba modality is not limited by the size of the input and can capture longer-range context dependencies. Second, this paper introduces a graph neural network for local modeling. Compared with CNN, the graph neural network updates node information in a nearest-neighbor manner and can locally model irregular features in remote sensing images instead of a receptive field with a fixed convolutional kernel size. Finally, this paper proposes a parallel fusion module that, while fully considering the characteristics of different modalities, uses parallel cross and concatMamba blocks to establish long-range dependencies between cross-modalities and limits the parameters within an acceptable range, and on this basis, uses graph neural network technology to enable deep fusion of modalities.

[0029] As Figure 1 shown, a U-shaped multi-modal fusion segmentation method combining a graph neural network and Mamba includes the following steps:

[0030] S1. Obtain a remote sensing image set Dataset (RGB_Dataset, DSM_Dataset), and take a subset

[0031] Dataset_part (RGB_Dataset_part, DSM_Dataset_part) of Dataset, label Dataset_part to obtain a class label set Dataset_part_Label. After cutting the remote sensing image set Dataset_part and the class label set Dataset_part_Label into fixed-size tokens (such as 1024x1024), divide the cut data set into a training set train_data and a validation set val_data according to a certain ratio.

[0032] As Figure 2 and Figure 3As shown in the figure, S2. Construct a multi-modal fusion segmentation network. The main body of this network follows a U-shaped architecture and consists of an encoder, a fusion layer, a decoder, and a segmentation head. The encoding part adopts a combination of Mamba and Graph. Among them, the first three layers of the encoder are composed of Mamba to capture global information, and the last layer of the encoder is composed of Graph to capture irregular local details. The fusion layer consists of carefully designed parallel fusion modules, fully considering the characteristics between different modalities. It uses parallel cross and concatMamba blocks to establish long-range dependencies between cross-modalities and limits the parameters within an acceptable range. On this basis, graph neural network technology is used to enable deep fusion of modalities. The decoder uses multiple upsampling modules to restore the hidden fusion features for the final segmentation process. After the output result of the decoder is processed by two layers of fully connected (Mlp), the final output is obtained.

[0033] The encoding part adopts a combination of Mamba and Graph. Among them, the first three layers of the encoder are composed of Mamba to capture global information. Mamba is inspired by continuous systems and maps a one-dimensional function x(t)→y(t) through the hidden state h(t). This system uses A as the evolution parameter, B and C as projection parameters, and D as the skip connection as shown in the following formula:

[0034] h′(t) = A h(t) + Bx(t)

[0035] y(t) = C h(t) + Dx(t)

[0036] where, A ∈ R N×N , B.C ∈ R N , D ∈ R 1 .

[0037] The image gray level is discrete information. Mamba discretizes the above continuous system as shown in the following formula:

[0038]

[0039] y t = Ch t + Dx t

[0040] where, Δ represents the time scale parameter, and A and B are the discretized A and B.

[0041] The last layer of the encoder is composed of Graph to capture irregular local details. For the feature representation F ∈ R input from the Mamba encoder to the Graph encoder C×H×W , we first divide it into N parts, and use Transformer ∈ R for each part D, construct the nodes V of the graph, connect the K nearest neighbors of each part as the edges E of the graph, and obtain the graph Graph=(V, E). The graph convolutional layer can exchange information between nodes by aggregating the features from its neighboring nodes, as shown in the following formula:

[0042] x i = h(x i , g(x i , Neightbor(x i ), W agg ), W update )

[0043] g(·)= [x i , max(Neightbor(x i ) - x i )]

[0044] h(·)= g(·)W update

[0045] Among them, Neightbor(x i ) represents the set of K nearest neighbor nodes of x i , and W agg , W update are the adaptive learning weights for aggregation and update respectively.

[0046] The fusion layer consists of carefully designed parallel fusion modules, fully considering the characteristics between different modalities. It uses parallel cross and concatMamba blocks to establish long-range dependencies between cross-modalities and limits the parameters within an acceptable range, and on this basis, uses graph neural network technology to enable deep fusion of modalities.

[0047] First, the features RGB i , DSM i ∈R Channel×H×W

[0048] are first split into 4 groups on the channel The 4 groups of Fusion i share the weights of the crossMamba block (using the same group). The operation of crossMamba is to exchange the C of RGB and DSM in the encoder Mamba detailed above (the Mamba encoder is described in detail above), and calculate y t as shown in the following formula:

[0049]

[0050] Secondly, the four groups of Fusioni for RGB and DSM obtained from the crossMamba block are shuffled and concatenated in the channel dimension and then fed into the graph convolutional layer (which is described in detail in the above figure encoder). Finally, the output of the graph convolutional layer is fed into the concatMamba block. The operation of the concatMamba block is to concatenate the input data in the channel dimension, duplicate it in reverse order, and then feed it into the Mamba operation in the form of two channels (which is described in detail in the above Mamba encoder). The result of the operation will be restored to the original order and directly added bit by bit.

[0051] The decoder restores the hidden fusion features through the use of multiple upsampling modules for the final segmentation process; after the output result of the decoder is processed by two layers of fully connected (MLP), the final output is obtained; the feature representation dimension of each layer of the decoder is shown in the following formula:

[0052]

[0053] Among them, H×W is the size of the input remote sensing image, and I∈(1,4).

[0054] When the decoder fuses the Fusion i features, it is shown in the following formula:

[0055] Decoder i =ConvBnRelu(W1×Upsample(Decoder i-1 )+W2×Fusion i )

[0056] Among them, Decoder i represents the decoder feature representation of the i-th layer, W1 and W2 represent adaptive weights, W1 + W2 = 1, ConvBnRelu represents (Conv 3×3 +BN+ReLu), and Upsample represents bilinear interpolation upsampling.

[0057] S301. The training set (RGB and DSM remote sensing images) cut into token sizes obtained from Dataset in S1 is sequentially input into the Mamba (the first three layers) and Graph (the last three layers) encoders, and the current feature representation is retained before each downsampling, and four different-resolution feature representations of RGB and DSM, namely RGB1, RGB2, RGB3, RGB4 and DSM1, DSM2, DSM3, DSM4, are obtained in turn.

[0058] S302. Take the feature representations RGB4 and DSM4 output by the last layer of the encoder (Graph encoder) as the inputs of the parallel fusion module to obtain Fusion4, and so on to obtain Fusion1, Fusion2, and Fusion3.

[0059] S303. The first layer of the decoder takes Fusion4 ∈ R C*H / 16*W / 16 as the input, where H*W is the size of the original input image and C is the number of channels. The number of channels in the first layer of the decoder corresponds to the number of channels in the last layer of the encoder and so on The first layer of the decoder consists of residual blocks, specifically Conv 3x3 +BN+ReLU)+X. Starting from the second layer of the decoder, after the decoder incorporates the Fusion i features, it is then input into the residual block.

[0060] S3. Balance and enhance the categories of the training set and validation set obtained from Dataset_part in S1, input them into S2 to construct the training of the multimodal fusion segmentation network, and take the parameter weights (.pth,.checkpoint files) with the highest accuracy in the validation stage; divide the entire remote sensing image dataset Dataset into token sizes, input them into the multimodal fusion segmentation network after loading the parameter weights, and obtain the semantic segmentation results of the interpretation of the ground objects in the remote sensing images.

[0061] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0062] It should be understood that the detailed description of the technical solutions of the present invention with the help of the preferred embodiments is illustrative rather than restrictive. Those of ordinary skill in the art can modify the technical solutions recorded in each embodiment on the basis of reading the specification of the present invention, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A U-shaped multimodal fusion segmentation method combining a graph neural network and a Mamba model, characterized in that: The method described includes: Obtain a remote sensing image set, annotate to obtain a class label set, and after cutting the remote sensing image set and the class label set into a fixed size, divide the cut data set into a training set and a validation set; Construct a multi-modal fusion segmentation network, the main architecture of which consists of an encoder, a fusion layer, a decoder, and a segmentation head; Perform class balancing and enhancement on the training set and the validation set, input them into the multi-modal fusion segmentation network for training, and take the parameter weights at the highest accuracy; Cut and input the entire set of remote sensing images into the multi-modal fusion segmentation network to obtain the semantic segmentation results of the interpretation of the ground objects of the remote sensing images. The encoding part of the constructed multi-modal fusion segmentation network adopts the Mamba-Graph combination method; The decoder restores the hidden fusion features through the use of multiple upsampling modules for the final segmentation process; After the output result of the decoder is processed by a two-layer fully connected MLP in the segmentation head, the final output is obtained; The construction of the multi-modal fusion segmentation network described includes the following steps: (1) Sequentially input the training set including RGB and DSM remote sensing images into the encoder, retain the current feature representation before each downsampling, and obtain 4 different resolution feature representations of RGB and DSM respectively; (2) Use the feature representations RGB4 and DSM4 output by the last layer of the encoder as the input of the parallel fusion module to obtain Fusion4, and so on to obtain Fusion1, Fusion2, and Fusion3; (3) The first layer of the decoder takes Fusion4 ∈ R C*H / 16*W / 16 as input, where H * W is the size of the original input image, C is the number of channels, and the number of channels in the first layer of the decoder corresponds to the number of channels in the last layer of the encoder and so on The first layer of the decoder consists of residual blocks, specifically Conv 3x3 + BN + ReLU) + X. From the second layer of the decoder onwards, the decoder incorporates Fusion i features and then inputs them into the residual blocks; The first three layers of the encoder are composed of Mamba, and the last layer of the encoder is composed of Graph; The fusion layer is composed of parallel fusion modules.

2. The U-shaped multi-modal fusion segmentation method combining a graph neural network and a Mamba model according to claim 1, characterized in that: The training set includes RGB and DSM remote sensing images.

3. The U-shaped multimodal fusion segmentation method combining a graph neural network and a Mamba model according to claim 1, characterized in that: The ratio of the training set to the validation set is 7:3.

Citation Information

Patent Citations

  • Coding and decoding structure-based multi-modal remote sensing image semantic segmentation method

    CN115984701A