Hybrid prompt architecture fusion method for multi-modal semantic segmentation

By introducing the methods of pre-trained model reuse and dynamic prompt learning in the multimodal semantic segmentation model, a multi-subspace alignment and hybrid prompt fusion architecture is constructed, which solves the problems of high model complexity and data scarcity, and achieves efficient multimodal information fusion and accurate semantic segmentation.

CN120599417AActive Publication Date: 2025-09-05BEIJING INST OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510549133.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-09-05
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

Existing multimodal semantic segmentation models are highly complex and data-scarce, resulting in difficult deployment and insufficient segmentation performance in resource-constrained environments.

Method used

Through the collaborative optimization of pre-trained model reuse and dynamic prompt learning, a multimodal embedding module based on the RGB pre-trained model is constructed, and multi-subspace alignment and hybrid prompt fusion methods are introduced to achieve efficient fusion of cross-modal features and parameter reduction.

Benefits of technology

While reducing the complexity of the model, it significantly improves the semantic segmentation performance, reduces the computing resource requirements, and improves the model's generalization ability and segmentation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599417A_ABST
    Figure CN120599417A_ABST
Patent Text Reader

Abstract

The invention discloses a mixed prompt architecture fusion method for multi-modal semantic segmentation, and belongs to the field of computer vision. The implementation method comprises the following steps: constructing a multi-modal embedding module based on an RGB pre-training model; initial prompts output by the multi-mode embedding module and RGB features input by the backbone network are projected to a low-rank subspace through linear mapping to complete feature alignment, a hybrid matrix is introduced to fuse the RGB features of the low-rank subspace with prompt information, and new prompt information is fused with RGB features of backbone network codes; a lightweight multi-subspace alignment and mixing prompt module is introduced; a multi-resolution self-attention encoder of a backbone network is used for encoding RGB image features, auxiliary image information generates an initial prompt through a multi-mode embedding module, and the initial prompt information and the RGB image features are fused through a multi-subspace alignment and mixed prompt module to form new prompt information and RGB image feature fusion. And fusing the semantic information of the RGB image and the semantic information of the auxiliary mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multimodal fusion method, in particular to a hybrid prompt architecture fusion method for multimodal semantic segmentation, and belongs to the field of computer vision. Background Art

[0002] Semantic segmentation, a fundamental task in computer vision, aims to assign a predefined semantic label to each pixel in an input image. In recent years, significant progress has been made in semantic segmentation technology with the release of large-scale RGB benchmark datasets such as Cityscapes and ADE20K. These datasets provide crucial support for the training and evaluation of segmentation models. However, segmentation methods that rely solely on RGB images often perform poorly in complex environmental conditions such as low light and inclement weather. This is primarily due to the fact that visual information can be severely occluded or distorted, resulting in reduced segmentation accuracy. To overcome these limitations, researchers have proposed multimodal semantic segmentation methods that enhance segmentation performance by incorporating auxiliary sensor data, such as depth images and thermal infrared images. For example, depth images provide precise spatial location information, while thermal infrared images can capture target features that are invisible to visible light in conditions such as low light or haze. This fusion of multimodal data not only enhances the model's ability to identify instance boundaries but also strengthens its understanding of contextual information, significantly improving segmentation results. Consequently, multimodal segmentation methods have surpassed traditional RGB-only segmentation methods in performance, opening up new directions for the research and application of semantic segmentation.

[0003] However, the application of multimodal semantic segmentation technology still faces two major challenges. 1. Increased model complexity. In order to process additional modal data, it is usually necessary to design independent backbone network branches for each modality to extract features, and introduce complex fusion modules to achieve the integration of multimodal features. This architectural expansion not only significantly increases the number of model parameters, but also greatly increases the computational complexity, making it difficult to deploy in resource-constrained environments (such as mobile devices or embedded systems). 2. Scarcity of multimodal data. Compared with RGB datasets, multimodal segmentation datasets are usually smaller in size, which limits the training effect and generalization ability of the model. Summary of the Invention

[0004] In response to the problems of high complexity and data scarcity in existing multimodal semantic segmentation models, the purpose of the present invention is to provide a hybrid prompt architecture fusion method for multimodal semantic segmentation. Through the collaborative optimization of pre-trained model reuse and dynamic prompt learning, efficient fusion of cross-modal features and parameter simplification are achieved. Without changing the underlying architecture of the pre-trained RGB segmentation model, auxiliary modality data can be efficiently integrated, thereby significantly improving segmentation performance while reducing model complexity.

[0005] The main purpose of the present invention is achieved through the following technical solutions:

[0006] The present invention discloses a hybrid prompt architecture fusion method for multimodal semantic segmentation, which constructs a multimodal embedding module based on the RGB pre-trained model to initialize the embedding information of the auxiliary modality, introduces a multi-subspace aligned hybrid prompt fusion method to generate multi-level prompt information, realize multimodal information fusion, and improve model segmentation accuracy. At the same time, the lightweight hybrid prompt architecture has lower model complexity than the traditional dual-stream fusion architecture, reducing the computing power requirements for model training and inference.

[0007] The present invention discloses a hybrid hint architecture fusion method for multimodal semantic segmentation, comprising the following steps:

[0008] Step 1: Build a multimodal embedding module based on the RGB pre-trained model to generate high-quality initial prompt information from the input auxiliary images such as depth images and thermal infrared images, effectively encode the auxiliary modality semantic information, avoid interference from redundant information of the original auxiliary images, and improve the adaptability of different modal auxiliary information to the RGB feature space.

[0009] Given auxiliary image information such as depth image, thermal infrared image, etc. Among them C f , H and W represent the number of channels, height and width of the image respectively, and the enhanced image is obtained by using channel alignment transformation and data enhancement method Among them C in Represents the number of channels of the RGB image. The initial block of the pre-trained ResNet-50 is used as the multimodal embedding network, and then the initial hint e0 is extracted by aligning it with the RGB image features through the mapping layer. The extraction expression of the initial hint e0 is shown in formula (1):

[0010]

[0011] Among them, x aux Indicates auxiliary modal input information, As the initial block of the pre-trained ResNet-50, the shallow architecture of the pre-trained ResNet-50 network is used to extract the features of the auxiliary modality information, so that the extracted auxiliary modality features match the shallow information distribution of the RGB backbone network. proj As the mapping layer, a single linear layer is used to map the extracted features to a specific feature space, which is consistent with the dimension of the RGB backbone model, to achieve high-quality initial prompt information generation.

[0012] Step 2: For the initial prompt output by the multimodal embedding module and the RGB features input by the Segformer backbone network, linear mapping is used to project them into a low-rank subspace to complete feature alignment. A mixing matrix is ​​introduced to fuse the RGB features of the low-rank subspace with the prompt information to generate new prompt information. The new prompt information is fused with the RGB features encoded by the Segformer backbone network. Without modifying the Segformer backbone network architecture, a lightweight multi-subspace alignment and mixed prompt module is introduced to provide feature information of auxiliary modalities to the backbone model, effectively improving the semantic segmentation performance.

[0013] Step 2.1: Project the initial prompt output by the multimodal embedding module and the RGB features input by the Segformer backbone network to complete the multi-subspace alignment. Given an RGB image Among them C in , H and W represent the number of RGB channels, height and width of the image respectively, and the enhanced image is obtained by using the data enhancement method of random horizontal flipping and perturbation addition. The enhanced data is normalized and input into the embedding module, divided into blocks of size 4×4, and the divided blocks are used as the input of the Segformer backbone network.

[0014] The RGB features of the Segformer backbone network input and the auxiliary modality features in the prompt branch exist in different feature spaces. These mismatched feature spaces need to be aligned. A low-rank subspace alignment method is introduced. Through the linear projection method, the RGB modality features and the auxiliary modality prompt information are aligned into a low-rank subspace. The aligned features are fused by addition. The multi-subspace alignment expression is shown in Equation (2):

[0015]

[0016] Where, i∈{1,2,…,N} represents the index of the layer, P i represents the i-th multi-subspace hybrid prompt module, e i Indicates the L i Layer prompt information, h i Indicates the L i The output hidden state of the layer, is the downsampling module, is the upsampling module. In addition, let d be the original feature space, r<<d, and r is the dimension of the low-rank aligned subspace. Using a smaller value of r allows parameterization of the hint module.

[0017] In order to improve the efficiency of information utilization, multiple pairs of down-projection and up-projection modules are introduced. Each pair of projections aligns the RGB modality and the auxiliary modality in a unique subspace, thereby improving the utilization of useful information. However, this method will significantly increase the number of parameters. Inspired by the low-rank adaptive fine-tuning method of large language models, even with extremely low ranks, good performance will still be retained because the pre-trained neural network model operates in low intrinsic dimensions. At the same time, the pre-trained Segformer segmentation model is also in low intrinsic dimensions. Therefore, by reducing the rank of each subspace to maintain the parameter efficiency of the overall model, the RGB modality features and the prompt information of the auxiliary modality are projected into multiple subspaces, as shown in Equations (3) and (4):

[0018] [W rgb,1 h i-1 ,W rgb,2 h i-1 ,...,W rgb,n h i-1 ], (3)

[0019] [W f,1 h i-1 ,W f,2 h i-1 ,...,W f,n h i-1 ], (4)

[0020] in, n represents the number of subspaces used. Dividing the original rank r by n ensures that the introduction of multiple subspaces does not bring additional parameters.

[0021] Step 2.2: Introduce a mixing matrix to fuse the RGB features of the low-rank subspace with the prompt information to generate new prompt information. The new prompt information is fused with the RGB features encoded by the Segformer backbone network. Without modifying the Segformer backbone network architecture, a lightweight multi-subspace alignment and mixed prompt module is introduced to provide auxiliary modality feature information to the backbone model, effectively improving the semantic segmentation performance.

[0022] In the low-rank subspace alignment process, a hybrid hint generation method is introduced to enable the RGB features and the auxiliary modal features in the hint branch to interact with each other and improve the expressiveness of the newly generated auxiliary hint information. In formulas (3) and (4), the RGB modal features and the hint information of the auxiliary modality are projected into multiple subspaces, and E rgb,n h i-1 Expressed as Introducing a mixing matrix Used to exchange information between subspaces. The mixed features The expression is shown in formula (5):

[0023]

[0024] Among them, M i,j Represents the value of the corresponding element in the mixing matrix M. Similarly, we get the prompt e i-1 Low-rank representation of The information of RGB modality is injected into the low-rank cue by addition and the It is back-projected back to the original feature space to generate a new hint e i . Backbone layer L i The RGB features are fused with the new hints by addition. The expression is shown in formula (6):

[0025]

[0026] Step 3: Use the multi-resolution self-attention encoder of the Segformer backbone network to encode RGB image features. Auxiliary image information such as depth images and infrared images is used to generate initial prompts through the multimodal embedding module. The initial prompt information and RGB image features are fused through the multi-subspace alignment and hybrid prompt module to form new prompt information and RGB image feature fusion, efficiently fusing the semantic information of RGB images and auxiliary modalities, improving the model inference speed and achieving improved semantic segmentation performance.

[0027] Given an RGB image After data pre-processing with blocks of size 4×4, the multi-resolution self-attention encoder of the Segformer backbone network takes the divided blocks as input and obtains the original image resolution at {1 / 4, 1 / 8, 1 / 16, 1 / 32}. The feature expression of the extracted image is shown in formula (7):

[0028]

[0029] in, is the RGB image block after data pre-processing, is a multi-resolution self-attention encoder, θ represents the parameters of the encoder, and f rgb Represents the extracted image features, corresponding to the hidden states h at different levels i The initial prompt information e0 can be obtained from formula (1). The obtained e0 is used as the initial prompt and the embedded original RGB input h0 is passed through the prompt module to generate a new prompt e1. After e1 passes through the model layer, the output hidden state h1 is obtained. The expressions are shown in formulas (8) and (9):

[0030] e1=P(h0,e0), (8)

[0031] h1=L(h0,e1), (9)

[0032] Among them, P and L represent the multi-subspace alignment and hybrid prompt module and the multi-resolution self-attention encoder layer, respectively. For different resolutions, the input feature is the prompt information e of the previous layer. i-1 and hidden state h i-1 , the input features are reduced to a low-rank subspace through linear projection and divided into multiple subspaces. The subspace features of the RGB modality are exchanged and fused using a mixing matrix. The fused features are combined through addition operations and restored to the original feature dimensions through linear projection to generate the final prompt e i , prompt information i With the previous hidden state h i-1 The addition is input to the encoder layer. The expressions are shown in equations (10), (11), (12), (13), and (14):

[0033]

[0034] h i ←L i (h i-1 +e i ), (14)

[0035] Where W x and W rgb is the mapping layer, reshape is the transformation operation, flatten is the flattening operation, is the intermediate state of the hidden state, W up For the upsampling layer, the mixing matrix M acts as a router, assigning different importance to each subspace through different weight mixing. At the same time, the reparameterization of the mixed hint module enables the model to utilize a multi-branch architecture, further promoting more effective propagation of auxiliary information, helping the model achieve better semantic segmentation performance.

[0036] The obtained hidden state, i.e., the fused features, is input into the multi-layer perceptron encoder. The multi-level feature channel dimensions are unified through the MLP layer. After upsampling, the channels are spliced. The spliced ​​features are processed by the MLP layer to obtain the reconstructed pixel-level segmentation mask. The expression is shown in Equation (15):

[0037]

[0038] In the formula, Linear is the linear MLP module, Up is the upsampling module, and Concat is the connection operation. Represents the segmentation result corresponding to the r fusion feature. At the same time, for the training loss function of the semantic segmentation task, the cross entropy loss function is used, and the expression is shown in formula (16):

[0039]

[0040] Among them, N is the total number of pixels in the image, M is the total number of segmentation categories, and y k,c is the true label of the k-th pixel in category c, p k,c The model predicts the probability that the pixel belongs to category c. The above-mentioned multimodal semantic segmentation hybrid hint architecture can achieve efficient fusion of multiple auxiliary modalities such as infrared images and depth images with RGB images, improving semantic segmentation performance while further reducing the complexity and inference cost of the multimodal segmentation model.

[0041] Beneficial effects

[0042] 1. The present invention discloses a hybrid prompt architecture fusion method for multimodal semantic segmentation. It uses a multimodal embedding module based on an RGB pre-trained model to generate high-quality initial prompt information from input auxiliary images such as depth images and thermal infrared images. Compared with traditional prompt initialization structures, only a small number of parameters are required. It can effectively encode auxiliary modal semantic information, avoid interference from redundant information of the original auxiliary image, and improve the adaptability of auxiliary information of different modalities to the RGB feature space.

[0043] 2. The present invention discloses a hybrid prompt architecture fusion method for multimodal semantic segmentation, which uses a hybrid prompt framework to convert auxiliary modalities into prompt information and fuses the prompt information with RGB image features. Compared with the traditional two-stream embedding fusion strategy, it can effectively reduce model complexity and reduce inference costs. At the same time, it can adaptively fuse multiple auxiliary modal information to further improve semantic segmentation accuracy.

[0044] 3. The present invention discloses a hybrid prompt architecture fusion method for multimodal semantic segmentation, which introduces a lightweight multi-subspace alignment and hybrid prompt module, projects the modal information into multiple low-rank subspace alignments, and uses a hybrid matrix to further enhance the representation ability of generated prompt information. Compared with the traditional semantic model prompt architecture, it can effectively reduce the computing power requirements for prompt information generation, and at the same time can handle the generation of multiple auxiliary modal prompts, thereby improving the model generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 Schematic diagram of a multimodal prompt segmentation framework based on multi-subspace alignment and hybrid prompting in this embodiment;

[0046] Figure 2 This is a schematic diagram of the multi-subspace alignment and hybrid prompt module in this embodiment;

[0047] Figure 3 This is a flow chart of a multimodal prompt segmentation framework based on multi-subspace alignment and hybrid prompting in this embodiment; DETAILED DESCRIPTION

[0048] The present invention will be described in detail below with reference to the accompanying drawings and embodiments, and the technical problems solved by the technical solution of the present invention and the beneficial effects thereof will be discussed. It should be noted that the described embodiments are intended to facilitate understanding of the present invention and do not have any limiting effect on the present invention.

[0049] Example 1

[0050] This embodiment discloses a hybrid prompt architecture fusion method for multimodal semantic segmentation, which is applied to the Segformer model. The specific steps are as follows:

[0051] Step 1: Build a multimodal embedding module based on the RGB pre-trained model to generate high-quality initial prompt information from the input depth image, effectively encode the auxiliary modality semantic information, avoid interference from redundant information of the original auxiliary image, and improve the adaptability of different modal auxiliary information to the RGB feature space.

[0052] like Figure 1 As shown in the figure, the backbone network model architecture used in this example is the Segformer-B5 model. The Mix Transformer encoder (MiT), pre-trained on the ADE20K dataset, is used as the backbone network encoder, and the decoder architecture is an MLP decoder with a dimension of 512. The MiT backbone network encoder has a multi-scale feature extraction architecture, with features at four different scales in stages 1 to 4. The dataset is NYU Depth V2, which contains 1,449 RGB-D samples covering 40 categories. The resolution of all RGB and depth images is 480×640.

[0053] Given depth image information Where H and W represent the number of channels, height, and width of the image, respectively. The enhanced image is obtained using channel alignment transformation and data enhancement method. Among them C in Indicates the number of channels of RGB image. Figure 2 As shown in the figure, the design uses ResNet50 as the feature extractor, loads pre-trained weights for initialization, selects the feature map obtained in the first stage of the ResNet50 model, and aligns the feature map with the feature dimension of the first stage of the backbone network through a single linear mapping layer. The extraction expression of the initial hint e0 is shown in formula (1):

[0054]

[0055] Among them, x aux Represents the input information of the auxiliary modality depth image, As the initial block of the pre-trained ResNet-50, the shallow architecture of the pre-trained ResNet-50 network is used to extract the features of the auxiliary modality information, so that the extracted auxiliary modality features match the shallow information distribution of the RGB backbone network. proj As the mapping layer, a single linear layer is used to map the extracted features to a specific feature space, which is consistent with the dimension of the RGB backbone model, to achieve high-quality initial prompt information generation.

[0056] We also further investigated the design of the pretrained feature extractor and analyzed the impact of initializing the extractor with pretrained weights on performance. The results, shown in Table 1, show that using a randomly initialized ResNet50 cue extractor yields an average Intersection-over-Union (IoU) of 59.5%, while initializing with pretrained weights improves to 60.1%. This demonstrates the effectiveness of using a pretrained RGB model to extract initial cues. The pretrained backbone network enables the model to extract more meaningful representations from the auxiliary modality, ultimately achieving better segmentation results.

[0057] Table 1: Impact of initializing the ResNet50 hint extractor with pre-trained weights

[0058] Feature extractor initialization mIoU random 59.5% Pre-training 60.1%

[0059] We then evaluated the impact of using different feature extractor architectures in the multimodal embedding module. The results are shown in Table 2. Since the backbone network uses the MiT encoder architecture, we evaluated the effectiveness of various MiT encoder variants as feature extractors. The results for MiT-B1 show that deeper extractors lead to lower average intersection-over-union (IoU), while gradually reducing the depth improves performance, with the best performance achieved using only the first stage. ResNet50 outperformed MiT-B1, demonstrating the effectiveness of convolutional models in extracting initial cues. Increasing the size of the MiT extractor did not lead to further improvements, further reinforcing the advantage of the convolutional extractor. These findings suggest that pre-trained convolutional cue extractors that focus on early features are more beneficial for cue segmentation.

[0060] Table 2: Impact of different hint extractor architectures and the number of layers used.

[0061] Feature Extractor Type Number of feature layers used mIoU MiT-B1 Stage {1,2,3,4} 59.2% MiT-B1 Stage {1,2,3} 59.5% MiT-B1 Phase {1,2} 59.6% MiT-B1 Phase {1} 59.7% MiT-B2 Phase {1} 59.6% MiT-B4 Phase {1} 59.2% MiT-B5 Phase {1} 59.3% ResNet50 Stage {1,2,3,4} 58.9% ResNet50 Stage {1,2,3} 59.5% ResNet50 Phase {1,2} 59.9% ResNet50 Phase {1} 60.1%

[0062] Step 2: For the initial prompt output by the multimodal embedding module and the RGB features input by the Segformer backbone network, linear mapping is used to project them into a low-rank subspace to complete feature alignment. A mixing matrix is ​​introduced to fuse the RGB features of the low-rank subspace with the prompt information to generate new prompt information. The new prompt information is fused with the RGB features encoded by the Segformer backbone network. Without modifying the Segformer backbone network architecture, a lightweight multi-subspace alignment and mixed prompt module is introduced to provide feature information of auxiliary modalities to the backbone model, effectively improving the semantic segmentation performance.

[0063] Step 2.1: Project the initial prompt output by the multimodal embedding module and the RGB features input by the Segformer backbone network to complete the multi-subspace alignment. Given an RGB image Among them C in , H and W represent the number of RGB channels, height and width of the image respectively, and the enhanced image is obtained by using the data enhancement method of random horizontal flipping and perturbation addition. The enhanced data is normalized and input into the embedding module, divided into blocks of size 4×4, and the divided blocks are used as the input of the Segformer backbone network.

[0064] Through the linear projection method, the RGB modality features and the auxiliary modality hint information are aligned into a low-rank subspace, and the aligned features are fused through the addition operation. The multi-subspace alignment expression is shown in formula (2):

[0065]

[0066] Where i∈{1,2,…,N} represents the index of the layer. i represents the i-th multi-subspace hybrid prompt module. e i Indicates the L i Layer prompt information, h i Indicates the L i The output hidden state of the layer, is the downsampling module, is the upsampling module. In addition, let d be the original feature space, r<<d, and r is the dimension of the low-rank aligned subspace. Using a smaller value of r allows parameterization of the hint module.

[0067] In order to improve the efficiency of information utilization, multiple pairs of down-projection and up-projection modules are introduced. Each pair of projections aligns the RGB modality and the auxiliary modality in a unique subspace, thereby improving the utilization of useful information. By reducing the rank of each subspace to maintain the parameter efficiency of the overall model, the RGB modality features and the clue information of the auxiliary modality are projected into multiple subspaces, as shown in Equations (3) and (4):

[0068] [W rgb,1 h i-1 ,W rgb,2 h i-1 ,...,W rgb,n h i-1 ], (3)

[0069] [W f,1 h i-1 ,W f,2 h i-1 ,...,W f,n h i-1 ], (4)

[0070] in, n represents the number of subspaces used. Dividing the original rank r by n ensures that the introduction of multiple subspaces does not bring additional parameters.

[0071] Step 2.2: Introduce a mixing matrix to fuse the RGB features of the low-rank subspace with the prompt information to generate new prompt information. The new prompt information is fused with the RGB features encoded by the Segformer backbone network. Without modifying the Segformer backbone network architecture, a lightweight multi-subspace alignment and mixed prompt module is introduced to provide auxiliary modality feature information to the backbone model, effectively improving the semantic segmentation performance.

[0072] In the low-rank subspace alignment process, a hybrid hint generation method is introduced to enable the RGB features and the deep modal features in the hint branch to interact with each other and improve the expressiveness of the newly generated auxiliary hint information. In formulas (3) and (4), the RGB modal features and the hint information of the auxiliary modality are projected into multiple subspaces, and W rgb,n h i-1 Expressed as Introducing a mixing matrix Used to exchange information between subspaces. The mixed features The expression is shown in formula (5):

[0073]

[0074] Among them, M i,j Represents the value of the corresponding element in the mixing matrix M. Similarly, we get the prompt e i-1 Low-rank representation of The information of RGB modality is injected into the low-rank cue by addition and the It is back-projected back to the original feature space to generate a new hint e i . Backbone layer L i The RGB features are fused with the new hints by addition. The expression is shown in formula (6):

[0075]

[0076] like Figure 2 As shown in the figure, in the multi-subspace alignment and hybrid cueing module, the input features are the cue information and hidden state of the previous layer. A linear projection layer is used to reduce the input features to a low-rank subspace and partition them into four subspaces. Subsequently, a mixing matrix is ​​used to exchange and fuse the subspace features of the RGB modality. The fused features are combined through addition and restored to the original feature dimensions through linear projection to generate the final cue.

[0077] We investigate the effects of different multi-subspace cue blending methods when applied to either the RGB or depth modality. The results, shown in Table 3, show that the model achieves an average Intersection-over-Union (IoU) of 59.3% without multi-subspace blending. Introducing multi-subspace blending to the RGB modality significantly improves performance to 60.1%, demonstrating the effectiveness of optimizing RGB cues through multi-subspace alignment. In contrast, applying multi-subspace blending only to the auxiliary modality yields only a modest improvement of 59.7%. Enabling multi-subspace blending in both modalities simultaneously does not yield further improvements, but instead causes a slight decrease in mIoU to 59.8%, slightly lower than the setting optimizing only the RGB modality. These results suggest that enhancing the RGB modality through multi-subspace blending plays a more critical role in improving segmentation performance. In contrast, applying the same strategy to the auxiliary modality yields more limited improvements, while simultaneously applying it to both modalities may introduce redundant or conflicting information. Therefore, multi-subspace blending is prioritized for the RGB cue to achieve the best results.

[0078] Table 3: Impact of different hint extractor architectures and the number of layers used.

[0079] method mIoU No mixing 59.3% RGB modal mixing 60.1% Auxiliary mode mixing 59.7% RGB+auxiliary mode dual mixing 59.8%

[0080] Step 3: Use the multi-resolution self-attention encoder of the Segformer backbone network to encode RGB image features. Auxiliary image information such as depth images and infrared images is used to generate initial prompts through the multimodal embedding module. The initial prompt information and RGB image features are fused through the multi-subspace alignment and hybrid prompt module to form new prompt information and RGB image feature fusion, efficiently fusing the semantic information of RGB images and auxiliary modalities, improving the model inference speed and achieving improved semantic segmentation performance.

[0081] Process such as Figure 3 As shown, given an RGB image After data pre-processing with blocks of size 4×4, the multi-resolution self-attention encoder of the Segformer backbone network takes the divided blocks as input and obtains the original image resolution at {1 / 4, 1 / 8, 1 / 16, 1 / 32}. The feature expression of the extracted image is shown in formula (7):

[0082]

[0083] in, is the RGB image block after data pre-processing, is a multi-resolution self-attention encoder, θ represents the parameters of the encoder, and f rgb Represents the extracted image features, corresponding to the hidden states h at different levels i The initial prompt information e0 can be obtained from formula (1). The obtained e0 is used as the initial prompt and the embedded original RGB input h0 is passed through the prompt module to generate a new prompt e1. After e1 passes through the model layer, the output hidden state h1 is obtained. The expressions are shown in formulas (8) and (9):

[0084] e1=P(h0,e0), (8)

[0085] h1=L(h0,e1), (9)

[0086] Among them, P and L represent the multi-subspace alignment and hybrid prompt module and the multi-resolution self-attention encoder layer, respectively. For different resolutions, the input feature is the prompt information e of the previous layer. i-1 and hidden state h i-1 , the input features are reduced to a low-rank subspace by linear projection and divided into multiple subspaces.

[0087] The input data is normalized, and data augmentation is performed using random flipping and random scaling within a range of [0.5, 1.75]. The auxiliary modality data is uniformly converted to a three-channel feature representation. The backbone model uses a pre-trained Segformer-B5 model to extract RGB image information. The auxiliary modality is projected into the feature space via a multimodal embedding module to initialize the cue information. Subsequently, four hybrid cue modules fuse these cues at multiple scales, effectively integrating depth information into the different layers of the RGB backbone network. The decoder output passes through a linear mapping layer to align the dimensions with the actual number of predicted categories. Training is performed for 500 iterations, with an initial learning rate of 0.04 and a polynomial decay scheme. SGD is used as the optimizer with a weight decay of 5e-4. Data augmentation techniques such as random flipping and random cropping are used during training. The backbone network parameters are frozen throughout, and only the parameters of the multimodal embedding module, the multi-subspace mixing module, and the final output linear mapping layer are trained.

[0088] During inference, the multi-scale inference enhancement results are used to compare with other methods, and the mean intersection over union (mIoU) is used to evaluate the accuracy of the segmentation results. The expressions are shown in formulas (10) and (11):

[0089]

[0090] Here, TP, FP, and FN represent the number of pixels correctly predicted, falsely reported, and missed, respectively, and N is the total number of categories. The multi-scale inference test results are shown in Table 4. The depth image modality can effectively enhance the RGB modality feature representation. In terms of multi-modal fusion effect, the method improves 4.3mIoU compared to the two-stream semantic segmentation architecture and 1.9mIoU compared to the hint learning semantic segmentation method. Furthermore, compared to the traditional two-way feature extraction architecture, the auxiliary modality hint mixing module has negligible parameters compared to the backbone architecture, requiring only half the number of parameters.

[0091] Table 4: Performance comparison of different modality segmentation methods.

[0092] Model mIoU Unimodal Segmentation Architecture 52.3% Two-stream semantic segmentation architecture 56.9% Tips for learning semantic segmentation methods 59.3% Multi-subspace hybrid hinting method 61.2%

[0093] At the same time, further ablation experiments are conducted on key hyperparameters. The ratio controls the dimension of the intermediate features and effectively determines the rank of the low-rank subspace. A larger ratio corresponds to a smaller rank and fewer trainable parameters. The results are shown in Table 5. When the dimensionality reduction ratio is increased from 1 to 4, mIoU increases from 59.4% to 60.1%, indicating that moderate rank reduction can improve feature efficiency. However, when the ratio is further increased to 8, the performance drops slightly to 59.6%, indicating that a too small rank may cause information loss and affect performance. The optimal trade-off point is reached at a ratio of 4, which is selected as the final parameter configuration.

[0094] Table 5: Ablation study of rank reduction ratio.

[0095]

[0096] We investigate the impact of the number of subspaces n used to blend cues. Unlike the rank reduction ratio, this parameter does not affect the model's parameter efficiency. The results, shown in Table 6, show that increasing the number of subspaces from 1 to 4 gradually improves performance, reaching a peak mIoU of 60.1%. However, performance degrades when the number of subspaces is further increased to 8, indicating that overly complex cue blending may introduce unnecessary redundancy.

[0097] Table 6: Ablation experiments on the number of mixing subspaces.

[0098]

[0099]

[0100] Therefore, the hybrid prompt architecture fusion method for multimodal semantic segmentation disclosed in the present invention constructs a multimodal embedding module based on the RGB pre-training model to initialize the embedding information of the auxiliary modality, introduces a multi-subspace aligned hybrid prompt fusion method, generates multi-level prompt information, realizes multimodal information fusion, effectively improves the model accuracy, and reduces the model calculation complexity and training cost, providing theoretical support for the development of multimodal information fusion models.

[0101] The above specific description further illustrates the purpose, technical solutions and beneficial effects of the invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A hybrid hint architecture fusion method for multimodal semantic segmentation, characterized by: The following steps are included: Step 1: Build a multimodal embedding module based on the RGB pre-trained model to generate high-quality initial prompt information from the input auxiliary image, effectively encode the auxiliary modality semantic information, avoid interference from redundant information in the original auxiliary image, and improve the adaptability of auxiliary information of different modalities to the RGB feature space; Step 2: For the initial prompt output by the multimodal embedding module and the RGB features input by the Segformer backbone network, linear mapping is used to project them into a low-rank subspace to complete feature alignment. A mixing matrix is ​​introduced to fuse the RGB features of the low-rank subspace with the prompt information to generate new prompt information. The new prompt information is fused with the RGB features encoded by the Segformer backbone network. Without modifying the Segformer backbone network architecture, a lightweight multi-subspace alignment and mixed prompt module is introduced to provide auxiliary modality feature information to the backbone model, effectively improving the semantic segmentation performance. Step 3: Use the multi-resolution self-attention encoder of the Segformer backbone network to encode RGB image features. The auxiliary image information generates initial prompts through the multimodal embedding module. The initial prompt information and RGB image features are fused through the multi-subspace alignment and hybrid prompt module to form new prompt information and RGB image feature fusion. The semantic information of the RGB image and the auxiliary modality are integrated to improve the model inference speed and achieve improved semantic segmentation performance.

2. The hybrid hint architecture fusion method for multimodal semantic segmentation according to claim 1, characterized in that: The implementation method of step one is: Given auxiliary image information such as depth image, thermal infrared image, etc. Among them C f , H and W represent the number of channels, height and width of the image respectively, and the enhanced image is obtained by using channel alignment transformation and data enhancement method Among them C in Respectively represent the number of channels of RGB images; the initial block of the pre-trained ResNet-50 is used as the multimodal embedding network, and then the initial prompt e0 is extracted by aligning with the RGB image features through the mapping layer; the extraction expression of the initial prompt e0 is shown in formula (1): Among them, x aux Indicates auxiliary modal input information, As the initial block of the pre-trained ResNet-50, the shallow architecture of the pre-trained ResNet-50 network is used to extract the features of the auxiliary modality information, so that the extracted auxiliary modality features match the shallow information distribution of the RGB backbone network. proj As the mapping layer, a single linear layer is used to map the extracted features to a specific feature space, which is consistent with the dimension of the RGB backbone model, to achieve high-quality initial prompt information generation.

3. The hybrid hint architecture fusion method for multimodal semantic segmentation according to claim 2, characterized in that: The implementation method of step 2 is: Step 2.1: Project the initial prompt output by the multimodal embedding module and the RGB features input by the Segformer backbone network to complete the multi-subspace alignment; given an RGB image Among them C in , H and W represent the number of RGB channels, height and width of the image respectively, and the enhanced image is obtained by using the data enhancement method of random horizontal flipping and perturbation addition. The enhanced data is normalized and input into the embedding module, divided into blocks of size 4×4, and the divided blocks are used as the input of the Segformer backbone network; The RGB features of the Segformer backbone network input and the auxiliary modality features in the prompt branch exist in different feature spaces. These mismatched feature spaces need to be aligned. A low-rank subspace alignment method is introduced to align the RGB modality features and the auxiliary modality prompt information into a low-rank subspace through a linear projection method. The aligned features are then fused through an addition operation. To improve information utilization efficiency, multiple pairs of down-projection and up-projection modules are introduced. Each pair of projections aligns the RGB modality and the auxiliary modality in a unique subspace, thereby improving the utilization of useful information. By reducing the rank of each subspace to maintain the parameter efficiency of the overall model, the RGB modality features and the clue information of the auxiliary modality are projected into multiple subspaces; Step 2.2: Introduce a mixing matrix to fuse the RGB features of the low-rank subspace with the prompt information to generate new prompt information. The new prompt information is fused with the RGB features encoded by the Segformer backbone network. Without modifying the Segformer backbone network architecture, a lightweight multi-subspace alignment and mixed prompt module is introduced to provide auxiliary modality feature information to the backbone model, effectively improving the semantic segmentation performance.

4. The hybrid hint architecture fusion method for multimodal semantic segmentation according to claim 3, characterized in that: In step 2.1, the multi-subspace alignment expression is shown in formula (2): Where i∈{1,2,…,N} represents the index of the layer; P i represents the i-th multi-subspace hybrid prompt module; h i Indicates the L i The output hidden state of the layer, is the downsampling module, is an upsampling module; in addition, d is the original feature space, r<<d, and r is the dimension of the low-rank alignment subspace.

5. The hybrid hint architecture fusion method for multimodal semantic segmentation according to claim 4, characterized in that: In step 2.1, the RGB modality features and the prompt information of the auxiliary modality are projected into multiple subspaces, as shown in equations (3) and (4): [W rgb,1 h i-1 ,W rgb,2 h i-1 ,...,W rgb,n h i-1 ], (3) [W f,1 h i-1 ,W f,2 h i-1 ,...,W f,n h i-1 ], (4) in, n represents the number of subspaces used. Dividing the original rank r by n ensures that the introduction of multiple subspaces does not bring additional parameters.

6. The hybrid hint architecture fusion method for multimodal semantic segmentation according to claim 5, characterized in that: The implementation method of step 2.2 is: In the low-rank subspace alignment process, a hybrid hint generation method is introduced to enable the RGB features and the auxiliary modal features in the hint branch to interact with each other and improve the expressiveness of the newly generated auxiliary hint information. In formulas (3) and (4), the RGB modal features and the hint information of the auxiliary modality are projected into multiple subspaces, and E rgb,n h i-1 Expressed as Introducing a mixing matrix Used to exchange information between subspaces; the mixed features The expression is shown in formula (5): Among them, M i,j Represents the value of the corresponding element in the mixing matrix M; similarly, we get the prompt e i-1 Low-rank representation of The information of RGB modality is injected into the low-rank cue by addition and the It is back-projected back to the original feature space to generate a new hint e i ; Backbone layer L i The RGB features are fused with the new hint by addition; the expression is shown in formula (6):

7. The hybrid hint architecture fusion method for multimodal semantic segmentation according to claim 6, characterized in that: The implementation method of step three is: Given an RGB image After data pre-processing with blocks of size 4×4, the multi-resolution self-attention encoder of the Segformer backbone network takes the divided blocks as input and obtains the original image resolution at {1 / 4, 1 / 8, 1 / 16, 1 / 32}. The feature expression of the extracted image is shown in formula (7): in, is the RGB image block after data pre-processing, is a multi-resolution self-attention encoder, θ represents the parameters of the encoder, and f rgb Represents the extracted image features, corresponding to the hidden states h at different levels i The initial prompt information e0 can be obtained from formula (1). The obtained e0 is used as the initial prompt and the embedded original RGB input h0 is passed through the prompt module to generate a new prompt e1. After e1 passes through the model layer, the output hidden state h1 is obtained. The expressions are shown in formulas (8) and (9): e1=P(h0,e0), (8) h1=L(h0,e1), (9) where P and L represent the multi-subspace alignment and hybrid hint module and the multi-resolution self-attention encoder layer, respectively. For different resolutions, the input feature is the hint information e of the previous layer. i-1 and hidden state h i-1 , the input features are reduced to a low-rank subspace through linear projection and divided into multiple subspaces; the subspace features of the RGB modality are exchanged and fused using a mixing matrix, the fused features are combined through addition operations and restored to the original feature dimensions through linear projection to generate the final prompt e i , prompt information i With the previous hidden state h i-1 The addition is input to the encoder layer; the expressions are shown in equations (10), (11), (12), (13), and (14): h i ←L i (h i-1 +e i ), (14) Where W x and W rgb is the mapping layer, reshape is the transformation operation, flatten is the flattening operation, is the intermediate state of the hidden state, W up For the upsampling layer, the mixing matrix M acts as a router, assigning different importance to each subspace through different weight mixing; at the same time, the reparameterization of the hybrid hint module enables the model to utilize a multi-branch architecture, further promoting more effective propagation of auxiliary information and helping the model achieve better semantic segmentation performance; The obtained hidden state, i.e., the fused feature, is input into the multi-layer perceptron encoder. The multi-level feature channel dimensions are unified through the MLP layer. After upsampling, the channels are spliced. The spliced ​​features are processed by the MLP layer to obtain the reconstructed pixel-level segmentation mask. The expression is shown in Equation (15): In the formula, Linear is the linear MLP module, Up is the upsampling module, and Concat is the connection operation. Represents the segmentation result corresponding to the r fusion feature; The above-mentioned multimodal semantic segmentation hybrid prompt architecture can achieve efficient fusion of multiple auxiliary modalities such as infrared images and depth images with RGB images, improve semantic segmentation performance, and at the same time further reduce the complexity and inference cost of the multimodal segmentation model.

8. The hybrid hint architecture fusion method for multimodal semantic segmentation according to claim 7, characterized in that: In step three, For the training loss function of the semantic segmentation task, the cross entropy loss function is used, and the expression is shown in formula (16): Among them, N is the total number of pixels in the image, M is the total number of segmentation categories, and y k,c is the true label of the k-th pixel in category c, p k,c It represents the probability that the model predicts that the pixel belongs to category c.

9. The hybrid hint architecture fusion method for multimodal semantic segmentation according to claim 1, 2, 3, 4, 5, 6, 7 or 8, characterized in that: The input auxiliary images include depth images and thermal infrared images.

Citation Information

Patent Citations

  • Sea-land port segmentation method based on space and semantic alignment fusion

    CN118691827A

  • Semantic segmentation model and segmentation method for high-resolution remote sensing image

    CN119206229A

  • Lidar point cloud segmentation method and apparatus, device, and storage medium

    WO2024021194A1

  • Three-dimensional lidar point cloud semantic segmentation method and apparatus based on deep learning

    WO2024130776A1

  • AU2020103905A4