Construction method of multi-modal ViT visual model structure considering modal missing situation
By designing a structure that considers the modal missing situation in the multimodal ViT visual model, using CameraViT and LidarViT to detect and replace missing modal features, the problem of performance degradation in the absence of modality is solved, and higher accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510124639.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-26
AI Technical Summary
The traditional multimodal ViT visual model does not fully consider the modal missing situation in structural design, resulting in significantly reducing the accuracy and robustness of the model in the case of modal missing, which may cause safety risks.
A multimodal ViT visual model structure that considers the modal missing situation is proposed. Through the single-modal branch structure design, including CameraViT and LidarViT, the missing mode is detected and projected to replace it, ensuring that the model can maintain high performance when the modal is missing.
In the modal absence situation, the model can maintain high accuracy and robustness, reduce safety hazards caused by modal absence, and improve the overall performance of the model.
Smart Images

Figure CN120071079A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multimodal computer vision, and relates to a multimodal ViT vision model, and particularly to a construction method for a multimodal ViT vision model structure considering the situation of modality loss. Background Art
[0002] Today, with the rapid development of computer technology, computer vision technology has developed to a relatively mature stage and is widely used in various tasks. Among them, the multimodal vision model can obtain richer and more comprehensive information from data of different modalities compared with the unimodal model, and use the features between different modalities to complement each other, so that the output result has higher accuracy and robustness.
[0003] In the traditional multimodal ViT vision model, due to the limited ability of data augmentation to simulate the situation of modality loss during the model training process, and the situation of modality loss has not been considered in the model structure design, once a certain modality is completely lost due to situations such as sensor offline or data transmission error, the accuracy and robustness of the model will drop sharply, and in tasks with high safety requirements such as autonomous driving, it may cause serious safety hazards. Summary of the Invention
[0004] In order to solve the problem that the traditional multimodal ViT vision model does not fully consider the situation of modality loss in the structure design, resulting in a significant reduction in the accuracy and robustness of the model in the case of modality loss, which may cause safety risks, the present invention provides a construction method for a multimodal ViT vision model structure considering the situation of modality loss. The model constructed by this method can still maintain high performance in the situation of modality loss, which is of great significance for ensuring the accuracy and robustness of the model.
[0005] The object of the present invention is achieved by the following technical solutions:
[0006] A construction method for a multimodal ViT vision model structure considering the situation of modality loss, comprising the following steps:
[0007] Step 1. Design of unimodal branch structure:
[0008] Step 11. Design of CameraViT structure:
[0009] CameraViT is used to divide an image into patches for learning, and the specific steps are as follows:
[0010] (1) For each image with length H I and width W I , divide it into N C patches with length νcH , a patch with width ν cW , that is
[0011] (2) Expand each patch into a one-dimensional feature of length ν cH ×ν cW ×3 and input it into the MLP:
[0012]
[0013] Among them, is the feature input to the MLP, is the result output by the MLP, and the lengths D c of both are consistent;
[0014] (3) Add a learnable embedding E c and a positional embedding E posc that retains position information to generate the embedded sequence feature Camera Token, which is expressed as:
[0015]
[0016] Among them, N is the number of patches, and x class is the trainable parameter for feature classification;
[0017] (4) Z 0 is used as the original feature and input into the CameraViT encoder. The finally output image feature H c is expressed as:
[0018]
[0019] Among them, LN represents the LayerNormalization layer normalization operation, and L c represents the depth of the Transformer model;
[0020] Steps one and two, LidarViT structure design:
[0021] LidarViT is used to process the point cloud input from the lidar. The input is processed by dividing the original point cloud in 3D space into voxels. The specific steps are as follows:
[0022] (1) Divide the point cloud space with length H lH ×ν lW ×ν lD into N l ×W l ×D l using a cube with side lengths of ν lA voxel, the feature set X of the voxel containing t points p It is expressed as:
[0023]
[0024] Wherein, Represents the laser points within the voxel, x i , y i , z i Represents the 3D coordinates of the i-th laser point, R 3 Represents three-dimensional Euclidean space;
[0025] (2) Calculate the centroid (c x , c y , c z ) of all points in the voxel by the average value, and add the offset of each point relative to the centroid as a supplement, the feature set of the voxel Is updated to:
[0026]
[0027] Wherein, Represents the set of the 3D coordinates of the i-th laser point and its offset supplement, R 6 Represents six-dimensional Euclidean space;
[0028] (3) Transform the updated feature set into a one-dimensional vector with a length of ν lH ×ν lW ×ν lD Then for each encoded feature i:
[0029]
[0030] As the point-level feature flag is transformed into the feature space f i In, perform max pooling on each element on f i To obtain the local fusion feature Through the connection of f i And To obtain the connection feature
[0031]
[0032] (4) Add the learnable embedding E l And the position embedding E posl , for the embedding sequence feature LidarToken with a length of D l Is expressed as:
[0033]
[0034] Among them, M is the number of voxels;
[0035] (5)Y 0 As the original feature, it is input into the LidarViT encoder, and the finally output point cloud feature H l is expressed as:
[0036]
[0037] Among them, L l is the depth of the Transformer model;
[0038] Step 2: Detection of missing modalities:
[0039] To detect the missing modality, it is necessary to evaluate the importance of the feature sequence output by the encoder, and screen the features with zero or less information through a scoring function:
[0040]
[0041] Among them, MSA (Multi-head self-attention) represents the multi-head self-attention layer, and MLP (Multilayer Perceptron) represents the multi-layer perceptron. represents the feature output by the l-th layer in the m-th modality encoder, represents the score of, represents the output of the MSA layer; this scoring function dynamically represents the importance of the feature sequence of the l-th layer of the m-th modality. Once the importance score is lower than the threshold, inter-modal projection will be performed to replace the corresponding modality feature;
[0042] Step 3: Projection of missing modalities:
[0043] Step 3-1: Projection of missing modalities:
[0044] (1) Define the projection of a single feature as:
[0045]
[0046] Among them, x m [n] represents the n-th layer feature of the m-th modality to be projected, and x m' [n] represents the n'-th layer feature of the m'-th modality for projection, and h is a multi-layer perceptron with a relatively shallow depth;
[0047] (2) If all the features in a modality are considered, the modality projection is represented as the concatenation of the feature sequence projections:
[0048]
[0049] (3) For the feature Pruning in the feature sequence dynamically detects low importance scores from all features and replaces unimportant features with projections of other embeddings. This dynamic process is performed before each encoding layer. Therefore Updated to:
[0050]
[0051] where Ⅱ is the judgment condition indicator, and the mask vector it outputs θ is the threshold; the operator ⊙ represents element-wise multiplication;
[0052] The loss function is expressed as:
[0053]
[0054] where L m represents the feature loss of the m-th modality, uses the first norm to mark whether inter-modal projection substitution has been performed, and λ is a hyperparameter for evaluating the losses of different modalities;
[0055] Step 3-2. Alignment of the projected positions:
[0056] Since feature substitution will destroy the position information in the original features, to alleviate this problem, the position embeddings of the original feature sequence are retained and the position embeddings in the subsequent layers are frozen. Finally, the position embeddings are substituted before the task head. Whether the original feature sequence has been substituted or not, the model still retains the original position information;
[0057] Step 4. Multi-modal fusion ViT:
[0058] Step 4-1. Use an MLP to connect the image feature H c and the point cloud feature H l to obtain N m fusion feature encodings X m :
[0059] X m = MLP(H c ∪H L )
[0060] Step 4-2. Add learnable embeddings E m and position embeddings E m to the feature sequence of length D posm . The fused feature embedding sequence FeatureFusionToken is expressed as:
[0061]
[0062] Step Four Three. The final output feature H m Output from the L m th layer of the Transformer encoder:
[0063]
[0064] Step Four Five. By adding different task heads, complete multimodal instance segmentation, object detection, and object tracking vision tasks.
[0065] Compared with the prior art, the present invention has the following advantages:
[0066] Aiming at the limited simulation ability of data augmentation for modal missing situations during the model training process of traditional ViT multimodal vision models, and the failure to consider the situation of modal missing in the model structure design, the present invention proposes a multimodal ViT vision model structure considering the situation of modal missing. The present invention properly handles the situation of modal missing, can still maintain high performance and stable output in the situation of modal missing, improves the accuracy and robustness of the model, and reduces potential safety hazards caused by modal missing. Description of the Drawings
[0067] Figure 1 It is the overall flowchart for constructing a multimodal ViT vision model structure considering the situation of modal missing.
[0068] Figure 2 It is the structural design diagram of a multimodal ViT vision model structure considering the situation of modal missing. Detailed Embodiments
[0069] The technical solutions of the present invention will be further described below in conjunction with the drawings, but are not limited thereto. Any modifications or equivalent replacements to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention shall be covered within the protection scope of the present invention.
[0070] The present invention provides a construction method for a multimodal ViT vision model structure considering the situation of modal missing. The method learns the features of objects in different modalities and the projection relationships between features in branches, introduces a feature scoring mechanism to screen out missing modalities, and finally performs projection transformation of features between modalities to replace the features of objects in the missing modalities, improving the accuracy and robustness of the model. As Figure 1 and Figure 2 shown, the specific steps are as follows:
[0071] Step One. Design of the single-modal branch structure:
[0072] Step One One. Design of the CameraViT model structure:
[0073] The CameraViT model is used to segment an image into patches for learning, and the specific steps are as follows:
[0074] (1) For each image with length H I and width W I : Segment it into N C patches with length ν cH and width ν cW , that is
[0075] (2) Unfold each patch into a one-dimensional feature of length ν cH ×ν cW ×3 and input it into the MLP:
[0076]
[0077] Where is the feature input to the MLP, is the result output by the MLP, and the length D c of both is the same.
[0078] (3) Add a learnable embedding E c and a positional embedding E posc that preserves positional information to the feature sequence. The generated embedded sequence feature Camera Token is expressed as:
[0079]
[0080] Where N is the number of patches, and x class is the trainable parameter for feature classification.
[0081] (4) Use Z 0 as the original feature and input it into the CameraViT encoder. The finally output image feature H c can be expressed as:
[0082]
[0083] Where LN represents the layer normalization operation of the LayerNormalization layer, and L c represents the depth of the Transformer model.
[0084] The CameraViT encoder is a Transformer encoder composed of alternating MSA layers and MLP blocks. Its hyperparameters are obtained through pre-training. Layer normalization is applied before each MLP block, and a residual connection is applied after each MLP block. The image feature H c2D vision tasks can be completed by adding different task headers and have the ability to connect with other features from different modalities to perform multi-modal tasks.
[0085] Steps 1 and 2: LidarViT Structure Design:
[0086] LidarViT is used to process the point cloud input from lidar. The input is processed by dividing the original point cloud in 3D space into voxels. Due to the sparsity of the point cloud, voxels that do not contain points need to be removed first to reduce the computational burden. The number of non-empty voxels will be used as the input length of the model. Randomly retaining N points from voxels containing more than N points reduces the computational burden and also reduces the bias that may be caused by highly inconsistent point cloud densities. The specific steps are as follows:
[0087] (1) Divide the point cloud space with dimensions H lH ×ν lW ×ν lD into N l ×W l ×D l voxels using a cube with side lengths ν l . The feature set of a voxel containing t points is represented as:
[0088]
[0089] where x i , y i , z i represent the 3D coordinates of each point.
[0090] (2) Calculate the centroid (c x , c y , c z ) of all points in the voxel by taking the average, and add the offset of each point relative to the centroid as a supplement. The feature set of the voxel is updated to:
[0091]
[0092] (3) Transform the updated feature set into a one-dimensional vector with length ν lH ×ν lW ×ν lD . Then for each encoded feature i:
[0093]
[0094] is transformed to the feature space f iAmong them, the FCN consists of a linear layer, a Batch Normalization layer, and a Sigmoid Linear Units layer. For f i Perform max pooling on each element of f to obtain local fusion features Through f i And The connection with can obtain the connection features
[0095]
[0096] (4) Add the learnable embedding E l And the position embedding E posl , for the embedding sequence feature LidarToken of length D l Can be expressed as:
[0097]
[0098] Among them, M is the number of voxels.
[0099] (5) Y 0 As the original feature, it is input into the LidarViT encoder, and the finally output point cloud feature H l Is expressed as:
[0100]
[0101] Among them, L l Is the depth of the Transformer model.
[0102] The LidarViT encoder is a Transformer encoder composed of alternating MSA layers and MLP blocks. Its hyperparameters are obtained by pre-training. Layer normalization is applied before each MLP block, and residual connections are applied after each MLP block. The point cloud feature H l Can complete 3D vision tasks by adding different task heads and has the ability to connect with other features from different modalities to perform multi-modal tasks.
[0103] Step 2: Detection of missing modalities:
[0104] In order to detect the missing modalities, it is necessary to evaluate the importance of the feature sequence output by the encoder and screen out the features with zero or less information through the scoring function:
[0105]
[0106] Among them, Represents the feature output by the l-th layer in the m-th modality encoder.
[0107] This scoring function dynamically represents the importance of the feature sequence of the m-th modality at the l-th layer. Once the importance score is lower than the threshold (usually set to 0.01), an inter-modal projection will be performed to replace the corresponding modal feature.
[0108] Step 3. Projection of missing modalities:
[0109] Step 3-1. Projection of missing modalities:
[0110] (1) Define the single feature projection as:
[0111]
[0112] where x m [n] represents the n-th layer feature of the m-th modality to be projected, and x m' [n] represents the n'-th layer feature of the m'-th modality for projection, and h is a shallow multi-layer perceptron.
[0113] (2) If all features in a modality are considered, the modality projection can be represented as the concatenation of feature sequence projections:
[0114]
[0115] (3) For the feature Pruning in the feature sequence dynamically detects low importance scores from all features and replaces unimportant features with projections of other embeddings. This dynamic process is performed before each encoding layer. Therefore is updated to:
[0116]
[0117] where Ⅱ is a judgment condition indicator, and its output mask vector θ is a threshold generally set to 0.01; the operator ⊙ represents element-wise multiplication;
[0118] The loss function can be expressed as:
[0119]
[0120] where L m represents the feature loss of the m-th modality, The first norm is used to mark whether an inter-modal projection replacement has been performed, and λ is a hyperparameter for evaluating the losses of different modalities.
[0121] Step 3-2. Alignment of positions after projection:
[0122] Since feature substitution can destroy the position information in the original features, to alleviate this problem, the method retains the original feature sequence position embedding and freezes the position embedding in subsequent layers, and finally substitutes the position embedding before the task head. Whether the original feature sequence is substituted or not, the model still retains the original position information.
[0123] Step 4: Multimodal Fusion ViT:
[0124] Step 4-1: Use an MLP to connect the image feature H c and the point cloud feature H l to obtain N m fusion feature encodings X m :
[0125] X m = MLP(H c ∪H L )
[0126] Step 4-2: Add the learnable embedding E m and the position embedding E m to the fusion feature sequence of length D posm . The fused feature embedding sequence FeatureFusionToken can be expressed as:
[0127]
[0128] Step 4-3: The finally output feature H m is output by the L m th layer of the multimodal fusion ViT encoder:
[0129]
[0130] The multimodal fusion ViT encoder is a Transformer encoder composed of alternating MSA layers and MLP blocks. Its hyperparameters are obtained by pre-training. Layer normalization is applied before each MLP block, and a residual connection is applied after each MLP block.
[0131] Step 4-5: By adding different task heads, visual tasks such as multimodal instance segmentation, object detection, and object tracking can be completed. The task head is used as the output layer to obtain the output of a specific task, and its general structure consists of an FCN and a SoftMax. Different hyperparameters are trained according to specific tasks such as classification, detection, and segmentation and specific application scenarios. In addition, to improve the model performance and reduce the training cost, CameraViT and LidarViT are pre-trained in advance, and then the fusion model is fine-tuned according to specific tasks.
[0132] Example 1:
[0133] In this embodiment, the method of the present invention was verified by applying the object detection task head to the multi-modal public dataset KITTI. Three common types of labels were selected: Car, Pedestrian, and cyclist, and a total of 7518 frames of data were used for the experiment. In this embodiment, data of a certain modality was randomly removed from all the data according to a percentage, and mAP was used as the evaluation index. The implementation results are shown in Table 1:
[0134] Table 1 Implementation Results in the KITTI Dataset
[0135]
[0136] The results show that compared with the traditional multi-modal object detection, the present invention has significantly higher accuracy when there is a modality missing. As the proportion of the missing modality data increases, the accuracy advantage of the present invention becomes more obvious, proving that the structural design of the present invention for the modality missing scenario is effective.
[0137] Example 2:
[0138] In this embodiment, the method of the present invention was verified by applying the semantic segmentation task head to the multi-modal public dataset Nuscenes. Four common types of labels were selected: Car, Pedestrian, bicycle, and motorcycle, and a total of 404 frames of data were used for the experiment. In this embodiment, data of a certain modality was randomly removed from all the data according to a percentage, and mIoU was used as the evaluation index. The implementation results are shown in Table 2:
[0139] Table 2 Implementation Results in the Nuscenes Dataset
[0140]
[0141] The results show that as the proportion of the missing modality increases, the advantage of the present invention over the traditional multi-modal semantic segmentation method gradually increases, proving that the structural design of the present invention for the modality missing scenario is effective.
[0142] It can be seen from the above embodiments that the method of the present invention can still maintain high performance and stable output in the multi-modal computer vision field compared with the traditional multi-modal ViT vision model in the case of modality missing, improving the accuracy and robustness of the model.
Claims
1. A method for constructing a multimodal ViT visual model structure considering the modality missing scenario, characterized in that The method comprises the following steps: Step 1: Single-mode branch structure design: Step 1. CameraViT structure design: CameraViT is used to segment images into patches for learning; Step 1 and 2: LidarViT structure design: LidarViT is used to process point cloud input from LiDAR by dividing the raw point cloud in 3D space into voxels. Step 2: Detection of missing modes: In order to detect the absence of modality, it is necessary to evaluate the importance of the feature sequence output by the encoder and filter out features with zero or less information through the scoring function: Among them, MSA stands for multi-head self-attention layer, MLP stands for multi-layer perceptron, represents the features of the output of the lth layer in the mth modal encoder, express Ratings, Represents the output of the MSA layer; the scoring function dynamically represents the importance of the feature sequence of the lth layer of the mth modality. Once the importance score is lower than the threshold, inter-modal projection will be performed to replace the corresponding modal features; Step 3: Projection of missing modes: Step 31: Projection of missing modes: (1) Define a single feature projection as: Among them, x m [n] represents the n-th layer feature of the m-th mode being projected, x m' [n] represents the n′th layer feature of the m′th modality to be projected, and h is a multi-layer perceptron with a shallow depth; (2) If all features in a modality are considered, the modality projection is represented as the concatenation of feature sequence projections: (3) For features Pruning in the feature sequence dynamically detects low importance scores from all features and replaces the unimportant features with other embedded projections. This dynamic process is performed before each encoding layer, so Updated to: Among them, Ⅱ is the judgment condition indicator, and its output mask vector θ is the threshold; the operator ⊙ represents element-by-element multiplication; The loss function is expressed as: Among them, L m represents the feature loss of the mth mode, The first norm is used to mark whether inter-modal projection substitution is performed, and λ is a hyperparameter for evaluating different modal losses; Step 32: Position alignment after projection: Since feature substitution will destroy the position information in the original features, in order to alleviate this problem, the original feature sequence position embedding is retained, and the position embedding in the subsequent layers is frozen. Finally, the position embedding is replaced before the task head. Regardless of whether the original feature sequence is replaced, the model still retains the original position information; Step 4: Multimodal Fusion ViT: Step 4.
1. Use an MLP to connect image features H c And point cloud features H l To obtain N m Fusion feature encoding X m : X m =MLP(H c ∪H L ) Step 4.2: When the length is D m The feature sequence of m and position embedding E posm , the fused feature embedding sequence feature Fusion Token is expressed as: Step 4.3: Final output feature H m By the Lth ... m Layer output: Step 4 and 5: Complete multimodal instance segmentation, target detection, and target tracking visual tasks by adding different task heads.
2. The method for constructing a multimodal ViT visual model structure considering a modality missing scenario according to claim 1 is characterized in that The specific steps of the steps are as follows: (1) Set each card to be H I , width is W I Image Split into N C The length is ν cH , width is ν cW Patch, that is (2) Expand each patch to a length of ν cH ×ν cW ×3 one-dimensional features and input into MLP: in, is the feature input to the MLP, is the result of MLP output, and the length of both is D c Stay consistent; (3) Add a learnable embedding E to the feature sequence c and the position embedding E that preserves the position information posc , the generated embedded sequence feature Camera Token is expressed as: Where N is the number of patches, x class It is a trainable parameter for feature classification; (4) Z0 is used as the original feature and input into the CameraViT encoder. The final output image feature H c It is expressed as: Hc=LN(Z Lc ) Among them, LN represents the LayerNormalization layer normalization operation, L c Indicates the depth of the Transformer model.
3. The method for constructing a multimodal ViT visual model structure considering a modality missing scenario according to claim 1 is characterized in that The CameraViT encoder is a Transformer encoder consisting of alternating MSA layers and MLP blocks, whose hyperparameters are obtained by pre-training. Layer normalization is applied before each MLP block and residual connection is applied after each MLP block.
4. The method for constructing a multimodal ViT visual model structure considering a modality missing scenario according to claim 1, characterized in that The specific steps of step one and step two are as follows: (1) Through the length, width and height of ν lH ×ν lW ×ν lD The cube will be H in length, width and height. l ×W l ×D l The point cloud space is divided into N l voxels, a feature set X of voxels containing t points p It is expressed as: in, represents the laser point within the voxel, x i ,y i ,z i represents the 3D coordinates of the i-th laser point, R 3 Represents three-dimensional Euclidean space; (2) Calculate the center of mass of all points in the voxel by averaging (c x ,c y ,c z ), and add the offset of each point relative to the center of mass as a supplement, the feature set of voxel Updated to: in, represents the set of 3D coordinates of the i-th laser point and its offset complement, R 6 represents six-dimensional Euclidean space; (3) Transform the updated feature set into a length of ν lH ×ν lW ×ν lD A one-dimensional vector of , then for each encoded feature i: As point-level feature markers, they are transformed into feature space f through FCN i In the i Perform maximum pooling on each element to obtain local fusion features By f i and The connection feature is obtained by connecting (4) Add learnable embedding E to the feature sequence l and position embedding E posl , for a length of D l The embedded sequence feature LidarToken is expressed as: Where M is the number of voxels; (5) Y0 is used as the original feature and input into the LidarViT encoder. The final output point cloud feature H l It is expressed as: Among them, L l is the depth of the Transformer model.
5. The method for constructing a multimodal ViT visual model structure considering a modality missing scenario according to claim 4 is characterized in that The LidarViT encoder is a Transformer encoder consisting of alternating MSA layers and MLP blocks, whose hyperparameters are obtained by pre-training. Layer normalization is applied before each MLP block and residual connection is applied after each MLP block.
6. The method for constructing a multimodal ViT visual model structure considering a modality missing scenario according to claim 4, characterized in that The FCN consists of a linear layer, a BatchNormalization layer, and a Sigmoid LinearUnits layer.
7. The method for constructing a multimodal ViT visual model structure considering a modality missing scenario according to claim 1, characterized in that The multimodal fusion ViT encoder is a Transformer encoder composed of alternating MSA layers and MLP blocks, whose hyperparameters are obtained by pre-training. Layer normalization is applied before each MLP block, and residual connection is applied after each MLP block.
Citation Information
Patent Citations
Multi-modal image aesthetic quality evaluation method for solving modal missing problem
CN114549850A
Multi-modal emotion recognition method and system for modal missing scene
CN116933051A
Multi-modal Transform semantic segmentation algorithm for coping with RGB-D modal deficiency
CN117671265A
Multi-modal sentiment analysis method combining pre-training model and self-attention block
CN118898046A
Modal information completion method, apparatus, and device
US20230206121A1