Multimodal Knowledge Fusion Method, Storage Medium and Application Based on Knowledge Transfer
The knowledge transfer method aligns 2D image features with 4D point cloud features to enhance texture information in 4D point clouds, addressing sparsity and alignment issues, thereby improving 4D scene understanding tasks efficiently.
Patent Information
- Application Number
- CN202311424648.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-10-31
AI Technical Summary
The existing four-dimensional point cloud representation framework lacks texture information and cannot effectively solve the four-dimensional scene understanding task. The two-dimensional image-assisted multimodal fusion method increases the computational overhead and time asymmetry problems.
Through knowledge migration in the training stage, the texture information of the two-dimensional image is migrated to the four-dimensional point cloud, and the multi-modal attention mechanism of attention occlusion is used for time alignment. Only point cloud mode is input in the implementation stage to reduce computing resource overhead.
It improves the performance of the four-dimensional scene understanding task, while avoiding additional computing resource consumption, solving the problem of sparsity and texture lack of point cloud data.
Smart Images

Figure CN117671438B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a multimodal knowledge fusion method, a storage medium, and an application based on knowledge transfer. Background Art
[0002] The field of four-dimensional (4D) point cloud understanding, which aims to analyze dynamic three-dimensional (3D) point cloud sequences, is developing rapidly. Different from traditional picture and video tasks, 4D point cloud sequences directly obtain spatial geometric information from the three-dimensional space and time information from the video sequence, which is more conducive to real-world interaction. These attributes are very important for 4D scene understanding such as action recognition and semantic segmentation.
[0003] Existing 4D point cloud representation frameworks all extend common 3D point cloud models to 4D and introduce additional temporal representation modules to achieve cross-temporal feature interaction. However, due to the sparsity and irregularity of 4D point clouds and the lack of texture information, it is impossible to well solve 4D scene understanding tasks only through the representation learning of the 4D point cloud modality.
[0004] Two-dimensional image-assisted multimodal fusion is a feasible way to supplement the rich texture information of two-dimensional images for 4D point clouds to enhance 4D point cloud representation. However, introducing two-dimensional image data inevitably introduces additional network design and computational overhead. And there is a problem of temporal information misalignment between two-dimensional images and 3D point clouds, which brings instability to multimodal fusion technology.
[0005] For traditional 4D point cloud representation frameworks, they lack texture information and cannot solve the irregularity and sparsity of point cloud data, and cannot well solve 4D scene understanding tasks such as action recognition and semantic segmentation. Although the method of two-dimensional image-assisted multimodal fusion supplements texture information for 4D point clouds, the introduced two-dimensional images bring additional computational overhead and network design, posing challenges to online 4D scene understanding tasks. The misalignment of two-dimensional images and 4D point clouds in the time dimension affects the performance of multimodal fusion technology.
[0006] In summary, there is currently a lack of a multimodal knowledge fusion method that can fuse two-dimensional images while reducing the computational resource overhead in the implementation stage. Summary of the Invention
[0007] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a multimodal knowledge fusion method, a storage medium, and an application based on knowledge transfer to achieve the reduction of computational resource overhead in the implementation stage while fusing two-dimensional images.
[0008] The purpose of the present invention can be achieved through the following technical solutions:
[0009] One aspect of the present invention provides a multimodal knowledge fusion method based on knowledge transfer, including the following steps:
[0010] Training stage:
[0011] Obtain the image sequence for training and its temporal gradient information, as well as point cloud data, and obtain image features, temporal features, and point cloud sequence representations through encoding;
[0012] For the image features and the temporal features, obtain the two-dimensional image representation after time alignment through image consistency processing;
[0013] For the two-dimensional image representation and the point cloud sequence representation, transfer the knowledge of the two-dimensional image representation to the point cloud sequence representation through a multimodal attention mechanism with attention occlusion to obtain a fusion feature, and perform training based on the fusion feature;
[0014] Implementation stage:
[0015] Obtain the point cloud data for prediction and perform encoding, obtain a new fusion feature through a multimodal attention mechanism with attention occlusion, and execute a preset four-dimensional scene understanding task based on the new fusion feature.
[0016] As a preferred technical solution, the process of obtaining the two-dimensional image representation includes:
[0017] Use a temporal sliding window to extract and perform secondary encoding on the image features and the temporal features respectively to obtain the secondarily encoded image features and the secondarily encoded temporal features;
[0018] Based on the secondarily encoded image features and the secondarily encoded temporal features, obtain the time correlation feature through fusion;
[0019] Based on the time correlation feature and the secondarily encoded image features, obtain the two-dimensional image representation through fusion.
[0020] As a preferred technical solution, in the training stage, the secondarily encoded image features and the secondarily encoded temporal features achieve time alignment through contrastive learning in the time dimension.
[0021] As a preferred technical solution, the time correlation feature is obtained through fusion by a fully connected network.
[0022] As a preferred technical solution, the two-dimensional image representation is obtained through fusion based on the attention mechanism.
[0023] As a preferred technical solution, in the training stage, the attention occlusion is configured such that the point cloud sequence representation cannot utilize the information of the two-dimensional image representation sequence representation.
[0024] As a preferred technical solution, the point cloud data is four-dimensional point cloud data.
[0025] As a preferred technical solution, the image sequence and its temporal gradient information obtain image features and temporal features through their respective image encoders, and the point cloud data obtains a point cloud sequence representation through a point cloud encoder.
[0026] Another aspect of the present invention provides an application of the foregoing multimodal knowledge fusion method based on knowledge transfer, and the foregoing method is applied to a four-dimensional scene understanding task, and the four-dimensional scene understanding task includes action recognition and / or semantic recognition.
[0027] Another aspect of the present invention provides a computer-readable storage medium, including one or more programs for execution by one or more processors of an electronic device, and the one or more programs include instructions for executing the foregoing multimodal knowledge fusion method based on knowledge transfer.
[0028] Compared with the prior art, the present invention has the following advantages:
[0029] In the training stage of the present invention, first, a two-dimensional image representation after time alignment is obtained by performing consistency processing on the image features and the temporal feature images, and then the knowledge of the two-dimensional image representation is transferred to the point cloud sequence representation through a multimodal attention mechanism with attention occlusion, realizing the transfer of two-dimensional image knowledge to point cloud sequence knowledge. In the implementation stage, only a single point cloud modality is input, without relying on picture representations, and finally, while improving the performance of the four-dimensional scene understanding task, no additional computing resources are added. Description of the Drawings
[0030] Figure 1 It is a schematic diagram of the basic framework for 4D point cloud video understanding based on multimodal knowledge transfer in the embodiment;
[0031] Figure 2 It is a schematic diagram of the image consistency module in the basic framework of the embodiment;
[0032] Figure 3 It is a schematic diagram for comparing the solution of the present application with the prior art;
[0033] Figure 4 It is a schematic diagram of the input data in the embodiment;
[0034] Figure 5 It is a schematic diagram of the attention occlusion module in the basic framework of the embodiment;
[0035] Figure 6 It is a schematic diagram of the visualization of action segmentation in the embodiment. Detailed Embodiments
[0036] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0037] Embodiment 1
[0038] In view of the problems existing in the foregoing prior art, this embodiment provides a multi-modal knowledge fusion method based on knowledge transfer. The method provides a four-dimensional point cloud understanding framework based on multi-modal knowledge transfer. In the training stage of this application, the texture information contained in the two-dimensional image is added to the four-dimensional point cloud through knowledge transfer, and an attention fusion module with control of information leakage between multi-modalities is introduced to achieve single-modal input of the point cloud in the implementation stage.
[0039] In addition, we also provide a series of multi-modal time alignment strategies, including additional temporal gradient input, sliding window design in the time dimension, and contrast learning strategy in the time dimension, to ensure knowledge transfer between the four-dimensional point cloud and the two-dimensional image at the same moment.
[0040] See Figure 3 showing a comparison between the proposed patent solution and the prior art. Through knowledge transfer in the training stage, this method can obtain better performance with only single-modal point cloud input in the implementation stage.
[0041] As Figure 1 shown is a schematic diagram of the four-dimensional point cloud understanding framework of this embodiment, which is divided into a training stage and an implementation stage.
[0042] Training stage: The input data in the training stage is a four-dimensional point cloud, a two-dimensional image (see Figure 4 (a) in Figure 4 ) and its temporal gradient (see (b) in ). The four-dimensional point cloud obtains the characteristic features of the point cloud sequence through the point cloud encoder. The two-dimensional picture and its temporal gradient obtain the features of the picture and time through their respective image encoders, and obtain the image representation through the two-dimensional image consistency module. The image representation and the point cloud sequence representation are fused through the multi-modal attention mechanism network to achieve knowledge transfer from the image to the point cloud sequence.
[0043] Implementation stage: The input data in the implementation stage is only a four-dimensional point cloud. The multi-modal attention mechanism with an attention occlusion module is used to add additional texture and temporal information to the point cloud data.
[0044] Among them, the output features of the multimodal attention mechanism enter the point cloud output layer and the image output layer respectively to obtain predictions. The point cloud and the image share the same label and perform backpropagation, using the same method as traditional deep learning tasks.
[0045] See Figure 2 For the schematic diagram of the two-dimensional image consistency module, the input two-dimensional picture and its temporal gradient pass through a sliding window in the time series, stack multiple time frames together, and after passing through the encoder, fuse their respective features through a fully connected network to obtain time correlation features. Finally, the time correlation features and the two-dimensional picture are fused through an attention mechanism to obtain the two-dimensional image representation. Among them, between the features of the two-dimensional image and the temporal gradient, temporal alignment is achieved using contrastive learning in the time dimension.
[0046] See Figure 5 , in the training stage, the point cloud sequence representation cannot utilize the information of the picture representation, while the picture representation can have the complete input modality information (point cloud representation and picture representation). This way, while transferring picture knowledge to the point cloud sequence in the training stage, the point cloud sequence does not depend on the picture representation in the implementation stage.
[0047] To verify the effectiveness of this method, in multiple four-dimensional scene understanding tasks, quantitative experiments are carried out on two main datasets using the method proposed in this embodiment.
[0048] See Table 1 for the results of quantitative experiments on the action segmentation task using the method of this embodiment and the P4Transformer method on datasets such as CVPR and HOI4D.
[0049] Table 1 Experimental results under the action segmentation task
[0050]
[0051] See Table 2 for the experimental results of the semantic segmentation task using the method of the embodiment and the P4Transformer method.
[0052] Table 2 Experimental results under the semantic segmentation task
[0053]
[0054] See Figure 6 For Figure 4 the schematic diagram of the visualization result under the action segmentation task when the data in
[0055] The above results show that the method of this embodiment achieves better results than the related technical methods used for comparison in the experiment.
[0056] Embodiment 2
[0057] This embodiment provides an application of the multi-modal knowledge fusion method based on knowledge transfer in Embodiment 1. The aforementioned method can be applied to various four-dimensional scene understanding tasks including action segmentation, semantic segmentation, action segmentation, etc. Except for the adopted understanding framework (i.e., model structure) and data processing within the model, all are implemented using existing technologies.
[0058] Embodiment 3
[0059] This embodiment provides a computer-readable storage medium, including a program for execution by a processor of an electronic device. The aforementioned program includes instructions for executing the multi-modal knowledge fusion method based on knowledge transfer as described in Embodiment 1.
[0060] Through knowledge transfer in the training stage, the present invention introduces two-dimensional image assistance in the four-dimensional point cloud representation framework. By using two-dimensional images and temporal gradients, additional texture and temporal information are added to the point cloud data. This technology can supplement texture information for the point cloud while still maintaining a separate point cloud modality input during the implementation stage, ultimately achieving the improvement of the performance of four-dimensional scene understanding tasks without increasing additional computing resources. In addition, the present invention adopts an algorithm design of time and motion consistency. By setting a temporal sliding window and multi-frame temporal contrast learning between multi-modalities, the temporal information between input modalities is aligned, solving the technical problem of alignment in the time dimension between multi-modal data.
[0061] The present invention does not depend on the structure of the point cloud encoder and can improve the performance of most mainstream point cloud encoders in four-dimensional scene understanding tasks. While improving the task performance, it does not increase the computational overhead during the model implementation stage.
[0062] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A multi-modal knowledge fusion method based on knowledge transfer, characterized in that, It includes the following steps: Training stage: Obtain the image sequence for training and its temporal gradient information, as well as the point cloud data, and obtain the image features, temporal features, and point cloud sequence representation through encoding; For the image features and the temporal features, obtain the two-dimensional image representation after time alignment through image consistency processing; For the two-dimensional image representation and the point cloud sequence representation, transfer the knowledge of the two-dimensional image representation to the point cloud sequence representation through a multi-modal attention mechanism with attention occlusion, obtain the fused features, and perform training based on the fused features; Implementation stage: Obtain the point cloud data for prediction and perform encoding, obtain the new fused features through a multi-modal attention mechanism with attention occlusion, and perform a preset four-dimensional scene understanding task based on the new fused features, The obtaining process of the two-dimensional image representation includes: Use a temporal sliding window to extract and perform secondary encoding on the image features and the temporal features respectively, to obtain the image features after secondary encoding and the temporal features after secondary encoding; Based on the image features after secondary encoding and the temporal features after secondary encoding, obtain the time correlation features through fusion; Based on the time correlation features and the image features after secondary encoding, obtain the two-dimensional image representation, In the training stage, the image features after secondary encoding and the temporal features after secondary encoding achieve time alignment through contrastive learning in the time dimension, The point cloud data is four-dimensional point cloud data, In the training stage, the attention occlusion is configured to prevent the point cloud sequence representation from using the information of the picture representation. The picture representation has complete input modality information, that is, the point cloud representation and the picture representation. While realizing the transfer of picture knowledge to the point cloud sequence, the point cloud sequence does not depend on the picture representation in the implementation stage.
2. The multimodal knowledge fusion method based on knowledge transfer according to claim 1, wherein The time correlation features are obtained through fusion by a fully connected network.
3. A multimodal knowledge fusion method based on knowledge transfer according to claim 1, characterized in that, The two-dimensional image representation is obtained through fusion based on the attention mechanism.
4. A multimodal knowledge fusion method based on knowledge transfer according to claim 1, characterized in that The image sequence and its temporal gradient information obtain the image features and the temporal features through their respective image encoders, and the point cloud data obtains the point cloud sequence representation through a point cloud encoder.
5. Application of a multimodal knowledge fusion method based on knowledge transfer as described in any one of claims 1-4, characterized in that Applied in a four-dimensional scene understanding task, the four-dimensional scene understanding task includes action recognition and / or semantic recognition.
6. A computer-readable storage medium, characterized in that, It includes one or more programs executed by one or more processors of a power supply device, and the one or more programs include instructions for executing the multi-modal knowledge fusion method based on knowledge transfer according to any one of claims 1-4.
Citation Information
Patent Citations
Multi-mode fusion method based on multiple views and image segmentation in three-dimensional target detection
CN113052066A
Image-laser radar data fusion method based on mixed attention mechanism
CN114398937A