A Multimodal Data Fusion Method and System
Through the cross-modal coordinated attention mechanism and adaptive fusion weight, the problems of information loss and high computing cost in multimodal fusion of three-dimensional point clouds and two-dimensional images are solved, and deep collaboration and efficient computing are achieved.
Patent Information
- Application Number
- CN202510174259.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The existing multimodal fusion technology of three-dimensional point cloud and two-dimensional image has problems such as information loss, insufficient deep synergy, insufficient modal complementary information, and high computational costs.
The cross-modal collaborative attention mechanism is used to perform multi-scale feature extraction and adaptive fusion weight calculation, dynamically adjust the weight allocation between modes, and optimize the calculation efficiency through multi-scale dynamic fusion.
Effectively reduce information loss, enhance deep synergy between modes, balance modal contribution, reduce computing complexity, and improve real-time and resource efficiency.
Smart Images

Figure CN119693759B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of point cloud data processing, and particularly to a multi-modal data fusion method and system. Background Art
[0002] The multi-modal data fusion task of three-dimensional point cloud and two-dimensional image is mainly applied to fields such as autonomous driving and environmental perception, aiming to combine the three-dimensional point cloud collected by lidar and the two-dimensional image data captured by a camera to improve the performance of tasks such as object detection and semantic segmentation. Existing technical solutions can be divided into four modes. Early fusion realizes the spatial alignment of the two modalities at the raw data level through an external calibration matrix. For example, Frustum PointNet projects the region of interest in the image into the point cloud for three-dimensional object detection. The mid-term fusion method performs fusion after multi-modal feature extraction. For example, PointPainting and FusionPainting project two-dimensional semantic information onto the point cloud and then process it using a three-dimensional network. The late fusion method processes the multi-modal branches separately and then fuses the outputs, retaining the independent processing process of each modality. The multi-phase fusion combines the advantages of early, mid-term, and late fusion and designs a multi-stage fusion module. For example, PMF fuses by projecting the point cloud onto the image plane, and 2D3DNet jointly optimizes the 2D segmentation and 3D point cloud model. These technical solutions provide different implementation paths for the fusion of three-dimensional point cloud and two-dimensional image.
[0003] Existing three-dimensional point cloud and two-dimensional image multi-modal fusion technologies still have some objective drawbacks in practical applications. Early fusion methods are prone to information loss due to the sparsity of the point cloud during the projection process, especially in the far-distance area, and are limited by the field of view of the camera. Although the mid-term fusion method processes at the feature level, it usually ignores the deep cooperation between different modalities, or due to the resolution difference, the information of a certain modality dominates, affecting the overall fusion effect. Although late fusion retains the independent processing of each modality, it fails to fully utilize the complementary information between modalities, resulting in limited fusion effect. In addition, although the multi-phase fusion method attempts to combine the advantages of multiple fusion stages, it has a complex design, high computational cost, and may increase the complexity of the model, affecting real-time performance and resource efficiency. Summary of the Invention
[0004] In view of this, the present invention proposes a multi-modal data fusion method and system, which reduces information loss during the point cloud projection process through a cross-modal collaborative attention mechanism and enhances the deep collaboration between different modalities. Through multi-scale feature extraction and adaptive fusion weights, the dominant problem caused by the resolution difference of different modalities is effectively solved, and at the same time, the weight allocation between modalities is dynamically adjusted, improving the adaptability and flexibility of the model. In addition, this method optimizes the overall computing efficiency by intelligently allocating computing resources, solving the problem of excessive computing burden in the prior art.
[0005] To achieve the above object, a multi-modal data fusion method is provided. The multi-modal data includes three-dimensional point cloud data and two-dimensional image data, and the method includes the following steps:
[0006] S1: Perform multi-scale feature extraction on the multi-modal data to obtain three-dimensional point cloud data features and two-dimensional image data features;
[0007] S2: Dynamically calculate the mutual relationship between the three-dimensional point cloud features and the image data features based on a cross-modal collaborative attention mechanism;
[0008] S3: Perform dynamic feature reconstruction on the image features and the point cloud features;
[0009] S4: Perform multi-modal fusion on the reconstructed image features and point cloud features to obtain the final fusion features 。
[0010] Preferably, performing multi-scale feature extraction on the multi-modal data includes performing feature extraction on the two-dimensional image data and performing feature extraction on the three-dimensional point cloud data.
[0011] Preferably, performing feature extraction on the two-dimensional image data specifically includes: using a convolutional neural network model (CNN) to extract multi-scale image features of the two-dimensional image data. After the two-dimensional image data passes through different convolutional layers of the convolutional neural network, a series of feature maps are generated, and the convolutional neural network model captures the local information and global information of the two-dimensional image data respectively; the image features at each scale are represented as , where n is the number of feature vectors, d is the dimension of the features, and s is the index of multi-scale processing, used to distinguish the features extracted at different scales.
[0012] Preferably, performing feature extraction on the three-dimensional point cloud data specifically includes: using the point cloud processing network of PointNet++ to perform hierarchical sampling and feature extraction on the three-dimensional point cloud data, generating geometric features of the point cloud at different scales, and the point cloud features at each scale are represented as , , where m is the number of points, d is the feature dimension, and s is the index of multi-scale processing, which is used to distinguish features extracted at different scales.
[0013] Preferably, in the S2, the specific method for dynamically calculating the correlation between the three-dimensional point cloud features and the image data features based on the cross-modal collaborative attention mechanism is as follows:
[0014] S2.1: Map the image features to the query vector ;
[0015] S2.2: Map the point cloud features to the key vector ;
[0016] S2.3: Map the point cloud features to the value vector ;
[0017] S2.4: Calculate the correlation between the image features and the point cloud features;
[0018] S2.5: Use the attention weight matrix to perform weighted summation on the value vector to obtain the image features combined with point cloud information .
[0019] Preferably, the S3 is specifically: Dynamically reconstruct the features of the image features and the point cloud features according to the attention weight matrix , that is is the image features after dynamic feature reconstruction, is the point cloud features after dynamic feature reconstruction.
[0020] Preferably, the S4 is specifically:
[0021] S4.1: Calculate the adaptive fusion weight;
[0022] Introduce an adaptive weight generation module to calculate the weights of each modality at different scales according to the reconstructed image features and point cloud features , where is the MLP neural network module. The input of the MLP neural network module is the reconstructed image features and point cloud features, and the output of the MLP neural network module is the fusion weight;
[0023] S4.2: Perform multi-modal fusion on the reconstructed image features and point cloud features based on the adaptive fusion weight to obtain the final fusion features ;
[0024] Among them, the S4.2 is expressed by the formula as:
[0025] 。
[0026] Preferably, the S2.1 is specifically: mapping the image feature to a query vector ; that is ; where is a learnable matrix for mapping the image feature to the query vector, is the dimension of the query vector after mapping; the S2.2 is specifically: mapping the image feature to a key vector ; that is ; where is a learnable matrix for mapping the point cloud feature to the key vector, is the dimension of the query vector after mapping; the S2.3 is specifically: mapping the image feature to a value vector ; that is ; where is a learnable matrix for mapping the point cloud feature to the value vector, is the dimension of the value vector after mapping.
[0027] Preferably, in the S2.4, the mutual relationship between the image feature and the point cloud feature is characterized by an attention weight matrix; based on the query vector and the key vector calculate the attention weight matrix;
[0028] The calculating the attention weight matrix based on the query vector and the key vector is specifically: using the dot product attention mechanism, calculate the dot product of the query vector and the key vector , and normalize the dot product through the softmax function to obtain the attention weight matrix ;
[0029] where , 。
[0030] According to another aspect of the present invention, there is provided a multimodal data fusion system, which adopts the above-mentioned multimodal data fusion method, and the system includes:
[0031] A feature extraction module, configured to perform multi-scale feature extraction on the multimodal data to obtain three-dimensional point cloud data features and two-dimensional image data features;
[0032] A cross-modal collaborative attention mechanism module for dynamically calculating the mutual relationship between the three-dimensional point cloud features and the image data features based on the cross-modal collaborative attention mechanism;
[0033] A dynamic feature reconstruction module for dynamically reconstructing the image features and the point cloud features;
[0034] A multi-modal fusion module for multi-modal fusion of the reconstructed image features and point cloud features to obtain the final fused features
[0035] The advantages and beneficial effects of the present invention are:
[0036] Through the dynamic cross-modal attention mechanism, the present invention accurately captures the mutual relationship between image and point cloud features, avoiding information loss caused by inaccurate projection, especially performing well in long-distance regions; through the adaptive attention mechanism, the present invention enables the image and point cloud features to dynamically interact, achieving deep-level modal collaboration and ensuring that all scale features of the two modalities can be effectively utilized; through the adaptive weight generation mechanism, the present invention balances the contributions of different modalities in fusion, avoiding the problem of information imbalance; at the same time, the present invention optimizes the fusion process, reducing the computational complexity through multi-scale dynamic fusion and improving the real-time performance and resource efficiency. Description of the Drawings
[0037] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the present invention or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0038] Figure 1 It is a flowchart of a multi-modal data fusion method provided by an embodiment of the present invention;
[0039] Figure 2 It is a schematic diagram of a multi-modal data fusion system provided by an embodiment of the present invention. Detailed Embodiments
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0041] Appendix Figure 1 shows a flowchart of a multi-modal data fusion method, as shown in the appendix Figure 1 shown, a multi-modal data fusion method, wherein the multi-modal data includes three-dimensional point cloud data and two-dimensional image data, and the method includes the following steps:
[0042] S1: Perform multi-scale feature extraction on the multi-modal data to obtain three-dimensional point cloud data features and two-dimensional image data features;
[0043] Among them, performing multi-scale feature extraction on the multi-modal data includes performing feature extraction on the two-dimensional image data and performing feature extraction on the three-dimensional point cloud data;
[0044] The convolutional neural network model (CNN) is a neural network model widely used in the field of image recognition and processing. Its basic structure includes an input layer, a convolutional layer, a pooling layer, and a fully connected layer. Among them, the input layer is used to receive the original image data and convert it into a form that the neural network can process. The convolutional layer extracts the features of the image through convolutional operations. The convolutional operation convolves the image with a set of convolutional kernels to obtain a feature map. A convolutional kernel is a small matrix that multiplies and sums with the image element by element in a sliding window manner to obtain each pixel value in the feature map. The convolutional operation can preserve the spatial relationship and local features of the image. Through different convolutional kernels, the CNN can learn different features, such as edges, textures, and shapes, etc. The pooling layer is used to reduce the size of the feature map and improve the calculation efficiency; common pooling operations include max pooling and average pooling. The purpose of the pooling operation is to reduce the size of the feature map, improve the calculation efficiency, and have a certain translational invariance. Through the pooling operation, the CNN can abstract the details of the image to better capture the overall features of the image. The fully connected layer is used to map the feature map to the output layer to achieve image classification or regression. Each neuron in the fully connected layer is connected to all neurons in the previous layer and discriminates different classes by learning weight parameters. The fully connected layer plays a decision-making role in the CNN. Through the learned weight parameters, the information of the feature map can be converted into a prediction of the image class;
[0045] Therefore, the specific method for performing feature extraction on the two-dimensional image data is: using a convolutional neural network model (CNN) to extract multi-scale image features of the two-dimensional image data. After the two-dimensional image data passes through different convolutional layers of the convolutional neural network, a series of feature maps are generated. The convolutional neural network model captures the local information and global information of the two-dimensional image data respectively; the image features of each scale are represented as , , where n is the number of feature vectors, d is the dimension of the features, and s is the index of multi-scale processing, which is used to distinguish features extracted at different scales.
[0046] PointNet++ is a network model used to process 3D point cloud data in the field of deep learning. PointNet++ is an improvement over PointNet, which addresses some limitations of PointNet when dealing with point cloud data. The main improvements of PointNet++ are as follows:
[0047] Local feature extraction: PointNet++ introduces sampling and grouping, considering local neighborhood features, while PointNet only considers single-point features and cannot represent local structures well.
[0048] Hierarchical structure: PointNet++ adopts a hierarchical structure, which can effectively extract local features of different regions according to different receptive field sizes. In contrast, the global feature of PointNet is directly obtained by max pooling, which is prone to information loss.
[0049] Rotation invariance: PointNet++ uses local relative coordinates for feature extraction and eliminates the TNet network, while PointNet uses TNet to ensure the rotation invariance of point cloud features.
[0050] Sample unevenness problem: PointNet++ proposes multi-scale method MSG and multi-level method MRG to solve the sample unevenness problem, while PointNet does not handle it.
[0051] Segmentation network structure: PointNet++ adopts an Encoder-Decoder structure, and features are connected through skip link concatenation, while PointNet directly integrates global feature and local embedding features.
[0052] Therefore, the feature extraction of 3D point cloud data is specifically as follows: Use the point cloud processing network of PointNet++ to perform hierarchical sampling and feature extraction on 3D point cloud data, generate geometric features of the point cloud at different scales, and the point cloud features at each scale are represented as , , where m is the number of points, d is the feature dimension, and s is the index of multi-scale processing, which is used to distinguish features extracted at different scales;
[0053] S2: Dynamically calculate the mutual relationship between the 3D point cloud features and the image data features based on the cross-modal collaborative attention mechanism;
[0054] Among them, the specific method for dynamically calculating the correlation between the 3D point cloud features and the image data features based on the cross-modal collaborative attention mechanism is as follows:
[0055] S2.1: Map the image features to a query vector ;
[0056] Among them, the specific method of S2.1 is: Map the image features to a query vector through matrix multiplication; that is ; where is a learnable matrix for mapping the image features to the query vector, is the dimension of the query vector after mapping;
[0057] S2.2: Map the point cloud features to a key vector ;
[0058] Among them, the specific method of S2.2 is: Map the image features to a key vector through matrix multiplication; that is ; where is a learnable matrix for mapping the point cloud features to the key vector, is the dimension of the query vector after mapping;
[0059] S2.3: Map the point cloud features to a value vector ;
[0060] Among them, the specific method of S2.3 is: Map the image features to a value vector through matrix multiplication; that is ; where is a learnable matrix for mapping the point cloud features to the value vector, is the dimension of the value vector after mapping;
[0061] S2.4: Calculate the correlation between the image features and the point cloud features;
[0062] In this step, the correlation between the image features and the point cloud features is characterized by an attention weight matrix; specifically, based on the query vector and the key vector calculate the attention weight matrix;
[0063] Furthermore, the method based on the query vector And the key vector Specifically, calculating the attention weight matrix is as follows:
[0064] Using the dot - product attention mechanism, calculate the dot - product of the query vector and the key vector , and normalize the dot - product through the softmax function to obtain the attention weight matrix ;
[0065] Wherein, , ;
[0066] S2.5: Use the attention weight matrix to perform weighted summation on the value vector to obtain the image features combined with point cloud information ;
[0067] Wherein, the formula of S2.5 is: ;
[0068] Through the above steps, the image features after fusing point cloud information are generated . These features not only contain the information of the original image but also combine the point cloud feature information related to it.
[0069] S3: Perform dynamic feature reconstruction on the image features and the point cloud features;
[0070] Wherein, S3 specifically is: According to the attention weight matrix perform dynamic feature reconstruction on the image features and the point cloud features, that is is the image feature after dynamic feature reconstruction, is the point cloud feature after dynamic feature reconstruction; Through the attention mechanism, the image features can be reconstructed by combining point cloud information, and the point cloud features will also absorb image information. During the feature reconstruction process, the complementary information between modalities is fully utilized, thereby reducing information loss and the limitations of single - modality features.
[0071] S4: Perform multi - modal fusion on the reconstructed image features and the point cloud features to obtain the final fused features ;
[0072] Wherein, S4 specifically is:
[0073] S4.1: Calculate the adaptive fusion weight;
[0074] Wherein, in order to balance the contributions of multi - modal features, an adaptive weight generation module is introduced to calculate the weights of each modality at different scales according to the reconstructed image features and the point cloud features , where It is an MLP neural network module. The input of the MLP neural network module is the reconstructed image features and the point cloud features, and the output of the MLP neural network module is the fusion weight;
[0075] Through adaptive adjustment, features of different scales can be flexibly assigned different weights according to the task requirements, preventing a certain modality from dominating the fusion process.
[0076] S4.2: Perform multimodal fusion on the reconstructed image features and the point cloud features based on the adaptive fusion weight to obtain the final fusion features ;
[0077] Among them, the S4.2 is expressed by the formula:
[0078] The fused features represent the deep interaction and complementary information between the image and the point cloud at different scales. Multiscale fusion ensures the integrity of the fused features globally and locally, avoids information loss, and makes the features more expressive.
[0079] Through the above solution, this embodiment has the following advantages:
[0080] 1) Avoid information loss: Existing early fusion methods are prone to information loss under point cloud sparsity and field of view limitations. This technology accurately captures the mutual relationship between image and point cloud features through a dynamic cross-modal attention mechanism, avoiding information loss caused by inaccurate projection, especially performing well in long-distance regions.
[0081] 2) Enhance modality collaboration: Existing mid-term fusion methods lack deep collaboration between different modalities. This technology enables dynamic interaction between image and point cloud features through an adaptive attention mechanism, realizes deep modality collaboration, and ensures that features at all scales of both modalities can be effectively utilized.
[0082] 3) Balance the modality dominance effect: In existing methods, due to resolution differences, a certain modality may dominate, affecting the fusion effect. This technology balances the contributions of different modalities in the fusion through an adaptive weight generation mechanism, avoiding the problem of information imbalance.
[0083] 4) Simplify calculation and improve real-time performance: Although the multi-phase fusion technology is comprehensive, it has a complex design and high computational cost. This technology optimizes the fusion process, reduces the computational complexity through multiscale dynamic fusion, and improves the real-time performance and resource efficiency.
[0084] Embodiment 2. This embodiment includes a multimodal data fusion system, which Figure 2 shows a structural diagram of a multimodal data fusion system, as shown in the appendix Figure 2As shown, the system adopts the multi-modal data fusion method of Embodiment 1, and the system includes:
[0085] A feature extraction module, configured to perform multi-scale feature extraction on the multi-modal data to obtain three-dimensional point cloud data features and two-dimensional image data features;
[0086] A cross-modal collaborative attention mechanism module, configured to dynamically calculate the mutual relationship between the three-dimensional point cloud features and the image data features based on the cross-modal collaborative attention mechanism;
[0087] A dynamic feature reconstruction module, configured to perform dynamic feature reconstruction on the image features and the point cloud features;
[0088] A multi-modal fusion module, configured to perform multi-modal fusion on the reconstructed image features and point cloud features to obtain final fusion features
[0089] Embodiment 3. This embodiment includes a computer-readable storage medium, on which a data processing program is stored, and the data processing program is executed by a processor to implement the multi-modal data fusion method of Embodiment 1.
[0090] Those skilled in the art should understand that the embodiments herein can be provided as methods, apparatuses (devices), or computer program products. Therefore, the embodiments herein can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Including but not limited to RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0091] This article is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (devices), and computer program products according to the embodiments herein. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate for implementing in the process Figure 1 a process or multiple processes and / or blocks Figure 1Apparatus for the functions specified in one or more boxes.
[0092] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device implements the process Figure 1 One process or multiple processes and / or boxes Figure 1 Steps for the functions specified in one or more boxes.
[0093] The above-described embodiments and / or implementation manners are only used to illustrate the preferred embodiments and / or implementation manners for implementing the technology of the present invention, and do not impose any formal restrictions on the implementation manners of the technology of the present invention. Any person skilled in the art, without departing from the scope of the technical means disclosed in the content of the present invention, can make some changes or modifications to other equivalent embodiments, but should still be regarded as the same technology or embodiment as the present invention in essence. Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described here, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A method for fusion of multimodal data, wherein the multimodal data includes three-dimensional point cloud data and two-dimensional image data, characterized in that: The method comprises the following steps: S1: performing multi-scale feature extraction on the multimodal data to obtain three-dimensional point cloud data features and two-dimensional image data features; S2: Dynamically calculating the relationship between the three-dimensional point cloud features and the image data features based on a cross-modal collaborative attention mechanism; the dynamic calculation of the relationship between the three-dimensional point cloud features and the image data features based on a cross-modal collaborative attention mechanism is specifically: S2.1: The image features Mapping to query vector ; s is the index of multi-scale processing, which is used to distinguish features extracted at different scales; S2.2: The point cloud features Mapping to key vector ; S2.3: The point cloud features Mapping to a vector of values ; S2.4: Calculate the relationship between the image features and the point cloud features; In S2.4, the relationship between the image features and the point cloud features is represented by an attention weight matrix; Based on the query vector and key vector Calculating the attention weight matrix; The query vector and key vector The calculation of the attention weight matrix is specifically as follows: using the dot product attention mechanism, the calculation based on the query vector and key vector The dot product is normalized by the softmax function to obtain the attention weight matrix ; in, , ; n is the number of feature vectors, m is the number of points; d is the feature dimension; S3: Dynamically reconstruct the image features and the point cloud features; S3 specifically comprises: according to the attention weight matrix Dynamic feature reconstruction is performed on the image features and the point cloud features, that is, , , is the image feature reconstructed by dynamic features, Point cloud features reconstructed from dynamic features; S4: Perform multimodal fusion on the reconstructed image features and the point cloud features to obtain the final fusion features ; The S4 is specifically: S4.1: Calculate adaptive fusion weights; An adaptive weight generation module is introduced to calculate the weight of each modality at different scales based on the reconstructed image features and point cloud features. ,in is an MLP neural network module, the input of the MLP neural network module is the reconstructed image features and the point cloud features, and the output of the MLP neural network module is the fusion weight; S4.2: Based on the adaptive fusion weight, the reconstructed image features and the point cloud features are multimodally fused to obtain the final fusion features. ; Wherein, the S4.2 is expressed by the formula: 。 2. A multimodal data fusion method according to claim 1, characterized in that: Performing multi-scale feature extraction on the multimodal data includes performing feature extraction on the two-dimensional image data and performing feature extraction on the three-dimensional point cloud data.
3. A multimodal data fusion method according to claim 2, characterized in that: The feature extraction of the two-dimensional image data is specifically performed as follows: a convolutional neural network model (CNN) is used to extract multi-scale image features of the two-dimensional image data. After the two-dimensional image data passes through different convolutional layers of the convolutional neural network, a series of feature maps are generated. The convolutional neural network model captures the local information and global information of the two-dimensional image data respectively; the image features of each scale are represented as , , d is the feature dimension.
4. A multimodal data fusion method according to claim 2 or 3, characterized in that: The feature extraction of 3D point cloud data is as follows: using the point cloud processing network of PointNet++, performing hierarchical sampling and feature extraction on the 3D point cloud data, generating geometric features of point clouds at different scales, and representing the point cloud features of each scale as follows: , , d is the feature dimension.
5. The multimodal data fusion method according to claim 1, characterized in that: The S2.1 is specifically: transforming the image features by matrix multiplication Mapping to query vector ;Right now ;in, is a learnable matrix for mapping the image features to a query vector, is the dimension of the query vector after mapping; S2.2 is specifically: transforming the image features into Mapping to key vector ;Right now ;in, is a learnable matrix for mapping the point cloud features to key vectors, is the dimension of the query vector after mapping; S2.3 is specifically: transforming the image features into Mapping to a vector of values ;Right now ;in, is a learnable matrix for mapping the point cloud features to a value vector, d is the feature dimension, is the dimension of the mapped value vector.
6. A multimodal data fusion system for three-dimensional point cloud data processing, characterized in that: The system adopts a multimodal data fusion method according to any one of claims 1 to 5, and the system comprises: A feature extraction module, used to perform multi-scale feature extraction on the multimodal data to obtain three-dimensional point cloud data features and two-dimensional image data features; A cross-modal collaborative attention mechanism module, used for dynamically calculating the relationship between the three-dimensional point cloud features and the image data features based on the cross-modal collaborative attention mechanism; A dynamic feature reconstruction module, used for dynamically reconstructing the image features and the point cloud features; The multimodal fusion module is used to perform multimodal fusion on the reconstructed image features and the point cloud features to obtain final fused features.
Citation Information
Patent Citations
Three-dimensional target detection method based on multi-modal fusion and deformable attention
CN117975436A
4D millimeter wave radar and visual adaptive fusion target identification system
CN118155174A
Multi-modal feature interaction 3D multi-target tracking method based on deep learning
CN118447354A