Multimodal data fusion processing method based on multi-stage fusion strategy
By constructing a multimodal data fusion method with a multi-stage fusion strategy, the problems of modal weight solidification and insufficient adaptability in traditional methods are solved, efficient multimodal data fusion is achieved, and information utilization and semantic expression accuracy are improved.
Patent Information
- Application Number
- CN202511056928.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-30
AI Technical Summary
Traditional multimodal data fusion methods have problems such as fixed modal weights, low information utilization, and insufficient adaptability to complex scenarios.
A three-layer fusion architecture is constructed at the data level, feature level, and decision level. Data preprocessing, feature fusion, and decision optimization are performed through a multi-stage fusion strategy. Adaptive image block allocation, attention module, jump fusion method, and mean aggregation algorithm are used to enhance the point cloud color representation capability and semantic expression accuracy.
It significantly improves the information utilization and semantic expression accuracy in complex scenarios, enhances the model's robustness to occlusion and lighting changes, and achieves efficient fusion and complementarity of multimodal data.
Smart Images

Figure CN120563991B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a multimodal data fusion processing method based on a multi-stage fusion strategy. Background Art
[0002] Multimodality refers to the use of information from multiple different forms or perceptual channels for expression, communication, and understanding, typically including visual, auditory, textual, tactile, and other sensory input and output methods. In computer science, artificial intelligence, and machine learning, multimodal technology refers to the integration of data from different modalities, such as images, text, audio, and video, to enhance a model's understanding and reasoning capabilities. This integration improves the completeness and accuracy of information, as each modality can provide unique information for specific tasks. This multimodal technology is widely used in scenarios such as human-computer interaction, autonomous driving, and medical diagnosis, demonstrating its strong application potential.
[0003] Multimodal image classification further expands this scope. It is not limited to a single image type, but combines multiple image modalities such as RGB images, infrared images, and depth images to obtain more detailed information about the observed object. In addition, multimodal object detection technology uses the comprehensive information of this data to detect and locate specific targets in images or videos. For example, its application in autonomous driving and night vision monitoring effectively improves recognition accuracy and system response speed.
[0004] However, in the current multimodal data fusion, traditional methods generally use fixed models such as preset and big data experience to set modal weights, resulting in low information utilization. In addition, the adaptability of multimodal data fusion methods is seriously insufficient in complex scenarios.
[0005] In response to the above technical problems, this application proposes a solution. Summary of the Invention
[0006] In the present invention, by constructing a three-layer fusion architecture of data level, feature level and decision level, targeted algorithm processing is performed in the data preprocessing stage, feature fusion stage and decision optimization stage, thereby enhancing the color representation ability of the point cloud, dynamically quantifying the importance and credibility of each feature element, realizing key information enhancement and interference suppression, strengthening the semantic expression ability while retaining detailed information, significantly improving the information utilization and semantic expression accuracy in complex scenarios, solving the problems of modal weight solidification, low information utilization and insufficient adaptability to complex scenarios in traditional methods, and proposing a multimodal data fusion processing method based on a multi-stage fusion strategy.
[0007] The purpose of the present invention can be achieved through the following technical solutions:
[0008] The multimodal data fusion processing method based on the multi-stage fusion strategy includes the following steps:
[0009] Step 1: Collect the original data, perform modal judgment on the collected data, and classify it into three-dimensional data and two-dimensional data;
[0010] Step 2: Extract geometric features from 3D data, extract texture information from 2D data, and dynamically match the geometric features and texture information to construct a composite data structure containing spatial geometric features and texture information;
[0011] Step 3: Introduce the attention module, generate a one-dimensional attention map through encoding and decoding the composite data structure, and perform credibility and importance assessment of feature elements based on the one-dimensional attention map;
[0012] Step 4: Fuse the cross-modal algorithms through the jump fusion algorithm, record the fusion results, extract the abstract features through the mean aggregation algorithm, and record the extracted features;
[0013] Step 5: Generate the final result of data fusion based on the features obtained in step 4;
[0014] Step 6: Perform reliability judgment on the final feature fusion result, and generate a result optimization signal or a result reliability signal based on the reliability judgment result;
[0015] Step 7: Record the key parameters in steps 2 to 4 based on the reliability judgment results and generate corresponding reminders.
[0016] As a preferred embodiment of the present invention, the device further includes a multimodal data collection port, which is used to acquire three-dimensional data and two-dimensional data, and determine the data type to obtain single-channel usable data;
[0017] a data fusion processing module, wherein the data fusion processing module obtains single-channel available data, performs algorithm adaptive matching based on the single-channel available data, obtains three-dimensional geometric information and two-dimensional texture information, and spatially matches the three-dimensional geometric information and the two-dimensional texture information to generate composite structure data;
[0018] A feature fusion processing module, which obtains composite structure data, encodes and decodes the composite structure data through the element attention module, generates a one-dimensional attention map, obtains feature points, and evaluates the importance and credibility of the feature points;
[0019] The decision fusion module performs cross-modal feature interaction through a jump fusion algorithm to optimize the features, and then extracts abstract features through a mean aggregation algorithm to fuse the original features with the deep features.
[0020] As a preferred embodiment of the present invention, the method further includes a result judgment module, which obtains the final feature fusion result through the decision fusion module, performs credibility judgment on the feature fusion result, generates result reliability, compares the result reliability with a set database, and generates a result reliability signal or a result optimization signal according to the comparison result;
[0021] A dynamic recording module obtains key parameters during the operation of the data processing module, the feature fusion processing module and the decision fusion module, and issues abnormal reminders for key parameters when generating result optimization signals.
[0022] As a preferred embodiment of the present invention, the single-channel available data acquired by the data fusion processing module includes three-dimensional point cloud data and two-dimensional RGB color image data;
[0023] The data fusion processing module obtains the three-dimensional geometric information by creating a rectangular coordinate system with a selected origin, and performing coordinate generation and model construction of the 3D point cloud data through the set origin;
[0024] After acquiring the two-dimensional RGB color image data, the data fusion processing module sequentially performs algorithm filtering, denoising, grayscale conversion, and contrast adjustment on the two-dimensional RGB color image data to obtain a temporary image. The data fusion processing module then constructs a matrix of the frequency of occurrence of grayscale values at specific distances and directions in the temporary image, and calculates characteristic quantities such as energy, entropy, and contrast to obtain the roughness and regularity of the texture, which are recorded as two-dimensional texture information.
[0025] As a preferred embodiment of the present invention, the method in which the data fusion processing module spatially matches the three-dimensional geometric information and the two-dimensional texture information is as follows:
[0026] The point cloud model in the 3D geometric information is used as the basic source point cloud. For each point in the basic source point cloud, the corresponding point with the closest Euclidean distance is searched in the target point cloud to generate the initial matching point pair and achieve spatial coordinate alignment.
[0027] Through the point pair relationship after registration, each 3D point is projected onto a 2D plane through an algorithm to generate an irregular image block;
[0028] The data fusion processing module adds two-dimensional texture features of corresponding positions on the irregular image blocks to generate composite structure data with both spatial rectangular coordinate system coordinates and texture description.
[0029] As a preferred embodiment of the present invention, when generating a one-dimensional attention map, the feature fusion processing module first obtains a preset feature type, and the feature fusion processing module compares the feature type in the composite structure data. If the comparison similarity is greater than the set benchmark, it is determined to be a high-attention area. If the comparison similarity is less than the set benchmark, it is determined to be a non-attention area. The high-attention area and the non-attention area after the determination are completed are normalized to generate a one-dimensional attention map, wherein the one-dimensional attention map is a weight value sequence.
[0030] As a preferred embodiment of the present invention, the feature fusion processing module filters the values in the one-dimensional attention map, records the part with a weight less than the set value as low-credibility, low-importance data, and records the part with a weight greater than the set value as high-credibility data, high-importance data.
[0031] As a preferred embodiment of the present invention, when the decision fusion module is running, the mean aggregation method is as follows: extracting features from the initially acquired multimodal data respectively, and fusing all feature results extracted under the same modality to obtain a feature mean;
[0032] The jump fusion method is to fuse the features extracted from the composite structure data with the features extracted from the single modality image;
[0033] The decision fusion module fuses the features obtained by the jump fusion algorithm with the feature mean of mean aggregation to obtain the final feature fusion result.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] In the present invention, in the data preprocessing stage, an adaptive image block allocation algorithm is adopted to dynamically match geometric and texture information based on the spatial correspondence between three-dimensional point clouds and two-dimensional images. Irregularly shaped image blocks are assigned to each point cloud data through the nearest neighbor registration point principle, and a unified data structure containing spatial coordinates and texture information is constructed, which significantly enhances the color representation ability of the point cloud. In the feature fusion stage, an element attention module is designed, and a one-dimensional attention map is generated using an encoding-decoding structure. The importance and credibility of each feature element are dynamically quantified to achieve key information enhancement and interference suppression, effectively improving the model's robustness to occlusion and illumination changes. In the decision optimization stage, a jump fusion method is used to achieve interactive optimization of cross-modal intermediate features. Abstract features are extracted through mean aggregation and fully connected layer iterations, and jump connections are introduced to fuse original features with deep features, enhancing semantic expression capabilities while retaining detailed information. By constructing a three-layer fusion architecture at the data level, feature level, and decision level, and utilizing a multi-stage collaborative mechanism, this method achieves efficient fusion and complementarity between 3D point clouds and RGB images, significantly improving information utilization and semantic expression accuracy in complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] To facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings.
[0037] Figure 1 is a system block diagram of the present invention;
[0038] Figure 2 It is a system flow chart of the present invention. DETAILED DESCRIPTION
[0039] The following is a clear and complete description of the technical solutions of the present invention in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0040] Example 1: Please refer to Figure 1 - Figure 2 As shown, the multimodal data fusion processing method based on the multi-stage fusion strategy includes the following steps:
[0041] Step 1: Collect the original data, perform modal judgment on the collected data, and classify it into three-dimensional data and two-dimensional data;
[0042] Step 2: Extract geometric features from 3D data, extract texture information from 2D data, and dynamically match the geometric features and texture information to construct a composite data structure containing spatial geometric features and texture information;
[0043] Step 3: Introduce the attention module, generate a one-dimensional attention map through encoding and decoding the composite data structure, and perform credibility and importance assessment of feature elements based on the one-dimensional attention map;
[0044] Step 4: Fuse the cross-modal algorithms through the jump fusion algorithm, record the fusion results, extract the abstract features through the mean aggregation algorithm, and record the extracted features;
[0045] Step 5: Generate the final result of data fusion based on the features obtained in step 4;
[0046] Step 6: Perform reliability judgment on the final feature fusion result, and generate a result optimization signal or a result reliability signal based on the reliability judgment result;
[0047] Step 7: Record the key parameters in steps 2 to 4 based on the reliability judgment results and generate corresponding reminders.
[0048] Example 2: Please refer to Figure 1 - Figure 2 As shown in the figure, the multimodal data fusion processing method based on the multi-stage fusion strategy includes a multimodal data collection port, a data fusion processing module, a feature fusion processing module, a decision fusion module, a result judgment module and a dynamic recording module. The multimodal data collection port is used to acquire three-dimensional data and two-dimensional data, and judge the data type to obtain single-channel available data;
[0049] The data fusion processing module obtains single-channel available data, which includes 3D point cloud data and 2D RGB color image data. It performs algorithm adaptive matching based on the single-channel available data to obtain 3D geometric information and 2D texture information. The data fusion processing module obtains 3D geometric information by creating a rectangular coordinate system with a selected origin, and generates coordinates of the 3D point cloud data and constructs a model based on the set origin.
[0050] After the data fusion processing module obtains the two-dimensional RGB color image data, it performs algorithm filtering, denoising, grayscale conversion, and contrast adjustment on the two-dimensional RGB color image data in sequence to obtain a temporary image. The data fusion processing module then constructs a matrix of the frequency of occurrence of grayscale values at specific distances and directions in the temporary image, and calculates characteristic quantities such as energy, entropy, and contrast to obtain the roughness and regularity of the texture, which is recorded as two-dimensional texture information.
[0051] The data fusion processing module spatially matches the three-dimensional geometric information and the two-dimensional texture information to generate composite structure data. The specific method of spatially matching the three-dimensional geometric information and the two-dimensional texture information is as follows:
[0052] The point cloud model in the 3D geometric information is used as the basic source point cloud. For each point in the basic source point cloud, the corresponding point with the closest Euclidean distance is searched in the target point cloud to generate the initial matching point pair and achieve spatial coordinate alignment.
[0053] Through the point pair relationship after registration, each 3D point is projected onto a 2D plane using the differentiable Gaussian splatting technique to generate irregular image blocks.
[0054] The data fusion processing module adds two-dimensional texture features of corresponding positions on the irregular image blocks to generate composite structure data with both spatial rectangular coordinate system coordinates and texture description.
[0055] The feature fusion processing module obtains the composite structure data, encodes and decodes the composite structure data through the element attention module, and generates a one-dimensional attention map. When generating the one-dimensional attention map, the feature fusion processing module first obtains the preset feature type. The feature fusion processing module compares the feature type in the composite structure data. If the comparison similarity is greater than the set benchmark, it is determined to be a high-attention area. If the comparison similarity is less than the set benchmark, it is determined to be a non-attention area. The high-attention area and the non-attention area after the determination are completed are normalized to generate a one-dimensional attention map, where the one-dimensional attention map is a weight value sequence;
[0056] The feature fusion processing module filters the values in the one-dimensional attention map, records the parts with weights less than the set value as low-credibility and low-importance data, and records the parts with weights greater than the set value as high-credibility and high-importance data, completing the importance and credibility evaluation of feature points;
[0057] The decision fusion module uses the jump fusion algorithm to perform cross-modal feature interaction and optimize the features. It then uses the mean aggregation algorithm to extract abstract features and fuse the original features with the deep features.
[0058] When the decision fusion module is running, the mean aggregation method is as follows: the initially acquired multimodal data is subjected to feature extraction respectively, and all feature results extracted under the same modality are fused to obtain the feature mean;
[0059] The jump fusion method is to fuse the features extracted from the composite structure data with the features extracted from the single modality image, that is, to combine the original data features and the features after deep processing in layers;
[0060] The decision fusion module fuses the features obtained by the jump fusion algorithm with the feature mean of mean aggregation to obtain the final feature fusion result.
[0061] The result judgment module obtains the final feature fusion result through the decision fusion module, and imports the final feature fusion result through the manually preset evaluation model to judge the feature clarity and feature missing degree of the final feature fusion result, thereby generating the credibility of the feature fusion result. The credibility is directly proportional to the feature clarity and inversely proportional to the feature missing degree. The credibility judgment of the feature fusion result is completed, and the result reliability is generated. The result reliability is compared with the set database. If the result reliability is less than the set data standard, a result optimization signal is generated. If the result reliability is greater than or equal to the set data standard, a data reliability signal is generated.
[0062] The dynamic recording module obtains the key parameters of the data processing module, feature fusion processing module and decision fusion module during operation, and issues abnormal reminders for the key parameters when generating result optimization signals. The key parameters include the composite structure data generation process of the data fusion processing module, the element attention module operation parameters of the feature fusion processing module, and the jump fusion parameters and mean aggregation parameters of the decision fusion module.
[0063] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to specific embodiments. Obviously, many modifications and variations are possible based on the contents of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A multimodal data fusion processing method based on a multi-stage fusion strategy, characterized in that: The following steps are involved: Step 1: Collect the original data, perform modal judgment on the collected data, and classify it into three-dimensional data and two-dimensional data; Step 2: Extract geometric features from 3D data, extract texture information from 2D data, and dynamically match the geometric features and texture information to construct a composite data structure containing spatial geometric features and texture information; Step 3: Introduce the attention module, generate a one-dimensional attention map through encoding and decoding the composite data structure, and perform credibility and importance assessment of feature elements based on the one-dimensional attention map; Step 4: Fuse the cross-modal algorithms through the jump fusion algorithm, record the fusion results, extract the abstract features through the mean aggregation algorithm, and record the extracted features; Step 5: Generate the final result of data fusion based on the features obtained in step 4; Step 6: Perform reliability judgment on the final feature fusion result, and generate a result optimization signal or a result reliability signal based on the reliability judgment result; Step 7: Record the key parameters in steps 2 to 4 based on the reliability judgment results and generate corresponding reminders.
2. The multimodal data fusion processing method based on the multi-stage fusion strategy according to claim 1 is characterized in that: It also includes a multimodal data collection port, which is used to acquire three-dimensional data and two-dimensional data, and determine the data type to obtain single-channel available data; a data fusion processing module, wherein the data fusion processing module obtains single-channel available data, performs algorithm adaptive matching based on the single-channel available data, obtains three-dimensional geometric information and two-dimensional texture information, and spatially matches the three-dimensional geometric information and the two-dimensional texture information to generate composite structure data; A feature fusion processing module, which obtains composite structure data, encodes and decodes the composite structure data through the element attention module, generates a one-dimensional attention map, obtains feature points, and evaluates the importance and credibility of the feature points; The decision fusion module performs cross-modal feature interaction through a jump fusion algorithm to optimize the features, and then extracts abstract features through a mean aggregation algorithm to fuse the original features with the deep features.
3. The multimodal data fusion processing method based on the multi-stage fusion strategy according to claim 2 is characterized in that: The system also includes a result judgment module, which obtains the final feature fusion result through the decision fusion module, performs credibility judgment on the feature fusion result, generates result reliability, compares the result reliability with a set database, and generates a result reliability signal or a result optimization signal according to the comparison result; A dynamic recording module obtains key parameters during the operation of the data processing module, the feature fusion processing module and the decision fusion module, and issues abnormal reminders for key parameters when generating result optimization signals.
4. The multimodal data fusion processing method based on the multi-stage fusion strategy according to claim 2 is characterized in that: The single-channel available data acquired by the data fusion processing module includes three-dimensional point cloud data and two-dimensional RGB color image data; The data fusion processing module obtains the three-dimensional geometric information by creating a rectangular coordinate system with a selected origin, and performing coordinate generation and model construction of the 3D point cloud data through the set origin; After acquiring the two-dimensional RGB color image data, the data fusion processing module sequentially performs algorithm filtering, denoising, grayscale conversion, and contrast adjustment on the two-dimensional RGB color image data to obtain a temporary image. The data fusion processing module then constructs a matrix of the frequency of occurrence of grayscale values at specific distances and directions in the temporary image, and calculates characteristic quantities such as energy, entropy, and contrast to obtain the roughness and regularity of the texture, which are recorded as two-dimensional texture information.
5. The multimodal data fusion processing method based on the multi-stage fusion strategy according to claim 2 is characterized in that: The method for the data fusion processing module to spatially match the three-dimensional geometric information and the two-dimensional texture information is: The point cloud model in the 3D geometric information is used as the basic source point cloud. For each point in the basic source point cloud, the corresponding point with the closest Euclidean distance is searched in the target point cloud to generate the initial matching point pair and achieve spatial coordinate alignment. Through the point pair relationship after registration, each 3D point is projected onto a 2D plane through an algorithm to generate an irregular image block; The data fusion processing module adds two-dimensional texture features of corresponding positions on the irregular image blocks to generate composite structure data with both spatial rectangular coordinate system coordinates and texture description.
6. The multimodal data fusion processing method based on the multi-stage fusion strategy according to claim 2 is characterized in that: When generating a one-dimensional attention map, the feature fusion processing module first obtains a preset feature type, and compares the feature type in the composite structure data. If the comparison similarity is greater than a set benchmark, it is determined to be a high-attention area. If the comparison similarity is less than the set benchmark, it is determined to be a non-attention area. The high-attention area and the non-attention area after the determination are completed are normalized to generate a one-dimensional attention map, wherein the one-dimensional attention map is a weight value sequence.
7. The multimodal data fusion processing method based on the multi-stage fusion strategy according to claim 2 is characterized in that: The feature fusion processing module filters the values in the one-dimensional attention map, records the part with a weight less than the set value as low-credibility, low-importance data, and records the part with a weight greater than the set value as high-credibility data, high-importance data.
8. The multimodal data fusion processing method based on the multi-stage fusion strategy according to claim 2 is characterized in that: When the decision fusion module is running, the mean aggregation method is as follows: feature extraction is performed on the initially acquired multimodal data respectively, and all feature results extracted under the same modality are fused to obtain the feature mean; The jump fusion method is to fuse the features extracted from the composite structure data with the features extracted from the single modality image; The decision fusion module fuses the features obtained by the jump fusion algorithm with the feature mean of mean aggregation to obtain the final feature fusion result.
Citation Information
Patent Citations
Flotation froth image segmentation method and device based on multi-modal data fusion
CN116258719A
Unmanned aerial vehicle image scene semantic segmentation method based on depth prior
CN118314347A