Multi-modal multi-source heterogeneous data fusion method
By setting up an independent transformation model for multimodal data and performing dynamic weight allocation, the problem of insufficient semantic extraction capability of multimodal data in the prior art is solved, and efficient and accurate multimodal data fusion is achieved.
Patent Information
- Application Number
- CN202510462841.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-25
AI Technical Summary
The existing technology mainly relies on Word2Vec for field vectorization, and does not design special feature extraction models for multimodal data such as images and videos, resulting in limited semantic extraction capabilities of non-text data and difficult to achieve cross-modal semantic alignment.
A multimodal multi-source heterogeneous data fusion method is designed, by setting an independent transformation model for each modal, marking the same data multiplexed shared feature vector, modal-specific conversion and dimensionality reduction processing are performed on the differentiated data, and dynamic weight allocation is used to generate joint characterization.
It improves the efficiency and accuracy of multimodal data processing, reduces calculation amount and noise interference, and enhances the reliability and accuracy of the fusion results.
Smart Images

Figure CN120372540A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and specifically to a multi-modal multi-source heterogeneous data fusion method. Background Art
[0002] With the rapid development of technologies such as the Internet of Things and social media, the ways of generating data have become more diverse, including various modalities such as text, images, videos, and audio. Moreover, these data often come from different devices or platforms and have different formats and structures. There are usually rich associations and information among these data. However, due to different modalities and representation methods, unified processing and fusion are required. The knowledge fusion of multi-source heterogeneous multi-modal data is a complex and important research topic, which involves extracting and integrating information from data of different sources, formats, and modalities to construct a more comprehensive and accurate knowledge graph or intelligent system. In the real world, data often comes from different sources and modalities, including various forms such as text, images, and audio. Most existing fusion methods can only handle limited modality combinations and highly rely on manual processing, lacking effective fusion means for the increasingly complex multi-modal data fusion. Additionally, due to diverse data sources and uneven data quality, it is difficult to ensure consistency, resulting in the reliability of the fusion results being affected.
[0003] According to the disclosed patent 202411455575.2, a multi-modal multi-source heterogeneous data fusion method, belonging to the field of data processing technology, includes the following steps: S1, configure a data pipeline from the original database to the target database, and extract all source data from the original database to the target database through the data pipeline to obtain a source-attached data set; S2, perform structuring and cleaning operations on the source-attached data set to obtain a first data set; S3, dump the first data set to obtain a second data set; S4, perform data fusion processing on the second data set, including data mapping and conversion, and data standardization, to obtain a third data set; S5, for the third data set, perform feature extraction and processing for different types of data to obtain a final data set; S6, perform persistent storage on the final data set. The multi-modal multi-source heterogeneous data fusion method provided by the present invention can effectively solve the data quality problem in multi-modal multi-source heterogeneous data fusion and improve the consistency and reliability of the fusion results.
[0004] However, in the process of use, traditional multi-modal multi-source heterogeneous data fusion methods mainly rely on Word2Vec for field vectorization and do not design dedicated feature extraction models (such as CNN, Transformer) for multi-modal data (such as images, videos), resulting in limited semantic extraction ability for non-text data. For example, images are only processed through file links without extracting visual features (such as object detection, pixel details), making it difficult to achieve cross-modal semantic alignment. Therefore, new technical solutions need to be designed to solve this problem. Summary of the Invention
[0005] The object of the present invention is to overcome the deficiencies of the prior art, meet the actual needs, and provide a multi-modal multi-source heterogeneous data fusion method to solve the technical problems of current traditional multi-modal multi-source heterogeneous data fusion methods that mainly rely on Word2Vec for field vectorization and do not design dedicated feature extraction models (such as CNN, Transformer) for multi-modal data (such as images, videos), resulting in limited semantic extraction ability for non-text data. For example, images are only processed through file links without extracting visual features (such as object detection, pixel details), making it difficult to achieve cross-modal semantic alignment.
[0006] To achieve the object of the present invention, the technical solution adopted by the present invention is as follows: Design a multi-modal multi-source heterogeneous data fusion method, including the following steps:
[0007] S1. Set independent conversion models for each modal data, and convert text, image, video, and audio modal data into intermediate feature vectors respectively;
[0008] S2. Mark the same data and different data in different modalities, where the same data is cross-modal shared semantic information, and the different data is modality-specific information;
[0009] S3. Reuse the shared feature vectors for the same data, and perform modality-specific conversion and dimensionality reduction processing on the different data;
[0010] S4. Input the multi-modal feature vectors with unified dimensions into the fusion model to generate a joint representation through dynamic weight allocation;
[0011] S5. Execute the target multi-modal task based on the joint representation, including cross-modal retrieval or multi-modal generation.
[0012] Preferably, in step S1, the BERT or GPT model is used to extract semantic vectors for the text modality, the ResNet or VGG is used to extract features for the image modality, the 3D CNN is used to process the key frame sequence for the video modality, and the Mel spectrogram or Wav2Vec model is used to convert acoustic features for the audio modality.
[0013] Preferably, in step S2, the cross-modal alignment model CLIP is used for semantic matching to mark the same data, and the different data is distinguished by calculating the feature similarity, and the similarity threshold is set to 0.7 - 0.9.
[0014] Preferably, in step S3, the principal component analysis (PCA) is used to reduce the dimension of the different data to 30% - 50% of the original feature dimension, and it is mapped to the unified dimension range of 512 - 1024 dimensions through a fully connected layer.
[0015] Preferably, in the step S4, the fusion model dynamically allocates weights for each modality by using an attention mechanism, where the weight ratio of the same data is 60%-80%, and the weight ratio of the different data is 20%-40%.
[0016] Preferably, the attention mechanism includes a cross-modal Transformer model, and the number of multi-head attention layers is 4-8 layers, and the hidden layer dimension is 512-768.
[0017] Preferably, the step S3 further includes caching the feature vectors of the same data, and directly calling them when the cache hits, reducing the number of repeated calculations.
[0018] Preferably, in the step S4, when the number of input modalities exceeds 3, a graph neural network (GNN) is used to construct a modality node relationship graph, and the fusion is realized through node aggregation and update.
[0019] Preferably, a lightweight classifier is connected to the output end of the fusion model, and the classifier is a 2-3 layer fully connected network, and the activation function uses GELU or Swish.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0021] 1. By setting independent conversion models for each modality data (such as text, image, video, audio), the present invention can process these data in parallel. Different modality data can be processed by their respective models simultaneously without waiting for the processing of other modality data to complete. This parallel processing ability can shorten the overall processing time. At the same time, each modality data has its unique characteristics and processing requirements. By setting independent models for each modality, these models can be more specialized and optimized for specific modality data, thereby improving processing efficiency and accuracy. Moreover, the independent models can be optimized and trained separately without considering the influence of other modality data, which helps to improve the performance of each model and reduce the complexity and errors that may be introduced by the integration of multi-modal data.
[0022] 2. By marking the same data, the present invention can identify the semantic information shared across modalities and directly reuse these shared feature vectors. In subsequent processing steps, these data do not need to be feature-extracted or converted again, thereby improving the overall processing efficiency, avoiding repeated calculation of features for the same data in different modalities, reducing the amount of calculation. At the same time, performing modality-specific conversion on different data can perform targeted processing according to the characteristics of different modality data, thereby better extracting modality-specific information. Through modality-specific conversion, information irrelevant or redundant to the current modality can be removed, reducing noise interference and improving data quality.
[0023] 3. By dynamically allocating the weights of each modality, the present invention can adaptively adjust the fusion strategy according to the data characteristics, thereby improving the accuracy and reliability of the fusion result. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a schematic flowchart of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0025] The present invention will be further described below in conjunction with the drawings and embodiments:
[0026] A multi-modal multi-source heterogeneous data fusion method, see Figure 1 , including the following steps:
[0027] S1. Set up independent conversion models for each modality data, and convert text, image, video and audio modality data into intermediate feature vectors respectively;
[0028] S2. Mark the same data and different data in different modalities. The same data is cross-modal shared semantic information, and the different data is modality-specific information;
[0029] S3. Reuse the shared feature vectors for the same data, and perform modality-specific conversion and dimensionality reduction processing on the different data;
[0030] S4. Input the multi-modal feature vectors with unified dimensions into the fusion model, and generate a joint representation through dynamic weight allocation;
[0031] S5. Perform the target multi-modal task based on the joint representation, including cross-modal retrieval or multi-modal generation.
[0032] Specifically, see Figure 1 , in step S1, the BERT or GPT model is used to extract semantic vectors for the text modality, the ResNet or VGG is used to extract features for the image modality, the 3D CNN is used to process the key frame sequence for the video modality, and the Mel spectrogram or Wav2Vec model is used to convert acoustic features for the audio modality.
[0033] More specifically, see Figure 1 , in step S2, the cross-modal alignment model CLIP is used to perform semantic matching to mark the same data, and the different data is distinguished by calculating the feature similarity, and the similarity threshold is set to 0.7-0.9.
[0034] Further, see Figure 1 , in step S3, the principal component analysis (PCA) is used to reduce the dimension of the different data to 30%-50% of the original feature dimension, and it is mapped to the unified dimension range of 512-1024 dimensions through the fully connected layer.
[0035] Even further, see Figure 1, in step S4, the fusion model uses an attention mechanism to dynamically allocate weights for each modality, where the weight ratio of the same data is 60%-80%, and the weight ratio of different data is 20%-40%.
[0036] It should be noted that, referring to Figure 1 , the attention mechanism includes a cross-modal Transformer model, whose number of multi-head attention layers is 4-8 layers, and the hidden layer dimension is 512-768.
[0037] It should be noted that, referring to Figure 1 , step S3 also includes caching the feature vectors of the same data and directly calling them when the cache hits, reducing the number of repeated calculations.
[0038] It should be introduced that, referring to Figure 1 , in step S4, when the number of input modalities exceeds 3, a graph neural network (GNN) is used to construct a modality node relationship graph, and the fusion is achieved through node aggregation and update.
[0039] It should be noted that, referring to Figure 1 , it also includes connecting a lightweight classifier at the output end of the fusion model. The classifier is a 2-3 layer fully connected network, and the activation function uses GELU or Swish.
[0040] Example 1
[0041] Multi-modal Q&A system
[0042] Scenario: Question-answering task based on images and texts (such as "the number of cats in the picture")
[0043] Steps:
[0044] S1. Feature extraction:
[0045] Text: Use BERT to extract the entity vector in the question (such as "cat").
[0046] Image: Use ResNet-50 to extract object detection features (detect the position and number of "cats").
[0047] S2. Data marking: Align the text "cat" with the cat area in the image through CLIP (similarity threshold 0.85).
[0048] S3. Fusion processing:
[0049] The same data ("cat" semantics) is directly reused.
[0050] The image details (pixels, background) are reduced to 40% of the original dimension through PCA (512→205 dimensions).
[0051] S4. Dynamic Fusion: The cross-modal Transformer (6-layer attention) assigns weights (70% for the same data and 30% for the different data).
[0052] S5. Output: The lightweight classifier (2-layer fully connected + GELU) generates answers with a confidence threshold of 0.9.
[0053] Effect: The answer accuracy is increased by 12% (compared with the traditional Word2Vec method), and the inference speed is increased by 30%.
[0054] Example 2
[0055] Cross-modal Retrieval
[0056] Scenario: Retrieve relevant images according to text descriptions (such as "beach sunset")
[0057] Steps:
[0058] S1. Feature Extraction:
[0059] Text: GPT-3 generates semantic vectors for "beach sunset".
[0060] Image: VGG-16 extracts global features and local details (such as sky tone, wave texture).
[0061] S2. Data Alignment: CLIP calculates the similarity between text and image (threshold 0.8) and marks the shared semantics ("beach", "sunset").
[0062] S3. Dimensionality Reduction and Optimization: The image background details are compressed to 35% of the original dimension (4096 → 1434 dimensions) through PCA.
[0063] S4. Fusion Model: The attention mechanism (8-layer multi-head) dynamically weights (75% for the same data and 25% for the different data).
[0064] S5. Retrieval Output: GNN constructs an image-text relationship graph and returns the top-5 relevant images.
[0065] Effect: The retrieval accuracy (mAP) reaches 92%, which is 18% higher than the traditional method.
[0066] Example 3
[0067] Video Sentiment Analysis
[0068] Scenario: Analyze the sentiment tendency of movie clips (such as "sad" or "happy")
[0069] Steps:
[0070] S1. Feature Extraction:
[0071] Video: 3D CNN processes the key frame sequence to extract facial expression and action features.
[0072] Audio: Wav2Vec extracts speech emotion features (such as intonation, rhythm).
[0073] S2. Data labeling: Align the crying scenes in the video with the low tones in the audio (similarity threshold 0.75).
[0074] S3. Dimensionality reduction processing: Reduce the video action details to 30% (1024→307 dimensions), and compress the audio background noise to 20%.
[0075] S4. Dynamic fusion: Transformer (768-dimensional hidden layer) assigns weights (60% for video, 40% for audio).
[0076] S5. Classification output: A lightweight classifier (3-layer fully connected + Swish) outputs emotion labels.
[0077] Effect: The F1-score of emotion classification reaches 89%, and the calculation time is reduced by 40%.
[0078] Example 4
[0079] Multimodal advertisement generation
[0080] Scenario: Generate advertisement copy based on product images and descriptions
[0081] Steps:
[0082] S1. Feature extraction:
[0083] Image: ResNet-101 extracts product features (such as color, shape).
[0084] Text: BERT encodes product descriptions (such as "high-end", "portable").
[0085] S2. Semantic alignment: CLIP matches image features with text keywords (similarity threshold 0.7).
[0086] S3. Dimensionality reduction and optimization: Reduce the complex texture of the image to 45% (2048→922 dimensions).
[0087] S4. Fusion and generation: A cross-modal Transformer (4-layer attention) fuses features to generate coherent copy. S5. Post-processing: Filter out low-quality generation results with a confidence threshold of 0.88.
[0088] Effect: The correlation score between the generated copy and the image is increased by 25%, and the satisfaction rate of manual evaluation reaches 85%.
[0089] Example 5
[0090] Medical Multimodal Diagnosis
[0091] Scenario: Combining CT images and pathological texts to assist in disease diagnosis
[0092] Steps:
[0093] S1. Feature extraction:
[0094] Image: 3D ResNet extracts CT lesion features (such as nodule size and density).
[0095] Text: BioBERT parses key indicators in the pathological report (such as "malignant" and "differentiation degree").
[0096] S2. Data alignment: CLIP aligns image features with text terms (similarity threshold 0.9).
[0097] S3. Dimensionality reduction processing: The details of the CT image are reduced to 50% (4096 → 2048 dimensions).
[0098] S4. Dynamic fusion: GNN constructs an image-text association graph, and nodes aggregate and update features.
[0099] S5. Diagnostic output: A lightweight classifier outputs the disease probability, and the confidence threshold is 0.95.
[0100] Effect: The diagnostic accuracy rate is increased by 15%, and the false positive rate is reduced by 8%.
[0101] Comparative Example 1
[0102] Traditional field mapping method (based on Word2Vec)
[0103] Scenario: Cross-modal retrieval
[0104] Steps:
[0105] S1. Feature extraction:
[0106] Text: Word2Vec generates word vectors.
[0107] Image: Only stores the file link and does not extract visual features.
[0108] S2. Field mapping: Matches image labels based on text similarity (cosine distance).
[0109] S3. Fusion processing: Simply concatenates the text and image labels without dynamic weight assignment.
[0110] Defects:
[0111] Image semantics relies on manual labels, missing key features (such as the "sunset hue" in "beach sunset"). The retrieval accuracy (mAP) is only 74%, 18% lower than that of Example 2.
[0112] Comparative Example 2
[0113] Static fusion strategy (without attention mechanism)
[0114] Scenario: Video emotion analysis (video + audio)
[0115] Steps:
[0116] S1. Feature extraction:
[0117] Video: 3D CNN extracts action features.
[0118] Audio: MFCC extracts acoustic features.
[0119] S2. Fusion processing: Concatenate features with fixed weights (50% for video and 50% for audio).
[0120] S3. Classification output: Direct classification by a fully connected network.
[0121] Defects:
[0122] It is impossible to dynamically adjust the modal contribution (e.g., the audio weight should be higher in a sad scene).
[0123] The F1-score of emotion classification is only 72%, 17% lower than that of Example 3.
[0124] The specific comparison data is as follows:
[0125]
[0126]
[0127]
[0128] In summary, the embodiments improve the accuracy and efficiency of multi-modal tasks through modal-specific models, dynamic fusion strategies, and dimensionality reduction optimization. The comparative examples expose the deficiencies of traditional methods in feature extraction, fusion flexibility, etc., further verifying the technical advantages of the present invention.
[0129] In addition, the components designed in the present invention are all common standard components or components known to those skilled in the art. Their structures and principles can all be learned by those skilled in the art through technical manuals or through conventional experimental methods. Those skilled in the art can fully implement them without further elaboration. The content protected by the present invention does not involve improvements to internal structures and methods.
[0130] The embodiments disclosed in the present invention are preferred embodiments, but not limited thereto. Those of ordinary skill in the art can easily understand the spirit of the present invention based on the above embodiments and make different extensions and changes. As long as they do not depart from the spirit of the present invention, they are within the protection scope of the present invention.
Claims
1. A multi-modal multi-source heterogeneous data fusion method, characterized in that, It includes the following steps: S1. Set up independent conversion models for each modality data, and convert text, image, video, and audio modality data into intermediate feature vectors respectively; S2. Mark the same data and different data in different modalities. The same data is cross-modal shared semantic information, and the different data is modality-specific information; S3. Reuse the shared feature vectors for the same data, and perform modality-specific conversion and dimensionality reduction processing on the different data; S4. Input the multi-modal feature vectors with unified dimensions into the fusion model, and generate a joint representation through dynamic weight allocation; S5. Execute the target multi-modal task based on the joint representation, including cross-modal retrieval or multi-modal generation.
2. The multimodal multi-source heterogeneous data fusion method according to claim 1, characterized in that, In step S1, the BERT or GPT model is used to extract semantic vectors for the text modality, the ResNet or VGG is used to extract features for the image modality, the 3D CNN is used to process the key frame sequence for the video modality, and the Mel spectrogram or Wav2Vec model is used to convert acoustic features for the audio modality.
3. The multimodal multi-source heterogeneous data fusion method according to claim 1, characterized in that In step S2, the cross-modal alignment model CLIP is used to perform semantic matching to mark the same data, and the different data is distinguished by calculating the feature similarity. The similarity threshold is set to 0.7 - 0.
9.
4. The multimodal multi-source heterogeneous data fusion method according to claim 1, wherein In step S3, the principal component analysis (PCA) is used to reduce the dimensionality of the different data to 30% - 50% of the original feature dimension, and it is mapped to the unified dimension range of 512 - 1024 dimensions through a fully connected layer.
5. The multimodal multi-source heterogeneous data fusion method according to claim 1, characterized in that In step S4, the fusion model uses the attention mechanism to dynamically allocate weights for each modality. Among them, the weight ratio of the same data is 60% - 80%, and the weight ratio of the different data is 20% - 40%.
6. The multimodal multi-source heterogeneous data fusion method according to claim 5, wherein The attention mechanism includes a cross-modal Transformer model, and the number of multi-head attention layers is 4 - 8 layers, and the hidden layer dimension is 512 - 768.
7. The multimodal multi-source heterogeneous data fusion method according to claim 1, wherein Step S3 also includes caching the feature vectors of the same data, and directly calling them when the cache hits, reducing the number of repeated calculations.
8. The multimodal multi-source heterogeneous data fusion method according to claim 1, wherein In step S4, when the number of input modalities exceeds 3, a graph neural network (GNN) is used to construct a modality node relationship graph, and the fusion is achieved through node aggregation and update.
9. The multimodal multi-source heterogeneous data fusion method according to claim 1, wherein It also includes connecting a lightweight classifier at the output end of the fusion model. The classifier is a 2 - 3 layer fully connected network, and the activation function uses GELU or Swish.
Citation Information
Patent Citations
A multi-modal, multi-source, heterogeneous data fusion method
CN118981453B
Cited By
Multi-source heterogeneous data fusion method and system based on edge calculation
CN120541795A
Multi-source heterogeneous data fusion method and system based on edge computing
CN120541795B
Feature fusion and engineering processing method of multi-modal data
CN122490448A
Feature fusion and engineering processing method of multi-modal data
CN122490448B