Massive multi-source and multi-modal data fusion method
By building a multimodal fusion model, the ViT, BERT and HuBERT models of the Transformer architecture are used to extract features, and combined with the deep cross attention mechanism and adaptive fusion layer, the problems of shallow fusion levels, high computational complexity and information loss in massive multi-source multimodal data fusion are solved, efficient and intelligent data fusion is achieved, and data accuracy and system robustness are improved.
Patent Information
- Application Number
- CN202510759252.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When the prior art deals with the fusion of massive multi-source multi-modal data, there are problems such as shallow fusion levels, high computational complexity, information loss and redundancy, and graph structure construction relying on manual rules. It is difficult to effectively mine and utilize the deep semantic correlation and complementarity of multi-modal data, and lack scalability and flexibility.
A multimodal fusion model is constructed, and the image, text and audio features are extracted using the ViT, BERT and HuBERT models of the Transformer architecture are respectively used to extract images, text and audio features, combined with the deep cross attention mechanism and adaptive fusion layer, and optimize the model structure through multimodal data augmentation and joint training to achieve efficient and intelligent data fusion.
It improves the effect and application value of data fusion, enhances the ability to understand cross-modal semantics, reduces the computational complexity, ensures the accuracy and completeness of the data, improves the robustness and adaptability of the system, and adapts to different application scenarios and data characteristics.
Smart Images

Figure CN120277619A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing and fusion, and specifically relates to a method for fusion of massive multi-source and multi-modal data. Background Art
[0002] With the rapid development of information technology, massive multi-source and multi-modal data has become the norm in modern society. These data are widely present in many fields such as the Internet, the Internet of Things, and social media, covering various forms such as text, images, audio, and video. How to effectively integrate these data from different sources and different modalities to explore their potential value is an important research topic in the current field of big data processing and artificial intelligence.
[0003] The massive multi-source multi-modal data fusion technology aims to integrate data of multiple modalities into a unified framework, and improve the accuracy and efficiency of data analysis through cross-modal information interaction and complementarity; as the core link of information processing and knowledge discovery, this technology has shown important application value in many fields such as intelligent analysis, Internet of Things, and social media monitoring. However, although a variety of fusion technologies have been applied to this field, most of them are based on traditional machine learning methods or statistical models. These methods face many challenges when processing massive, multi-source, and multi-modal data, such as the heterogeneity and inconsistency between different modal data, which exacerbates the difficulty of fusion, the huge differences caused by the exponential growth of data volume, and the loss of information during the fusion process.
[0004] In recent years, the rise of deep learning technology has provided new ideas for the fusion of massive multi-source and multi-modal data, especially the successful application of Transformer models and their variants, such as BERT (Bidirectional Encoder Based on Transformer) and GPT (Generative Pre-trained Transformer) in the field of natural language processing, and the breakthrough of ViT (Vision Transformer) in the field of image recognition, which has made cross-modal deep fusion possible. However, most of these technologies focus on multi-modal fusion of a single modality or a specific field, and are still insufficient for the fusion of massive, multi-source, and heterogeneous multi-modal data, which limits the comprehensiveness and depth of data fusion.
[0005] Traditional feature fusion methods extract features of each modality through models such as CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), or Transformer, and perform simple concatenation or weighted fusion; however, such methods ignore the deep semantic associations between modalities and have a shallow fusion level. The attention mechanism method dynamically adjusts the modality weights by introducing the attention mechanism, such as text-image alignment in VQA (Visual Question Answering), but in the face of massive data, the computational complexity of this method increases, and the fusion effect is difficult to fully guarantee. The graph neural network method regards multi-modal data as graph nodes and represents the associations between modalities through edges, but the construction of the graph structure depends on artificial rules and is difficult to flexibly handle the complexity of massive multi-source data.
[0006] In summary, when dealing with the fusion of massive, diverse, and multi-modal data, existing technologies mainly focus on feature extraction, model construction, and fusion strategies. These technologies mainly include feature-level fusion, decision-level fusion, model-level fusion, graph neural network methods, etc., and usually aim to extract useful information from data of different modalities and fuse this information in some way to support subsequent analysis and decision-making. However, these technologies also have some significant drawbacks as follows: Shallow fusion level: Most of the existing feature fusion methods stay at the simple concatenation or weighting at the feature level, ignoring the deep semantic associations and complementarities between different modalities. This shallow fusion method is difficult to fully explore and utilize the rich information in multi-modal data, limiting the effect and application value of data fusion.
[0007] High computational complexity: Although the attention mechanism-based method improves the information selectivity and pertinence in the fusion process to a certain extent, in the face of massive data, its computational complexity increases significantly. Whether it is feature extraction, model training, or the fusion process, a large amount of computational resources are required, resulting in low processing efficiency; this not only increases the consumption of computational resources but also affects the real-time performance and response speed of data processing.
[0008] Information loss and redundancy: In the feature extraction and fusion process, some important information may be lost, affecting the comprehensiveness of subsequent fusion; at the same time, there may be redundant information between different modalities, which will increase the processing complexity and interference, affect the fusion effect, and reduce the performance and accuracy of the model.
[0009] Graph structure construction depends on humans: The graph neural network method represents the relationships between multi-modal data by constructing a graph structure, but this construction process often depends on artificially designed rules or heuristic algorithms. This method lacks adaptability and is difficult to flexibly handle the complexity and diversity of massive multi-source data. At the same time, it also increases the difficulty and cost of model construction.
[0010] Lack of scalability and flexibility: When dealing with different application scenarios and data characteristics, existing technologies often lack sufficient scalability and flexibility. As the data volume continues to grow and application scenarios continue to change, existing technologies may not be able to effectively adapt to these changes, resulting in a decline in the fusion effect or an inability to meet actual requirements.
[0011] Therefore, when exploring the profound field of massive multi-source multi-modal data fusion technology and facing a complex and ever-changing data processing environment, it is particularly important to propose a comprehensive technical solution. Summary of the Invention
[0012] To solve the technical problems in massive multi-source multi-modal data fusion, such as shallow fusion level, high computational complexity, complex model design, information loss and redundancy, and insufficient cross-modal interaction, the present invention provides a method for massive multi-source multi-modal data fusion. By constructing an efficient, intelligent, and adaptive data fusion framework, innovative fusion strategies, and optimized model structures, more efficient, intelligent, and comprehensive data fusion is achieved.
[0013] A method for massive multi-source multi-modal data fusion includes the following steps: (1) Obtain a massive dataset with multiple sources and rich multi-modal information (such as text, images, audio, etc.), and preprocess the data therein; (2) Construct a multi-modal fusion model, including: A visual processing module that extracts visual features from image data based on the ViT model; A text processing module that extracts semantic features from text data based on the BERT model; A sound processing module that extracts sound features from audio data based on the pre-trained HuBERT (Hidden-unit BERT, BERT through hidden unit mask prediction) model; A multi-modal fusion module that fuses the extracted visual features, semantic features, and sound features under the Joint Architecture framework; (3) Perform data augmentation on the preprocessed massive dataset, and use this dataset to train the multi-modal fusion model; (4) Use the fusion features generated by the trained multi-modal fusion model to complete corresponding application scenario tasks.
[0014] Further, the data preprocessing in step (1) includes data integration, data cleaning, and standardization. For text data, data cleaning includes processing such as missing value filling, outlier correction, stop word removal, and duplicate item removal; for audio data, data cleaning includes denoising processing; for image data, data cleaning includes processing such as denoising, image cropping, scaling, and rotation; for image data and audio data, standardization uses min-max normalization processing; for text data, standardization uses the word embedding method to convert the vocabulary in the text into a vector representation of a fixed dimension. This step aims to uniformly convert data from different sources and in different formats into a form suitable for subsequent processing, ensuring the consistency and comparability of the data.
[0015] Further, the visual processing module first adjusts the input image to the fixed size required by the ViT model, and then divides the adjusted image into a series of small blocks of a fixed size. Each small block is flattened into a vector as a token in the sequence. These tokens not only contain the color and texture information of the image but also retain the information of the spatial relationship through their positions in the sequence; furthermore, a special classification token is added at the beginning of the sequence, and position encoding or position embedding is added at the end of the sequence; the processed token sequence is input into a Transformer encoder cascaded by multiple encoders. Each encoder layer contains a self-attention mechanism and a feed-forward neural network. When processing each token, the self-attention mechanism can consider the information of all other tokens in the sequence, thereby capturing the complex spatial dependence relationships in the image; the feed-forward neural network further performs a non-linear transformation on the output of the self-attention mechanism to enhance the expressive ability of the ViT model; the output after being processed by the Transformer encoder is a high-dimensional feature vector, which encodes the key visual information of the input image.
[0016] Further, the text processing module first tokenizes the text using the WordPiece algorithm to obtain a token sequence, then adds special markers [CLS] and [SEP] at the beginning and end of the token sequence respectively, and assigns a position embedding to each token in the sequence; the processed token sequence is input into a pre-trained BERT model. The BERT model processes the input token sequence through a self-attention mechanism and multiple layers of Transformer encoders, captures the complex relationships between tokens, and extracts deep semantic features; the self-attention mechanism allows the model to consider other tokens in the entire sequence when processing each token, thereby capturing the global context information of the text; the BERT model finally outputs a high-dimensional feature vector, which encodes the semantic and syntactic information of the input text.
[0017] Further, the sound processing module first initializes and configures the pre-training parameters of the HuBERT model (including the number of layers of the Transformer encoder, embedding dimension, attention mechanism, etc.), then uses the k-means clustering method to cluster audio frames, assigns a pseudo-label to each audio frame, and uses the HuBERT model trained in the previous stage to predict more refined pseudo-labels for the model training in the next stage; the audio frames and their pseudo-labels are used as inputs to train the HuBERT model to learn the intrinsic representation of audio data through a masked prediction task, where the masked prediction task is to randomly mask some audio frames and train the model to predict the pseudo-labels of these masked frames; finally, new audio data is input into the HuBERT model completed by the above pre-training to output a high-dimensional feature vector, which encodes the deep features of the input audio (including not only basic acoustic attributes such as pitch, rhythm, timbre, and duration, but also may include higher-level semantic information such as the emotional color of speech and the characteristics of environmental sounds).
[0018] In the process of multi-modal data fusion, the data sources are mainly images, texts, sounds, etc. Each modality carries unique and rich information. To effectively integrate this information and improve the performance of the model, it is necessary to efficiently extract key features from each modality. In view of the complexity of multi-modal data, the present invention adopts a modular design principle to construct the overall model. Each of the above modules is independently designed and optimized for its specific task. This design not only improves the scalability and maintainability of the model, but also facilitates customized development for the data characteristics of specific modalities; through the Transformer model customized for different modalities of data, the present invention can more efficiently and accurately extract key features from data such as images, sounds, and texts, laying a solid foundation for subsequent data fusion and joint modeling.
[0019] To further reduce the computational complexity and improve the model performance, the multi-modal fusion model also introduces a feature screening technology, and uses PCA (Principal Component Analysis), LDA (Linear Discriminant Analysis) or a model-based feature importance evaluation method to screen the extracted features, so as to screen out the most representative feature vectors, eliminate redundant information, and retain key features.
[0020] To establish a close connection between multimodal data, the multimodal fusion module feeds visual features, semantic features, and audio features into the adaptive fusion layer for fusion after appropriate alignment and transformation. The adaptive fusion layer uses a deep cross-attention mechanism to calculate the cross-attention results between the visual, semantic, and audio features, and dynamically adjusts the weights of the attention mechanism to allocate the contribution degrees of different modal information. Finally, the cross-attention results between the modal features are fused, and the fusion strategy is automatically adjusted according to different application scenarios and task requirements, such as feature-level fusion, decision-level fusion, model-level fusion, weighted fusion, splicing fusion, or graph neural network-based fusion. This mechanism enhances the model's ability to capture key information, reduces information loss and redundancy, continuously optimizes the attention distribution during the training process, realizes more accurate and efficient multimodal fusion, and its adaptive ability ensures the excellent performance of the model in various situations and promotes the effective integration of different modal information.
[0021] Furthermore, before performing step (3), transfer learning technology is used to pre-train the multimodal fusion model, and the pre-trained model parameters are used as initialization parameters to accelerate the training process of the model in step (3). For the visual processing module, the ViT model of this module is trained on ImageNet (a large-scale image database). For the text processing module, the BERT model of this module is trained on a large-scale text corpus. For the audio processing module, the HuBERT model of this module is trained on LibriSpeech (a free speech corpus). At the same time, a parameter sharing mechanism is adopted in the multimodal fusion model, allowing different modal network layers to share parameters to a certain extent, promoting cross-modal information exchange and complementarity. This design enhances the generalization ability of the model, reduces the dependence on the amount of data, and helps to discover potential associations between different modalities, providing a strong starting point for the multimodal data fusion task.
[0022] To improve the generalization ability and robustness of the model, a multimodal data augmentation strategy is implemented. In step (3), for image data, data augmentation applications include image transformation operations such as image rotation, cropping, and color transformation to generate diverse image samples. For text data, data augmentation uses methods such as synonym replacement, back-translation, and sentence recombination to enrich text expressions. For audio data, data augmentation uses methods such as noise addition, audio speed change, volume adjustment, audio cropping, and splicing to improve the model's adaptability to complex sound environments. In addition, data augmentation operations also generate new text-image sample pairs, text-audio sample pairs, and audio-image sample pairs in combination with multimodal characteristics, ensuring the semantic consistency of these sample pairs while increasing data diversity. This multimodal data augmentation method not only enhances the model's adaptability to complex data distributions but also promotes the interaction and fusion between different modalities.
[0023] To further address the problem of insufficient data and improve the model performance, in step (3), when performing data augmentation on the massive dataset, a generative adversarial network (GAN) and conditional data augmentation techniques are also introduced. The generative adversarial network generates high-quality fake data through adversarial training between a generator and a discriminator. The generated fake data is close to the real data in terms of feature distribution. Using the fake data generated by GAN to further expand the dataset can effectively improve the generalization ability and robustness of the model. During the training process, the parameters and structure of GAN are continuously adjusted to optimize the quality of the generated data, further promoting the development of multi-modal data fusion technology. The conditional data augmentation technique uses the data of one modality to guide the data augmentation process of another modality. For example, in the case of corresponding text and images, the image can be enhanced specifically according to the text content (such as adding objects related to the text description, changing the background while maintaining consistency with the text, etc.). Conditional data augmentation can ensure that the relevance between different modality data is not destroyed, while increasing the diversity and complexity of the data, which helps to improve the generalization ability of the model.
[0024] Furthermore, in step (3), the training process of the multi-modal fusion model adopts the methods of joint training and knowledge distillation: First, for a specific task, a multi-modal joint training framework is designed, which can process data from different modalities simultaneously. Then, one or more single-modal models that have been trained on a large dataset are used as teacher models, and the knowledge of the teacher models is transferred to the multi-modal joint training framework through knowledge distillation. This method can overcome the domain adaptation problems that may be encountered when directly using pre-trained models for fine-tuning. Joint training can make full use of the complementarity between multi-modal data, while knowledge distillation can retain the powerful capabilities of the teacher models and reduce the risk of overfitting to new data.
[0025] The present invention can deeply explore the deep semantic associations and complementarities between different modality data, improve the effect and application value of data fusion, optimize the feature extraction algorithm to reduce information loss, and use dimensionality reduction and feature selection techniques to remove redundant information. The fusion strategy design of the present invention aims to make full use of complementarity, improve the fusion effect and model performance. By optimizing the algorithm and model design, the computational complexity is reduced, and the efficiency and real-time performance of processing massive data are improved. In addition, the present invention realizes the self-adaptability of graph structure construction to enhance the model's processing ability for complex and changeable data and the scalability and flexibility of the technology, so as to better adapt to the changes in different application scenarios and data characteristics, thereby ensuring the continuous optimization of the data fusion effect and meeting the actual requirements.
[0026] In addition, the present invention can utilize the complementary characteristics of different data modalities to improve the perception and judgment capabilities of the system. Especially in the fields of autonomous driving, intelligent healthcare, sentiment analysis, human-computer interaction, etc., it can demonstrate great potential. The present invention aims to achieve efficient, intelligent, and comprehensive multi-modal data fusion through refined dataset processing, innovative model architecture design, flexible multi-modal fusion strategies, and efficient data augmentation methods to address the current technical bottlenecks.
[0027] Therefore, the innovation and beneficial technical effects of the technical solution of the present invention are mainly reflected in the following aspects: 1. Data processing efficiency and cost.
[0028] Through the multi-modal data fusion technology, the present invention can simultaneously process information from different sensors or data sources, significantly improving the parallelism and efficiency of data processing. Compared with the prior art, the present invention can integrate and analyze massive multi-source multi-modal data faster, reducing the data processing time. In addition, the present invention can reduce the data processing cost. The multi-modal data fusion technology reduces the excessive dependence on a single data source or single-modal data by optimizing the data processing process, thereby reducing the costs of data acquisition, storage, and processing. Moreover, by improving the automation degree of data processing, the present invention also reduces the human input.
[0029] 2. Data accuracy and integrity.
[0030] The multi-modal data fusion technology of the present invention can utilize the complementarity and redundancy between different modal data, and perform fusion processing on the data through various data models (such as deep learning models), effectively improving the data accuracy. Compared with the prior art, the present invention can more accurately reflect the actual situation, reducing misjudgment and missed judgment, improving data accuracy, and at the same time ensuring data integrity. During the data fusion process, the present invention can ensure the integrity and consistency of different modal data, avoiding information loss caused by data loss or damage, which is helpful for subsequent data analysis and decision-making.
[0031] 3. Cross-modal semantic understanding and cognitive reasoning.
[0032] Enhanced cross-modal semantic understanding ability; compared with the prior art, the present invention pays more attention to the research on cross-modal semantic understanding and cognitive reasoning. Through advanced technologies such as deep learning, the present invention can better understand the semantic relationships between different modal data, realizing a higher level of information fusion and reasoning.
[0033] Improving system robustness: The multi-modal data fusion technology of the present invention can utilize the redundancy between multiple modal data to improve the fault tolerance and robustness of the system. Even if the data of a certain modality has problems or is missing, the system can still make accurate judgments and decisions using the data of other modalities.
[0034] 4. Application scenario expansion.
[0035] The technology of massive multi-source multi-modal data fusion has broad application prospects. The multi-modal data fusion technology has broad application prospects in multiple fields, such as intelligent transportation, smart home, medical diagnosis, etc. Compared with the existing technology, the present invention can better meet the needs of different fields and provide more comprehensive and accurate information support. With the continuous development of artificial intelligence technology, the multi-modal data fusion technology will continue to progress and improve. As an innovative achievement in this field, the present invention is expected to promote the further development and application expansion of related technologies and drive technological innovation.
[0036] 5. Privacy protection and security.
[0037] In the process of data processing and fusion, the present invention pays attention to privacy protection and data security. By adopting advanced encryption technologies and data isolation measures, the security and privacy of user data are ensured; compared with the existing technology, the present invention has more advantages in privacy protection and enhanced privacy protection capabilities. Brief description of the drawings
[0038] Figure 1 It is a schematic flow diagram of the method for massive multi-source multi-modal data fusion of the present invention. Detailed implementation manners
[0039] In order to describe the present invention more specifically, the technical solutions of the present invention will be described in detail below in conjunction with the drawings and specific implementation manners.
[0040] To address the complex challenges in the technology of massive multi-source multi-modal data fusion, the present invention innovatively designs a comprehensive fusion framework. This framework first ensures the quality and consistency of the input data through a multi-source data screening and cleaning mechanism; then uses a strategy that combines the attention mechanism in deep learning with GNN (graph neural network) to achieve efficient extraction and fusion of cross-modal features. In this process, the present invention introduces an adaptive weight allocation algorithm to dynamically adjust the contribution degrees of different modal data and ensure the accuracy and robustness of the fusion results. To address the problems of large data scale and uneven distribution, the present invention adopts a distributed computing architecture and data augmentation technology, and generates inter-modal data based on GAN (generative adversarial network), effectively expanding the training data set and improving the generalization ability of the model in scarce data scenarios. The specific execution process is as Figure 1 shown: (1)Dataset preparation and preprocessing.
[0041] Before performing multi-modal data fusion, it is first necessary to preprocess the data to ensure the compatibility and consistency of different modal data, which usually includes steps such as data cleaning, standardization and normalization, and feature extraction.
[0042] 1-1 Data Integration When integrating multi-source and multi-modal data, it is first necessary to clarify the data sources and data types. The data types used in this embodiment are mainly in the field of intelligent healthcare. The data sources include the hospital's electronic medical record system (text data), medical imaging system (image data), and voice diagnosis for patients (sound data); when integrating these data, problems such as inconsistent data formats, timestamp alignment, and patient privacy protection need to be solved.
[0043] 1-2 Data Cleaning Data cleaning is a crucial step to ensure data quality. For text data, in addition to removing duplicates and missing values, it is also necessary to perform word segmentation, part-of-speech tagging, stop word removal, etc.; for image data, in addition to denoising, it is also necessary to perform operations such as image cropping, scaling, and rotation to ensure that the data input into the model has a unified size and format. Removing outliers and missing values is to ensure data quality. For missing values, methods such as mean filling, ignoring, or model prediction can be used for processing; for outliers, they are processed by ignoring, replacing, or correcting after identification.
[0044] 1-3 Standardization and Normalization Standardization and normalization are crucial steps in data preprocessing, aiming to unify data from different sources or scales into the same dimension and range so that subsequent algorithm models can process and analyze more effectively. Standardization methods need to be applied to various types of data, including image data (pixel value normalization, color space conversion to improve processing efficiency), text data (reducing lexical diversity through stemming and lemmatization to promote the consistency of text analysis), and sound data (using peak normalization and loudness normalization to optimize sound quality).
[0045] Image data normalization is a processing method that scales the pixel values of an image to a specific range. The main purpose is to eliminate the dimensional differences between images and improve the accuracy and efficiency of image processing. In this embodiment, min-max normalization is adopted, that is, by calculating the minimum and maximum values in the image data, a linear transformation is performed on each pixel value to scale it to the specified range [0, 1]. Text data normalization mainly involves the processing after text vectorization to eliminate the differences in feature representations of different texts. In this embodiment, the word embedding method is used to convert the words in the text into vector representations of a fixed dimension, and these vectors can capture the semantic relationships between words. The purpose of sound data normalization is to reduce the differences between different samples and map the data of the sound signal to the specified amplitude range. In this embodiment, the min-max normalization method is also used to normalize the sound data. By performing a linear transformation on the amplitude of the sound signal, it is scaled to the specified range [0, 1] to eliminate the amplitude differences between different samples.
[0046] 1-4 Feature Extraction For the feature extraction of different modality data, this invention adopts a preprocessing model. Each module deeply depends on the Transformer architecture and its variants to make full use of its advantages in processing sequence data. They are all based on powerful pre-trained models and achieve efficient and accurate feature extraction for visual, text, and sound data through deep learning and self-attention mechanisms, laying a solid foundation for the fusion and understanding of multi-modal data. Each data module uses its respective selected feature extraction method to obtain the desired output results from the image, text, and sound data sets respectively. f v 、 f t 、 f a 。In terms of vision, the image to be processed is input into the corresponding model, and the model will output the feature vector of the image; in terms of text, the selected method is used to convert the text into a feature vector; in terms of sound extraction, the selected method is used to extract the feature vector from the sound signal.
[0047] To further reduce the computational complexity and improve the model performance, this embodiment introduces feature screening techniques such as PCA (Principal Component Analysis), LDA (Linear Discriminant Analysis), or model-based feature importance evaluation to screen out the most representative feature subset, eliminate redundant information, and retain key features.
[0048] (2)Model Construction and Pre-training.
[0049] 2-1 Module Construction When building a multi-modal fusion model, it is necessary to construct a visual processing module, a text processing module, an audio processing module, and a multi-modal fusion module respectively. The visual, text, and audio processing modules all use the Transformer network structure and corresponding different training models, while the multi-modal fusion module needs to design a mechanism to effectively fuse features from different modalities.
[0050] The visual processing module focuses on the in-depth extraction of image features. It adopts the ViT model based on Transformer to extract visual features, including but not limited to color, texture, shape, spatial relationship, and dynamic changes, etc. The ViT model extracts the global and local features of the image by dividing the image into a series of small patches, regarding these patches as tokens in the sequence, and then processing these sequences through the standard Transformer encoder. This way not only retains the powerful ability of Transformer in processing sequence data but also enables the model to capture more complex spatial dependency relationships in the image, which can be used for tasks such as image classification, object recognition, scene understanding, and action recognition. The output of the visual processing module is usually a high-dimensional feature vector that encodes the key visual information of the input image or video. The main implementation process is as follows: First, the input image is adjusted to the fixed size (224×224 pixels) required by the ViT model, and then the adjusted image is divided into a series of fixed-size patches. Each patch is usually flattened into a vector and used as a token in the sequence. These tokens not only contain the color and texture information of the image but also retain the information of spatial relationships through their positions in the sequence.
[0051] To enable the model to understand the start and end of the sequence and the global information of the image, the present invention adds a special classification token to the beginning of the sequence and adds positional encoding or learned position embeddings at the end of the sequence to retain the spatial position information of each token.
[0052] The token sequence obtained from preprocessing is input into a standard Transformer encoder, which is composed of multiple stacked encoder layers. Each encoder layer contains a self-attention mechanism and a feed-forward neural network. The self-attention mechanism allows the model to consider the information of all other tokens in the sequence when processing each token, thereby capturing the complex spatial dependencies in the image. This mechanism is the most core part of the Transformer when processing sequence data and is also the key for ViT to effectively extract global and local features of the image. Additionally, the feed-forward neural network further performs a non-linear transformation on the output of the self-attention mechanism to enhance the model's expressive power.
[0053] After being processed by the Transformer encoder, the initial classification token is updated to contain the global information of the image. The output of this classification token is used as the feature representation of the entire image. By adding an additional fully connected layer to the output of this classification token, tasks such as image classification are performed. For other tasks (such as object recognition, scene understanding, etc.), further processing of the feature vector or combination with the outputs of other modules is required. Finally, the visual processing module outputs a high-dimensional feature vector that encodes the key visual information of the input image or video and can be used for subsequent multi-modal data fusion, classification, recognition, etc. tasks.
[0054] The text processing module is responsible for the understanding and representation of text information and uses the BERT model based on Transformer. In the field of text feature extraction, many difficult problems in natural language processing tasks such as machine translation, text classification, sentiment analysis, etc. can be effectively solved by using Transformer and its variant BERT model. The BERT model captures the complex relationships between words in the text through the self-attention mechanism, thereby extracting deep semantic features. Transformer and its BERT model provide strong support for text feature extraction in multi-modal data fusion. By using the text processing module to extract semantic and syntactic features from the text, these features include the meaning of words, sentence structure, context relationships, etc., to make full use of the powerful representation ability of the pre-trained model and play a key role in tasks such as language translation, sentiment analysis, text summarization, etc. The main implementation process is as follows: Text cleaning: Remove the noise in the text, including HTML tags, special characters, extra spaces, etc.
[0055] Tokenization: Split the text into units that the model can process. BERT usually uses the WordPiece algorithm for tokenization, which combines lexical and character-level information.
[0056] Add special tokens: Add special tokens ([CLS] and [SEP]) at the beginning and end of the text sequence to indicate the start and end of the sequence, as well as possible sentence separation.
[0057] Position encoding: Assign a position embedding to each token to preserve its position information in the sequence, as the Transformer model itself does not directly handle the order of the sequence.
[0058] Load the pre-trained BERT model, including its weights and configuration. The BERT model has been pre-trained on large-scale text data and can capture rich language knowledge and context relationships. Configure the output layer of the model according to the specific task (such as text classification, sentiment analysis, language translation, etc.). For example, in a text classification task, a fully connected layer needs to be added on top of the BERT output layer to map the hidden states of BERT to class labels.
[0059] Input the preprocessed text sequence (including token embeddings, position embeddings, etc.) into the BERT model. The BERT model processes the input sequence through the self-attention mechanism and multiple layers of Transformer encoders, captures the complex relationships between words, and extracts deep semantic features; the self-attention mechanism allows the model to consider other tokens in the entire sequence when processing each token, thereby capturing the global context information of the text.
[0060] Extract the corresponding features from the output of the BERT model according to the task requirements. In a text classification task, the hidden state of the [CLS] token is usually used as the representation of the entire sequence and passed to the classification layer for classification. In a sentiment analysis task, attention needs to be paid to the output of specific parts of the sentence; in a text summarization task, the hidden states of the entire sequence need to be utilized to generate the summary. The output layer converts the extracted features into the format required by the task, such as class probabilities, sentiment scores, or summary text.
[0061] Post-process the output of the model, convert the class probabilities into specific class labels, and format the generated text summary. Use evaluation metrics to evaluate the performance of the model on the test set, and then adjust the model configuration, training parameters, or data preprocessing steps according to the evaluation results to further improve the model performance.
[0062] The sound processing module focuses on the in-depth feature extraction of audio signals and adopts the cutting-edge HuBERT model based on Transformer. The HuBERT model aims to mine and learn the intrinsic representation of audio from a large amount of unlabeled audio data through self-supervised learning, so as to effectively extract rich sound features. These features widely cover basic acoustic attributes such as pitch, rhythm, timbre, and duration, as well as higher-level features such as the emotional color of speech and the characteristics of environmental sounds. The innovation of the HuBERT model lies in its use of a pre-training strategy. First, it learns the low-level representation of audio by predicting the pseudo-labels of audio frames (audio feature vectors based on k-means clustering), and then gradually refines these representations in higher-level iterations to approach more high-level semantic information. This step-by-step self-supervised learning method enables HuBERT to capture the complex temporal dynamics and spatial structures in audio, thus more accurately parsing and understanding sound signals. The sound processing module can recognize and understand speech commands, distinguish different speakers, analyze the emotional state of speech, etc. These features are crucial for applications such as automatic speech recognition, speech synthesis, and sound event detection. The main implementation process is as follows: Configure the parameters of the HuBERT model, including the number of layers of the Transformer encoder, embedding dimension, attention mechanism, etc.
[0063] Initial stage: Using an unsupervised method, in this embodiment, the k-means clustering method is used to cluster audio frames and assign a pseudo-label to each audio frame.
[0064] Iterative stage: Use the model trained in the previous stage to predict more refined pseudo-labels for the next stage of training. This process is recursive, gradually improving the accuracy and usefulness of the pseudo-labels.
[0065] Take the audio frames and their pseudo-labels as input, and train the HuBERT model to learn the intrinsic representation of audio through a masked prediction task. The masked prediction task means randomly covering some audio frames and training the model to predict the pseudo-labels of these covered frames. As the training progresses, gradually increase the difficulty of the prediction task, such as covering longer audio segments or introducing more noise and interference.
[0066] After the training is completed, the HuBERT model has learned the low-level and high-level feature representations of audio signals. For new audio inputs, the pre-trained HuBERT model can be used for inference to obtain the in-depth feature representation of audio frames. These features not only include basic acoustic attributes such as pitch, rhythm, timbre, and duration, but may also contain higher-level semantic information such as the emotional color of speech and the characteristics of environmental sounds.
[0067] The multi-modal fusion module is a core component in multi-modal learning. It aims to effectively integrate information from different modalities so that the model can comprehensively utilize this information for decision-making. In this implementation, the multi-module fusion framework adopts a joint architecture, which maps the feature representations of different modalities f u into a shared semantic subspace to enable the fusion of multi-modal features. Under this framework, each modality first extracts features through its respective pre-trained model (ViT for vision, BERT for text, HuBERT for sound). Then, after appropriate alignment and transformation, these features are fed into a fusion layer for fusion. The fusion layer uses the cross-modal attention mechanism method to achieve the interaction and fusion between different modality features. The main implementation process is as follows: Design the fusion architecture: Select an appropriate fusion architecture, the joint architecture, according to the task requirements.
[0068] Implement the fusion algorithm: Adopt the attention mechanism. The fusion method based on the attention mechanism can dynamically adjust different modality features to enhance the interaction between different modality features.
[0069] Output the joint representation: Use the fused features as the output of the multi-modal model for subsequent task processing.
[0070] 2-2 Initialization of the pre-trained model Pre-training is an important means to accelerate model training and improve model performance. In the implementation of the massive multi-source multi-modal data fusion technology, using pre-trained models can significantly accelerate the model training process and improve model performance. First, it is necessary to load the pre-trained models required for the visual processing module and the text processing module. These models are usually trained on large-scale datasets and have good generalization ability. Second, use the weights of the pre-trained models as initial parameters to initialize the network parameters of the visual processing module, the text processing module, and the multi-modal fusion module. Through pre-training, the model already has a certain generalization ability, so it can converge to the optimal solution faster during the training process. The training parameters set can be the learning rate, batch size (Batch Size), number of training epochs (Epochs), etc. In addition, pre-training can also reduce the model's dependence on training data and reduce the risk of overfitting.
[0071] For the visual processing module, a CNN model pre-trained on large datasets such as ImageNet can be used; for the text processing module, BERT or GPT models pre-trained on a large amount of text data can be used.
[0072] (3)Implementation of the multi-modal fusion strategy.
[0073] 3-1 Fusion model and method Feature-level fusion: The features extracted from different data sources are concatenated, weighted, or combined to obtain a more comprehensive and information-rich feature vector. This fusion method aims to enhance the model's understanding and expression ability of data by combining complementary information from multiple feature sets. For example, in image processing, image features and text features can be concatenated.
[0074] Decision-level fusion: The decisions from different data sources or different models are integrated to obtain a more reliable decision result. This fusion method aims to reduce the uncertainty or error that may be brought by a single decision by integrating the advantages of multiple independent decisions. The decision-level fusion method used in this embodiment is the weighted average method. Different weights are given according to the reliability or confidence of different prediction results, and then weighted averaging is performed to obtain the final decision result.
[0075] Model fusion: The prediction results of multiple models are integrated to obtain a more accurate final prediction. The model fusion method used in this invention is Bagging (Bootstrap Aggregating). Bagging is a bootstrap aggregation algorithm and a parallel ensemble learning method. Its basic idea is to randomly draw multiple sample sets of the same size but allowing duplicates from the original dataset by the bootstrap sampling method. Subsequently, multiple different base models (such as decision trees, neural networks, etc.) are trained in parallel using these sample sets. Finally, the prediction results of each base model are aggregated by averaging (for regression problems) or voting (for classification problems), thereby reducing the variance of the base model and improving the generalization ability of the overall model.
[0076] 3-2 Deep cross-attention mechanism The deep cross-attention mechanism is a method to enhance feature representation by calculating the correlation between different inputs (such as text, images, sounds, etc.). It usually includes three main components: query Q, key K, and value V. These components can be feature vectors or matrices from different modalities and allow features of different modalities to pay attention to each other during the fusion process, thereby capturing richer information. This mechanism is usually implemented through an attention layer, which can calculate the similarity between features of different modalities and assign weights according to the similarity. In this embodiment, three different feature sequences are involved, namely image features, text features, and sound features. The cross-attention between each pair of features (such as text and sound, sound and image, image and text) is calculated. Taking the cross-attention between text and sound as an example, the query of the text feature Q text and the key of the sound feature K audio and value V audio :
[0077]
[0078]
[0079] Wherein: and are the key and value weight matrices of the voice features respectively.
[0080] Calculate the attention score and apply it to the voice features:
[0081] Wherein: Q is the query matrix, whose dimension is m × d k , m is the number of queries, d k is the dimension of each query; K is the key matrix, whose dimension is n × d k , n is the number of keys; V is the value matrix, whose dimension is n × d k , d k is the dimension of each value (in some cases d v = d k ). is the scaling factor, which is used to prevent the dot product result from becoming very large when d k is very large, resulting in the disappearance of the softmax function's steps; is the calculated attention score (normalized dot product), which is used to measure the relationship between text features and voice features; is the calculated attention weight, which is a m × n matrix, representing the correlation strength between each query and each key.
[0082] Then, fuse the calculated attention-weighted features. For example, the cross-attention results of text and voice and the cross-attention results of voice and image can be fused by addition or concatenation:
[0083]
[0084] The fusion results are further combined to obtain the final feature representation:
[0085] The subsequent processing is to input the fused features X fused into the subsequent fully connected layer or other task-related layers for final prediction or classification:
[0086] 3-3 Adaptive Fusion Layer The adaptive fusion layer dynamically adjusts the fusion weights according to the task requirements, which is achieved by introducing learnable parameters. In the adaptive fusion layer, a series of learnable parameters are first initialized. These parameters usually start with random values and are updated through the backpropagation algorithm during the training process. The number of parameters depends on the feature dimensions to be fused and the fusion strategy; the adaptive fusion layer can handle the complex relationships between different modalities more flexibly and improve the fusion effect. In addition, the fusion strategy used in this embodiment is weighted fusion, that is, each feature vector is scaled according to its corresponding learnable weight and then added or concatenated. The value of the weight determines the contribution degree of different feature vectors in the fusion result.
[0087] 3-4 Multi-task Learning To make full use of the rich information in multi-modal data, a multi-task learning framework can be adopted. Under this framework, the model simultaneously learns multiple related tasks and improves the learning efficiency by sharing the representation layer. In the field of intelligent healthcare, a model can be trained to predict the disease type (main task) and the disease severity (auxiliary task) simultaneously, and use the prediction result of the disease severity to assist the prediction of the disease type.
[0088] 3-5 Experimental Design To verify the effectiveness of the multi-modal fusion strategy, a series of experiments need to be designed. The experimental design should include the comparison of different fusion strategies (early fusion, late fusion, and hybrid fusion), the comparison of different model architectures (CNN+LSTM (Long Short-Term Memory Network) and Transformer), and the performance evaluation on different datasets.
[0089] (4)Model Training and Optimization.
[0090] 4-1 Loss Function Design Design appropriate loss functions according to the task type (such as classification, regression, ranking, etc.). For multi-modal fusion tasks, the design of the loss function needs to comprehensively reflect the contributions of all modalities and tasks, and it may be necessary to design a composite loss function that can consider multiple modalities and multiple tasks simultaneously. When there are two main tasks (Task A and Task B), corresponding to two modalities (Modality 1 and Modality 2) respectively, a composite loss function is constructed in the form of a weighted sum:
[0091] where: L A,Modality1 and L A,Modality2 respectively represent the losses of Person A in a single modality, which may be the categorical cross-entropy loss (for classification tasks) or the mean squared error loss (for regression tasks); L B,CombinedModalities represents the loss of Task B after fusing multi-modal information, α , β , γ are hyperparameters used to balance the importance of different tasks and modalities.
[0092] 4-2 Optimization Algorithm Selection Selecting an appropriate optimization algorithm is crucial for model training. Commonly used optimization algorithms include SGD (Stochastic Gradient Descent), Adam (Adaptive Moment Estimation), RMSprop (Root Mean Square Propagation), etc.; during training, parameters such as the learning rate and momentum can be adjusted according to the convergence speed and training stability of the model.
[0093] 4-3 Hyperparameter Tuning Hyperparameter tuning is an important means to improve model performance. Optimal hyperparameter combinations can be found through methods such as grid search, random search, Bayesian optimization, etc.; during the tuning process, attention should be paid to the generalization ability of the model to avoid overfitting or underfitting.
[0094] (5)Model Evaluation and Performance Analysis.
[0095] 5-1 Evaluation Metric Selection Select appropriate evaluation metrics according to the task requirements. For classification tasks, commonly used evaluation metrics include accuracy, precision, recall, F1-score, etc.; for regression tasks, metrics such as mean squared error (MSE), root mean squared error (RMSE) can be used.
[0096] 5-2 Performance Analysis Analyze the performance of the model under different datasets and different parameter settings. By comparing the experimental results, the impacts of different fusion strategies, model architectures, optimization algorithms, etc. on the model performance can be understood. At the same time, the error types of the model can also be analyzed for subsequent improvement.
[0097] 5-3 Case Analysis Select specific cases for analysis to demonstrate the effects and existing problems of the model in actual applications. For example, in the field of intelligent healthcare in this embodiment, indicators such as the accuracy rate and recall rate of the model in predicting a certain disease can be analyzed, and possible misjudgment reasons and improvement directions can be explored, as shown in Table 1 specifically: Table 1
[0098] (6)Deployment and Application.
[0099] 6-1 Model Deployment Deploy the trained model to the actual production environment, which includes converting the model into a format suitable for deployment (such as ONNX (Open Neural Network Exchange), TensorRT (TensorRT), etc.), configuring the model server (such as TensorFlow Serving, TorchServe, etc.), and conducting performance optimization and stability testing.
[0100] 6-2 Application Scenarios Determine the applicable application scenarios according to the characteristics and performance of the model. For example, in the field of intelligent healthcare, the multi-modal fusion model can be used in scenarios such as disease diagnosis and treatment plan recommendation; in the field of intelligent security, it can be used in scenarios such as face recognition and abnormal behavior detection.
[0101] 6-3 Continuous Iteration and Optimization With the continuous accumulation of data and the continuous development of technology, it is necessary to continuously iterate and optimize the model, which includes collecting new data to train the model, improving the model architecture and fusion strategy, and optimizing the running efficiency and stability of the model.
[0102] 6-4 Challenges and Solutions Various challenges may be encountered during the deployment and application processes, such as data privacy protection, model interpretability, cross-domain adaptability issues, etc. To address these challenges, the solutions adopted in this invention are to protect data privacy by using differential privacy technology, develop a model with strong interpretability, and use technologies such as transfer learning to improve the cross-domain adaptability of the model.
[0103] The above description of the embodiments is provided to enable those of ordinary skill in the art to understand and apply the present invention. It is obvious that those skilled in the art can easily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without creative efforts. Therefore, the present invention is not limited to the above embodiments, and all improvements and modifications made by those skilled in the art based on the disclosure of the present invention should fall within the protection scope of the present invention.
Claims
1. A method for fusing massive multi-source and multi-modal data, characterized in that, It includes the following steps: (1) Obtain a massive dataset with multiple sources and rich multimodal information, and preprocess the data therein; (2) Construct a multimodal fusion model, including: A visual processing module that extracts visual features from image data based on the ViT model; A text processing module that extracts semantic features from text data based on the BERT model; A sound processing module that extracts sound features from audio data based on the pre-trained HuBERT model; A multimodal fusion module that fuses the extracted visual features, semantic features, and sound features under the Joint Architecture framework; (3) Augment the preprocessed massive dataset and use this dataset to train the multimodal fusion model; (4) Use the fusion features generated by the trained multimodal fusion model to complete the corresponding application scenario tasks.
2. The method for fusing a large amount of multi-source and multi-modal data according to claim 1, characterized in that: The data preprocessing in step (1) includes data integration, data cleaning, and standardization. For text data, data cleaning includes processing such as missing value filling, outlier correction, stop word removal, and duplicate item removal; for audio data, data cleaning includes denoising processing; for image data, data cleaning includes processing such as denoising, image cropping, scaling, and rotation in addition to denoising; for image data and audio data, standardization uses min-max normalization processing; for text data, standardization uses a word embedding method to convert the words in the text into vector representations of a fixed dimension.
3. A method for fusing a large amount of multi-source and multi-modal data according to claim 1, characterized in that: The visual processing module first adjusts the input image to the fixed size required by the ViT model, and then divides the adjusted image into a series of small blocks of a fixed size. Each small block is flattened into a vector as a token in the sequence. These tokens not only contain the color and texture information of the image but also retain the information of the spatial relationship through their positions in the sequence; furthermore, a special classification token is added at the beginning of the sequence, and position encoding or position embedding is added at the end of the sequence; the processed token sequence is input into a Transformer encoder cascaded by multiple encoders. Each encoder layer contains a self-attention mechanism and a feed-forward neural network. When processing each token, the self-attention mechanism can take into account the information of all other tokens in the sequence, thereby capturing the complex spatial dependencies in the image; the feed-forward neural network further performs a non-linear transformation on the output of the self-attention mechanism to enhance the expressive power of the ViT model; the output after being processed by the Transformer encoder is a high-dimensional feature vector that encodes the key visual information of the input image.
4. A method for fusing massive multi-source multi-modal data according to claim 1, characterized in that: The text processing module first tokenizes the text using the WordPiece algorithm to obtain a token sequence, then adds special markers [CLS] and [SEP] at the beginning and end of the token sequence respectively, and assigns a position embedding to each token in the sequence; the processed token sequence is input into a pre-trained BERT model, and the BERT model processes the input token sequence through the self-attention mechanism and multiple layers of Transformer encoders, captures the complex relationships between tokens, and extracts deep semantic features; The self-attention mechanism allows the model to consider other tokens in the entire sequence when processing each token, thus capturing the global context information of the text; the BERT model finally outputs a high-dimensional feature vector, which encodes the semantic and syntactic information of the input text.
5. A method for fusing a large amount of multi-source and multi-modal data according to claim 1, characterized in that: The sound processing module first initializes the pre-trained parameters of the HuBERT model, then uses the k-means clustering method to cluster the audio frames, assigns a pseudo-label to each audio frame, and uses the HuBERT model trained in the previous stage to predict a more refined pseudo-label for the next stage of model training; the audio frames and their pseudo-labels are used as inputs to train the HuBERT model to learn the intrinsic representation of audio data through the masked prediction task, which randomly masks some audio frames and trains the model to predict the pseudo-labels of these masked frames; finally, new audio data is input into the HuBERT model pre-trained as above to output a high-dimensional feature vector, which encodes the deep features of the input audio.
6. A method for fusing a large amount of multi-source and multi-modal data according to claim 1, characterized in that: The multi-modal fusion model also introduces a feature screening technology, and uses PCA, LDA or a model-based feature importance evaluation method to screen the extracted features.
7. A method for fusing a large amount of multi-source and multi-modal data according to claim 1, characterized in that: The multi-modal fusion module aligns and transforms the visual features, semantic features and sound features appropriately and then sends them into the adaptive fusion layer for fusion. The adaptive fusion layer uses the deep cross-attention mechanism to calculate the cross-attention results between the visual, semantic and sound features, and dynamically adjusts the weights of the attention mechanism to allocate the contribution degrees of different modal information. Finally, the cross-attention results between the features of each modality are fused, and the fusion strategy is automatically adjusted according to different application scenarios and task requirements, such as feature-level fusion, decision-level fusion, model-level fusion, weighted fusion, concatenation fusion or graph neural network-based fusion.
8. A method for fusing massive multi-source and multi-modal data according to claim 1, characterized in that: Before performing step (3), the multi-modal fusion model is pre-trained using transfer learning technology, and the pre-trained model parameters are used as initialization parameters to accelerate the training process of the model in step (3); For the visual processing module, train the ViT model of this module on ImageNet; for the text processing module, train the BERT model of this module on a large-scale text corpus; for the sound processing module, train the HuBERT model of this module on LibriSpeech; at the same time, adopt a parameter sharing mechanism in the multimodal fusion model, allowing network layers of different modalities to share parameters to a certain extent to promote the exchange and complementarity of cross-modal information.
9. A method for fusing a large amount of multi-source and multi-modal data according to claim 1, characterized in that: In step (3), for image data, data augmentation applications include image transformation operations such as image rotation, cropping, and color transformation to generate diverse image samples; for text data, data augmentation uses methods such as synonym replacement, back translation, and sentence recombination to enrich text expressions; for audio data, data augmentation uses methods such as noise addition, audio speed change, volume adjustment, audio cropping and splicing to improve the model's adaptability to complex sound environments; in addition, data augmentation operations also generate new text-image sample pairs, text-audio sample pairs, and audio-image sample pairs in combination with multimodal characteristics to ensure the semantic consistency of these sample pairs while increasing data diversity; In step (3), the data augmentation of the massive dataset also introduces a generative adversarial network and conditional data augmentation technology. The generative adversarial network generates high-quality pseudo data through adversarial training between the generator and the discriminator, and the generated pseudo data is close to the real data in terms of feature distribution; The conditional data augmentation technology is to use the data of one modality to guide the data augmentation process of another modality.
10. A method for fusing a large amount of multi-source and multi-modal data according to claim 1, characterized in that: In step (3), the training process of the multimodal fusion model adopts the methods of joint training and knowledge distillation: first, for a specific task, design a multimodal joint training framework that can process data from different modalities simultaneously; then use one or more single-modal models that have been trained on a large dataset as teacher models, and transfer the knowledge of the teacher models to the multimodal joint training framework through knowledge distillation.
Citation Information
Patent Citations
Text association type short video multi-mode emotion recognition method and system
CN117636196A
Multi-modal model and method for fusing characters, images and audios
CN118861988A
Ambient sound event detection method based on multi-modal data fusion
CN119446154A
Cited By
Human body walking perception index analysis method, system and device and storage medium
CN120473184A
Data processing method based on Internet of Things multi-source information
CN120492815A
CVOCA feature extraction method and system based on multi-modal large model
CN120596902A
Cvoca feature extraction method and system based on multi-modal large model
CN120596902B
Multi-modal fusion deep learning analysis method and system
CN120763881A