A Multimodal Data Augmentation Method and System Based on an AI Training Platform
Through multimodal data fusion and deep learning pipelines, the AI training platform has solved the problem of low automation degree and bias in multimodal data processing and model learning in multimodal data processing, and achieved efficient multimodal data processing and model generalization capabilities.
Patent Information
- Application Number
- CN202510121495.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-01-26
AI Technical Summary
The existing AI training platform has many user manual modifications and low degree of automation when processing multimodal data, and single modal data leads to biased model learning, affecting model generalization capabilities and data processing efficiency.
Develop a multimodal data fusion pipeline for data standardization, timing synchronization and alignment, and feature extraction. Generative adversarial networks are used to generate high-quality synthetic data, and feature-level and decision-level fusion is carried out through multimodal deep learning pipelines, and a multimodal Transformer model is integrated for comprehensive decision-making.
It has improved the automation level of the AI training platform, reduced the cost of manual labeling, improved the generalization ability and data processing efficiency of the model, and enhanced the automation level of multimodal training tasks.
Smart Images

Figure CN119557630B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal data processing, and specifically to a multimodal data enhancement method and system based on an AI training platform. Background Art
[0002] With the rapid development of artificial intelligence technology, AI training platforms are increasingly being used in various industries. They have gradually evolved from traditional single-modality processing (such as image processing, text processing, etc.) to more complex multi-modal data fusion and processing. The main function of these platforms is to provide a complete set of tool chains for training, testing, and deploying deep learning models, and integrate functions such as data processing, algorithm development, model training, and reasoning.
[0003] However, as the scale of AI models grows and the application scenarios become more complex, traditional AI R&D platforms also face some challenges: in multimodal processing scenarios, AI models require appropriate data processing strategies to enable AI models to deeply mine the features of multimodal data and learn patterns in multimodal scenarios. How to achieve multimodal data fusion by combining data from different modalities (such as images, text, sensor data, etc.) to make up for the limitations of single modality data and improve the generalization ability of the model and data processing efficiency is a problem that needs to be solved at present. In addition, when facing multimodal data processing tasks, existing AI training platforms generally have the problem of frequent manual modifications by users and low degree of platform training automation. Summary of the invention
[0004] The technical task of the present invention is to address the above shortcomings and provide a multimodal data enhancement method and system based on an AI training platform to solve the problem of model learning bias caused by single modality data in the prior art, improve the generalization ability of the model and data processing efficiency, and enhance the degree of automation of the AI training platform.
[0005] The technical solution adopted by the present invention to solve its technical problem is:
[0006] A multimodal data enhancement method based on an AI training platform, the implementation of the method includes:
[0007] First, a multimodal data fusion pipeline is developed on the AI training platform. Data of different modalities (such as images, text, sensor data, etc.) are collected through the data acquisition module and preprocessed. In the preprocessing stage, the multimodal data fusion pipeline is used to extract features from data from different modalities and perform joint modeling at a specific level.
[0008] Then, a data augmentation pipeline is used to automatically generate high-quality synthetic data;
[0009] Finally, a multi-modal deep learning pipeline is developed to fully learn and understand the deep features of different modal data.
[0010] Furthermore, the multi-modal data fusion pipeline is used to process data sources of different modalities and synchronize, align, and fuse the data according to the actual application scenario.
[0011] The multi-modal data fusion pipeline includes the following steps:
[0012] (1) Data standardization: Standardize the multi-modal data (images, text, sensor data, speech) collected from different data sources.
[0013] (2) Temporal synchronization and alignment: Synchronize and align the time series data (sensor data) and non-time series data (images, text) to ensure effective fusion of multi-modal data within the same time period.
[0014] (3) Data feature extraction: The multi-modal feature extraction model includes a convolutional neural network (CNN), a recurrent neural network (RNN), and a Transformer.
[0015] Furthermore, standardizing the multi-modal data collected from different data sources includes:
[0016] For image data: Scale and normalize it, supporting two normalization methods: The first is to adjust the image pixel values from the range of 0 - 255 to 0 - 1; the second is to normalize according to specific mean and standard deviation values, using the normalization values of the ImageNet dataset: mean [0.485, 0.456, 0.406], standard deviation [0.229, 0.224, 0.225].
[0017] For text data: Tokenize and encode it, use natural language processing toolkits (SpaCy, NLTK) for tokenization, and use models such as BERT and Word2Vec to convert the tokenized text into word embedding vectors.
[0018] For time series data: Perform interpolation and alignment processing to ensure that all data can be analyzed on the same time axis or in the same dimension.
[0019] The above-mentioned temporal synchronization and alignment are specifically implemented as follows:
[0020] Slice the entire time series into multiple time periods, each time period containing data of multiple modalities, and missing time points can be filled by interpolation.
[0021] The above-mentioned data feature extraction specifically includes:
[0022] Image data feature extraction: Use the pre-trained convolutional neural network (CNN) ResNet model to extract high-level features from images; the pre-trained model has been trained on a large image dataset and can be directly applied to feature extraction; after the last convolutional layer, use global average pooling to convert the feature map into a fixed-length feature vector for further multi-modal fusion;
[0023] Text data feature extraction: Use the pre-trained language model Transformer to encode text data and extract semantic feature vectors of the text; through pre-training and fine-tuning, the model can understand the semantics and structure of the text; then, convert each tokenized word or sub-word into a fixed-length vector, and the output feature vector can represent the context relationship and semantic information in the text;
[0024] Sensor data feature extraction: Use the recurrent neural network (RNN) to extract time series features from sensor data; for time series data, the recurrent neural network (RNN) can capture time dependencies and is suitable for feature extraction of time series data.
[0025] Furthermore, the data augmentation pipeline uses a generative adversarial network (GANs) generation model to generate new samples using existing multi-modal data:
[0026] Train the GAN model with existing multi-modal data as input; the generator generates new samples, while the discriminator evaluates the authenticity of the generated samples; through continuous iterative training, the generator can learn the distribution of real data and generate new samples similar to real samples; for example, new images, texts, or time series can be generated to expand the dataset; the generated samples will be input into the model as additional training data to improve the generalization ability of the model.
[0027] Furthermore, to optimize the generation effect of GANs, regularly update the parameters of the generation model and re-train it using newly collected multi-modal data to improve the diversity and quality of the generated samples; through a continuous feedback mechanism, continuously optimize the effects of feature extraction and data augmentation.
[0028] Furthermore, the multi-modal deep learning pipeline includes a feature-level fusion strategy and a decision-level fusion strategy;
[0029] The feature-level fusion strategy is to first extract features from the data of each modality, align the features of different modalities in a certain space or dimension, and then input them into a joint model; the advantage of feature-level fusion is that it can retain the detailed information of each modality and can extract the correlation between modalities by sharing the network structure in deep learning; therefore, this method integrates feature-level fusion;
[0030] The decision-level fusion strategy is to process the data of each modality separately and then make a comprehensive decision on their respective output results; therefore, this method integrates the decision-level fusion strategy.
[0031] Furthermore, the feature-level fusion strategy includes:
[0032] Feature alignment: Create a shared latent space based on a fully connected neural network to enable features of different modalities to be aligned within the same space; optimize the alignment effect through a loss function (minimizing the reconstruction error) to ensure the maximization of the correlation between modalities; and, during the alignment process, introduce a regularization technique (L2 regularization) to prevent overfitting;
[0033] Deep fusion model: Design a multi-modal joint model, a multi-modal Transformer, embed this deep fusion model as a built-in model into the AI training platform, and at the same time support users to customize the model; use the aligned features as input for further processing; the specific network architecture of the multi-modal joint model is as follows:
[0034] The input layer of the multi-modal Transformer receives the aligned feature vectors, undergoes non-linear transformation through multiple fully connected layers, and finally outputs the prediction results; in order to retain the original feature information, skip connections are introduced between different layers to enhance the learning ability of the model; the self-attention layer is used to enable the model to focus on the important parts of the features in different modalities, so that the model can adaptively adjust the weights of the features of each modality;
[0035] The decision-level fusion strategy includes:
[0036] Independent model training: Design independent models for each modality: use CNN for images, BERT for text, and LSTM for other data; embed the trained models as built-in basic models into the AI training platform, and at the same time users can choose to customize the basic models; each model is independently trained on the modality data of its own to capture modality-specific features; after training, each independent model makes predictions on the input data and generates corresponding output results (such as classification probabilities);
[0037] Weighted voting: Set the weights of each modality model and combine the outputs of the models; the weights are adjusted based on the performance of the models on the validation set. For example, a model with a higher accuracy rate accounts for a greater proportion in the final decision.
[0038] The multi-modal data processing method with the feature-level fusion strategy and the decision-level fusion strategy not only helps to improve the performance and generalization ability of the model in theory, but also can reduce costs, improve efficiency, and enhance the automation degree of users' multi-modal task training on the AI platform in practical applications.
[0039] The present invention also claims to protect a multi-modal data augmentation system based on an AI training platform, including:
[0040] A multi-modal data fusion pipeline for processing data sources of different modalities and synchronizing, aligning, and fusing data according to the actual application scenario;
[0041] A data augmentation pipeline for generating models through generative adversarial networks (GANs), and automatically generating new samples for the user's learning tasks using existing multi-modal data to improve the generalization and robustness of model training:
[0042] A multi-modal deep learning pipeline, including a feature-level fusion strategy and a decision-level fusion strategy, for fully learning and understanding the deep features of different-modal data and assisting the user in conveniently training a multi-modal learning model;
[0043] The system specifically realizes multi-modal data augmentation through the above method.
[0044] The present invention also claims to protect a multi-modal data augmentation device based on an AI training platform, characterized by including: at least one memory and at least one processor;
[0045] The at least one memory is used for storing machine-readable programs;
[0046] The at least one processor is used for calling the machine-readable program to implement the above method.
[0047] The present invention also claims to protect a computer-readable medium, characterized in that computer instructions are stored on the computer-readable medium, and when the computer instructions are executed by a processor, the above method is implemented.
[0048] Compared with the prior art, a multi-modal data augmentation method and system based on an AI training platform of the present invention have the following beneficial effects:
[0049] 1. Improve the automation level of the AI training platform: The data fusion channel can automatically fuse features of different modalities, comprehensively capture information in the data automatically, reduce the bias caused by a single modality, improve the adaptability to new data, and further improve the automation level of multi-modal training tasks.
[0050] 2. Improve the optimization of multi-modal data training on the AI training platform: In the decision fusion strategy, independent models are trained for specific modalities respectively, which can give full play to their respective advantages. After comprehensive decision-making, the accuracy of the final prediction can be improved, assisting the user to better train and optimize the model.
[0051] 3. Reduce the cost of data annotation: Through intelligent feature extraction and fusion, the need for manual annotation and data screening is reduced, and the time cost in the data preparation stage is lowered.
[0052] 4. Improve the efficiency of data processing: The integration strategy can automatically complete multiple tasks during the data processing process, saving the time for multimodal data processing in the case of large-scale datasets and improving the training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a diagram of a multimodal data fusion pipeline module provided by an embodiment of the present invention;
[0054] Figure 2 is a diagram of a multimodal deep learning pipeline structure provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0055] The present invention will be further described below in conjunction with specific embodiments.
[0056] An embodiment of the present invention provides a multimodal data augmentation method based on an AI training platform. The implementation of this method includes:
[0057] First, develop a multimodal data fusion pipeline on the AI training platform. Collect data of different modalities (such as images, texts, sensor data, etc.) through a data acquisition module and perform preprocessing; in the preprocessing stage, use the multimodal data fusion pipeline to extract features from data of different modalities and perform joint modeling at a specific level;
[0058] Then, adopt a data augmentation pipeline to automatically generate high-quality synthetic data;
[0059] Finally, develop a multimodal deep learning pipeline to fully learn and understand the deep features of data of different modalities.
[0060] This method comprehensively improves the multi-modal data processing ability of the AI training platform in multiple aspects. In terms of optimizing the multi-modal data training of the AI training platform, this method effectively makes up for the limitations of single-modal data through intelligent processing technology, improving the generalization ability of the model and the data processing efficiency. In terms of reducing the data standard cost, this method reduces the manual input while improving the data processing efficiency by realizing efficient feature extraction and data augmentation. In terms of improving the data processing efficiency, this method deeply mines the feature information of multi-modal data through a hybrid strategy, providing strong support for the application of the AI training platform in a multi-modal environment. Combining the above characteristics, this method can automatically and comprehensively capture the information in the data, improve the data processing ability, and greatly improve the automation degree of multi-modal training tasks. For users, training multi-modal tasks on the AI platform will reduce the interaction cost, and the efficiency of training and optimizing the multi-modal learning task model will be greatly improved.
[0061] The specific implementation of this method is as follows:
[0062] 1. Develop a multi-modal data fusion pipeline. This pipeline is used to process data sources of different modalities and synchronize, align, and fuse the data according to the actual application scenario. It includes the following three steps:
[0063] (1) Data standardization: Standardize the multi-modal data (images, text, sensor data, speech) collected from different data sources.
[0064] For image data: Perform scaling and normalization, supporting two normalization methods: The first is to adjust the image pixel values from the range of 0-255 to 0-1; the second is to normalize according to specific mean and standard deviation values, using the normalization values of the ImageNet dataset: mean [0.485, 0.456, 0.406], standard deviation [0.229, 0.224, 0.225].
[0065] For text data: Perform word segmentation and encoding. Use natural language processing toolkits (SpaCy, NLTK) for word segmentation, and use models such as BERT and Word2Vec to convert the segmented text into word embedding vectors.
[0066] For time series data: Perform interpolation and alignment processing to ensure that all data can be analyzed on the same time axis or the same dimension.
[0067] (2) Temporal synchronization and alignment: Synchronize and align time series data (sensor data) and non-time series data (images, text) to ensure the effective fusion of multi-modal data within the same time period. The specific approach is to slice the entire time series into multiple time periods, each time period containing data of multiple modalities, and missing time points can be filled by interpolation.
[0068] Collect different modalities of data from multiple data sources through the data acquisition module to ensure that each time slice can contain multi-modal data. This will be described in detail below:
[0069] Time series data: For time series data with clear timestamps (such as sensor data, audio data), the data will be timestamp-aligned to ensure that the data can be organized in chronological order.
[0070] Text data: First, the sources of text data can be basically divided into the following types. First, the text data comes from real-time collection (such as conversation records), then a timestamp can be assigned to each group of conversations during collection. Second, the text data is some log records, then the log will contain the time when the text was generated; third, the text data comes from crawling network news or comments using a crawler software, then the timestamp will also be recorded during crawling. Therefore, the text data will record its timestamp during collection, so as to ensure that it can be organized according to time.
[0071] Image data: Image data is usually collected by cameras or other imaging devices and has clear timestamps. The data can be organized in chronological order according to the timestamps.
[0072] Therefore, multi-modal data can obtain its timestamp information during the data acquisition process. As long as appropriate experimental data is selected, it can be ensured that each time slice can contain multi-modal data.
[0073] For missing time points, in the time series, some data points with timestamps are missing, which are discovered by using the timestamp interval check method. During data preprocessing, check the interval between each timestamp and the previous timestamp. If the time interval between a certain timestamp and the previous timestamp does not conform to the predetermined sampling interval, it means that the data at that time point is missing.
[0074] (3) Data feature extraction: The multi-modal feature extraction model includes convolutional neural network (CNN), recurrent neural network (RNN) and Transformer.
[0075] Image data feature extraction: Use the pre-trained convolutional neural network (CNN) ResNet model to extract high-level features from images. The pre-trained model has been trained on a large image dataset and can be directly applied to feature extraction. After the last convolutional layer, use global average pooling to convert the feature map into a fixed-length feature vector for further multi-modal fusion.
[0076] In image processing tasks, high-level features refer to features that are extracted from the original image data, more abstract, and have a higher semantic level. These features can represent the complex information of the image, such as objects, scenes, textures, and other semantic contents, and can help machine learning models better understand the image content. For the extraction of high-level features using the ResNet network in this method, the following will be specifically elaborated:
[0077] In deep learning, image data is processed into a multi-dimensional array with dimensions (batch size, height, width, number of channels). First, the image data is normalized, and the data dimensions remain unchanged during this process. Then, the ResNet network is used, and through a series of operations such as convolution and pooling, a high-dimensional vector is obtained. This vector is the high-level feature of the image, with dimensions (batch_size, 2048) and 2048 dimensions.
[0078] Text data feature extraction: Use the pre-trained language model Transformer to encode the text data and extract the semantic feature vector of the text. Through pre-training and fine-tuning, the model can understand the semantics and structure of the text. Then, each tokenized word or sub-word is converted into a fixed-length vector, and the output feature vector can represent the context relationship and semantic information in the text.
[0079] Sensor data feature extraction: Use the recurrent neural network RNN to extract the time series features in the sensor data. For time series data, the recurrent neural network RNN can capture time-dependent relationships and is suitable for feature extraction of time series data.
[0080] 2. Develop a multi-modal data augmentation pipeline. The data augmentation pipeline generates models through generative adversarial networks (GANs) and uses existing multi-modal data to generate new samples.
[0081] Train the GAN model with existing multi-modal data as input. The generator generates new samples, while the discriminator evaluates the authenticity of the generated samples. Through continuous iterative training, the generator can learn the distribution of real data and generate new samples similar to real samples. For example, new images, texts, or time series can be generated to expand the dataset. The generated samples will be used as additional training data input into the model to improve the generalization ability of the model.
[0082] The generative adversarial network mainly consists of two opposing parts: the generator and the discriminator. They play against each other through adversarial training. The following is a detailed introduction:
[0083] The task of the generator is to generate highly realistic data. The generator takes some random noise (usually a low-dimensional vector, such as random values drawn from a normal distribution) as input, and after being processed by a neural network, outputs fake data.
[0084] The task of the discriminator is to determine whether the input data is "real" or "generated". It is a binary classifier that outputs a value representing the authenticity of the input data. Usually, it outputs a probability, representing the probability that the data is real data. The discriminator and the generator are alternately trained to reach the balance point of the GAN network, at which time the data generated by the generator is realistic enough.
[0085] The discriminator evaluates the authenticity of the generated samples as follows:
[0086] During the training process of the GAN, the discriminator uses the binary cross-entropy loss function to measure the gap between the prediction result and the true label. The goal of the loss function is to maximize the probability that the real data belongs to the "true" class and maximize the probability that the generated data belongs to the "false" class. The goal of optimizing the discriminator is to minimize the above loss function. Through the backpropagation algorithm, the parameters (weights and biases) of the discriminator are optimized so that the discriminator can accurately distinguish real data and generated data.
[0087] To optimize the generation effect of GANs, the parameters of the generation model are updated regularly, and the newly collected multi-modal data is used for retraining to improve the diversity and quality of the generated samples. Through a continuous feedback mechanism, the effects of feature extraction and data augmentation are continuously optimized.
[0088] 3. Develop a multi-modal deep learning pipeline. The multi-modal deep learning pipeline includes a feature-level fusion strategy and a decision-level fusion strategy.
[0089] The feature-level fusion strategy is to first extract features from the data of each modality, align the features of different modalities in a certain space or dimension, and then input them into a joint model. The advantage of feature-level fusion is that it can retain the detailed information of each modality, and in deep learning, the correlation between modalities can be extracted by sharing the network structure. Therefore, this method integrates feature-level fusion. The feature-level fusion strategy includes:
[0090] Feature alignment: Create a shared latent space based on a fully connected neural network so that the features of different modalities can be aligned in the same space. The alignment effect is optimized through a loss function (minimizing the reconstruction error) to ensure the maximization of the correlation between modalities. And, during the alignment process, regularization techniques (L2 regularization) are introduced to prevent overfitting.
[0091] Deep Fusion Model: Design a multimodal joint model, a multimodal Transformer, embed this deep fusion model as a built-in model into the AI training platform, and at the same time support users to customize the model. Use the aligned features as input for further processing. The specific network architecture of the multimodal joint model is as follows:
[0092] The input layer of the multimodal Transformer receives the aligned feature vectors, performs non-linear transformations through multiple fully connected layers, and finally outputs the prediction results. In order to retain the original feature information, skip connections are introduced between different layers to improve the learning ability of the model. The self-attention layer can help the model focus on the important parts of the features in different modalities, enabling the model to adaptively adjust the weights of each modality feature.
[0093] The decision-level fusion strategy is to process the data of each modality separately and then make a comprehensive decision on their respective output results. Therefore, this method integrates the decision-level fusion strategy. The decision-level fusion strategy includes:
[0094] Independent Model Training: Design independent models for each modality. Use CNN for images, BERT for text, and LSTM for other data. Embed the trained models as built-in basic models into the AI training platform, and at the same time users can choose to customize the basic models. Each model is independently trained on the modality data of its own to capture modality-specific features. After training, each independent model makes predictions on the input data and generates corresponding output results (such as classification probabilities).
[0095] Weighted Voting: Set the weights of each modality model and combine the outputs of the models. The weights are adjusted based on the performance of the models on the validation set. For example, a model with a higher accuracy rate accounts for a greater proportion in the final decision.
[0096] The multimodal data processing method with feature-level fusion strategy and decision-level fusion strategy not only theoretically helps to improve the performance and generalization ability of the model, but also can reduce costs, improve efficiency, and enhance the user experience in practical applications. Generally speaking, this method has broad application prospects and potential, and is of great significance to promoting the development and application of multimodal data processing technology.
[0097] The joint modeling at a specific level described in this method refers to the fusion of data of different modalities (images, text, time series data) at a specific feature level during the multimodal data processing. The feature level mentioned in this method refers to the features processed by multimodal feature extraction. These features are no longer direct pixel or word vectors, but data that reflects their semantic information.
[0098] A joint model refers to a model that combines data from multiple modalities (such as images, text, time series) for training. The joint model mentioned in this method refers to a multi-modal Transformer model. This multi-modal model is built based on the Transformer infrastructure, and its network structure is as follows: First is the multi-modal feature input layer, which unifies the features of different modalities to the same dimension through a fully connected layer. For image features, a 2048-dimensional feature vector is converted to 768 dimensions; for text features, the text is processed by BERT to obtain 768-dimensional features; for time series features, the 256-dimensional features obtained through RNN processing are converted to 768 dimensions (the common and effective dimension setting size of the Bert model). Then comes the 12-layer ENCODER structure of the Transformer model. Each ENCODER layer includes a self-attention mechanism, a feed-forward neural network, and residual connections and layer normalization. Next, a fusion layer is set up to combine the features of different modalities after being encoded by the Encoder using the self-attention mechanism to obtain a joint multi-modal representation, so as to make full use of the complementary information between different modalities and highlight more relevant and important information. Next, based on the fused multi-modal representation, the 12-layer Decoder layer generates the output required for the task. Among them, the structure of the Decoder layer is similar to that of the Encoder, and it is also based on self-attention and a feed-forward neural network. And, in this method, skip connections (such as residual connections) are introduced between different layers of the Encoder, so that the gradient can be directly passed to the bottom layer, alleviating the optimization problem of deep networks. At the same time, the skip connections can also fuse features at different levels to obtain a richer representation. Finally, there is a 3-layer fully connected output layer with 128 nodes, 512 nodes, 256 nodes, and 3 nodes (taking the classification task as an example). Among them, the activation function of the network layer uses softmax.
[0099] An embodiment of the present invention also provides a multi-modal data augmentation system based on an AI training platform, including:
[0100] A multi-modal data fusion pipeline for processing data sources of different modalities and synchronizing, aligning, and fusing the data according to the actual application scenario;
[0101] A data augmentation pipeline for generating new samples using the existing multi-modal data through a generative adversarial network (GANs) generation model:
[0102] A multi-modal deep learning pipeline, including a feature-level fusion strategy and a decision-level fusion strategy, for fully learning and understanding the deep features of different modality data;
[0103] This system specifically implements multi-modal data augmentation through the multi-modal data augmentation method based on the AI training platform described in the above embodiments.
[0104] First, develop a multi-modal data fusion pipeline on the AI training platform. Collect data of different modalities (such as images, texts, sensor data, etc.) through the data acquisition module and perform preprocessing. In the preprocessing stage, use the multi-modal data fusion pipeline to extract features from data of different modalities and perform joint modeling at a specific level.
[0105] Then, adopt a data augmentation pipeline to automatically generate high-quality synthetic data.
[0106] Finally, develop a multi-modal deep learning pipeline to fully learn and understand the deep features of data of different modalities.
[0107] The multi-modal data fusion pipeline includes the following three steps:
[0108] (1) Data standardization: Standardize the multi-modal data (images, texts, sensor data, voices) collected from different data sources.
[0109] For image data: Perform scaling and normalization, supporting two normalization methods: The first is to adjust the image pixel values from the range of 0-255 to 0-1; the second is to normalize according to specific mean and standard deviation, using the normalization values of the ImageNet dataset: mean [0.485, 0.456, 0.406], standard deviation [0.229, 0.224, 0.225].
[0110] For text data: Perform word segmentation and encoding, use natural language processing toolkits (SpaCy, NLTK) for word segmentation, and use models such as BERT and Word2Vec to convert the segmented text into word embedding vectors.
[0111] For time series data: Perform interpolation and alignment processing to ensure that all data can be analyzed on the same time axis or the same dimension.
[0112] (2) Temporal synchronization and alignment: Synchronize and align time series data (sensor data) and non-time series data (images, texts) to ensure the effective fusion of multi-modal data within the same time period. The specific approach is to slice the entire time series into multiple time periods, each time period containing data of multiple modalities, and missing time points can be filled by interpolation.
[0113] (3) Data feature extraction: The multi-modal feature extraction model includes convolutional neural network (CNN), recurrent neural network (RNN) and Transformer.
[0114] Image data feature extraction: Use the pre-trained convolutional neural network (CNN) ResNet model to extract high-level features from images. The pre-trained model has been trained on a large image dataset and can be directly applied to feature extraction. After the last convolutional layer, global average pooling is used to convert the feature map into a fixed-length feature vector, which is further used for multimodal fusion.
[0115] Text data feature extraction: Use the pre-trained language model Transformer to encode the text data and extract the semantic feature vectors of the text. Through pre-training and fine-tuning, the model can understand the semantics and structure of the text. Then, each tokenized word or sub-word is converted into a fixed-length vector, and the output feature vector can represent the context relationship and semantic information in the text.
[0116] Sensor data feature extraction: Use the recurrent neural network RNN to extract the time series features in the sensor data. For time series data, the recurrent neural network RNN can capture the time-dependent relationships and is suitable for the feature extraction of time series data.
[0117] For the multimodal data augmentation pipeline, train a GAN model using existing multimodal data as input. The generator generates new samples, while the discriminator evaluates the authenticity of the generated samples. Through continuous iterative training, the generator can learn the distribution of real data and generate new samples similar to the real samples. For example, new images, texts, or time series can be generated to expand the dataset. The generated samples will be used as additional training data and input into the model to improve the generalization ability of the model.
[0118] To optimize the generation effect of GANs, regularly update the parameters of the generation model and re-train it using newly collected multimodal data to improve the diversity and quality of the generated samples. Through a continuous feedback mechanism, continuously optimize the effects of feature extraction and data augmentation.
[0119] The multimodal deep learning pipeline includes a feature-level fusion strategy and a decision-level fusion strategy:
[0120] The feature-level fusion strategy is to first extract the features of the data in each modality, align the features of different modalities in a certain space or dimension, and then input them into a joint model. The advantage of feature-level fusion is that it can retain the detailed information of each modality and can extract the correlations between modalities through sharing network structures in deep learning. Therefore, this method integrates feature-level fusion. The feature-level fusion strategy includes:
[0121] Feature Alignment: Create a shared latent space based on a fully connected neural network to enable features of different modalities to be aligned within the same space. Optimize the alignment effect through a loss function (minimizing the reconstruction error) to ensure the maximization of the correlation between modalities. Moreover, during the alignment process, introduce a regularization technique (L2 regularization) to prevent overfitting.
[0122] Deep Fusion Model: Design a multi-modal joint model, a multi-modal Transformer, embed this deep fusion model as a built-in model into the AI training platform, and at the same time support users to customize the model. Use the aligned features as input for further processing. The specific network architecture of the multi-modal joint model is as follows:
[0123] The input layer of the multi-modal Transformer receives the aligned feature vectors, performs non-linear transformations through multiple fully connected layers, and finally outputs the prediction results. To retain the original feature information, skip connections are introduced between different layers to enhance the learning ability of the model. The self-attention layer can help the model focus on the important parts of the features in different modalities, enabling the model to adaptively adjust the weights of the features of each modality.
[0124] The decision-level fusion strategy is to process the data of each modality separately and then make a comprehensive decision on the respective output results. Therefore, this method integrates the decision-level fusion strategy. The decision-level fusion strategy includes:
[0125] Independent Model Training: Design independent models for each modality. Use CNN for images, BERT for text, and LSTM for other data. Embed the trained models as built-in basic models into the AI training platform, and at the same time users can choose to customize the basic models. Each model is independently trained on the modality data of its own to capture modality-specific features. After training, each independent model makes predictions on the input data to generate corresponding output results (such as classification probabilities).
[0126] Weighted Voting: Set the weights of each modality model and combine the outputs of the models. The weights are adjusted based on the performance of the models on the validation set. For example, a model with a higher accuracy rate accounts for a greater proportion in the final decision.
[0127] An embodiment of the present invention also provides a multi-modal data augmentation device based on an AI training platform, which is characterized by including: at least one memory and at least one processor;
[0128] The at least one memory is used to store machine-readable programs;
[0129] The at least one processor is used to call the machine-readable programs to implement the multi-modal data augmentation method based on the AI training platform described in the above embodiments.
[0130] An embodiment of the present invention also provides a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor is caused to execute the multi-modal data augmentation method based on the AI training platform described in the above embodiments. Specifically, a system or device equipped with a storage medium can be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer (or CPU or MPU) of the system or device is caused to read and execute the program code stored in the storage medium.
[0131] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.
[0132] Embodiments of the storage medium for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code can be downloaded from a server computer via a communication network.
[0133] In addition, it should be clear that not only can the functions of any one of the above embodiments be achieved by executing the program code read by the computer, but also by causing an operating system or the like operating on the computer based on the instructions of the program code to complete part or all of the actual operations.
[0134] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion unit is caused to execute part and all of the actual operations, thereby implementing the functions of any one of the above embodiments.
[0135] The present invention has been described in detail above with reference to the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above-mentioned multiple embodiments, those skilled in the art can know that more embodiments of the present invention can be obtained by combining the code review means in the above different embodiments, and these embodiments are also within the protection scope of the present invention.
Claims
1. A multi-modal data augmentation method based on an AI training platform, characterized in that, The implementation of this method includes: First, develop a multi-modal data fusion pipeline on the AI training platform. Collect data of different modalities through the data acquisition module and perform preprocessing. In the preprocessing stage, use the multi-modal data fusion pipeline to extract features from data of different modalities and perform joint modeling at a specific level; Then, adopt a data augmentation pipeline to automatically generate high-quality synthetic data; Finally, develop a multi-modal deep learning pipeline to fully learn and understand the deep features of different modality data; The multi-modal data fusion pipeline normalizes the multi-modal data collected from different data sources, including: For image data: perform scaling and normalization, supporting two normalization methods: The first is to adjust the image pixel values from the range of 0-255 to 0-1; the second is to normalize according to specific mean and standard deviation, using the normalization values of the ImageNet dataset: mean [0.485, 0.456, 0.406], standard deviation [0.229, 0.224, 0.225]; For text data: perform word segmentation and encoding. Use the natural language processing toolkit for word segmentation, and use the BERT and Word2Vec models to convert the segmented text into word embedding vectors; For time series data: perform interpolation and alignment processing to ensure that all data can be analyzed on the same time axis or the same dimension; Time series synchronization and alignment are specifically implemented as follows: Slice the entire time series into multiple time periods. Each time period contains data of multiple modalities, and missing time points can be filled by interpolation; Data feature extraction specifically includes: Image data feature extraction: Use the pre-trained convolutional neural network ResNet model to extract high-level features from images; after the convolutional layer, use global average pooling to convert the feature map into a fixed-length feature vector for further multi-modal fusion; Text data feature extraction: Use the pre-trained language model Transformer to encode text data and extract semantic feature vectors of the text; then, convert each segmented word or sub-word into a fixed-length vector, and the output feature vector can represent the context relationship and semantic information in the text; Sensor data feature extraction: Use a recurrent neural network to extract time series features in sensor data; The data augmentation pipeline generates new samples using the existing multi-modal data through the generative adversarial network generation model: Train the GAN model, using the existing multi-modal data as input; the generator generates new samples, while the discriminator evaluates the authenticity of the generated samples; through continuous iterative training, the generator can learn the distribution of real data and generate new samples similar to real samples; the generated samples will be input into the model as additional training data; The multi-modal deep learning pipeline includes a feature-level fusion strategy and a decision-level fusion strategy; The feature-level fusion strategy is to first extract features from the data of each modality, align the features of different modalities in a specified space or dimension, and then input them into a joint model; it includes feature alignment and a deep fusion model; The feature alignment: Create a shared latent space based on a fully connected neural network to enable features of different modalities to be aligned within the same space; Optimize the alignment effect through a loss function to ensure the maximization of the correlation between modalities; And, during the alignment process, introduce regularization techniques to prevent overfitting; The deep fusion model: Design a multi-modal joint model, a multi-modal Transformer, embed this deep fusion model as a built-in model into the AI training platform, and at the same time support users to customize the model; Use the aligned features as input for further processing; The specific network architecture of the multi-modal joint model is as follows: The input layer of the multi-modal Transformer receives the aligned feature vectors, undergoes non-linear transformation through multiple fully connected layers, and finally outputs the prediction results; Skip connections are introduced between different layers; The self-attention layer is used to enable the model to focus on the important parts of the features in different modalities, so that the model can adaptively adjust the weights of the features of each modality; The decision-level fusion strategy is to make a comprehensive decision on the output results after processing the data of each modality separately, including independent model training and weighted voting; The independent model training: Design independent models for each modality: use CNN for images, BERT for text, and LSTM for other data; Embed the trained models as built-in basic models into the AI training platform, and at the same time users can choose to customize the basic models; Each model is independently trained on the modality data of its own to capture modality-specific features; After training, each independent model makes predictions on the input data and generates corresponding output results; The weighted voting: Set the weights of each modality model and combine the outputs of the models; The weights are adjusted based on the performance of the models on the validation set.
2. The multimodal data augmentation method based on an AI training platform according to claim 1, wherein The multi-modal data fusion pipeline is used to process data sources of different modalities and synchronize, align, and fuse the data according to the actual application scenario; The multi-modal data fusion pipeline includes the following steps: (1) Data standardization: Standardize the multi-modal data collected from different data sources; (2) Temporal synchronization and alignment: Synchronize and align time series data and non-time series data to ensure the effective fusion of multi-modal data within the same time period; (3) Data feature extraction: The multi-modal feature extraction model includes a convolutional neural network, a recurrent neural network, and a Transformer.
3. A multimodal data augmentation method based on an AI training platform according to claim 1, characterized in that, Regularly update the parameters of the generation model and retrain it with newly collected multi-modal data to improve the diversity and quality of the generated samples; Through a continuous feedback mechanism, continuously optimize the effects of feature extraction and data augmentation.
4. A multi-modal data augmentation system based on an AI training platform, characterized in that, Include: The multi-modal data fusion pipeline is used to process data sources of different modalities and synchronize, align, and fuse the data according to the actual application scenario; The data augmentation pipeline is used to generate new samples by using the existing multi-modal data through a generative adversarial network generation model: The multi-modal deep learning pipeline includes a feature-level fusion strategy and a decision-level fusion strategy to fully learn and understand the deep features of different modality data; The system specifically implements multimodal data augmentation through the method described in any one of claims 1 to 3.
5. A multi-modal data augmentation device based on an AI training platform, characterized in that, It includes: At least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program to implement the method described in any one of claims 1 to 3.
6. A computer-readable medium, characterized in that, Computer instructions are stored on the computer-readable medium, and when the computer instructions are executed by a processor, the method described in any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
Intelligent data enhancement method and device based on generative adversarial network and multi-modal data, and medium
CN118606715A
Quality safety optimization method and system based on multi-modal vertical large model technology
CN119130268A