Multi-modal data management method based on large model technology

Through the multimodal data governance method of large-modal technology, network crawlers and feature extraction models are used to build multimodal feature space, solving the problem of format, structure and semantic differences in multimodal data governance, and achieving efficient and intelligent data governance.

CN120296264APending Publication Date: 2025-07-11SHOUXIN TONGLIAN (BEIJING) DATA TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510229146.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing multimodal data governance methods are difficult to uniformly process different types of multimodal data, and there is noise, missing values and inconsistencies, and rely on manual operation efficiency and high probability of errors.

Method used

The multimodal data governance method based on big model technology is adopted to collect data through network crawlers, deduplication, error correction and data conversion are performed, and feature vectors are extracted using models such as BERT, convolutional neural network and WaveNet, and feature space is constructed by combining attention mechanisms and multimodal Transformer models, and abnormal detection and storage are used for large models.

Benefits of technology

It realizes accurate capture and unified processing of multimodal data, improves data quality and governance efficiency, reduces manual intervention, and enhances intelligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296264A_ABST
    Figure CN120296264A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal data management method based on a large model technology, and relates to the technical field of combination of artificial intelligence and a large data technology, and the method comprises the following steps: collecting multi-modal data of a target body through a web crawler, processing the collected multi-modal data, extracting features in the processed multi-modal data, and storing the features in a database; a multi-modal feature vector is formed through combination of an attention mechanism and a weighted summation method, association between different modal features is learned by using a multi-modal auto-encoder and a multi-modal Transform model, a multi-modal feature space is constructed, and a multi-modal feature vector is obtained. According to the method, a web crawler acquisition technology, a multi-modal feature extraction and fusion technology and a modern information technology are closely combined, multi-modal data such as texts, images, audios and videos are accurately captured, and high-quality multi-modal feature vectors and a unified multi-modal feature space are obtained; and real-time and comprehensive monitoring of the whole process from acquisition to processing and storage of the multi-modal data is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of the combination of artificial intelligence and big data technologies, and specifically relates to a multi-modal data governance method based on large model technology. Background Art

[0002] In today's digital age, data has become the core asset of enterprises and organizations. The importance of data governance is self-evident, which is directly related to data quality, security and compliance, and further affects the accuracy and efficiency of data-driven decision-making. With the rapid development of information technology, a large amount of multi-modal data has emerged. Due to the significant differences in format, structure and semantics, traditional governance methods limited to single data types are difficult to cope with. At the same time, large model technologies, especially models based on the Transformer architecture, have flourished and demonstrated powerful potential in processing large-scale, multi-modal data, and can capture complex data relationships. However, when applying it to multi-modal data governance, many problems such as data integration and processing, security and privacy protection, and improvement of automation efficiency are faced. In this context, it is urgent to develop an effective multi-modal data governance method based on large model technology;

[0003] Although existing multi-modal data governance methods based on large model technology have made great progress, there are still some problems to be optimized. There are many types of multi-modal data, and there are huge differences in format, structure and semantics among different types of multi-modal data, which are difficult to uniformly process; secondly, multi-modal data may have noise, missing values and inconsistencies, affecting the reliability and effectiveness of data. Moreover, traditional data governance processes often rely on manual operations, with low efficiency and high error probability. Summary of the Invention

[0004] To achieve the above objectives, the present invention is realized through the following technical solutions: A multi-modal data governance method based on large model technology, including the following steps:

[0005] Step 1: Use web crawlers to collect multi-modal data of the target body to provide data support for subsequent steps;

[0006] Step 2: Perform preprocessing on the collected multi-modal data, including data deduplication, data error correction and removal of irrelevant data, and then perform data conversion and data annotation on the preprocessed multi-modal data;

[0007] Step 3: Extract features from the text data after data conversion and data annotation, learn the semantic representation of the text, and obtain text feature vectors to learn the semantic representation of the text;

[0008] Step 4: Extract features from the image data after data conversion and data annotation, learn the visual representation of the image, and obtain image feature vectors;

[0009] Step Five: Extract features from the audio data after data conversion and data annotation, learn the acoustic representation of the audio, and obtain the audio feature vector;

[0010] Step Six: Extract features from the video data after data conversion and data annotation, learn the spatio-temporal representation of the video, and obtain the video spatio-temporal feature vector;

[0011] Step Seven: Combine the attention mechanism with the weighted summation method to integrate the feature vectors of multi-modal data, form a multi-modal feature vector, use the multi-modal autoencoder and multi-modal Transformer model to learn the associations between different modal features, and construct a multi-modal feature space;

[0012] Step Eight: Utilize the anomaly detection ability of the large model to automatically identify and correct errors and inconsistencies in the data, evaluate the data quality through the large model, including accuracy, integrity, consistency, and timeliness, screen out the low-quality data, and store the multi-modal data using the object storage architecture;

[0013] Step Nine: Fine-tune and update the large model, continuously monitor the performance of the large model to ensure the stability and reliability of the model in actual applications, and integrate it into the existing data management system.

[0014] A further improvement of the technical solution of the present invention lies in: In the above Step One, the process of using a web crawler to collect multi-modal data of the target body to provide data support for subsequent steps includes:

[0015] The multi-modal data of the target body is text data, image data, audio data, and video data;

[0016] Through the web crawler, locate the text elements of the target body, identify the image HTML tags in the target body, search for the HTML5 tags of audio in the target body, and the video HTML5 tags in the target body;

[0017] For the located text elements, the web crawler extracts the text data within the HTML tags;

[0018] For the identified image HTML tags, the web crawler parses the source link of the image from the image HTML tags and obtains the image data through a network request;

[0019] For the searched HTML5 tags of audio, the web crawler extracts the link address of the audio file and obtains the audio data through downloading;

[0020] For the searched video HTML5 tags, the web crawler parses the link of the video file and obtains the video data through downloading;

[0021] A further improvement of the technical solution of the present invention lies in that: in the second step, the preprocessing of the collected multimodal data for data deduplication, data error correction, and removal of irrelevant data, and then the process of data conversion and data annotation of the preprocessed multimodal data includes:

[0022] Based on the hash algorithm, calculate the hash values of text data, image data, audio data, and video data respectively, and determine the duplication situation of multimodal data by comparing the hash values. When the hash values of two data are the same, it indicates that there are duplicate multimodal data, and then perform deduplication processing on the duplicate multimodal data;

[0023] Use the spelling check and grammar check tools and business term dictionaries in natural language processing to correct spelling mistakes and grammar errors in the text data; check the resolution and damage degree of the image data, repair the damaged images through image repair technology, and perform image enhancement processing on the image data with low resolution; detect the noise situation, abnormal volume situation, and incomplete clip situation of the audio data, and process the audio data with noise, abnormal volume audio data, and incomplete clip audio data through audio noise reduction, volume adjustment, and clip repair respectively;

[0024] Annotate and classify the collected multimodal data, classify and annotate the multimodal data through machine learning algorithms, divide the multimodal data related to the target object and the multimodal data not related to the target object, and remove the multimodal data not related to the target object according to the requirements of the target object.

[0025] Convert text data, image data, audio data, and video data into the data processing format of the large model through text tokenization, image scaling, audio sampling rate conversion, and video frame extraction respectively, and perform data annotation on the multimodal data after data conversion processing.

[0026] A further improvement of the technical solution of the present invention lies in that: in the third step, the process of feature extraction of the text data after data conversion and data annotation, learning the semantic representation of the text, and obtaining the text feature vector includes:

[0027] Select the BERT pre-trained language model. The BERT pre-trained language model is composed of Transformer layers. Add a dynamic adjustment layer to the BERT pre-trained language model. Input the text data after text tokenization into the BERT pre-trained language model. Based on the bidirectional encoding mechanism, the Transformer layer performs layer-by-layer encoding operations on the text. The dynamic adjustment layer monitors the encoding operation results of the text data in real time and adjusts the parameters of the BERT pre-trained language model. After the collaborative processing of the text data by the Transformer layer and the dynamic adjustment layer, the BERT pre-trained language model outputs the text feature vector extracted from the text data, and this text feature vector is the text semantics.

[0028] A further improvement of the technical solution of the present invention lies in that: in the fourth step, for the image data after data conversion and data annotation, feature extraction is performed to learn the visual representation of the image, and the process of obtaining the image feature vector includes:

[0029] Use a convolutional neural network to perform feature extraction on the image data. The convolutional neural network consists of a traditional convolutional layer, an activation function, a pooling layer, and a fully connected layer. Input the image data, and after convolutional operation, output an offset. Introduce this offset into the convolutional layer to become a deformable convolutional layer. Among them, the convolutional layer has a built-in convolutional kernel;

[0030] Input the image data after image scaling into the convolutional neural network. The traditional convolutional layer extracts the basic feature map of the image data; then input the basic feature map into the deformable convolutional layer. The convolutional kernel uses the offset to perform deformation and sampling operations on the basic feature data of the image, and extracts the detailed feature map of the image data; pass the detailed feature map through the activation function for non-linear transformation, learn the relationship between the image feature maps, perform dimensionality reduction operations through the pooling layer, refine the detailed feature map of the image data, and then the fully connected layer integrates the detailed feature map of the image data and outputs the final image feature vector. This image feature vector represents the visual representation of the image.

[0031] A further improvement of the technical solution of the present invention lies in that: in the fifth step, for the audio data after data conversion and data annotation, feature extraction is performed to learn the acoustic representation of the audio, and the process of obtaining the audio feature vector includes:

[0032] Extract the feature data in the audio data through the WaveNet model. The extraction process of the WaveNet model includes four steps: causal convolution, dilated convolution, gated activation unit non-linear transformation, and feature aggregation and output. Among them, the convolutional layer has a built-in convolutional kernel;

[0033] Input the audio data after audio sampling rate conversion into the WaveNet model and enter the causal convolutional layer. Based on the input of the audio data at past and current moments, calculate the output of the audio data at the current moment. As the number of layers of the causal convolutional layer increases, extract the local feature data of the audio data at different time scales;

[0034] Apply dilated convolution, insert blank elements between the elements of the convolutional kernel to increase the receptive field of the convolutional kernel, and extract the feature data of the audio data over a long time span;

[0035] After convolutional operation, use a connection gate activation unit that combines the Sigmoid activation function and the Tanh activation function to perform non-linear transformation on the local feature data of the audio data at different time scales and the feature data of the audio data over a long time span;

[0036] The local feature data of the audio data at different time scales and the feature data of the audio data at a long time span extracted after causal convolution, dilated convolution and nonlinear transformation are aggregated and output to obtain the final audio feature vector, which is the acoustic representation of the audio data.

[0037] A further improvement of the technical solution of the present invention is that in step 6, the process of extracting features from the video data after data conversion and data annotation, learning the spatiotemporal representation of the video, and obtaining the spatiotemporal feature vector of the video includes:

[0038] The convolutional neural network and the 3D-CNN model are combined to extract features from video data. Both the convolutional neural network and the 3D-CNN model are composed of convolutional layers, activation functions, pooling layers and fully connected layers. The video data after video frame extraction is input into the convolutional neural network. Through the collaborative processing of the convolutional layer, activation function and pooling layer, the spatial features of each frame image are integrated and output through the fully connected layer of the convolutional neural network to extract the spatial features of the video data. The 3D-CNN model is then used to perform convolution operations on the spatial and temporal dimensions of the spatial features of the video data extracted by the neural network to mine the spatiotemporal features in the video data. Finally, the extracted spatiotemporal features are integrated through the fully connected layer in the 3D-CNN model to output the video spatiotemporal feature vector, which is the spatiotemporal representation of the video.

[0039] A further improvement of the technical solution of the present invention is that in step seven, the process of acquiring the multimodal feature vector, the association between different modal features and the multimodal feature space includes:

[0040] By combining the attention mechanism with the weighted summation method, the feature data of multimodal data is integrated to form a multimodal feature vector. The multimodal autoencoder and multimodal Transformer model are used to learn the association between different modal features and construct a multimodal feature space.

[0041] Introduce the attention mechanism, initialize and set the learnable query vector through the dot product attention method, perform dot product operations on the text feature vector, image feature vector, audio feature vector and video spatiotemporal feature vector and the learnable query vector respectively, and obtain the attention scores of the text feature vector, image feature vector, audio feature vector and video spatiotemporal feature vector. The attention scores are the weights of the text feature vector, image feature vector, audio feature vector and video spatiotemporal feature vector.

[0042] The multimodal feature vector is obtained by weighted summation method. The weighted summation process is as follows:

[0043] V multi =w t ×V t +wi ×V i +w a ×V a +w p ×V p

[0044] Among them, V multi is a multi-modal feature vector, and V t , V i , V a and V p are the text feature vector, image feature vector, audio feature vector, and video spatio-temporal feature vector respectively, and w t , w i , w a and w p are the weights of the text feature vector, image feature vector, audio feature vector, and video spatio-temporal feature vector respectively;

[0045] By combining the multi-modal autoencoder and the multi-modal Transformer model, the associations between different modal features are learned, and a multi-modal feature space is constructed;

[0046] The multi-modal autoencoder consists of an encoder and a decoder. Among them, the encoder includes the text encoder of LSTM, the CNN image encoder, the WaveNet audio encoder, and the video encoder combining 3D-CNN and LSTM. In the encoding stage, the text feature vector is input into the text encoder of LSTM, and through calculation, the text feature vector is encoded into a low-dimensional text intermediate representation; the CNN image encoder is used to encode the image feature vector into an image intermediate representation; the WaveNet audio encoder is used to encode the audio feature vector into an audio intermediate representation; the video spatio-temporal feature vector is encoded by the video encoder combining 3D-CNN and LSTM. Among them, 3D-CNN extracts the features of the spatial dimension and the time dimension in the video spatio-temporal feature vector, and LSTM further processes the video spatio-temporal feature vector to capture the time series dependencies in the video and obtain the video intermediate representation;

[0047] Through the splicing operation, the text intermediate representation, the image intermediate representation, the audio intermediate representation, and the video intermediate representation are fused to obtain the multi-modal data intermediate vector, and then the multi-modal data intermediate vector is input into the decoder to reconstruct the multi-modal data intermediate vector, minimize the reconstruction error, and learn the associations between different modal features;

[0048] Taking the text feature vector, image feature vector, audio feature vector, and video spatio-temporal feature vector as input data, inputting them into the multi-modal Transformer model. The multi-modal Transformer model encodes and integrates the input data according to the associations between different modal features to obtain an input sequence, processes the input sequence through the multi-head attention mechanism to calculate the weight distribution of different attentions, uses a feed-forward neural network to perform a non-linear transformation on the input sequence processed by the multi-head attention mechanism, and then performs layer normalization and residual connection processing on the input sequence. The multi-modal Transformer model outputs a feature vector sequence that fuses the association information of different modalities. Based on the feature vector sequence that fuses the association information of different modalities, a multi-modal feature space is constructed.

[0049] A further improvement of the technical solution of the present invention lies in: in the eighth step, the process of using the anomaly detection ability of the large model to identify and correct errors and inconsistencies in the data and storing the multi-modal data using an object storage architecture includes:

[0050] The large model comprehensively scans the multi-modal data to identify anomalies in the multi-modal data. For anomalies in text data, the large model corrects text errors through semantic analysis technology and language generation technology; for anomalies in image data, the large model uses an image correction algorithm to improve image errors; for anomalies in audio data, the large model uses an audio processing algorithm to correct audio errors; for anomalies in video data, the large model corrects video errors through a video synthesis algorithm.

[0051] The corrected text data, image data, audio data, and video data are stored in object form using an object storage architecture. Each object contains the data itself, metadata of detailed data attributes, and a global identifier.

[0052] A further improvement of the technical solution of the present invention lies in: in the ninth step, the process of fine-tuning and updating the large model, real-time monitoring of the performance of the large model, and integrating it into the existing data management system includes:

[0053] New multi-modal data is regularly collected, and the large model is fine-tuned and updated using the new multi-modal data. With the help of performance metric evaluation means, the performance of the large model is monitored in real time. Among them, the monitored performance metrics include accuracy, recall, mean squared error, confusion matrix, memory occupancy, running time, and throughput, and the large model is integrated into the existing data management system to realize the application of the multi-modal data governance method based on the large model in the data management system.

[0054] Due to the adoption of the above technical solution, the technical progress achieved by the present invention compared with the prior art is as follows: For this multimodal data governance method based on large model technology, compared with the traditional multimodal data governance method based on large model technology, the web crawler acquisition technology, multimodal feature extraction and fusion technology in the method of the present invention are closely combined with modern information technology to accurately capture multimodal data such as text, images, audio, and video, obtain high-quality multimodal feature vectors and a unified multimodal feature space, achieving real-time and comprehensive monitoring of the entire process of multimodal data from acquisition to processing and storage. It solves the problems that there are a wide variety of multimodal data types, and there are huge differences in format, structure, and semantics among different types of multimodal data, making it difficult to uniformly process them. Secondly, multimodal data may have noise, missing values, and inconsistencies, affecting the reliability and effectiveness of the data. Moreover, the traditional data governance process often relies on manual operations, with low efficiency and a high probability of errors. This ensures that the method in the present invention can refine the dynamic monitoring standards of the multimodal data governance method based on large model technology within a more accurate range, making the monitored data more accurate indicators under the same conditions. The research and application of this method significantly enhance the degree of intelligence in the process of multimodal data governance based on large model technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0056] Figure 1 It is a flowchart of a multimodal data governance method based on large model technology of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0058] As Figure 1 shown, the present invention provides a multimodal data governance method based on large model technology, which consists of the following steps:

[0059] Step 1: Use a web crawler to collect multimodal data of the target object to provide data support for subsequent steps;

[0060] Step 2: Preprocess the collected multimodal data by removing duplicates, correcting errors, and removing irrelevant data, and then perform data conversion and data annotation on the preprocessed multimodal data;

[0061] Step 3: Extract features from the text data after data conversion and data annotation, learn the semantic representation of the text, and obtain the text feature vector to learn the semantic representation of the text;

[0062] Step 4: Extract features from the image data after data conversion and data annotation, learn the visual representation of the image, and obtain the image feature vector;

[0063] Step 5: Extract features from the audio data after data conversion and data annotation, learn the acoustic representation of the audio, and obtain the audio feature vector;

[0064] Step 6: Extract features from the video data after data conversion and data annotation, learn the spatio-temporal representation of the video, and obtain the video spatio-temporal feature vector;

[0065] Step 7: Combine the attention mechanism with the weighted summation method to integrate the feature vectors of the multimodal data, form the multimodal feature vector, use the multimodal autoencoder and the multimodal Transformer model to learn the association between different modal features, and construct the multimodal feature space;

[0066] Step 8: Utilize the anomaly detection ability of the large model to automatically identify and correct errors and inconsistencies in the data, evaluate the data quality through the large model, including accuracy, integrity, consistency, and timeliness, filter out the data with poor quality, and store the multimodal data using the object storage architecture;

[0067] Step 9: Fine-tune and update the large model, continuously monitor the performance of the large model to ensure the stability and reliability of the model in actual applications, and integrate it into the existing data management system.

[0068] In Step 1, the process of using the web crawler to collect the multimodal data of the target object and providing data support for the subsequent steps includes:

[0069] The multimodal data of the target object is text data, image data, audio data, and video data;

[0070] Through the web crawler, locate the text elements of the target object, identify the image HTML tags in the target object, find the HTML5 tags of the audio in the target object, and the video HTML5 tags in the target object;

[0071] For the located text elements, the web crawler extracts the text data within the HTML tags;

[0072] For the identified image HTML tags, the web crawler parses the source link of the image from the image HTML tags and obtains the image data through a network request;

[0073] For the found HTML5 tags of audio, the web crawler extracts the link address of the audio file and obtains the audio data through downloading;

[0074] For the found video HTML5 tags, the web crawler parses the link of the video file and obtains the video data through downloading;

[0075] In step two, the process of preprocessing the collected multimodal data by removing duplicate data, correcting data errors, and removing irrelevant data, and then performing data conversion and data annotation on the preprocessed multimodal data includes:

[0076] Based on the hash algorithm, calculate the hash values of the text data, image data, audio data, and video data respectively. By comparing the hash values, determine the duplication situation of the multimodal data. When the hash values of two data are the same, it indicates that there is duplicate multimodal data, and then perform deduplication processing on the duplicate multimodal data;

[0077] Use the spelling inspection, grammar checking tools, and business term dictionaries in natural language processing to perform data error correction processing on the text data for spelling mistakes and grammar errors; check the resolution and damage degree of the image data, repair the damaged images through image repair technology, and perform image enhancement processing on the image data with low resolution; detect the noise situation, abnormal volume situation, and incomplete clip situation of the audio data, and process the audio data with noise, abnormal volume, and incomplete clips through audio noise reduction, volume adjustment, and clip repair respectively;

[0078] Perform annotation and classification on the collected multimodal data, classify and annotate the multimodal data through machine learning algorithms, divide the multimodal data related to the target body and the multimodal data not related to the target body, and remove the irrelevant multimodal data according to the requirements of the target body.

[0079] Convert the text data, image data, audio data, and video data into the data processing format of the large model through text tokenization, image scaling, audio sampling rate conversion, and video frame extraction respectively, and perform data annotation on the multimodal data after data conversion processing.

[0080] In step three, the process of extracting features from the text data after data conversion and data annotation, learning the semantic representation of the text, and obtaining the text feature vector includes:

[0081] Select the BERT pre-trained language model. The BERT pre-trained language model consists of Transformer layers. A dynamic adjustment layer is added to the BERT pre-trained language model. The text data after text tokenization is input into the BERT pre-trained language model. Based on the bidirectional encoding mechanism, the Transformer layers perform layer-by-layer encoding operations on the text. The dynamic adjustment layer monitors the encoding operation results of the text data in real time and adjusts the parameters of the BERT pre-trained language model. After the collaborative processing of the text data by the Transformer layers and the dynamic adjustment layer, the BERT pre-trained language model outputs the text feature vectors extracted from the text data, and this text feature vector represents the text semantics.

[0082] In step four, the process of extracting features from the image data after data conversion and data annotation to learn the visual representation of the image and obtain the image feature vector includes:

[0083] Use a convolutional neural network to extract features from the image data. The convolutional neural network consists of traditional convolutional layers, activation functions, pooling layers, and fully connected layers. Input the image data and output the offset after convolutional operations. Introduce this offset into the convolutional layer to become a deformable convolutional layer. Among them, the convolutional layer has a built-in convolutional kernel;

[0084] Input the image data after image scaling into the convolutional neural network. The traditional convolutional layer extracts the basic feature map of the image data; then input the basic feature map into the deformable convolutional layer. The convolutional kernel uses the offset to perform deformation and sampling operations on the basic feature data of the image to extract the detailed feature map of the image data; pass the detailed feature map through the activation function for non-linear transformation to learn the relationship between the image feature maps, perform dimensionality reduction operations through the pooling layer to refine the detailed feature map of the image data, and then the fully connected layer integrates the detailed feature map of the image data and outputs the final image feature vector, and this image feature vector represents the visual representation of the image.

[0085] In step five, the process of extracting features from the audio data after data conversion and data annotation to learn the acoustic representation of the audio and obtain the audio feature vector includes:

[0086] Extract the feature data in the audio data through the WaveNet model. The WaveNet model extraction process includes four steps: causal convolution, dilated convolution, gated activation unit non-linear transformation, and feature aggregation and output. Among them, the convolutional layer has a built-in convolutional kernel;

[0087] Input the audio data after audio sampling rate conversion into the WaveNet model and enter the causal convolutional layer. Calculate the output of the current moment audio data based on the input of the audio data at the past moment and the current moment. As the number of causal convolutional layers increases, the local feature data of the audio data at different time scales is extracted;

[0088] Using dilated convolution, blank elements are inserted between the elements of the convolution kernel to increase the receptive field of the convolution kernel and extract the feature data of the audio data over a long time span;

[0089] After the convolution operation, a connection gate activation unit that combines the Sigmoid activation function and the Tanh activation function is used to perform a non-linear transformation on the local feature data of the audio data at different time scales and the feature data of the audio data over a long time span;

[0090] The local feature data of the audio data at different time scales and the feature data of the audio data over a long time span extracted after causal convolution, dilated convolution, and non-linear transformation are aggregated and output to obtain the final audio feature vector, which is the acoustic representation of the audio data.

[0091] In step six, the process of extracting features from the video data after data conversion and data annotation, learning the spatio-temporal representation of the video, and obtaining the video spatio-temporal feature vector includes:

[0092] Feature extraction is performed on the video data by combining a convolutional neural network and a 3D-CNN model. Both the convolutional neural network and the 3D-CNN model are composed of a convolutional layer, an activation function, a pooling layer, and a fully connected layer. The video data after video frame extraction is input into the convolutional neural network. Through the collaborative processing of the convolutional layer, the activation function, and the pooling layer, and then through the fully connected layer of the convolutional neural network, the spatial features of each frame of the image are integrated and output to extract the spatial features of the video data. Then, the 3D-CNN model is used to perform convolutional operations on the spatial features of the video data extracted by the neural network in the spatial and temporal dimensions to mine the spatio-temporal features of the video data. Finally, the spatio-temporal features extracted are integrated through the fully connected layer in the 3D-CNN model, and the video spatio-temporal feature vector is output, which is the spatio-temporal representation of the video.

[0093] In step seven, the process of obtaining the multi-modal feature vector, the associations between different modal features, and the multi-modal feature space includes:

[0094] By combining the attention mechanism and the weighted summation method, the feature data of the multi-modal data is integrated to form a multi-modal feature vector. The multi-modal autoencoder and the multi-modal Transformer model are used to learn the associations between different modal features and construct a multi-modal feature space

[0095] Introduce the attention mechanism. Through dot-product attention, initialize and set the learnable query vector, and perform dot-product operations between the learnable query vector and the text feature vector, image feature vector, audio feature vector, and video spatio-temporal feature vector respectively to obtain the attention scores of the text feature vector, image feature vector, audio feature vector, and video spatio-temporal feature vector. This attention score is the weight of the text feature vector, image feature vector, audio feature vector, and video spatio-temporal feature vector;

[0096] Obtain the multi-modal feature vector through the weighted summation method. The weighted summation process is as follows:

[0097] V multi = w t × V t + w i × V i + w a × V a + w p × V p

[0098] Among them, V multi is the multi-modal feature vector, V t , V i , V a and V p are the text feature vector, image feature vector, audio feature vector, and video spatio-temporal feature vector respectively, and w t , w i , w a and w p are the weights of the text feature vector, image feature vector, audio feature vector, and video spatio-temporal feature vector respectively;

[0099] Combine the multi-modal autoencoder and the multi-modal Transformer model to learn the associations between different modal features and construct a multi-modal feature space;

[0100] The multimodal autoencoder consists of an encoder and a decoder. Among them, the encoder includes a text encoder of LSTM, a CNN image encoder, a WaveNet audio encoder, and a video encoder that combines 3D-CNN and LSTM. In the encoding stage, the text feature vector is input into the text encoder of LSTM, and through calculation, the text feature vector is encoded into a low-dimensional text intermediate representation; the CNN image encoder is used to encode the image feature vector into an image intermediate representation; the WaveNet audio encoder is utilized to encode the audio feature vector into an audio intermediate representation; the video encoder that combines 3D-CNN and LSTM encodes the video spatio-temporal feature vector. Among them, 3D-CNN extracts the features of the spatial dimension and the time dimension in the video spatio-temporal feature vector, and LSTM further processes the video spatio-temporal feature vector to capture the time series dependencies in the video and obtain the video intermediate representation;

[0101] Through the splicing operation, the text intermediate representation, the image intermediate representation, the audio intermediate representation, and the video intermediate representation are fused to obtain the multimodal data intermediate vector, and then the multimodal data intermediate vector is input into the decoder to reconstruct the multimodal data intermediate vector, minimize the reconstruction error, and learn the associations between different modal features;

[0102] Taking the text feature vector, the image feature vector, the audio feature vector, and the video spatio-temporal feature vector as input data, input them into the multimodal Transformer model. The multimodal Transformer model encodes and integrates the input data according to the associations between different modal features to obtain the input sequence, processes the input sequence through the multi-head attention mechanism, calculates the weight distributions of different attentions, uses the feed-forward neural network to perform non-linear transformation on the input sequence processed by the multi-head attention mechanism, and then performs layer normalization and residual connection processing on the input sequence. The multimodal Transformer model outputs a feature vector sequence that fuses the association information of different modalities. Based on the feature vector sequence that fuses the association information of different modalities, a multimodal feature space is constructed.

[0103] In step eight, the process of using the anomaly detection ability of the large model to identify and correct errors and inconsistencies in the data and storing the multimodal data using the object storage architecture includes:

[0104] The large model comprehensively scans the multimodal data to identify anomalies in the multimodal data. For the anomalies in the text data, the large model corrects the text errors through semantic analysis technology and language generation technology; for the anomalies in the image data, the large model uses image correction algorithms to improve the image errors; for the anomalies in the audio data, the large model corrects the audio errors using audio processing algorithms; for the anomalies in the video data, the large model corrects the video errors through video synthesis algorithms;

[0105] The object storage architecture is adopted to store the corrected text data, image data, audio data, and video data in the form of objects. Each object contains the data itself, metadata of detailed data attributes, and a global identifier.

[0106] In Step Nine, the process of fine-tuning and updating the large model, real-time monitoring of the large model's performance, and integration into the existing data management system includes:

[0107] Regularly collect new multimodal data, use the new multimodal data to fine-tune and update the large model, and rely on performance metric evaluation means to real-time monitor the large model's performance. Among them, the monitored performance metrics include accuracy, recall, mean squared error, confusion matrix, memory occupancy, running time, and throughput. Then integrate the large model into the existing data management system to realize the application of the multimodal data governance method based on the large model in the data management system.

[0108] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A multi-modal data governance method based on large model technology, characterized in that: Including the following steps: Step 1: Use a web crawler to collect multimodal data of the target entity; Step 2: Perform preprocessing on the collected multimodal data, including data deduplication, data error correction, and removal of irrelevant data, and then perform data conversion and data annotation on the preprocessed multimodal data; Step 3: Extract features from the text data after data conversion and data annotation to obtain text feature vectors; Step 4: Extract features from the image data after data conversion and data annotation to obtain image feature vectors; Step 5: Extract features from the audio data after data conversion and data annotation to obtain audio feature vectors; Step 6: Extract features from the video data after data conversion and data annotation to obtain video spatio-temporal feature vectors; Step 7: Combine the attention mechanism with the weighted summation method to integrate the feature vectors of multimodal data, form multimodal feature vectors, use multimodal autoencoders and multimodal Transformer models to learn the associations between different modal features, and construct a multimodal feature space; Step 8: Utilize the anomaly detection ability of the large model to identify and correct errors and inconsistencies in the data, and store the multimodal data using an object storage architecture; Step 9: Fine-tune and update the large model, monitor the performance of the large model in real time, and integrate it into the existing data management system.

2. The multimodal data governance method based on large model technology according to claim 1, characterized in that: In the said Step 1, the process of using a web crawler to collect multimodal data of the target entity includes: The multimodal data of the target entity is text data, image data, audio data, and video data; Through the web crawler, locate the text elements of the target entity, identify the image HTML tags in the target entity, search for the HTML5 tags of audio in the target entity, and the video HTML5 tags in the target entity; For the located text elements, the web crawler extracts text data within the HTML tags; For the identified image HTML tags, the web crawler parses the source link of the image from the image HTML tags and obtains the image data through a network request; For the searched HTML5 tags of audio, the web crawler extracts the link address of the audio file and obtains the audio data through downloading; For the searched video HTML5 tags, the web crawler parses the link of the video file and obtains the video data through downloading.

3. The multimodal data governance method based on large model technology according to claim 2, wherein: In the said Step 2, the process of performing data deduplication, data error correction, and removal of irrelevant data on the collected multimodal data, and then performing data conversion and data annotation on the preprocessed multimodal data includes: Based on the hash algorithm, calculate the hash values of text data, image data, audio data, and video data respectively, determine the duplication situation of multimodal data by comparing the hash values. When the hash values of two data are the same, it indicates that there is duplicate multimodal data, and then perform deduplication processing on the duplicate multimodal data; Data error correction processing is performed on text data by using a spelling inspection and grammar checking tool business term dictionary in natural language processing to correct typos and grammar errors; the resolution and damage degree of image data are checked, and damaged images are repaired through image repair technology, and image enhancement processing is performed on image data with low resolution; the noise condition, abnormal volume condition, and incomplete clip condition of audio data are detected, and the audio data with noise, the audio data with abnormal volume, and the audio data with incomplete clips are processed respectively through audio noise reduction, volume adjustment, and clip repair. The collected multimodal data is labeled and classified. The multimodal data is classified and labeled through machine learning algorithms, and the multimodal data related to the target body and the multimodal data unrelated to the target body are divided. According to the requirements of the target body, the unrelated multimodal data is removed. The text data, image data, audio data, and video data are respectively converted into the data processing format of the large model through text tokenization, image scaling, audio sampling rate conversion, and video frame extraction, and the multimodal data after data conversion processing is data-labeled.

4. A multimodal data governance method based on large model technology according to claim 3, characterized in that: In step three, the process of extracting text feature vectors by extracting features from the text data after data conversion and data labeling includes: Select the BERT pre-trained language model. The BERT pre-trained language model is composed of Transformer layers. A dynamic adjustment layer is added to the BERT pre-trained language model. The text data after text tokenization is input into the BERT pre-trained language model. Based on the bidirectional encoding mechanism, the Transformer layer performs layer-by-layer encoding operations on the text. The dynamic adjustment layer monitors the encoding operation results of the text data in real time and adjusts the parameters of the BERT pre-trained language model. After the collaborative processing of the text data by the Transformer layer and the dynamic adjustment layer, the BERT pre-trained language model outputs the text feature vectors extracted from the text data, and the text feature vectors are the text semantics.

5. The multimodal data governance method based on large model technology according to claim 4, wherein: In step four, the process of extracting image feature vectors by extracting features from the image data after data conversion and data labeling includes: Use a convolutional neural network to extract features from the image data. The convolutional neural network is composed of a traditional convolutional layer, an activation function, a pooling layer, and a fully connected layer. The image data is input, and the offset is output after convolutional operations. The offset is introduced into the convolutional layer to become a deformable convolutional layer. Among them, the convolutional layer has a built-in convolutional kernel. The image data after image scaling processing is input into the convolutional neural network. The traditional convolutional layer extracts the basic feature map of the image data; then the basic feature map is input into the deformable convolutional layer, and the convolutional kernel uses the offset to perform deformation and sampling operations on the image basic feature data to extract the detailed feature map of the image data; the detailed feature map undergoes a non-linear change through the activation function to learn the relationship between the image feature maps, and a dimensionality reduction operation is performed through the pooling layer to refine the detailed feature map of the image data, and then the fully connected layer integrates the detailed feature map of the image data to output the final image feature vector, and the image feature vector represents the visual representation of the image.

6. A multimodal data governance method based on large model technology according to claim 5, characterized in that: In the fifth step, the process of extracting feature data from the audio data after data conversion and data annotation to obtain an audio feature vector includes: Extract the feature data in the audio data through the WaveNet model. The extraction process of the WaveNet model includes four steps: causal convolution, dilated convolution, gated activation unit non-linear transformation, and feature aggregation and output. Among them, the convolutional layer has a built-in convolutional kernel; Input the audio data after audio sampling rate conversion into the WaveNet model and enter the causal convolutional layer. Calculate the output of the audio data at the current moment based on the input of the audio data at the past moment and the current moment. As the number of layers of the causal convolutional layer increases, the local feature data of the audio data at different time scales is extracted; Apply dilated convolution to insert blank elements between the elements of the convolutional kernel to increase the receptive field of the convolutional kernel and extract the feature data of the audio data over a long time span; After the convolution operation, use the connection gated activation unit that combines the Sigmoid activation function and the Tanh activation function to perform non-linear transformation on the local feature data of the audio data at different time scales and the feature data of the audio data over a long time span; Aggregate and output the local feature data of the audio data at different time scales and the feature data of the audio data over a long time span extracted after causal convolution, dilated convolution, and non-linear transformation to obtain the final audio feature vector. This audio feature vector is the acoustic representation of the audio data.

7. A multimodal data governance method based on large model technology according to claim 6, characterized in that: In the sixth step, the process of extracting feature data from the video data after data conversion and data annotation to obtain a video spatio-temporal feature vector includes: Extract the feature data of the video data through the combination of a convolutional neural network and a 3D-CNN model. Both the convolutional neural network and the 3D-CNN model are composed of a convolutional layer, an activation function, a pooling layer, and a fully connected layer. Input the video data after video frame extraction into the convolutional neural network. Through the collaborative processing of the convolutional layer, the activation function, and the pooling layer, and then through the fully connected layer of the convolutional neural network, integrate and output the spatial features of each frame of image, and extract the spatial features of the video data. Then use the 3D-CNN model to perform convolutional operations on the spatial features of the video data extracted by the neural network in the spatial and temporal dimensions to mine the spatio-temporal features of the video data. Finally, integrate the spatio-temporal features extracted through the fully connected layer in the 3D-CNN model and output the video spatio-temporal feature vector. This spatio-temporal feature vector is the spatio-temporal representation of the video.

8. The multimodal data governance method based on large model technology according to claim 7, characterized in that: In the seventh step, the process of obtaining the multi-modal feature vector, the association between different modal features, and the multi-modal feature space includes: Integrate the feature data of the multi-modal data through the combination of the attention mechanism and the weighted summation method to form a multi-modal feature vector. Use the multi-modal autoencoder and the multi-modal Transformer model to learn the association between different modal features and construct a multi-modal feature space Introduce the attention mechanism. Through the dot product attention method, initialize and set the learnable query vector, and perform dot product operations on the text feature vector, image feature vector, audio feature vector, and video spatio-temporal feature vector with the learnable query vector respectively to obtain the attention scores of the text feature vector, image feature vector, audio feature vector, and video spatio-temporal feature vector. The attention score is the weight of the text feature vector, image feature vector, audio feature vector, and video spatio-temporal feature vector. Obtain the multi-modal feature vector through the weighted summation method. The weighted summation process is as follows: V multi = w t × V t + w i × V i + w a × V a + w p × V p Among them, V multi is a multi-modal feature vector, and V t , V i , V a and V p are the text feature vector, the image feature vector, the audio feature vector, and the video spatio-temporal feature vector respectively. w t, w i , w a and w p are the weights of the text feature vector, the image feature vector, the audio feature vector, and the video spatio-temporal feature vector respectively; Combine the multi-modal autoencoder and the multi-modal Transformer model to learn the associations between different modal features and construct a multi-modal feature space. The multi-modal autoencoder consists of an encoder and a decoder. Among them, the encoder includes an LSTM text encoder, a CNN image encoder, a WaveNet audio encoder, and a video encoder combining 3D-CNN and LSTM. In the encoding stage, input the text feature vector into the LSTM text encoder, and encode the text feature vector into a low-dimensional text intermediate representation through calculation; use the CNN image encoder to encode the image feature vector into an image intermediate representation; use the WaveNet audio encoder to encode the audio feature vector into an audio intermediate representation; encode the video spatio-temporal feature vector through the video encoder combining 3D-CNN and LSTM. Among them, 3D-CNN extracts the features of the spatial dimension and time dimension in the video spatio-temporal feature vector, and LSTM further processes the video spatio-temporal feature vector to capture the time series dependencies in the video and obtain the video intermediate representation. Through the concatenation operation, fuse the text intermediate representation, image intermediate representation, audio intermediate representation, and video intermediate representation to obtain the multi-modal data intermediate vector, and then input the multi-modal data intermediate vector into the decoder to reconstruct the multi-modal data intermediate vector, minimize the reconstruction error, and learn the associations between different modal features. Take the text feature vector, image feature vector, audio feature vector, and video spatio-temporal feature vector as input data and input them into the multi-modal Transformer model. The multi-modal Transformer model encodes and integrates the input data according to the associations between different modal features to obtain the input sequence. Process the input sequence through the multi-head attention mechanism, calculate the weight distribution of different attentions, perform non-linear transformation on the input sequence processed by the multi-head attention mechanism using the feed-forward neural network, and then perform layer normalization and residual connection processing on the input sequence. The multi-modal Transformer model outputs a feature vector sequence that fuses different modal association information. Based on the feature vector sequence that fuses different modal association information, construct a multi-modal feature space.

9. The multimodal data governance method based on large model technology according to claim 8, characterized in that: In the eighth step described above, the process of using the anomaly detection ability of the large model to identify and correct errors and inconsistencies in the data and storing the multi-modal data using the object storage architecture includes: The big model conducts a comprehensive scan of multimodal data to identify anomalies in the multimodal data. For anomalies in text data, the big model uses semantic analysis technology and language generation technology to correct text errors; for anomalies in image data, the big model uses image correction algorithms to improve image errors; for anomalies in audio data, the big model uses audio processing algorithms to correct audio errors; for anomalies in video data, the big model uses video synthesis algorithms to correct video errors; The object storage architecture is used to store the modified text data, image data, audio data and video data in the form of objects, each object containing the data itself, metadata of detailed data attributes and a global identifier.

10. A multimodal data governance method based on large model technology according to claim 9, characterized in that: In step nine, the process of fine-tuning and updating the large model, monitoring the performance of the large model in real time, and integrating it into the existing data management system includes: New multimodal data is collected regularly, and the large model is fine-tuned and updated using the new multimodal data. The performance of the large model is monitored in real time with the help of performance indicator evaluation methods. The monitoring performance indicators include accuracy, recall rate, mean square error, confusion matrix, memory usage, running time and throughput, and the large model is integrated into the existing data management system.

Citation Information

Cited By

  • AI agent behavior prediction method based on multi-modal data

    CN120873993A

  • Data management method, model, system, product and equipment based on large language model and adaptive knowledge graph

    CN121009083A

  • Neurology disease identification method and system based on multi-modal information

    CN121260423A