Multi-modal data fusion method and system based on large model
Through the self-attention and cross-attention mechanisms of the large model, combined with lightweight encoders and adapters, a fusion feature matrix of cross-modal semantic association is generated, which solves the problems of semantic association and high computational cost in traditional multimodal fusion methods and achieves lightweight deployment and efficient fusion.
Patent Information
- Application Number
- CN202510840258.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-06-23
AI Technical Summary
Traditional multimodal fusion methods have difficulty capturing cross-modal semantic associations, poor model generalization, high computational costs, and difficulty meeting real-time requirements.
Through the self-attention and cross-attention mechanisms of the large model, dynamic calculation of the semantic weights of multimodal data is achieved, and a fusion feature matrix is generated by combining lightweight encoders and adapters. It is then deployed to edge devices through knowledge distillation and pruning optimization.
It realizes the generation of fusion feature matrix of cross-modal semantic association, supports lightweight deployment, improves fusion efficiency and generalization capability, and is suitable for edge devices.
Smart Images

Figure CN120822172A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data fusion, and in particular to a large model-based multimodal data fusion method and system. Background Art
[0002] The prior art has the following defects:
[0003] Modal heterogeneity: Traditional multimodal fusion methods (such as feature splicing and weighted averaging) have difficulty capturing cross-modal semantic associations, resulting in insufficient information utilization.
[0004] Poor model generalization: Existing algorithms (such as CNN+Attention) rely on specific task design and are not adaptable to unknown modal combinations or complex scenarios.
[0005] High computational cost: When large models directly process multimodal data, the number of parameters is huge, making it difficult to meet real-time requirements (such as edge device deployment). Summary of the Invention
[0006] In response to the needs and shortcomings of current technological development, the present invention provides a multimodal data fusion method and system based on a large model, which improves fusion efficiency and generalization capability by breaking the semantic gap between modalities, while supporting lightweight deployment.
[0007] In a first aspect, the present invention provides a multimodal data fusion method based on a large model, and the technical solutions adopted to solve the above technical problems are as follows:
[0008] A multimodal data fusion method based on a large model comprises the following steps:
[0009] S1, receiving multimodal data, performing noise layer filtering and time-space alignment;
[0010] S2: Map the data output from step S1 to the latent space of the large model through lightweight encoders of each modality. After extracting initial features using a single-modality pre-trained model, the latent space vector is generated through adapter conversion for downstream cross-modal retrieval and multimodal fusion.
[0011] S3: Receive the latent space vectors of each modality and the task prompt words, extract the semantic association within the single modality through the self-attention mechanism of the large model, and then use the cross-attention mechanism to calculate the cross-modal semantic weights. Finally, dynamically aggregate to generate a fusion feature matrix to achieve semantic integration of multimodal information;
[0012] S4: Receive the fused feature matrix output by the large model as a soft label and learn cross-modal semantic mapping capabilities through knowledge distillation. The trained model is dynamically quantized and optimized through pruning, deployed to the edge or front-end environment, and the optimized feature matrix is input into the task-customized lightweight head network, which combines multimodal context information to output the final task result.
[0013] Optionally, the large model involved refers to a high-parameter neural network obtained through pre-training on large-scale cross-modal data. It has the ability to map multimodal information into a unified latent space and supports cross-modal semantic association and task reasoning through the attention mechanism. Its core features include:
[0014] Learning cross-modal semantic alignment in the pre-training phase;
[0015] Providing standardized latent space dimensions as a unified representation basis for multimodal features;
[0016] It supports dynamic adjustment of cross-modal attention weights through task prompt words to generate a fusion feature matrix.
[0017] Optionally, step S1 specifically includes:
[0018] S1.1. Access multimodal data, including video data, text data, and image data, to provide raw materials for subsequent processing;
[0019] S1.2. Using a layered noise filtering framework, we first perform intra-modal primary filtering. This involves applying the corresponding basic filtering algorithms based on the characteristics of different modal data to preliminarily screen out candidate clean data. This candidate clean data is then fed into an adaptive noise determination algorithm, which uses intelligent analysis and determination to further filter out noisy data and improve data purity.
[0020] S1.3. Match multimodal data based on the data's inherent timestamp information to ensure that data from different modalities correspond in time, allowing for synchronous presentation and analysis of changes in the time series of each modality.
[0021] S1.4. With the help of multimodal coordinate mapping technology, the coordinate information of different modal data in the spatial dimension is converted and matched to eliminate the spatial differences caused by data acquisition equipment and viewing angle factors, and achieve precise spatial alignment of multimodal data.
[0022] Preferably, step S1.2 is performed, and a layered noise filtering framework is used to first perform intra-modal primary filtering:
[0023] For video data, use any of the three filtering algorithms, including inter-frame filtering, median filtering / Gaussian filtering, and abnormal frame detection, to preliminarily screen out candidate clean data;
[0024] For text data, use any of the four filtering algorithms, including regular expression filtering, stop word filtering, text similarity detection, and language model error correction, to preliminarily screen out candidate clean data;
[0025] For image data, any one of the four filtering algorithms, including spatial filtering, deblurring algorithm, and abnormal area detection, is used to preliminarily screen out candidate clean data.
[0026] Optionally, step S2 specifically includes:
[0027] S2.1. Each modality’s lightweight encoder performs preliminary dimensionality reduction and feature compression on the data output from step S1, mapping it to the basic dimensions of the latent space that can be processed by the large model;
[0028] S2.2. Use the text encoder of the single-modal pre-trained model BERT or the cross-modal pre-trained model CLIP to perform semantic encoding on the mapped text data and extract the feature vector h containing the context association text ; Through the single-modal pre-trained visual model ResNet or ViT, the image data is processed by convolution operation and attention mechanism to generate a feature vector h containing spatial structure information image ;
[0029] S2.3. For the initial feature vectors of text and image, linear transformation is performed through learnable adapter parameters to generate the text latent vector Z adapted to the cross-modal space text and the image latent vector Z aligned with the text space image :
[0030] Z text =W text *h text +b text ,
[0031] Z image =W image *h image +b image ,
[0032] Where W text 、b text 、W image 、b image are the learnable parameters of the adapter.
[0033] In a second aspect, the present invention provides a multimodal data fusion system based on a large model, and the technical solutions adopted to solve the above technical problems are as follows:
[0034] A multimodal data fusion system based on a large model, comprising:
[0035] Data receiving and preprocessing module, used to receive multimodal data, perform noise layer filtering and time-space alignment;
[0036] The multimodal feature mapping and conversion module is used to map the preprocessed data to the latent space of the large model through lightweight encoders of each modality. After extracting the initial features using the single-modal pre-trained model, the latent space vector is generated through adapter conversion for downstream cross-modal retrieval and multimodal fusion.
[0037] The vector receiving and dynamic aggregation module is used to receive the latent space vectors of each modality and the task prompt words, extract the semantic associations within a single modality through the self-attention mechanism of the large model, and then use the cross-attention mechanism to calculate the cross-modal semantic weights. Finally, dynamic aggregation generates a fusion feature matrix to achieve semantic integration of multimodal information;
[0038] The lightweight model deployment module is used to receive the fused feature matrix output by the large model as soft labels and learn cross-modal semantic mapping capabilities through knowledge distillation. The trained model is dynamically quantized and optimized through pruning, deployed to the edge or front-end environment, and the optimized feature matrix is input into the task-customized lightweight head network, which combines multimodal context information to output the final task results.
[0039] Optionally, the large model involved refers to a high-parameter neural network obtained through pre-training on large-scale cross-modal data. It has the ability to map multimodal information into a unified latent space and supports cross-modal semantic association and task reasoning through the attention mechanism. Its core features include:
[0040] Learning cross-modal semantic alignment in the pre-training phase;
[0041] Providing standardized latent space dimensions as a unified representation basis for multimodal features;
[0042] It supports dynamic adjustment of cross-modal attention weights through task prompt words to generate a fusion feature matrix.
[0043] Optionally, the data receiving and preprocessing modules involved specifically include:
[0044] A multimodal data receiving unit is used to access multimodal data, including video data, text data, and image data, to provide raw materials for subsequent processing;
[0045] The layered noise filtering unit is used to perform primary filtering within the modality using a layered noise filtering framework. This involves applying the corresponding basic filtering algorithm based on the characteristics of different modal data to preliminarily screen out candidate clean data. The candidate clean data is then input into the adaptive noise determination algorithm, which further filters the noise data through intelligent analysis and determination to improve data purity.
[0046] The time alignment unit is used to match multimodal data based on the timestamp information of the data, ensuring that the data of different modalities correspond in the time dimension, so that the changes in the time series of each modal data can be presented and analyzed synchronously;
[0047] The spatial alignment unit is used to convert and match the coordinate information of different modal data in spatial dimensions with the help of multimodal coordinate mapping technology, eliminate the spatial differences caused by data acquisition equipment and viewing angle factors, and achieve precise spatial alignment of multimodal data.
[0048] Preferably, the layered noise filtering unit adopts a layered noise filtering framework and first performs intra-modal primary filtering:
[0049] For video data, use any of the three filtering algorithms, including inter-frame filtering, median filtering / Gaussian filtering, and abnormal frame detection, to preliminarily screen out candidate clean data;
[0050] For text data, use any of the four filtering algorithms, including regular expression filtering, stop word filtering, text similarity detection, and language model error correction, to preliminarily screen out candidate clean data;
[0051] For image data, any one of the four filtering algorithms, including spatial filtering, deblurring algorithm, and abnormal area detection, is used to preliminarily screen out candidate clean data.
[0052] Optionally, the multimodal feature mapping and conversion modules involved specifically include:
[0053] The data mapping unit is used to perform preliminary dimensionality reduction and feature compression on the preprocessed data using lightweight encoders of each modality, and then map it to the basic dimensions of the latent space that can be processed by the large model;
[0054] The semantic encoding unit is used to use the text encoder of the single-modal pre-training model BERT or the cross-modal pre-training model CLIP to semantically encode the mapped text data and extract the feature vector h containing the context association. text ;
[0055] The image processing unit is used to perform convolution operations and attention mechanism processing on image data through the single-modal pre-trained visual model ResNet or ViT to generate a feature vector h containing spatial structure information image ;
[0056] The linear transformation unit is used to perform linear transformation on the initial feature vectors of text and image through learnable adapter parameters to generate a text latent vector Z adapted to the cross-modal space. text and the image latent vector Z aligned with the text space image :
[0057] Z text =W text *h text +b text ,
[0058] Z image =W image *h image +b image ,
[0059] Where W text 、b text 、W image 、b image are the learnable parameters of the adapter.
[0060] The multimodal data fusion method and system based on a large model of the present invention have the following beneficial effects compared with the prior art:
[0061] 1. This invention uses the self-attention and cross-attention mechanisms of a large model to dynamically calculate the semantic weights of multimodal data such as images, text, and videos, addressing the modality gap problem and generating a fused feature matrix containing cross-modal semantic associations, supporting the joint understanding of multimodal information by downstream tasks.
[0062] 2. The lightweight head of the present invention can be customized for specific tasks without modifying the core multimodal processing process, supporting rapid business iteration; the standardized process from preprocessing to feature fusion can be applied to scenarios such as disaster emergency response, intelligent diagnosis, and industrial detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Attachment Figure 1 is a flow chart of a method according to embodiment 1 of the present invention;
[0064] Attachment Figure 2 This is a module connection block diagram of the second embodiment of the present invention. DETAILED DESCRIPTION
[0065] In order to make the technical solution, the technical problems solved and the technical effects of the present invention more clear, the technical solution of the present invention is clearly and completely described below in conjunction with specific embodiments.
[0066] Example 1:
[0067] Combined with attachment Figure 1 This embodiment proposes a multimodal data fusion method based on a large model, which includes the following steps:
[0068] S1. Receive multimodal data and perform noise layer filtering and spatiotemporal alignment, specifically including:
[0069] S1.1. Access multimodal data, including video data, text data, and image data, to provide raw materials for subsequent processing.
[0070] S1.2. Using a layered noise filtering framework, we first perform intra-modal primary filtering. That is, based on the characteristics of different modal data, we apply the corresponding basic filtering algorithm to preliminarily screen out candidate clean data. The candidate clean data is then input into the adaptive noise judgment algorithm, and through intelligent analysis and judgment, the noise data is further filtered to improve data purity.
[0071] In this step, when performing intra-modal primary filtering:
[0072] For video data, use any of the three filtering algorithms, including inter-frame filtering, median filtering / Gaussian filtering, and abnormal frame detection, to preliminarily screen out candidate clean data;
[0073] For text data, use any of the four filtering algorithms, including regular expression filtering, stop word filtering, text similarity detection, and language model error correction, to preliminarily screen out candidate clean data;
[0074] For image data, any one of the four filtering algorithms, including spatial filtering, deblurring algorithm, and abnormal area detection, is used to preliminarily screen out candidate clean data.
[0075] S1.3. Match multimodal data based on the timestamp information of the data to ensure that the data of different modalities correspond in the time dimension, so that the changes in the time series of each modal data can be presented and analyzed synchronously.
[0076] S1.4. With the help of multimodal coordinate mapping technology, the coordinate information of different modal data in the spatial dimension is converted and matched to eliminate the spatial differences caused by data acquisition equipment and viewing angle factors, and achieve precise spatial alignment of multimodal data.
[0077] S2: Map the data output from step S1 to the latent space of a large model (such as CLIP) through lightweight encoders for each modality. After extracting initial features using a single-modality pre-trained model, the latent space vector is generated through adapter conversion for downstream cross-modal retrieval and multimodal fusion. Specifically, it includes:
[0078] S2.1. Each modality's lightweight encoder (e.g., 3D convolutional network for video, FastText for text, MobileNet for images) performs preliminary dimensionality reduction and feature compression on the data output from step S1, mapping it to a base dimension of the latent space that can be processed by a large model (e.g., CLIP).
[0079] S2.2. Use the text encoder of the single-modal pre-trained model BERT or the cross-modal pre-trained model CLIP to perform semantic encoding on the mapped text data and extract the feature vector h containing the context association text ; Through the single-modal pre-trained visual model ResNet or ViT, the image data is processed by convolution operation and attention mechanism to generate a feature vector h containing spatial structure information image ;
[0080] S2.3. For the initial feature vectors of text and image, linear transformation is performed through learnable adapter parameters to generate the text latent vector Z adapted to the cross-modal space text and the image latent vector Z aligned with the text space image :
[0081] Z text =W text *h text +b text ,
[0082] Z image =W image *h image +b image ,
[0083] Where W text 、b text 、W image 、b image are the learnable parameters of the adapter.
[0084] S3: Receive the latent space vectors of each modality and the task prompt word, extract the semantic association within the single modality through the self-attention mechanism of a large model (such as BLIP-2), and then use the cross-attention mechanism to calculate the cross-modal semantic weights. Finally, dynamically aggregate to generate a fusion feature matrix to achieve semantic integration of multimodal information;
[0085] S4. Adopting the TinyBERT lightweight model architecture, it receives the fused feature matrix output by a large model (such as BLIP-2) as a soft label and learns cross-modal semantic mapping capabilities through knowledge distillation. The trained model is dynamically quantized and optimized through pruning, deployed to the edge or front-end environment, and the optimized feature matrix is input into the task-customized lightweight head network, which combines multimodal context information to output the final task result.
[0086] It should be noted that the large models involved refer to high-parameter neural networks obtained through pre-training on large-scale cross-modal data. They have the ability to map multimodal information into a unified latent space and support cross-modal semantic association and task reasoning through the attention mechanism. They are not limited to the CLIP large model and BLIP-2 large model listed above. The core features of the large models include:
[0087] Learning cross-modal semantic alignment in the pre-training phase;
[0088] Providing standardized latent space dimensions as a unified representation basis for multimodal features;
[0089] It supports dynamic adjustment of cross-modal attention weights through task prompt words to generate a fusion feature matrix.
[0090] Example 2:
[0091] Combined with attachment Figure 2 This embodiment proposes a multimodal data fusion system based on a large model, which includes:
[0092] Data receiving and preprocessing module, used to receive multimodal data, perform noise layer filtering and time-space alignment;
[0093] The multimodal feature mapping and conversion module is used to map the preprocessed data to the latent space of a large model (such as CLIP) through lightweight encoders of each modality. After extracting the initial features using a single-modality pre-trained model, the latent space vector is generated through adapter conversion for downstream cross-modal retrieval and multimodal fusion.
[0094] The vector receiving and dynamic aggregation module is used to receive the latent space vectors of each modality and the task prompt word, extract the semantic association within a single modality through the self-attention mechanism of a large model (such as BLIP-2), and then use the cross-attention mechanism to calculate the cross-modal semantic weights. Finally, dynamic aggregation generates a fusion feature matrix to achieve semantic integration of multimodal information;
[0095] The lightweight model deployment module uses the TinyBERT lightweight model architecture, receives the fused feature matrix output by a large model (such as BLIP-2) as a soft label, and learns cross-modal semantic mapping capabilities through knowledge distillation. The trained model is dynamically quantized and pruned, then deployed to the edge or front-end environment. The optimized feature matrix is input into a task-customized lightweight head network, which combines multimodal context information to output the final task results.
[0096] In this embodiment, the data receiving and preprocessing module specifically includes:
[0097] A multimodal data receiving unit is used to access multimodal data, including video data, text data, and image data, to provide raw materials for subsequent processing;
[0098] The layered noise filtering unit is used to perform primary filtering within the modality using a layered noise filtering framework. This involves applying the corresponding basic filtering algorithm based on the characteristics of different modal data to preliminarily screen out candidate clean data. The candidate clean data is then input into the adaptive noise determination algorithm, which further filters the noise data through intelligent analysis and determination to improve data purity.
[0099] The time alignment unit is used to match multimodal data based on the timestamp information of the data, ensuring that the data of different modalities correspond in the time dimension, so that the changes in the time series of each modal data can be presented and analyzed synchronously;
[0100] The spatial alignment unit is used to convert and match the coordinate information of different modal data in spatial dimensions with the help of multimodal coordinate mapping technology, eliminate the spatial differences caused by data acquisition equipment and viewing angle factors, and achieve precise spatial alignment of multimodal data.
[0101] It should be added that the layered noise filtering unit adopts a layered noise filtering framework and first performs intra-modal primary filtering:
[0102] For video data, use any of the three filtering algorithms, including inter-frame filtering, median filtering / Gaussian filtering, and abnormal frame detection, to preliminarily screen out candidate clean data;
[0103] For text data, use any of the four filtering algorithms, including regular expression filtering, stop word filtering, text similarity detection, and language model error correction, to preliminarily screen out candidate clean data;
[0104] For image data, any one of the four filtering algorithms, including spatial filtering, deblurring algorithm, and abnormal area detection, is used to preliminarily screen out candidate clean data.
[0105] In this embodiment, the multimodal feature mapping and conversion module specifically includes:
[0106] The data mapping unit is used to perform preliminary dimensionality reduction and feature compression on the preprocessed data using lightweight encoders of various modalities (such as 3D convolutional networks for video, FastText for text, and MobileNet for images), and then map it to the base dimensions of the latent space that can be processed by the large model;
[0107] The semantic encoding unit is used to use the text encoder of the single-modal pre-training model BERT or the cross-modal pre-training model CLIP to semantically encode the mapped text data and extract the feature vector h containing the context association. text ;
[0108] The image processing unit is used to perform convolution operations and attention mechanism processing on image data through the single-modal pre-trained visual model ResNet or ViT to generate a feature vector h containing spatial structure information image ;
[0109] The linear transformation unit is used to perform linear transformation on the initial feature vectors of text and image through learnable adapter parameters to generate a text latent vector Z adapted to the cross-modal space. text and the image latent vector Z aligned with the text space image :
[0110] Z text =W text *h text +b text ,
[0111] Z image =W image *h image +b image ,
[0112] Where W text 、b text 、W image 、b image are the learnable parameters of the adapter.
[0113] It should be noted that the large models involved refer to high-parameter neural networks obtained through pre-training on large-scale cross-modal data. They have the ability to map multimodal information into a unified latent space and support cross-modal semantic association and task reasoning through the attention mechanism. They are not limited to the CLIP large model and BLIP-2 large model listed above. The core features of the large models include:
[0114] Learning cross-modal semantic alignment in the pre-training phase;
[0115] Providing standardized latent space dimensions as a unified representation basis for multimodal features;
[0116] It supports dynamic adjustment of cross-modal attention weights through task prompt words to generate a fusion feature matrix.
[0117] For the first and second embodiments of this invention, the large model fusion features are used as soft labels to train lightweight models, which can increase the inference speed by 3-5 times while retaining more than 90% of the cross-modal capabilities; the model size is reduced by 40%-60% through model compression technology (such as INT8 quantization), which is suitable for edge device or browser deployment.
[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multimodal data fusion method based on a large model, characterized in that: The steps include: S1, receiving multimodal data, performing noise layer filtering and time-space alignment; S2: Map the data output from step S1 to the latent space of the large model through lightweight encoders of each modality. After extracting initial features using a single-modality pre-trained model, the latent space vector is generated through adapter conversion for downstream cross-modal retrieval and multimodal fusion. S3: Receive the latent space vectors of each modality and the task prompt words, extract the semantic association within the single modality through the self-attention mechanism of the large model, and then use the cross-attention mechanism to calculate the cross-modal semantic weights. Finally, dynamically aggregate to generate a fusion feature matrix to achieve semantic integration of multimodal information; S4, receiving the fused feature matrix output by the large model as soft labels, and learning cross-modal semantic mapping capabilities through knowledge distillation; The trained model is dynamically quantized and pruned, then deployed to the edge or front-end environment. The optimized feature matrix is input into a lightweight head network customized for the task, and the final task result is output in combination with multimodal context information.
2. The multimodal data fusion method based on a large model according to claim 1, characterized in that: The large model refers to a high-parameter neural network obtained through pre-training on large-scale cross-modal data. It has the ability to map multimodal information into a unified latent space and supports cross-modal semantic association and task reasoning through the attention mechanism. Its core features include: Learning cross-modal semantic alignment in the pre-training phase; Providing standardized latent space dimensions as a unified representation basis for multimodal features; It supports dynamic adjustment of cross-modal attention weights through task prompt words to generate a fusion feature matrix.
3. The multimodal data fusion method based on a large model according to claim 1, characterized in that: The step S1 specifically includes: S1.
1. Access multimodal data, including video data, text data, and image data, to provide raw materials for subsequent processing; S1.
2. Using a layered noise filtering framework, we first perform intra-modal primary filtering. This involves applying the corresponding basic filtering algorithms based on the characteristics of different modal data to preliminarily screen out candidate clean data. This candidate clean data is then fed into an adaptive noise determination algorithm, which uses intelligent analysis and determination to further filter out noisy data and improve data purity. S1.
3. Match multimodal data based on the data's inherent timestamp information to ensure that data from different modalities correspond in time, allowing for synchronous presentation and analysis of changes in the time series of each modality. S1.
4. With the help of multimodal coordinate mapping technology, the coordinate information of different modal data in the spatial dimension is converted and matched to eliminate the spatial differences caused by data acquisition equipment and viewing angle factors, and achieve precise spatial alignment of multimodal data.
4. The multimodal data fusion method based on a large model according to claim 3 is characterized in that: Execute step S1.2, using the layered noise filtering framework, and first perform intra-modal primary filtering: For video data, use any of the three filtering algorithms, including inter-frame filtering, median filtering / Gaussian filtering, and abnormal frame detection, to preliminarily screen out candidate clean data; For text data, use any of the four filtering algorithms, including regular expression filtering, stop word filtering, text similarity detection, and language model error correction, to preliminarily screen out candidate clean data; For image data, any one of the four filtering algorithms, including spatial filtering, deblurring algorithm, and abnormal area detection, is used to preliminarily screen out candidate clean data.
5. The multimodal data fusion method based on a large model according to claim 3 is characterized in that: The step S2 specifically includes: S2.
1. Each modality’s lightweight encoder performs preliminary dimensionality reduction and feature compression on the data output from step S1, mapping it to the basic dimensions of the latent space that can be processed by the large model; S2.
2. Use the text encoder of the single-modal pre-trained model BERT or the cross-modal pre-trained model CLIP to semantically encode the mapped text data and extract the feature vector h containing the context association text ; Through the single-modal pre-trained visual model ResNet or ViT, the image data is processed by convolution operation and attention mechanism to generate a feature vector h containing spatial structure information image ; S2.
3. For the initial feature vectors of text and image, linear transformation is performed through learnable adapter parameters to generate the text latent vector Z adapted to the cross-modal space text and the image latent vector Z aligned with the text space image : Z text =W text *h text +b text , Z image =W image *h image +b image , Where W text 、b text 、W image 、b image are the learnable parameters of the adapter.
6. A multimodal data fusion system based on a large model, characterized by: It includes: Data receiving and preprocessing module, used to receive multimodal data, perform noise layer filtering and time-space alignment; The multimodal feature mapping and conversion module is used to map the preprocessed data to the latent space of the large model through lightweight encoders of each modality. After extracting the initial features using the single-modal pre-trained model, the latent space vector is generated through adapter conversion for downstream cross-modal retrieval and multimodal fusion. The vector receiving and dynamic aggregation module is used to receive the latent space vectors of each modality and the task prompt words, extract the semantic associations within a single modality through the self-attention mechanism of the large model, and then use the cross-attention mechanism to calculate the cross-modal semantic weights. Finally, dynamic aggregation generates a fusion feature matrix to achieve semantic integration of multimodal information; A lightweight model deployment module receives the fused feature matrix output by the large model as soft labels and learns cross-modal semantic mapping capabilities through knowledge distillation. The trained model is dynamically quantized and pruned, then deployed to the edge or front-end environment. The optimized feature matrix is input into a lightweight head network customized for the task, and the final task result is output in combination with multimodal context information.
7. The multimodal data fusion system based on a large model according to claim 6, characterized in that: The large model refers to a high-parameter neural network obtained through pre-training on large-scale cross-modal data. It has the ability to map multimodal information into a unified latent space and supports cross-modal semantic association and task reasoning through the attention mechanism. Its core features include: Learning cross-modal semantic alignment in the pre-training phase; Providing standardized latent space dimensions as a unified representation basis for multimodal features; It supports dynamic adjustment of cross-modal attention weights through task prompt words to generate a fusion feature matrix.
8. The multimodal data fusion system based on a large model according to claim 6, characterized in that: The data receiving and preprocessing module specifically includes: A multimodal data receiving unit is used to access multimodal data, including video data, text data, and image data, to provide raw materials for subsequent processing; The layered noise filtering unit is used to perform primary filtering within the modality using a layered noise filtering framework. This involves applying the corresponding basic filtering algorithm based on the characteristics of different modal data to preliminarily screen out candidate clean data. The candidate clean data is then input into the adaptive noise determination algorithm, which further filters the noise data through intelligent analysis and determination to improve data purity. The time alignment unit is used to match multimodal data based on the timestamp information of the data, ensuring that the data of different modalities correspond in the time dimension, so that the changes in the time series of each modal data can be presented and analyzed synchronously; The spatial alignment unit is used to convert and match the coordinate information of different modal data in spatial dimensions with the help of multimodal coordinate mapping technology, eliminate the spatial differences caused by data acquisition equipment and viewing angle factors, and achieve precise spatial alignment of multimodal data.
9. The multimodal data fusion system based on a large model according to claim 8, characterized in that: The layered noise filtering unit adopts a layered noise filtering framework and first performs intra-modal primary filtering: For video data, use any of the three filtering algorithms, including inter-frame filtering, median filtering / Gaussian filtering, and abnormal frame detection, to preliminarily screen out candidate clean data; For text data, use any of the four filtering algorithms, including regular expression filtering, stop word filtering, text similarity detection, and language model error correction, to preliminarily screen out candidate clean data; For image data, any one of the four filtering algorithms, including spatial filtering, deblurring algorithm, and abnormal area detection, is used to preliminarily screen out candidate clean data.
10. The multimodal data fusion system based on a large model according to claim 8, characterized in that: The multimodal feature mapping and conversion module specifically includes: The data mapping unit is used to perform preliminary dimensionality reduction and feature compression on the preprocessed data using lightweight encoders of each modality, and then map it to the basic dimensions of the latent space that can be processed by the large model; The semantic encoding unit is used to use the text encoder of the single-modal pre-training model BERT or the cross-modal pre-training model CLIP to semantically encode the mapped text data and extract the feature vector h containing the context association. text ; The image processing unit is used to perform convolution operations and attention mechanism processing on image data through the single-modal pre-trained visual model ResNet or ViT to generate a feature vector h containing spatial structure information image ; The linear transformation unit is used to perform linear transformation on the initial feature vectors of text and image through learnable adapter parameters to generate a text latent vector Z adapted to the cross-modal space. text and the image latent vector Z aligned with the text space image : Z text =W text *h text +b text , Z image =W image *h image +b image , Where W text 、b text 、W image 、b image are the learnable parameters of the adapter.
Citation Information
Patent Citations
Dynamic gesture recognition method, system and equipment and medium
CN116524593A
Lightweight adaptive network learning method oriented to multi-mode and multi-task learning
CN116644316A
Emotional intention semantic association method, system and equipment based on implicit label reasoning
CN117828534A
Multi-modal graph fusion learning method and system based on Boolean multiplication
CN117935002A
Medical multi-mode hidden space alignment fusion method and system based on VQ-GCN
CN119475243A
Cited By
Lightweight multi-modal representation learning method based on multilayer attention mechanism
CN121051701A
Multi-modal large model data integration treatment system and method
CN121144855A
Software function automatic test and evaluation method based on multi-modal large model
CN121412134A
Multi-source heterogeneous data element intelligent fusion and operation system based on AI large model
CN121637426A
Model scheduling method and system for end-cloud collaboration
CN122332076A