Multi-modal model joint training method for single-modal data and multi-modal data

AIGC technology generates aligned multimodal data sets and performs feature fusion training, which solves the problem that single-modal data cannot use semantic information of other modalities, realizes effective training of multimodal models, and improves model accuracy.

CN120449085APending Publication Date: 2025-08-08TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510524566.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Single-modal data cannot utilize the rich semantic information present in other modes, resulting in hindering the overall understanding of complex scenarios, and multi-modal data fails to fully utilize its complementary advantages in practical applications.

Method used

AIGC technology is used to generate an aligned multimodal data set, and the client fusion model is trained on the local client through feature extraction and fusion splicing, and upload local gradient updates to the server for multimodal global model training.

Benefits of technology

Multimodal model training between clients of single modal data is realized, which solves the training bottleneck caused by data modal differences and improves the model accuracy by about 10%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120449085A_ABST
    Figure CN120449085A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent technology, in particular to a multi-modal model joint training method for single-modal data and multi-modal data. In order to solve the problem of training bottleneck caused by data modal difference, the invention provides a novel multi-modal model joint training method for single-modal and multi-modal data, which comprises the following steps of: 1) generating a corresponding aligned multi-modal data set from single-modal client training data by using an AIGC technology; 2) carrying out different feature extraction on all the multi-modal data sets; 3) fusing and splicing the extracted corresponding features, and training a client fusion model at a local client; and 4) the clients generate local gradient updates based on the client fusion model and upload the local gradient updates to the server, and the server aggregates the received gradient updates so as to train and form a multi-modal global model.According to the multi-modal model training method, multi-modal model training can be carried out even among clients with single-modal data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of embodied intelligence technology, and in particular to the field of intelligent model training technology, specifically a multimodal model joint training method for single and multimodal data. Background Art

[0002] With the rapid development of technologies such as embodied intelligence and large models, the importance of data resources has received increasing attention. In the real world, the diversity of data continues to increase, covering multiple modalities such as text, images, and sound. Multimodal data fusion technology significantly improves the performance of AI models by integrating aligned data from different modalities, making more significant progress than single-modality methods. This technology has been widely used in many fields such as the Internet of Things and natural language processing. However, the sensitivity and privacy issues of multimodal data, such as user personal information and commercial secrets, often limit its public disclosure, which in turn affects the effective use of multimodal data in model training. Therefore, multimodal federated learning (MMFL), as a new research direction, has shown great potential.

[0003] Existing approaches to building MMFL systems assume that all clients have aligned multimodal training data, but this contradicts real-world applications. In real-world applications, high-quality, multimodal aligned data is significantly scarce, and most collected data is single-modal. In the field of embodied intelligence, intelligent agents (such as smart vehicles, robots, and smart home devices) rely on information collected from multiple sensors (cameras, microphones, temperature sensors, accelerometers, etc.) to make decisions. These sensor devices typically capture data in a single modality and fail to leverage the rich semantic information present in other modalities. For example, vehicle cameras can only capture images or video data, onboard sensors can only collect wave data, and vehicle systems record passengers' historical instructions and feedback. This data is stored in the vehicle's control unit (ECU) in text form. Since single-modal data cannot be integrated and analyzed with the rich semantic information in other modalities, it hinders the overall understanding of complex scenes and may cause ambiguity in the interpretation of information. However, if the above-mentioned multiple modal data are integrated and analyzed, the system can more accurately help the driver fully understand the surrounding environment. Although multimodal data has potential complementarity, due to its scarcity, this advantage is often not fully utilized in practical applications. Summary of the Invention

[0004] In order to solve the problem that unimodal data cannot be integrated and analyzed with the rich semantic information in other modalities, thereby hindering the overall understanding of complex scenes, the present invention provides a new multimodal model joint training method for unimodal and multimodal data.

[0005] The present invention is achieved by adopting the following technical solutions:

[0006] A method for jointly training a multimodal global model of single and multimodal data, wherein the single and multimodal data include single-modal client training data and multimodal client training data. The specific training method includes the following steps: 1) using AIGC technology to generate a corresponding aligned multimodal dataset from each single-modal client training data, so that all single and multimodal clients have shared potential features to accommodate heterogeneous feature spaces and maintain cross-modal semantic consistency, thereby achieving feature alignment between single-modal and multimodal clients; 2) performing different feature extraction on all multimodal datasets; 3) fusing and splicing the extracted corresponding features, and training a client fusion model on the local client; 4) each client generates a local gradient update based on the client fusion model and uploads it to an aggregation server. The aggregation server uses an aggregation algorithm to aggregate the local gradient updates received from multiple clients to train a multimodal global model.

[0007] The beneficial effects of the present invention are as follows: The training method described in the present invention uses AIGC technology to generate aligned multimodal data for clients whose training data is a single modality in multimodal federated learning, thereby realizing multimodal model training even between clients with only single modality data, and solving the training bottleneck caused by data modality differences. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0009] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0010] Figure 1 This is a training diagram for a single text client;

[0011] Figure 2 This is a diagram of training for a single-image client;

[0012] Figure 3 Schematic diagram of multimodal feature fusion. DETAILED DESCRIPTION

[0013] In order to more clearly understand the above-mentioned objectives, features and advantages of the present invention, the scheme of the present invention will be further described below. It should be noted that, in the absence of conflict, the embodiments of the present invention and the features therein can be combined with each other.

[0014] In the description, it should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance. It should be noted that, unless otherwise expressly specified or limited, the terms "installed," "connected," and "connected" should be understood broadly. For example, they can refer to fixed connections, removable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal connections between two components. Those skilled in the art will understand the specific meanings of the above terms based on specific circumstances.

[0015] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present invention, rather than all the embodiments.

[0016] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0017] Assume that the client training data in the MMFL system includes both text and image modal data. For unimodal client training data, the training data only includes one of the two modalities (such as images captured by vehicle cameras or related text information, such as road image information). Multimodal client training data has aligned data of both image and text modalities. The specific training method includes the following steps:

[0018] 1) If Figure 1 As shown, generate image data: use AIGC technology to generate corresponding image data from client training data that only holds text data:

[0019] a. Text encoding: The input text description T is first converted into an embedding representation e in the latent vector space through a pre-trained text encoder (e.g., CLIP model) T ,

[0020] e T =TextEncoder(T)

[0021] b. Noise addition: The image x0 (e.g., an image sampled from the data distribution) is gradually noised.

[0022]

[0023] x t represents the image after t steps of noise diffusion, α t is a factor that controls the intensity of noise addition, ∈ t It is the noise in the standard normal distribution. In this process, the image gradually loses its original information and becomes pure noise;

[0024] c. Reverse process: given a text embedding e T and the current noise image x t , through the trained denoising network UNet, the noise is predicted and denoised step by step to restore the qualified image.

[0025]

[0026] ∈ θ (x t , t) is the noise predicted by the denoising neural network, the network uses the current image x t and time step t as input to predict the noise. t and is a parameter that controls the noise intensity during the diffusion process. t is the noise standard deviation of the diffusion process, z t is noise sampled from a standard normal distribution and is used to introduce randomness in the back-diffusion process. At each step, the denoising network ∈ θ The goal is to minimize the difference between the predicted noise and the actual noise, so that the image can be gradually restored to a clear image that matches the text description. As t decreases to 0, the model gradually removes the noise and generates the final image;

[0027] 2) If Figure 2 As shown, text data is generated: AIGC technology is used to generate corresponding text data from client training data that only holds image data. This usually involves image-to-text technology, which extracts image features through a visual model and converts them into descriptive text in combination with a language generation model, achieving cross-modal conversion from vision to language. The specific process is as follows:

[0028] a. Image feature extraction: First, a convolutional neural network or other visual feature extraction model is used to extract high-dimensional feature representation F from the image I.

[0029] F=ImageEncoder(I)

[0030] b. Feature encoding: The extracted visual features are further encoded into representations suitable for processing by large language models.

[0031] z t =Attention(z t-1 , F t )

[0032] z t is the hidden state at time step t, representing the information after the update at that time step, z t-1 is the hidden state of the previous time step, F t is the image feature extracted at time step t, and Attention represents the attention mechanism, which dynamically adjusts the degree of attention to different information based on the relationship between the current image feature and the previous hidden state.

[0033] At each time step, the attention mechanism updates the hidden state by calculating the relationship between the hidden state of the current time step and the input image features. In this way, the model can gradually fuse visual information and language information to provide richer contextual information for subsequent text generation tasks.

[0034] c. The decoder gradually generates natural language descriptions from the image features.

[0035]

[0036] Among them, w t |w 1:t-1 is the generated text sequence, P(w t |w 1:t-1 , F) represents the conditional probability distribution for predicting the word w given the previous generation 1;t-1 and visual features F, the current word w t probability;

[0037] 3) If Figure 3 As shown in Figure 1, image feature extraction: Image feature extraction is performed on the original image data in the single-modal client data and multi-modal client data, as well as the image data generated by the AIGC technology. In order to enhance the effectiveness of the model, image features are extracted in two different domains: spatial domain and frequency domain: First, the Fourier Transform (FT) algorithm is used to convert the image data into frequency domain information. Then, the original image is processed by the VGG-19 model to obtain the d-dimensional spatial domain feature R F ∈R d×1 , and then use a series of CNN convolutional networks to extract frequency domain information to obtain the frequency domain features R of the image s ∈R d×1 ;

[0038] 4) Text feature extraction: Extract text features from the original text data in the unimodal client data and multimodal client data, as well as the text data generated using AIGC technology. In order to extract text features, the pre-trained BERT model is used to obtain the text feature R T ∈R d×1 ;

[0039] 5) Multimodal feature fusion: The extracted text and its corresponding image features are fused and spliced, and trained on the local client to form the corresponding client fusion model;

[0040] 6) Each client generates a local gradient update based on the client fusion model and uploads it to the aggregation server. The aggregation server uses an aggregation algorithm to aggregate the local gradient updates received from multiple clients to train a multimodal global model.

[0041] Experiments on the MSCOCO dataset show that the model trained using this method improves model accuracy by about 10% compared to the model trained using only a single modality.

[0042] The above description is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Although detailed descriptions have been made with reference to the aforementioned embodiments, those skilled in the art should understand that they may still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents; and such modifications or replacements do not deviate from the essence of the corresponding technical solutions within the scope of the technical solutions of the embodiments, and they should all be included in the scope of protection of the claims.

Claims

1. A method for jointly training multimodal models of single and multimodal data, wherein: The single- and multimodal data include single-modal client training data and multimodal client training data, and are characterized by the following steps: 1) using AIGC technology to generate a corresponding aligned multimodal dataset for each single-modal client training data, so that all single- and multimodal clients have shared potential features to accommodate heterogeneous feature spaces and maintain cross-modal semantic consistency, thereby achieving feature alignment between single-modal and multimodal clients; 2) performing different feature extraction on all multimodal datasets; 3) fusing and splicing the extracted corresponding features, and training a client fusion model on the local client; 4) each client generates a local gradient update based on the client fusion model and uploads it to the aggregation server. The aggregation server uses an aggregation algorithm to aggregate the local gradient updates received from multiple clients to train a multimodal global model.