Multimodal Federated Learning Training Method and Device

By training the initial model on the server side and transmitting the global feature representation to the client for aggregation, the resource limitation problem in multimodal federated learning is solved, enabling larger-scale multimodal model training and task completion.

CN116386058BActive Publication Date: 2026-05-26TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TSINGHUA UNIVERSITY
Filing Date
2023-02-17
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing multimodal federated learning algorithms cannot effectively utilize the large amounts of unimodal and multimodal data on the client side to train multimodal models on the server side, due to limitations in the computing and storage resources of edge devices.

Method used

An initial model is trained on the server side. Shared data is input into the server-side model to generate a global feature representation, which is then transmitted to the client for training. The client generates a local feature representation based on the local data modality type. The training of the server-side model is completed by aggregating the global and local feature representations.

Benefits of technology

It enables the training of larger-scale multimodal models without being limited by client resources, and can effectively utilize client-side unimodal and multimodal data to complete a variety of multimodal tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116386058B_ABST
    Figure CN116386058B_ABST
Patent Text Reader

Abstract

This invention provides a multimodal federated learning training method and apparatus. The method includes: inputting shared data into an initial server-side model to obtain an output global feature representation, and transmitting the global feature representation to a client; receiving local feature representations generated by the client; aggregating the local feature representations transmitted by the client based on the global feature representation and the local feature representations to obtain an aggregated feature representation; and training the server-side model based on the aggregated feature representation to complete one round of model training. This method uses shared data for knowledge transfer between the server and the client, ensuring that the client's private data is not transmitted to the server, and enabling the server to effectively utilize a large amount of unimodal and multimodal data on the client to train the server-side model, thereby training a larger-scale multimodal model capable of performing various multimodal tasks.
Need to check novelty before this filing date? Find Prior Art