Multimodal Federated Learning Training Method and Device
By training the initial model on the server side and transmitting the global feature representation to the client for aggregation, the resource limitation problem in multimodal federated learning is solved, enabling larger-scale multimodal model training and task completion.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2023-02-17
- Publication Date
- 2026-05-26
AI Technical Summary
Existing multimodal federated learning algorithms cannot effectively utilize the large amounts of unimodal and multimodal data on the client side to train multimodal models on the server side, due to limitations in the computing and storage resources of edge devices.
An initial model is trained on the server side. Shared data is input into the server-side model to generate a global feature representation, which is then transmitted to the client for training. The client generates a local feature representation based on the local data modality type. The training of the server-side model is completed by aggregating the global and local feature representations.
It enables the training of larger-scale multimodal models without being limited by client resources, and can effectively utilize client-side unimodal and multimodal data to complete a variety of multimodal tasks.
Smart Images

Figure CN116386058B_ABST