Personalized multi-modal federated learning method for heterogeneous edge device

By designing a personalized multimodal federated learning method in multimodal federated learning, the performance degradation caused by modal heterogeneity is solved, achieving higher model accuracy and lower privacy and communication costs.

CN120218189APending Publication Date: 2025-06-27BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510296038.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When existing multimodal federated learning methods face modal quantitative heterogeneity, modal features focus on heterogeneity and modal data heterogeneity, their performance deteriorates and it is difficult to effectively solve the problems of data privacy and communication costs.

Method used

A personalized multimodal federated learning method is designed. By dividing the training of the multimodal model network into two stages, and initializing the model and collaborative graph on the cloud server, using a graph-based personalized multimodal FL strategy, dynamically update the collaborative graph and optimize the personalized model of the client to adapt to heterogeneity.

Benefits of technology

It significantly improves the accuracy of each customer model, solves the problems of modal quantitative heterogeneity, modal features focusing on heterogeneity and modal data heterogeneity, and reduces data privacy risks and communication costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218189A_ABST
    Figure CN120218189A_ABST
Patent Text Reader

Abstract

A personalized multi-modal federated learning method for heterogeneous edge devices belongs to the field of federated learning, and comprises the following steps: in a multi-modal federated learning scene, a cloud server initializes a model and a collaboration diagram, and sends the initial model to each client; in order to solve the problem of modal quantity heterogeneity, training of a multi-modal model network is divided into two stages, and in the stage 1, federation training of single-modal clients and multi-modal clients is carried out; and stage 2, carrying out multi-modal customer personalized federation training. In the stage 1, the multi-modal model is decoupled, so that single-modal customers and multi-modal customers can cooperate on their common modalities, and the generalization and robustness of the modal feature extraction network are improved. According to the invention, a personalized multi-modal FL based on the graph is designed, different modal emphasis and modal data heterogeneous levels can be accurately adapted, and personalized effectiveness and collaboration benefits are balanced by dynamically updating the collaboration graph on the cloud server and optimizing the personalized model of the client.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of federated learning, and particularly relates to a personalized multi-modal federated learning method for heterogeneous edge devices. Background Art

[0002] Multi-modal sensing is booming in the real world. Because in scenarios with complex tasks such as autonomous driving, health monitoring, and human-computer interaction, relying solely on a single sensor modality cannot well complete the tasks. For example, in the autonomous driving task, visual images contain rich color and texture information, but are affected by lighting, weather changes, and viewpoints, resulting in a decline in model performance or even inability to work properly due to environmental changes, and there is also a possibility of privacy leakage; depth information provides geometric information of elements, which is beneficial for determining object boundaries; lidar can provide accurate position space information, but lacks a description of the object appearance, is vulnerable to bad weather, and is expensive; thermal imaging is convenient for identifying different objects through infrared detection; polarization information and event data are good at dealing with specular reflections and dynamic real-world scenarios.

[0003] Inspired by the five senses of humans, people wonder if it is possible to combine and utilize information from multiple modalities to enhance the perception of the surrounding environment; and the data of each modality has its own advantages in different scenarios. An intuitive idea is to use these modalities to complement each other. By fusing the features of multiple modalities, more accurate analysis and appropriate responses to events can be made in dynamic scenarios. There have been a large number of autonomous driving works that use lidar, radar, and camera modalities for driving-related tasks such as vehicle detection, object detection, and path planning; compared with the original single-modal method, the accuracy has been greatly improved. However, most current research focuses on models deployed on centralized devices. This method uses a large amount of raw data for training and then distributes the model to each customer for use. However, collecting raw training data brings serious data privacy problems.

[0004] The protection of data privacy is becoming increasingly important, especially in the medical field, where privacy protection is often the most important factor determining whether users use the technology; at the same time, uploading edge data will consume a large amount of communication resources - which is often expensive. With the continuous improvement of the computing power of edge devices, training models distributively at the terminal has become a promising way. Federated learning (FL) provides a paradigm for distributed learning: the terminal uses local data to train a local model, and the customer only uploads the model parameters. The cloud server aggregates the collected models and then distributes the updated model to the customer, repeating this process. In this way, the data remains local, protecting data privacy while reducing communication costs. In addition, the environment where edge devices are located is usually dynamically changing, such as different weather / roads in the autonomous driving task, different contexts / languages in the emotion recognition task. Through federated collaboration, the adaptability of the model in a dynamically changing environment can be enhanced.

[0005] However, most existing federated learning methods use unimodal data to train unimodal models, such as image data MNIST, CIFAR10, CIFAR100, text data, and audio data. These methods cannot be directly applied to multimodal systems because of the following differences. First, due to different installation requirements of customers or different manufacturers, customers often have different combinations of sensors, resulting in different types and numbers of modalities for customers. Different modality nodes have significantly different model architectures. How can FL aggregate the node models? At the same time, sensors may drop out due to failures (such as power surges) or energy depletion. The absence of modalities will lead to a decline or even failure of the multimodal model performance, hindering the application of normal modality data. How should the design of FL cope with such changing scenarios? In addition, different customers have different modalities that they focus on training. For example, in an emotion recognition system, due to personal habits or the long-term environmental requirements of users, for users who are more inclined to express themselves by typing, the prediction accuracy of their multimodal model network depends more on the text modality feature extraction network that captures the user's language and expression methods, while the prediction effect of the multimodal model network of users who often send voice depends more on the voice modality feature extraction network that recognizes oral intentions and tones. Therefore, the multimodal model networks of different users have modalities that they focus more on training. Finally, in a multimodal FL system, there is also the same data heterogeneity problem as in unimodal FL. Because of different environments, the data of the same modality collected by different users are different. At this time, directly aggregating models by FedAvg will cause serious model non-convergence. Generally speaking, multimodal FL has more serious heterogeneity problems than unimodal FL, affecting the accuracy and convergence of federated learning.

[0006] Heterogeneity in federated learning is a key issue. The heterogeneity existing in multimodal FL can be summarized as multimodal heterogeneity, including modality number heterogeneity, modality feature emphasis heterogeneity, and modality data heterogeneity. These several types of heterogeneity are coupled with each other, aggravating the differences between the models of different customers, and thus affecting the performance of the models obtained by federated learning training. For the above-mentioned heterogeneity problems, there have been multiple related works.

[0007] 1. Multimodal federated learning for modality number heterogeneity;

[0008] In a multimodal FL system, due to different installation requirements or budget constraints, customers may have different modalities. For example, some customers may only have modality A, some customers may only have modality B, while others may have both A and B. The diversity of modality combinations leads to different model architectures on different nodes, making the aggregation process more complex. Generally speaking, there are mainly two methods to solve this problem:

[0009] (1) Some methods use the same auto-encoder feature extraction network for both single-modal nodes and multi-modal nodes, which can be directly aggregated due to the same network structure. However, the network on multi-modal nodes focuses on extracting shared features of different modalities, while the network on single-modal nodes focuses on extracting features of the corresponding modality. Directly aggregating these networks will degrade the performance of single-modal and multi-modal models.

[0010] (2) Another method uses the structure of a conventional feature extraction network + classification network, grouping customers with the same modality combination in the same group and only allowing customers in the same group to collaborate. However, it limits the collaboration of single-modal and multi-modal nodes on one of their same modalities.

[0011] In addition, sensors may malfunction due to faults (such as power surges) or power depletion, resulting in modality loss. In the case of a significant lack of modality information, customers may blindly attach importance to modalities with poor training effects. In this case, customers may try to improve the performance of defective modalities without realizing that the performance degradation is due to the lack of a specific modality, thus limiting the effective utilization of available data.

[0012] 2. For multi-modal federated learning with heterogeneous modal feature emphasis;

[0013] The multi-modal model networks of customers have different training intensities on different modalities because different modalities on each customer contribute differently to the improvement of task performance. Existing research has visualized the contributions of different modal features and trained a group of random forest classifiers specific to individuals and activities for three activities of 10 people. The results show that for the "exercise" activity, accelerometer data plays an important role in identifying the activities of user 2, while the prediction for user 8 relies more on GPS coordinates. Therefore, the corresponding feature extraction networks for different modalities have been trained to different degrees.

[0014] To address this heterogeneity, some methods introduce additional attention modules locally, enabling the model to focus on the characteristic features of each client. However, this method requires significant modification of the model structure, and the attention modules added after each convolutional layer increase the computational and storage burdens on resource-constrained devices; some methods cluster customers based on modal emphasis, and only customers within the same cluster are allowed to collaborate. However, not all multi-modal datasets exhibit obvious clustering characteristics, which limits the scalability of this method. In addition, the dataset does not exhibit the same appropriate number of classification clusters in each FL round, and if customers are misclassified due to the pre-set number of clusters, the model accuracy will decrease significantly.

[0015] 3. For multi-modal federated learning with heterogeneous modal data;

[0016] Due to environmental and individual differences, data of the same modality on different clients is usually heterogeneous. Therefore, different data distributions of the same modality lead to significant inconsistencies in the same feature extraction network. Directly averaging and aggregating these networks may seriously damage the performance of the model, and the present invention requires a more complex aggregation strategy to adapt to the heterogeneity of modal data.

[0017] In multi-modal FL, personalized design usually also follows these methods. Using multi-task learning to capture the relationships between multi-modal nodes, however, this method ignores the negative impact of aggregating different modal feature extraction networks; some methods use hypergraph neural networks to model complex client relationships and achieve personalized aggregation, but it requires a global consensus prototype booster on the cloud server to obtain general knowledge from the public dataset and then guide the training of the client model. This method relies on the public dataset on the cloud server side, and hypergraph computing is very expensive, bringing a large amount of computational overhead; some methods adopt multi-task federated learning, regarding the activity recognition of each user as a separate task. The local model is divided into a shared module and a personalized module - the shared module uses FedAvg for aggregation, while the personalized module remains local to alleviate data heterogeneity. Although this strategy is effective, it requires significant modification to the model structure, and the additional attention module imposes huge computational and storage requirements on resource-limited devices; in some methods, each client sends its modal feature extraction network to the cloud server, and then the cloud server aggregates these networks and extracts features from the public multi-modal dataset, and then sends the features for different clients back to the clients to assist in the personalized training of the local decoding network. However, this method requires a complete public dataset, which may be difficult to obtain in privacy-sensitive fields such as healthcare, and the performance of the model highly depends on the public dataset. In addition, sending features may pose a risk of privacy leakage. Summary of the Invention

[0018] Currently, multi-modal learning improves the robustness and accuracy of the model by integrating complementary information from multiple modalities, solving the limitations of single-modal methods in dealing with complex tasks. To address privacy and communication constraints, distributed multi-modal model training adopts federated learning (FL), which only uploads model parameters instead of the original data. However, in practical systems, due to the coupling of three types of modal heterogeneity on edge devices: heterogeneous number of modalities, heterogeneous emphasis on modal features, and heterogeneous modal data, traditional multi-modal FL methods will experience a significant performance decline.

[0019] Therefore, the present invention considers how to design a personalized federated learning strategy for these coupled heterogeneities in federated learning in a multi-modal dynamic scenario to solve the difficult problems studied above and promote the deployment of federated learning in multi-modal systems.

[0020] In summary, in view of the deficiencies of the prior art, in a multi-modal dynamic scenario, considering the three types of heterogeneity of the number of modalities, the emphasis on modal features, and the heterogeneity of modal data of the device, starting from maximizing the performance of the local model, a personalized multi-modal federated learning training strategy for high-performance heterogeneous devices is designed.

[0021] The technical solutions adopted by the present invention to solve the technical problems are as follows:

[0022] A personalized multi-modal federated learning method for heterogeneous edge devices provided by the present invention includes the following steps:

[0023] In a multi-modal federated learning scenario, the cloud server initializes the model and the collaboration graph, and sends the initial model to each client;

[0024] In stage 1, single-modal clients and multi-modal clients perform federated training: multiple single-modal FLs for each modality are run in parallel on the cloud server; the cloud server sends each single-modal model to the corresponding single-modal client and the multi-modal client having that modality; each client updates the model through local stochastic gradient descent, and each client sends the single-modal model network to the corresponding single-modal FL; the FedAvg aggregation algorithm is used for the aggregation operation; at the end of stage 1, the single-modal clients receive the aggregated modal feature extraction network and classification network, and the multi-modal clients only receive the modal feature extraction network as the starting point of stage 2; this step is repeated until convergence;

[0025] In stage 2, multi-modal clients perform personalized federated training: the multi-modal clients send the modal feature extraction network and classification network corresponding to each modality to the cloud server respectively; the cloud server updates each collaboration graph; the cloud server sends the personalized model aggregated according to the collaboration graph of the current round to the multi-modal clients; the multi-modal clients update the model locally according to the optimization equation until the model is uploaded to the cloud server again in the communication round; this step is repeated until convergence.

[0026] Further, the network training of multi-modal nodes or faulty nodes is decoupled, and a phased training method is adopted.

[0027] Further, the multi-modal federated learning scenario includes a cloud server and N clients; each client contains M (M≥2) different modalities locally; for any given client c i (1≤i≤N), its local training dataset is D i ={s:(X,y)}, where X={x k |k∈M i},s represents the training sample, y represents the label corresponding to the training sample, k represents a certain modality of the client data, and x k represents the training data of modality k, Specified customer c i Available modalities on customer c i Containing |M i |(1 ≤ |M i | ≤ M) modalities; the goal of multi-modal FL is to train a set of models {θ i (x k |k ∈ M i}|1 ≤ i ≤ N}, customized for customers with different modality combinations M i , where θ i represents the model of customer i (which may be single-modal or multi-modal, depending on the number of modalities the customer has).

[0028] Furthermore, for the single-modal model network, the single-modal customer c i (s)(1 ≤ i ≤ N s ) trains the single-modal model θ i (s) according to its local data modality k ∈ {1, 2,..., M i}; the single-modal model consists of a modality feature extraction network and a classification network g k (·):

[0029]

[0030] For the multi-modal model network, on the multi-modal customer c i with M i (m)(1 ≤ i ≤ N m ), data from each modality will be separately input into the modality-specific modality feature extraction network:

[0031]

[0032] where is the feature of this modality extracted by the modality feature extraction network , D k represents the dimension of the modality k feature, x k represents the training data of modality k; then direct feature concatenation is performed to fuse all modality features:

[0033]

[0034] where the symbol represents the concatenation operation;

[0035] The resulting fused feature is input into several fully connected layers to produce an output; the multi-modal classification network based on fusion in the multi-modal model network is represented as:

[0036] y output= g(h i ) (4)

[0037] where g(·) represents a multi-modal classification network; the multi-modal model representation of the multi-modal customer c i (m) with M i modalities is expressed as:

[0038]

[0039] where represents a set of single-modal feature extraction networks for extracting the feature representation vectors of M i modalities.

[0040] Furthermore, in stage 2, the multi-modal customer c i (m) uploads the set of single-modal feature extraction networks and the classification network g(·) to the cloud server respectively; the cloud server aggregates the models, the set N k represents the total number of customers with modality k, and the customer set represents all customers with modality k; the set N c represents the total number of multi-modal customers with the same modality combination, and the customer set represents all multi-modal customers with the same modality combination; a collaboration graph G k (V k , W k ) is introduced to model the collaboration relationship between customers, c (V c , W c ) models the collaboration relationship between customers, is the adjacency matrix of the collaboration graph for the modality feature extraction network corresponding to modality k, is the adjacency matrix of the collaboration graph for the classification network; the (i, j)-th element in the adjacency matrix represents the collaboration intensity between the i-th customer and the j-th customer on the corresponding model part.

[0041] Furthermore, the collaboration intensity is estimated by the similarity of model parameters, and the cosine similarity is used to measure the model similarity; the optimization equation for the adjacency matrix of the collaboration graph of the modality feature extraction network corresponding to modality k of the i-th customer is:

[0042]

[0043] where represents the collaboration intensity between customer i and customer j on the modality feature extraction network corresponding to modality k, represents the parameters of the modality feature extraction network corresponding to modality k of customer i, represents the parameters of the modality feature extraction network corresponding to modality k of customer j, and the set Nk represents the total number of all customers with modality k; the first constraint in formula (6) ensures that all collaboration intensities remain non - negative, and the second constraint sets the overall collaboration budget for all customers.

[0044] Furthermore, the multi - modal model tends to preferentially train the modality that contributes the most to the task; the modality features are mainly represented by distance vectors, which consist of the variations of each single - modality feature extraction network:

[0045]

[0046] where represents the variation of the parameters of the modality feature extraction network corresponding to modality 1 of customer i, represents modality M of customer i i corresponding to the variation of the parameters of the modality feature extraction network.

[0047] Even further, client i will send the distance vector to the cloud server The optimization equation of the collaboration graph adjacency matrix of the classification network of the i - th customer is:

[0048]

[0049] where represents the distance vector that client j will send to the cloud server, represents the collaboration intensity between customer i and customer j on the classification network, and the set N c represents the total number of all multi - modal customers with the same modality combination; the cosine similarity is restricted within the range of [- 1,1]; the first constraint in formula (8) ensures the non - negativity of all collaboration intensities, and the second constraint sets the overall collaboration budget for all customers.

[0050] Furthermore, the customer downloads the model from the cloud server:

[0051]

[0052] where represents the modality feature extraction network corresponding to modality k of customer i, represents the collaboration intensity between customer i and customer j on the modality feature extraction network corresponding to modality k, represents the parameters of the modality feature extraction network corresponding to modality k of customer j, θ i,g(·) represents the parameters of the classification network of customer i, represents the collaboration intensity between customer i and customer j on the classification network, θ j,g(·) represents the parameters of the classification network of customer j, represents the classification network of customer i, θ iThe model representing client i, set N k Denotes the total number of all clients with modality k, set N c Denotes the total number of all multimodal clients with the same modality combination, M i Denotes the number of modalities, and k denotes a certain modality on the client.

[0053] Furthermore, the client locally trains the model θ i And it is optimized through the following optimization equation:

[0054]

[0055] Where, H(θ i ) represents the client's locally trained model θ i The optimization objective of, F i (θ i ) represents the loss function of the client's local model θx, Represents the modality feature extraction network parameters corresponding to modality 1 of client i, Represents modality M of client i i The corresponding modality feature extraction network parameters, Represents the modality feature extraction network parameters corresponding to modality 1 of client j, Represents modality M of client j i The corresponding modality feature extraction network parameters, Represents the collaboration intensity between client i and client j on the classification network, Represents the collaboration intensity between client i and client j on the modality feature extraction network corresponding to modality 1, Represents the collaboration intensity between client i and client j on the modality feature extraction network corresponding to modality M i The corresponding modality feature extraction network, N1 represents the total number of clients with modality 1, Represents the total number of clients with modality M i The first term of Equation (12) minimizes the empirical loss of the local task performance.

[0056] The beneficial effects of the present invention are:

[0057] A personalized multimodal federated learning method for heterogeneous edge devices provided by the present invention, its main technical route lies in:

[0058] Firstly, to solve the problem of heterogeneous modality numbers, the present invention divides the training of the multimodal model network into two stages. In stage 1, the multimodal model is decoupled, enabling unimodal clients and multimodal clients to collaborate on their common modalities, improving the generalization and robustness of the modality feature extraction network.

[0059] Secondly, in order to precisely adapt to different modal emphases and modal data heterogeneity levels, the present invention designs a graph-based personalized multi-modal FL, which balances personalized utility and collaborative benefits by dynamically updating the collaborative graph on the cloud server and optimizing the personalized model of the client.

[0060] Finally, the present invention evaluates the effectiveness of the algorithm on the MHAD dataset.

[0061] Compared with the prior art, a personalized multi-modal federated learning method for heterogeneous edge devices provided by the present invention has the following advantages:

[0062] 1) The present invention jointly considers three types of heterogeneity in the multi-modal federated scenario, namely, the heterogeneity of the number of modalities, the heterogeneity of modal emphasis, and the heterogeneity of modal data, while most of the existing studies only consider a single heterogeneity problem.

[0063] 2) Considering the heterogeneity of the modal combinations (including the number) of devices, the present invention extends multi-modal federated learning to various modal combination situations and designs a multi-modal model decoupling strategy with high universality and low complexity.

[0064] 3) Considering the heterogeneity of modal emphasis of devices, which is a unique phenomenon in multi-modal model training. According to the feature fusion scheme, the present invention analyzes that this heterogeneity will affect the classification network part of the model, and updates the personalized collaborative graph of the classification network according to this index to enhance the collaboration of similar users. At the same time, in order to reduce the algorithm complexity, the present invention decouples the personalized FL problem to be independently solved on the cloud server side and the client side.

[0065] 4) Considering the heterogeneity of modal data of devices, according to the feature fusion scheme, the present invention analyzes that this heterogeneity will affect the modal feature extraction network of the model, and forms a personalized collaborative graph corresponding to the modal feature extraction network according to this index to enhance the collaboration of similar users, and at the same time, the local optimization equation prevents overfitting.

[0066] 5) The graph-based personalized method of the present invention embeds the personalized model as a node on the graph and embeds the collaborative weight as the connection weight on the graph. The model aggregates according to the graph, and by dynamically updating the collaborative graph, the model propagates information in the graph, promotes knowledge sharing while maintaining personalization, not only improves the learning efficiency, but also can prevent malicious attacks. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 It is a flowchart of a personalized multi-modal federated learning method for heterogeneous edge devices provided by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0068] The present invention will be further described in detail below with reference to the accompanying drawings.

[0069] A personalized multi-modal federated learning method for heterogeneous edge devices provided by the present invention proposes a new personalized multi-modal FL framework, HeteroPMMFL. Through model decoupling, HeteroPMMFL effectively solves the problem of heterogeneous modal numbers on different clients; designs a personalized collaboration graph to alleviate the heterogeneity of modal emphasis and modal data among different multi-modal users, and strengthens the cooperation of similar users. This framework is applicable to any multi-modal dataset and can be easily extended to various modal combinations. Experimental results show that the present invention can significantly improve the accuracy of each client model in the presence of three types of heterogeneity.

[0070] See Figure 1 As shown, a personalized multi-modal federated learning method for heterogeneous edge devices provided by the present invention has the following specific implementation process:

[0071] I. System model;

[0072] The present invention considers a multi-modal federated learning scenario consisting of a cloud server and N clients. Each client can contain at most M (M≥2) different modalities locally. For any given client c i (1≤i≤N), its local training dataset is denoted as D i ={s:(X,y)}, where X={x k |k∈M i}), s represents the training sample, y represents the label corresponding to the training sample, k represents a certain modality of the client data, and x k represents the training data of modality k, specifies the available modalities on client c i , and client c i contains |M i |(1≤|M i |≤M) modalities; the goal of multi-modal FL is to train a set of models θ i (x k |k∈M i )|1≤i≤N} tailored for clients with different modality combinations M i , and θ i represents the model of client i (which may be single-modal or multi-modal, depending on the number of modalities the client has).

[0073] II. Client model;

[0074] In clustered federated edge learning, the goal of the system is to learn multiple models to meet the heterogeneous data on the devices. The federated learning training process specifically includes the following steps:

[0075] The present invention adopts a feature-level fusion strategy. First, single-modal features are independently extracted, then different-modal features are fused, and finally the fused features are input into a classification network to obtain the final output. Compared with the result level, this method allows for deeper modal interactions; moreover, the feature-level fusion method can better adapt to changes in modal combinations. One only needs to replace the different modal feature extraction networks according to the changes in the modalities, which can be pre-trained backbones or other deep learning models, such as CNNs for images and LSTMs for images; while the data fusion method is to decorate one modality with another modality. Once the modal combination changes, this data fusion method needs to be completely changed because data of different modalities have different characteristics.

[0076] For the single-modal model network, the single-modal client c i (s) (1 ≤ i ≤ N s ) trains the single-modal model θ i (s) according to its local data modality k ∈ {1, 2,..., M i}. The single-modal model consists of a modal feature extraction network and a classification network g k (·) and is defined as:

[0077]

[0078] For the multi-modal model network, on the multi-modal client c i with M i (m) (1 ≤ i ≤ N m ), data from each modality will be separately input into the modality-specific modal feature extraction networks:

[0079]

[0080] where is the feature of this modality extracted by the modal feature extraction network , D k represents the dimension of the feature of modality k, and x k represents the training data of modality k.

[0081] The set of single-modal feature extraction networks is denoted as for extracting the feature representation vectors of M i modalities. For simplicity and efficiency, the present invention selects direct feature concatenation to fuse all modal features:

[0082]

[0083] where the symbol represents the concatenation operation.

[0084] The obtained fused features are input into several fully connected layers to generate an output. The fusion-based multi-modal classification network in the multi-modal model network can be expressed as:

[0085] y output = g(h i ) (4)

[0086] where g(·) represents the multi-modal classification network. The multi-modal model of the multi-modal client c i with M i modalities is represented as:

[0087]

[0088] III. Personalized multi-modal federated learning strategy for optimizing the client model;

[0089] 1. Algorithm overview;

[0090] Although the networks of multi-modal models and single-modal models are different, the feature extraction networks for the same modality are the same. Therefore, for modality number heterogeneity, the present invention collaborates single-modal clients and multi-modal clients by untying the training of the multi-modal feature extraction network and dividing the training of the classification network g(·) into two stages.

[0091] Meanwhile, the present invention proposes a personalized multi-modal FL method to balance the personalized utility and collaborative benefits. The personalized multi-modal FL method proposed by the present invention only needs to optimize the aggregation strategy, without introducing excessive complexity and model modification, protects data privacy, and does not require a public dataset. At the same time, it should be noted that due to the limited labeled data on each node, compared with single-modal models, multi-modal models are more prone to overfitting. Therefore, the design strategy needs to balance between personalization and generalization. The present invention decomposes the personalized FL problem into the cloud server side and the client side, and proposes an algorithm to solve this optimization challenge. As Figure 1 shown, the specific process of a personalized multi-modal FL method proposed by the present invention is described as follows:

[0092] a) The cloud server initializes the model and the collaboration graph, and sends the initial model to each client;

[0093] b) Phase 1 - Federated Training for Unimodal and Multimodal Clients: Multiple unimodal FLs for each modality are run in parallel on the cloud server. The cloud server sends the modality models to the unimodal clients and the multimodal clients that have that modality. Each client updates the model through local Stochastic Gradient Descent (SGD), and then each client sends the unimodal model network to the corresponding unimodal FL. At the end of Phase 1, the unimodal clients receive the aggregated modality feature extraction network and classification network, while the multimodal clients only receive the modality feature extraction network, which serves as the starting point for Phase 2. Among them, the aggregation algorithm uses the Federated Averaging Algorithm (FedAvg). This step is repeated until convergence.

[0094] c) Phase 2 - Personalized Federated Training for Multimodal Clients: The multimodal clients send the modality feature extraction network and classification network corresponding to each modality to the cloud server respectively; the cloud server updates each collaboration graph according to the optimization equations (Equations (6) and (8)) designed in the present invention; the cloud server sends the personalized model aggregated according to the collaboration graph of the current round to the multimodal clients. The multimodal clients then update the model locally according to the optimization equation (Equation (12)) designed in the present invention until the model is uploaded to the cloud server again in the communication round. This step is repeated until convergence.

[0095] 2. Multimodal Federated Learning Strategy for Collaborating Different Modality-Combined Clients;

[0096] In Phase 1, the present invention decouples the multimodal model into a unimodal feature extraction network and a unimodal classification network, and uploads the unimodal models to the unimodal FL systems running in parallel specific to the modality. Unimodal clients and multimodal clients can collaborate on any of their common modalities, while avoiding the following problem: when using the same auto-encoder model structure, the unimodal client network focuses on extracting unimodal features, while the multimodal client network focuses on extracting modality-shared features, resulting in poor performance in directly fusing the auto-encoder. At the end of Phase 1, the unimodal clients receive the aggregated modality feature extraction network and classification network, while the multimodal clients only receive the modality feature extraction network, which serves as the starting point for Phase 2. In Phase 2, only the multimodal clients participate, and a personalized multimodal FL method is designed to balance the utility of personalization and the benefits of collaboration.

[0097] 3. Personalized Multimodal Federated Learning Strategy Based on Personalized Collaboration Graph;

[0098] The specific implementation process is as follows:

[0099] (1) First is the client uploading the model and the cloud server aggregating the model;

[0100] Multimodal client ci (m) Upload the set of single-modal feature extraction networks and the classification network g(·) to the cloud server respectively.

[0101] For model aggregation on the cloud server, the set N k represents the total number of all customers with modality k, and the customer set represents all customers with modality k; the set N c represents the total number of all multi-modal customers with the same modality combination, and the customer set represents all multi-modal customers with the same modality combination. To model the collaboration relationship between customers, the present invention introduces a collaboration graph is the adjacency matrix of the collaboration graph of the modality feature extraction network corresponding to modality k, is the adjacency matrix of the collaboration graph of the classification network. The (i,j)-th element in the adjacency matrix represents the collaboration strength between the i-th customer and the j-th customer on the corresponding model part.

[0102] (2) For modality data heterogeneity, it is considered that the more similar the data distribution of this modality is, the more similar the feature extraction networks on this modality are, thus enhancing collaboration. Considering the inaccessibility of data, the present invention estimates the collaboration strength through model parameter similarity. The more similar the models are, the stronger the collaboration is. Then the optimization equation of the adjacency matrix of the collaboration graph of the modality feature extraction network corresponding to modality k of the i-th customer is specifically:

[0103]

[0104] where, represents the collaboration strength between customer i and customer j on the modality feature extraction network corresponding to modality k, represents the parameters of the modality feature extraction network corresponding to modality k of customer i, represents the parameters of the modality feature extraction network corresponding to modality k of customer j.

[0105] Equation (6) uses the collaboration graph to precisely define the collaboration strength, thereby enhancing collaboration between similar customers, promoting local data homogeneity, and being able to adapt to different levels of data heterogeneity, and filtering out irrelevant or potentially malicious customers by minimizing the collaboration weights. The second innovation is to use cosine similarity instead of the commonly used l2 distance to measure model similarity. This method has two major advantages: 1) It is not affected by the absolute scale of model parameters and provides a more accurate similarity measure; 2) It standardizes the similarity within a fixed range, simplifying hyperparameter tuning across different settings. The cosine similarity is restricted to the range of [-1,1]. The first constraint in Equation (6) can ensure that all collaboration strengths remain non-negative, and the second constraint sets the overall collaboration budget for all customers.

[0106] (3) In addition, the present invention emphasizes the heterogeneity of modal feature emphasis, and the multi-modal model tends to preferentially train the modality that contributes the most to the task. The modal feature emphasis is represented by a distance vector, which consists of the changes of each single-modal feature extraction network:

[0107]

[0108] Among them, represents the change in the parameters of the modal feature extraction network corresponding to modality 1 of customer i, represents modality M of customer i i corresponding to the change in the parameters of the modal feature extraction network.

[0109] Client i will send the distance vector A greater change in the modality-specific modal feature extraction network indicates that the customer gives priority to the training of this modality, thereby affecting the extracted features and subsequently affecting the training position focus of the classification network. Then, the collaborative graph adjacency matrix optimization equation of the classification network of the i-th customer is specifically:

[0110]

[0111] Among them, represents the distance vector that client j will send to the cloud server, represents the collaboration strength between customer i and customer j on the classification network.

[0112] Cosine similarity is restricted within the range of [-1, 1]. The first constraint in formula (8) can ensure the non-negativity of all collaboration strengths, and the second constraint sets the overall collaboration budget for all customers.

[0113] (4) Then the customer downloads the model from the cloud server:

[0114]

[0115]

[0116] Among them, represents the classification network parameters of customer j, represents the classification network of customer j.

[0117] (5) Then the client locally trains the model θ i and optimizes it through the following optimization equation:

[0118]

[0119] Among them, H(θ i ) represents the client locally training the model θi Optimization objective, F i (θ i ) represents the local model θ of the client i 's loss function, represents the modal feature extraction network parameters corresponding to modality 1 of client i, represents modality M of client i i corresponding modal feature extraction network parameters, represents the modal feature extraction network parameters corresponding to modality 1 of client j, represents modality M of client j i corresponding modal feature extraction network parameters, represents the collaboration intensity between client i and client j on the classification network, represents the collaboration intensity between client i and client j on the modal feature extraction network corresponding to modality 1, represents client i and client j in modality M i corresponding collaboration intensity on the modal feature extraction network, N1 represents the total number of clients with modality 1, represents having modality M i total number of clients.

[0120] The first term of Equation (12) minimizes the empirical loss of the local task performance. In addition, considering the limited labeled data, multi-modal models are more prone to overfitting than single-modal models. Therefore, the design strategy must balance between personalization and generalization. The design of the collaboration graph ensures personalization. To ensure generalization, the present invention has: 1) Both the modal feature extraction network and the classification network should participate in global aggregation, rather than only a part of the model participating in aggregation as in some existing studies; 2) The local training optimizes the regularization terms of the second term and subsequent terms of Equation (12), which can maximize the similarity between the local model and the aggregated model, prevent excessive deviation from the aggregated model, and reduce the risk of overfitting.

[0121] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A personalized multimodal federated learning method for heterogeneous edge devices, characterized in that: The following steps are involved: In the multimodal federated learning scenario, the cloud server initializes the model and collaboration graph and sends the initial model to each client; In phase 1, federated training of unimodal clients and multimodal clients is performed: multiple unimodal FLs for each modality are run in parallel on the cloud server; the cloud server sends each unimodal model to the corresponding unimodal client and the multimodal client that owns the modality; each client updates the model through local stochastic gradient descent, and each client sends the unimodal model network to the corresponding unimodal FL; the FedAvg aggregation algorithm is used for aggregation; at the end of phase 1, unimodal clients receive the aggregated modal feature extraction network and classification network, and multimodal clients only receive the modal feature extraction network as the starting point of phase 2; This step is repeated until convergence; Phase 2 performs personalized federated training for multimodal customers: multimodal customers send the modal feature extraction network and classification network corresponding to each modality to the cloud server respectively; The cloud server updates each collaboration graph; the cloud server sends the personalized model obtained by aggregating the collaboration graphs of the current round to the multimodal client; the multimodal client updates the model locally according to the optimization equation until the model is uploaded to the cloud server again in the communication round; This step is repeated until convergence.

2. According to claim 1, a personalized multimodal federated learning method for heterogeneous edge devices is characterized in that: The network training of multimodal nodes or faulty nodes is decoupled and a phased training approach is adopted.

3. According to claim 1, a personalized multimodal federated learning method for heterogeneous edge devices is characterized in that: The multimodal federated learning scenario includes a cloud server and N clients; each client contains M (M ≥ 2) different modalities locally; for any given client c i (1≤i≤N), whose local training data set is D i ={s:(X,y)}, where X = {x k |k∈M i }, s represents the training sample, y represents the label corresponding to the training sample, k represents a certain mode of customer data, x k represents the training data of mode k, Designated Customer i Available modalities on client c i Contains|M i |(1≤|M i |≤M) modalities; the goal of multimodal FL is to train a model {θ i (x k |k∈M i )|1≤i≤N}, for different modal combinations M i Tailor-made for customers, i Represents the model of customer i.

4. According to claim 1, a personalized multimodal federated learning method for heterogeneous edge devices is characterized in that: For a single-modal model network, the single-modal client c i (s)(1≤i≤N s ) according to its local data mode k∈{1,2,…,M i }Training unimodal model θ i (s); The single modal model is composed of a modal feature extraction network and classification network g k (·) composition: For a multimodal model network, in a network with M i Multimodal client c i (m)(1≤i≤N m ), the data from each modality will be input into the modality-specific modality feature extraction network separately: in, It is the modal feature extraction network The extracted features of this mode, D k represents the dimension of the modal k feature, x k Represents the training data of modality k; then direct feature concatenation is performed to fuse all modality features: Among them, the symbol ⊕ represents the splicing operation; The obtained fusion features are input into several fully connected layers to generate output; the fusion-based multimodal classification network in the multimodal model network is expressed as: y output =g(h i ) (4) Among them, g(·) represents a multimodal classification network; with M i Multimodal client c i The multimodal model of (m) is expressed as: in, Represents a collection of unimodal feature extraction networks, used to extract M i The feature representation vector of each mode.

5. The personalized multimodal federated learning method for heterogeneous edge devices according to claim 1, characterized in that: In stage 2, the multimodal client c i (m) Collect the unimodal feature extraction networks and classification network g(·) are uploaded to the cloud server respectively; the cloud server aggregates the model and sets N k represents the total number of customers with mode k, the customer set represents all customers with mode k; the set N c Represents the total number of multimodal customers with the same modality combination, the customer set Represents all multimodal customers with the same modality combination; introduces the collaboration graph G k (V k ,W k ),G c (V c ,W c ) to model the collaborative relationships between customers, is the collaboration graph adjacency matrix of the modality feature extraction network corresponding to modality k, is the adjacency matrix of the collaboration graph of the classification network; the (i, j)th element in the adjacency matrix represents the collaboration intensity between the i-th customer and the j-th customer on the corresponding model part.

6. The personalized multimodal federated learning method for heterogeneous edge devices according to claim 5, characterized in that: The collaboration intensity is estimated by the similarity of model parameters, and the model similarity is measured by cosine similarity. The optimization equation of the collaboration graph adjacency matrix of the modal feature extraction network corresponding to the modality k of the i-th customer is: in, represents the collaboration intensity between customers i and j on the modality feature extraction network corresponding to modality k, represents the modal feature extraction network parameters corresponding to the modality k of customer i, represents the modal feature extraction network parameters corresponding to the modality k of customer j, set N k represents the total number of customers with modality k; the first constraint in formula (6) ensures that all collaboration intensities remain non-negative, and the second constraint sets the overall collaboration budget for all customers.

7. The personalized multimodal federated learning method for heterogeneous edge devices according to claim 1, characterized in that: Multimodal models tend to prioritize training the modality that contributes most to the task; Modal features focus on distance vector representation, which is derived from the changes in each unimodal feature extraction network. composition: in, represents the change of the modal feature extraction network parameters corresponding to modality 1 of customer i, The mode M of customer i i Corresponding changes in modal feature extraction network parameters.

8. The personalized multimodal federated learning method for heterogeneous edge devices according to claim 7, characterized in that: Client i will send the distance vector to the cloud server The optimization equation of the collaboration graph adjacency matrix of the classification network of the i-th customer is: in, represents the distance vector that client j will send to the cloud server, represents the collaboration intensity between customer i and customer j on the classification network, and the set N c Represents the total number of multimodal customers with the same modality combination; cosine similarity is constrained to be in the range [-1,1]; the first constraint in formula (8) ensures the non-negativity of all collaboration intensities, and the second constraint sets the overall collaboration budget for all customers.

9. The personalized multimodal federated learning method for heterogeneous edge devices according to claim 1, characterized in that: Customers download the model from the cloud server: in, represents the modal feature extraction network corresponding to modality k of customer i, represents the collaboration intensity between customers i and j on the modality feature extraction network corresponding to modality k, represents the modal feature extraction network parameters corresponding to the modality k of customer j, θ i,g(·) represents the classification network parameters of customer i, represents the collaboration intensity between customer i and customer j on the classification network, θ j,g(·) represents the classification network parameters of customer j, represents the classification network of customer i, θ i Represents the model of customer i, set N k represents the total number of customers with mode k, set N c represents the total number of multimodal customers with the same modality combination, M i represents the number of modes, and k represents a certain mode on the customer.

10. The personalized multimodal federated learning method for heterogeneous edge devices according to claim 1, characterized in that: Client local training model θ i And optimize it through the following optimization equation: Among them, H(θ i ) represents the client local training model θ i The optimization goal is i (θ i ) represents the client local model θ i The loss function is represents the modal feature extraction network parameters corresponding to modality 1 of customer i, The mode M of customer i i The corresponding modal feature extraction network parameters, represents the modal feature extraction network parameters corresponding to modality 1 of customer j, represents the mode M of customer j i The corresponding modal feature extraction network parameters, represents the collaboration intensity between customers i and j on the classification network, represents the collaboration intensity between customer i and customer j on the modal feature extraction network corresponding to modality 1, Indicates that customers i and j are in mode M i The corresponding modality feature extracts the collaboration intensity on the network. N1 represents the total number of customers with modality 1. Indicates that it has mode M i The first term of (12) minimizes the empirical loss of local task performance.