Wireless multi-modal federated learning system for modal heterogeneous users

By optimizing user scheduling and bandwidth allocation in the wireless multimodal federated learning system, combined with the decision-level fusion model, the problem of inefficient training of modal heterogeneous users is solved, improving the classification accuracy of the model and saving energy overhead.

CN120529414AActive Publication Date: 2025-08-22SHANGHAI JIAOTONG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510817708.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-08-22
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

When facing modal heterogeneous users, the multimodal federated learning system in the existing wireless scenarios has problems such as increasing computational volume, complex communication process and inconsistent modal convergence speed, resulting in low model training efficiency and overfitting.

Method used

A wireless multimodal federated learning system is designed to optimize user scheduling and uplink bandwidth allocation through the collaborative work of servers and base stations, combining decision-level converged multimodal models, and to achieve efficient training of multimodal models using channel gain and energy constraints.

Benefits of technology

It improves the classification accuracy of multimodal models, saves 20% energy overhead, and adapts to a wide range of wireless multimodal federated learning application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120529414A_ABST
    Figure CN120529414A_ABST
Patent Text Reader

Abstract

The invention provides a wireless multi-modal federated learning system for modal heterogeneous users. The wireless multi-modal federated learning system comprises a server, a base station and users, the server is directly deployed at a base station side or is connected with the base station through an optical cable, the server broadcasts a global multi-modal model to a user through the base station, issues a user scheduling result and an uplink bandwidth used by the user at the same time, and performs global aggregation after receiving a local multi-modal model uploaded by the user to obtain a new global multi-modal model; the base station uses a wireless network to complete communication between the server and the user; and the scheduled user uses own data to update the received model, and uploads the obtained local multi-modal model to the base station. According to the method, wireless multi-modal federated learning under modal heterogeneity is considered, a wireless multi-modal federated learning system is designed, the energy consumption of federated learning is reduced through joint optimization of user scheduling and uplink power, and the performance of a multi-modal model on single-mode and multi-mode data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of wireless communications and artificial intelligence, and in particular to a wireless multimodal federated learning system for modally heterogeneous users. Background Art

[0002] Existing federated learning systems and methods in wireless scenarios usually only consider the case where user data is unimodal. When user data is multimodal, the modal heterogeneity between users will bring new problems to multimodal federated learning in wireless scenarios, so corresponding systems and solutions are needed.

[0003] Modal heterogeneity in multimodal federated learning manifests itself in the fact that different users possess different modal data. Due to the limitations of user devices, modal heterogeneity is quite common in actual application scenarios. Currently, in the research field of multimodal federated learning, there are many solutions to the problem of modal heterogeneity, including reconstruction of missing modal representations, knowledge distillation based on heterogeneous models, and modification of multimodal fusion methods. In wireless scenarios with fixed latency and limited bandwidth and computing resources, training the required reconstruction models and reconstructing the missing modal representations will introduce a lot of additional computational effort. Similarly, the learning process of knowledge distillation on proxy datasets will also introduce additional computational effort. In addition, the collection and sharing of proxy datasets will also complicate the original communication process.

[0004] In addition to the difficulties in training methods, the modal heterogeneity of multimodality also leads to significant differences in computing and communication for users with different modalities. Due to the huge differences in the structure and modal data of each modal sub-model in multimodality, inconsistent convergence speeds of each modality are inevitable during the training process. This problem means that the absolute balance strategy for each modality cannot guarantee the rapid convergence of the multimodal model. On the contrary, overtraining on the converged modality not only wastes computing power, but also may lead to overfitting of the modal data. Accordingly, the unconverged modality cannot be fully trained.

[0005] Patent application document CN116386058A discloses a multimodal federated learning training method and apparatus, wherein the method comprises: inputting shared data into an initial server-side model to obtain an output global feature representation, and transmitting the global feature representation to a client; receiving a local feature representation generated by the client; aggregating the local feature representation transmitted by the client based on the global feature representation and the local feature representation to obtain an aggregated feature representation; and training the server-side model based on the aggregated feature representation to complete one round of model training. However, this patent cannot completely solve the existing technical problems, nor can it meet the requirements of the present invention. Summary of the Invention

[0006] In view of the defects in the prior art, the purpose of the present invention is to provide a wireless multimodal federated learning system for modality heterogeneous users.

[0007] The wireless multimodal federated learning system for modality-heterogeneous users provided by the present invention includes:

[0008] The server is configured to determine the user scheduling result and uplink bandwidth allocation result for each communication round, generate and broadcast a global multimodal model, receive local multimodal models uploaded by scheduled users, and perform global aggregation to update the global multimodal model;

[0009] The base station uses a wireless network to complete communication between the server and the user, collects the channel gain of the user equipment, broadcasts the global multimodal model, user scheduling results, and uplink bandwidth allocation results issued by the server, receives the local multimodal model uploaded by the scheduled user, and transmits it to the server;

[0010] There are multiple users, each of which has heterogeneous multimodal data, that is, the modalities owned by different users are any subsets of the set of all modalities; the scheduled user updates the received global multimodal model using the local multimodal data to obtain a local multimodal model, and uploads the local multimodal model to the base station using the allocated uplink bandwidth;

[0011] The server is directly deployed on the base station side or connected to the base station via an optical cable.

[0012] Preferably, the server decides the user scheduling vector in the nth communication round And the uplink communication bandwidth allocation result Where U is the total number of users, and the participating users in the nth communication round constitute the set The global multimodal model of the server is composed of M global unimodal sub-models, that is,

[0013] The base station uses the wireless network to complete the communication between the user and the server, and collects the channel gain of user i in the nth communication round as

[0014] U users constitute a set All modalities owned by the user constitute a collection The set of users with modality m is The modality set of user i is The number of modes is The local dataset is Among them D i is the size of the dataset, x i,m,j is the eigenvector of the sample mode, yi,j is the label of the sample; the scheduled user After updating locally, upload the local multimodal model The set of participating users with modal data m in the nth communication round

[0015] Preferably, during the local update process, user i uses the local multimodal dataset to calculate the multimodal loss function as follows:

[0016]

[0017] Where L(·) is the loss function; Indicates that the feature x of mode m is k,m,j Input to the global unimodal sub-model The output result is obtained, and then the output results of each modality owned by user i are averaged, and the output result of the multimodal model is obtained by using the decision-level fusion in the multimodal model; each unimodal sub-model in the multimodal model uses the unimodal features to complete the reasoning task of machine learning, and the sum of the user's unimodal loss function is expressed as:

[0018]

[0019] For user-owned modals The unimodal loss function is calculated as:

[0020]

[0021] Due to the heterogeneity of modalities, for the modalities that users are missing The unimodal loss function is defined as the global loss, and its gradient is defined as the global unimodal gradient, that is:

[0022]

[0023] Adding the two together gives the user's local loss function:

[0024] H i (θ n-1 )=F i (θ n-1 )+G i (θ n-1 )

[0025] In the user's local update process, the number of local training cycles in each communication round is set to 1, and the batch gradient descent method is used to obtain the local gradient:

[0026] H i (θ n-1 )=F i (θn-1 )+G i (θ n-1 )

[0027] The local gradient is also composed of M local sub-gradients, namely:

[0028]

[0029] For user-owned modals The corresponding local sub-gradient is calculated as:

[0030]

[0031] For the user-missing modality, the local sub-gradient is defined as the global gradient Finally, the local update formula of user i is obtained:

[0032]

[0033] Preferably, the global multimodal model on the server is obtained by aggregating the local multimodal models uploaded by the scheduled users. Its aggregate form is:

[0034]

[0035] The aggregation weight is defined as When no scheduled user has mode m, that is, At this time, the parameters of the single-modal sub-model of this modality are kept the same as at the end of the previous communication round, that is:

[0036]

[0037] The goal of the entire wireless multimodal federated learning system is to minimize the sum of the multimodal loss function and the unimodal loss function, which is expressed as:

[0038]

[0039] Preferably, when user i performs local update, computational delay and computational energy overhead are generated. Assuming that all users have the same computing power, that is, the CPU frequency f and energy consumption coefficient α are the same; use β m represents the number of CPU cycles required to train a sample on modality m. According to the decision-level fusion method, user i in modality The number of CPU cycles required to train one sample is: The user's computing delay and computing energy cost are:

[0040]

[0041] Preferably, after completing the local update, user i generates communication delay and communication energy overhead in the process of uploading the local multimodal model. The uplink communication of multiple users is realized through frequency division multiple access. According to Shannon's formula, in the nth communication round, user The uplink rate is:

[0042]

[0043] in, is the bandwidth allocated to user i, p is the uplink communication power, N0 is the power spectral density of white noise, is the channel gain;

[0044] In an FDMA system, the total bandwidth allocated to participating users shall not exceed the total bandwidth of the base station used for wireless multi-mode federated services:

[0045]

[0046] According to their own modalities, different users upload single-modal sub-models of different modalities. Single-modal sub-models of the same modality have the same structure, so l m The data length corresponding to the single-mode sub-model of mode m is represented by the total data volume of the user uplink communication. Thus, it represents the uplink delay of user i in the nth communication round:

[0047]

[0048] Then the energy consumption of user i in the uplink communication process is given as:

[0049]

[0050] Preferably, for each participating user in the nth communication round, within the specified maximum delay T max The local model is uploaded within 10 seconds, so the user The latency constraints need to be met:

[0051] T i n,com +T i cmp ≤T max

[0052] Assume that the energy allocated by the user to a single communication round of wireless multimodal federated learning is E add , thus the residual energy of user i in the nth communication round is:

[0053]

[0054] The residual energy of each communication round is accumulated and used in the high-cost communication round. For the entire training process, it is only necessary that the total residual energy of user i is not negative, that is, there is a constraint:

[0055]

[0056] Under the constraints of bandwidth, delay and energy, the decision on access users and bandwidth allocation is made with the loss function H(θ N ) is the objective function, and the performance optimization problem of wireless multimodal federated learning is expressed as:

[0057]

[0058] Where N→+∞ means that enough communication rounds are trained until the multimodal model converges.

[0059] Preferably, based on the properties that the loss function H(θ) has γ smoothness and ρ-Lipschitz continuity, the global gradient of the single-modal sub-model has an upper bound And the difference between the local gradient and the global gradient of the unimodal sub-model has an upper bound Under the assumption that , the upper bound of the loss function is derived as:

[0060]

[0061] There is a form Where γ is the smoothness coefficient, ρ is the Lipschitz continuity coefficient, are the upper bound coefficient of the m-mode gradient and the m-mode gradient difference coefficient of user i, respectively. The ideal aggregation weight when all users are connected

[0062] Preferably, the constant term H(θ 0 ), it is split into the sum of the differences of a series of global loss functions In addition, for the long-term average energy constraint C5, we first construct a virtual queue about energy, and then give the virtual queue mean stability constraint C5′ equivalent to C5:

[0063]

[0064] After equivalent transformation of the objective function and C5, we obtain the standard structure that can be solved using the Lyapunov optimization method:

[0065]

[0066] stC1,C2,C3,C4,C5′

[0067] According to the Lyapunov optimization method, the Lyapunov drift function with penalty term is given:

[0068]

[0069] Substitute the upper bound of performance, and then amplify it through a series of inequalities, ignoring the constant term and a n Irrelevant terms, we get the upper bound for optimization and the transformed optimization problem:

[0070]

[0071] stC1,C2,C3,C4

[0072] Objective function J1(a n ,B n ) is the upper bound of the performance of multimodal federated learning, and the second is the energy cost of the current communication round. The Lyapunov penalty factor V>0 can balance the relationship between the two: when V→+∞, the optimization method tends to increase the energy cost to improve the performance of wireless multimodal federated learning, while when V→0, the optimization method strives to reduce the energy cost and ignores the performance of wireless multimodal federated learning. The specific setting of V is determined according to different energy cost and performance requirements.

[0073] Preferably, due to a n is the combinatorial optimization variable, B n is a continuous optimization variable, which is written as the equivalent main problem P4 and decomposed into the combinatorial optimization subproblem P4.1 and the continuous optimization subproblem P4.2, respectively:

[0074]

[0075] stC1,C2,C3,C4.

[0076]

[0077] stC1,C3.

[0078]

[0079] stC2,C3,C4.

[0080] in, represents the optimal solution of P4.2, and the objective function of P4.2 is:

[0081]

[0082] P4.2 becomes a convex problem through a series of transformations, and the optimal bandwidth is given by the KKT condition:

[0083]

[0084] The related variables satisfy the relationship:

[0085]

[0086] The remaining formulas involved in the numerical solution are as follows:

[0087]

[0088]

[0089] Using the above relationship, the optimal bandwidth vector is obtained;

[0090] P4.1 is solved by immune algorithm, assuming that the randomly generated primary antibody set is Where S is the total number of antibodies, Let g be the user access vector corresponding to the s-th antibody of generation 0, set the current generation g = 0, and construct the affinity function aff to measure the quality of the antibody according to the objective function in P4.1:

[0091]

[0092] in, ι>0 is the exponential coefficient for adjusting the dispersion of affinity. Currently, for infeasible solutions, the affinity function is set to 0;

[0093] The antibody concentration function den of the immune algorithm measures the similarity between the antibody and other antibodies in the set. Its calculation formula is:

[0094]

[0095] The similarity function sim(·) is measured by comparing the Hamming distance dis(·) of two antibodies with the distance threshold Dis:

[0096]

[0097] The affinity of the antibody and the antibody concentration together determine the degree of motivation when the antibody is selected:

[0098]

[0099] Among them, ε1 and ε2 represent the weights corresponding to the affinity function and the antibody concentration function, respectively. The S / μ antibodies with the largest incentive in the selection antibody set constitute the set. Finally, each selected antibody is cloned into μ times of the original one, and these antibodies are randomly mutated to obtain the mutated set Select Collection The highest affinity antibodies, and randomly generated S / μ antibodies constitute the next generation antibody collection Then, the above operation is repeated for the new antibody set until the maximum number of iterations is reached, and the optimal

[0100] Compared with the prior art, the present invention has the following beneficial effects:

[0101] (1) The wireless multimodal federated learning system of the present invention can adapt to situations where users have heterogeneous modalities and can ensure that the multimodal model and each unimodal sub-model can achieve convergence by training with users' heterogeneous modal data, thus being applicable to a wide range of wireless multimodal federated learning application scenarios.

[0102] (2) Compared with similar algorithms, the solution of the present invention can improve the accuracy of multimodal data classification tasks by up to 4.62%, and can improve the accuracy of unimodal data classification tasks by up to 2.79%.

[0103] (3) The solution of the present invention takes into account the energy constraints of user equipment. While achieving the above-mentioned performance advantages, it saves at least 20% of energy overhead compared to similar algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0104] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:

[0105] Figure 1 A wireless multimodal federated learning system for modality-heterogeneous users;

[0106] Figure 2 It is a multimodal model based on decision-level fusion;

[0107] Figure 3a and Figure 3b The effects of different Lyanov penalty factors on the multi-mode performance, single-mode performance and energy consumption of the CREMA-D dataset and IEMOCAP dataset are shown respectively;

[0108] Figure 4a to Figure 4d The comparison of the multi-mode performance, single-mode performance and energy consumption of the present invention and other solutions on the CREMA-D dataset is shown respectively;

[0109] Figure 5a to Figure 5d The following are the comparisons of the multi-mode performance, single-mode performance and energy consumption of the present invention and other solutions on the IEMOCAP dataset. DETAILED DESCRIPTION

[0110] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.

[0111] Example

[0112] The present invention can be applied in next-generation wireless communication networks to assist in the efficient training of multimodal models. In next-generation wireless communication networks, some terminal devices are equipped with one or more sensing devices such as cameras, microphones, and sensors to jointly collect data and train models. This embodiment can be deployed in a base station or edge server to build a multimodal collaborative system for use in multiple scenarios such as intelligent driving, telemedicine, and industrial Internet of Things. By efficiently training multimodal models, tasks can be assisted faster and more accurately, improving user experience and enhancing the intelligence level of the system.

[0113] The embodiment of the present invention discloses a wireless multimodal federated learning system, such as Figure 1 As shown, it includes U users, base stations, and servers. All users constitute the set All modes of data constitute a set Conversely, user i has a modality set of It represents the number of modalities of user i. The local data set of user i under this modality is expressed as in is the size of the dataset, x i,m,j is the eigenvector of mode m of sample j, y i,j is the label of sample j. All users with m modalities constitute the set

[0114] Similar to single-modal federated learning, the entire process of multimodal federated learning is also divided into N communication rounds. Each communication round consists of model delivery, local update, model upload, and global aggregation. In the nth communication round, the multimodal global model θ n-1 First, it is broadcast to all users. For the first communication round, define θ 0 is a randomly initialized multimodal model. It should be noted that, depending on the different modal fusion methods, the multimodal models used in multimodal learning are also different. In this invention, in order to adapt to the modal heterogeneity between users and avoid the redundant burden of calculation and communication, a multimodal model based on decision-level fusion is adopted. Its specific architecture is as follows: Figure 1 As shown, the global multimodal model is composed of M global single-modal sub-models, that is, During the training of the nth communication round, the global unimodal sub-model obtained in the n-1th communication round is According to the input feature vector x of user i in modality m i,m,j Calculate the output If some modalities are missing, no output of that modality will be generated. The output of each modality is then averaged using the fusion device to obtain the output of the multimodal model. Therefore, after the user receives the global multimodal model sent by the server, the multimodal loss calculated based on the local dataset is:

[0115]

[0116] Among them, L() is the loss function, which is generally in the form of cross entropy for classification tasks and mean square error for regression tasks; The global unimodal submodel obtained for the n-1th communication round y i,j is the sample label;

[0117] In addition to the multimodal loss function, each unimodal sub-model in the multimodal model can also use unimodal features to complete machine learning reasoning tasks. Therefore, the sum of the user's unimodal loss functions is expressed as:

[0118]

[0119] For user-owned modals The unimodal loss function can be calculated as:

[0120]

[0121] Among them, v m is the weight of the unimodal loss function, is the input feature vector.

[0122] Due to the heterogeneity of modalities, for the modalities that users are missing The unimodal loss function is defined as the global loss, and its gradient is defined as the global unimodal gradient That is:

[0123]

[0124] The above definition is to unify the form of multimodal loss function and unimodal loss function. The two are added together to obtain the user's local loss function form:

[0125] H i (θ n-1 )=F i (θ n-1 )+G i (θn-1 ).

[0126] During the user's local update process, the number of local training cycles in each communication round is set to 1. At the same time, the batch gradient descent method is used to add the multimodal gradient and the unimodal gradient to obtain the local gradient:

[0127]

[0128] Similar to the local model, the local gradient is also composed of M local sub-gradients, namely:

[0129]

[0130] For user-owned modals The corresponding local sub-gradient can be calculated as:

[0131]

[0132] For the user-missing modality, the local sub-gradient is also defined as the global gradient Finally, the local update formula of user i can be obtained:

[0133]

[0134] Where η is the learning rate set in advance. Although the above formula involves the model update of the missing mode, the calculation process is actually an identity about the global update

[0135] The goal of multimodal federated learning is to minimize the global loss function, which is the weighted sum of local loss functions. Therefore, the goal of the entire wireless multimodal federated learning system is to minimize the sum of the multimodal loss function and the unimodal loss function:

[0136]

[0137] User participation is expressed as participation vector To represent, the participating users in the nth communication round constitute a set The set of participating users with modal data m in the nth communication round

[0138] After the user completes local training, the local model will be uploaded. Figure 1 For a modal that the user has Its aggregate form is:

[0139]

[0140] The aggregation weight is defined as η is the learning rate set in advance.

[0141] Once no scheduled user has mode m, we have At this time, the parameters of the single-modal sub-model of this modality are kept the same as at the end of the previous communication round, that is:

[0142]

[0143] exist Figure 1 In the wireless multimodal federated learning architecture, at the beginning of each communication round, the server will first solve the performance optimization problem of wireless multimodal federated learning based on the user, channel and model conditions, thereby generating a user participation vector And the channel bandwidth used by users to upload local models in The 1-0 indicator variable indicating whether user i is scheduled and the allocated uplink bandwidth, and then the server sends a n , B n and multimodal global models, user The users do not participate in training and uploading in this communication round, and finally complete the global aggregation process with the participation of some users.

[0144] The communication process in wireless multimodal federated learning can be divided into two parts: base station downlink communication and user uplink communication. Communication between the base station and the server occurs via optical cables, offering high communication rates and negligible latency. Downlink communication relies on the base station's broadcast communication. Due to the base station's high downlink transmission power and ample downlink bandwidth, downlink latency is negligible compared to uplink latency. Similarly, since base stations generally have ample energy supplies, the focus is on the energy consumption of individual user terminals.

[0145] After the user completes the local update, the process of uploading the local multimodal model will generate communication delay and communication energy overhead. Since users have significant differences in the amount of data transmitted due to modal heterogeneity, the uplink communication of multiple users is implemented through Frequency Division Multiple Access (FDMA). According to Shannon's formula, in the nth communication round, the user The uplink rate is:

[0146]

[0147] in, is the bandwidth allocated to user i, p is the uplink communication power, N0 is the power spectral density of white noise, In an FDMA system, the total bandwidth allocated to participating users cannot exceed the total bandwidth B used by the base station for wireless multi-mode federation services. max :

[0148]

[0149] in, B is a 1-0 indicator variable indicating whether user i is scheduled; max is the total bandwidth used by the base station for wireless multimodal federated services;

[0150] Different users will upload unimodal sub-models of different modalities according to their own modalities, but unimodal sub-models of the same modality have the same structure, so l m It can represent the data length corresponding to the single-mode sub-model of mode m. The total data volume of the user uplink communication is The uplink delay of user i in the nth communication round can be expressed as:

[0151]

[0152] The energy consumption of user i during uplink communication can be given as:

[0153]

[0154] Local updates by users will cause computational delay and computational energy overhead. Assume that all users have the same computing power, that is, the CPU frequency f and energy consumption coefficient α are the same. Due to the different structures of the single-modal sub-models and the different input data characteristics of different modalities, the computing power required to complete the training on a sample is also different. m represents the number of CPU cycles required to train one sample on modality m. According to the decision-level fusion method, the computing power required to train multiple single-modality sub-models is additive, while the computing power of decision-level fusion is linearly related to the number of modalities: one addition of modal decisions requires β0 CPU cycles, The addition of modes requires CPU cycles. Although the local loss function includes multimodal local loss and unimodal local loss, the output of the unimodal sub-model has been obtained in the calculation of the multimodal local loss. Therefore, only the loss function of the unimodal sub-model output and the label needs to be calculated. This amount of calculation can be ignored compared to the amount of calculation of the neural network forward propagation. Similarly, the amount of calculation required for the unimodal local gradient can also be ignored. Thus, the user i in the modality The number of CPU cycles required to train one sample is Then we can express the user's computing delay and computing energy cost as follows:

[0155]

[0156] Among them, one addition of modal decision requires β0 CPU cycles;

[0157] For each participating user in the nth communication round, it is necessary to max The local model is uploaded within 10 seconds, so the user The latency constraints need to be met:

[0158] T i n,com +T i cmp ≤T max .

[0159] In addition to the time constraint, since users are generally wireless terminals with their own batteries and cannot afford excessive energy consumption, the energy allocated to a single communication round of wireless multimodal federated learning by users is set to E add , thus the residual energy of user i in the nth communication round is:

[0160]

[0161] Unlike delay, the residual energy of each communication round can be accumulated and used in high-cost communication rounds. For the entire training process, it is only necessary that the total residual energy of user i is not negative, that is, there is a constraint:

[0162]

[0163] Under the constraints of bandwidth, delay and energy, the decision on access users and bandwidth allocation is made with the loss function H(θ N ) is the objective function, and the performance optimization problem of wireless multimodal federated learning is expressed as:

[0164]

[0165] Here N→+∞ indicates that enough communication rounds are trained until the multimodal model converges.

[0166] Based on the fact that the loss function H(θ) has the properties of γ smoothness and ρ-Lipschitz continuity, the global gradient of the single-modal sub-model has an upper bound And the difference between the local gradient and the global gradient of the unimodal sub-model has an upper bound Under the assumption that , the upper bound of the loss function can be derived as:

[0167]

[0168] There is a form Where γ is the smoothness coefficient, ρ is the Lipschitz continuity coefficient, are the upper bound coefficient of the m-mode gradient and the m-mode gradient difference coefficient of user i, respectively. The above four coefficients can be obtained by estimation. in, is the m-modal gradient difference coefficient of user i; γ is the smoothness coefficient; ρ is the Lipschitz continuity coefficient; is the process variable; is the upper bound coefficient of the m-mode gradient; is the ideal aggregation weight when user i is fully connected; is the m-modal gradient difference coefficient of user i.

[0169] Subtract the constant term H(θ from the objective function of P1 0 ), it can be decomposed into the sum of the differences of a series of global loss functions In addition, for the long-term average energy constraint C5, first construct a virtual queue about energy And the virtual queue mean stability constraint C5′ equivalent to C5 is given:

[0170]

[0171] in, For virtual queues on energy

[0172] After equivalent transformation of the objective function and C5, we obtain the standard structure that can be solved using the Lyapunov optimization method:

[0173]

[0174] stC1,C2,C3,C4,C5′.

[0175] According to the Lyapunov optimization method, the Lyapunov drift function with penalty term can be given:

[0176]

[0177] After a series of inequalities, the constant terms are ignored and a n The irrelevant terms can be used to obtain the upper bound for optimization and the transformed optimization problem:

[0178]

[0179] stC1,C2,C3,C4.

[0180] Objective function J1(a n ,B n) is the upper bound of the performance of multimodal federated learning, and the second is the energy cost of the current communication round. The Lyapunov penalty factor V>0 can balance the relationship between the two: when V→+∞, the optimization method will tend to increase the energy cost to improve the performance of wireless multimodal federated learning, while when V→0, the optimization method will try to reduce the energy cost and ignore the performance of wireless multimodal federated learning. The specific setting of V can be determined according to different energy cost and performance requirements.

[0181] Due to a n is the combinatorial optimization variable, B n is a continuous optimization variable, which can be written as the equivalent main problem P4 and decomposed into the combinatorial optimization subproblem P4.1 and the continuous optimization subproblem P4.2, respectively:

[0182]

[0183] stC1,C2,C3,C4.

[0184]

[0185] stC1,C3.

[0186]

[0187] stC2,C3,C4.

[0188] in represents the optimal solution of P4.2, and the objective function of P4.2 is:

[0189]

[0190] P4.2 can be transformed into a convex problem through a series of transformations, and the optimal bandwidth can be given by using the KKT condition:

[0191]

[0192] The related variables satisfy the relationship:

[0193]

[0194] Although the above relationship cannot be given The closed form of , but can be solved numerically by Newton's iteration method. The remaining formulas involved are:

[0195]

[0196]

[0197] Using the above relationship, the optimal bandwidth vector can be obtained.

[0198] P4.1 can be solved by immune algorithm. Let the randomly generated primary antibody set be Where S is the total number of antibodies, is the user access vector corresponding to the s-th antibody in generation 0, and the current generation g is set to 0. According to the objective function in P4.1, the affinity function aff that measures the quality of the antibody can be constructed:

[0199]

[0200] in, ι>0 is the exponential coefficient for adjusting the dispersion of affinity; is the user access vector corresponding to the s-th antibody in the g-th generation; s is the antibody subscript; S is the total number of antibodies.

[0201] Currently, for infeasible solutions, the affinity function is set to 0. The antibody concentration function den of the immune algorithm measures the similarity between the antibody and other antibodies in the set. Its calculation formula is:

[0202]

[0203] The similarity function sim(·) is measured by comparing the Hamming distance dis(·) of two antibodies with the distance threshold Dis:

[0204]

[0205] The affinity of the antibody and the antibody concentration together determine the incentive for the antibody to be selected:

[0206]

[0207] Among them, ε1 and ε2 represent the weights corresponding to the affinity function and the antibody concentration function respectively; aff() is the affinity function; den() is the antibody concentration function.

[0208] Select S / μ antibodies with the highest motivation in the antibody set to form a set Finally, each selected antibody is cloned into μ times of the original one, and these antibodies are randomly mutated to obtain the mutated set Select Collection The highest affinity antibodies, and randomly generated S / μ antibodies constitute the next generation antibody collection Then, the above operation is repeated for the new antibody set until the maximum number of iterations is reached, thereby obtaining the optimal a n* .

[0209] The embodiment of the present invention discloses a wireless multimodal federated learning system. Figure 1 The basic structure of the present invention is described. Figure 2 Describes the basic model on which the invention is based, FIG3~ Figure 3b , Figure 4a to Figure 4d , Figure 5a to Figure 5d The performance of the invention is compared with that of other solutions.

[0210] Modal heterogeneity in multimodal federated learning manifests itself in the fact that different users possess varying modal data. Wireless multimodal federated learning is an emerging system that enables comprehensive training using data from various modalities across users. In wireless scenarios, latency is constant, and bandwidth and computing resources are limited. Modal heterogeneity in multimodal learning leads to significant differences in computing and communication between users with different modalities. The impact of modal heterogeneity on multimodal performance requires further characterization.

[0211] The present invention provides a wireless multimodal federated learning system (such as Figure 1 The system design incorporates the requirements of federated learning training performance and, within the constraints of communication and computing resources, improves the performance of multimodal federated learning through user selection and resource scheduling.

[0212] Users in the system possess multimodal data and differ in computing and communication capabilities. After updating the federated learning model using their own data, the user uploads it to the base station via the wireless network. The base station is configured with a server that receives the user-uploaded model and performs partial aggregation. Subsequently, the central server receives the model uploaded by the edge server and performs global aggregation, completing the federated learning task.

[0213] In view of the modal differences of user data in the system, the present invention aims to schedule user participation in the federated learning process, and at the same time comprehensively considers the performance of each modal model and the user's computing and communication capabilities, and designs a novel wireless multimodal federated learning system.

[0214] This paper adopts a multimodal model with decision-level fusion and constructs a single-modal loss function using the output of each modality, incorporating it into the final objective. This system design can adapt to heterogeneous user modalities and ensures that each modality can reach convergence.

[0215] This invention jointly optimizes user scheduling and uplink bandwidth allocation. User scheduling is based on real-time channel conditions, latency, and energy constraints, while also considering the training and convergence of each modality. This optimized design enables the system to optimally schedule appropriate users, improving the performance of wireless multimodal federated learning.

[0216] The wireless multimodal federated learning system of the present invention can adapt to situations where user modal heterogeneity exists, and can ensure that the multimodal model and each unimodal sub-model can achieve convergence by training with the user's modal heterogeneous data, and is suitable for a wide range of wireless multimodal federated learning application scenarios.

[0217] Those skilled in the art will appreciate that, in addition to implementing the system, device, and various modules provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same program in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, and the like by logically programming the method steps. Therefore, the system, device, and various modules provided by the present invention can be considered a hardware component, and the modules included therein for implementing various programs can also be considered structures within the hardware component; the modules for implementing various functions can also be considered both software programs for implementing the method and structures within the hardware component.

[0218] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.

Claims

1. A wireless multimodal federated learning system for modality-heterogeneous users, characterized by: include: The server is configured to determine the user scheduling result and uplink bandwidth allocation result for each communication round, generate and broadcast a global multimodal model, receive local multimodal models uploaded by scheduled users, and perform global aggregation to update the global multimodal model; The base station uses a wireless network to complete communication between the server and the user, collects the channel gain of the user equipment, broadcasts the global multimodal model, user scheduling results, and uplink bandwidth allocation results issued by the server, receives the local multimodal model uploaded by the scheduled user, and transmits it to the server; There are multiple users, each of which has heterogeneous multimodal data, that is, the modalities owned by different users are any subsets of the set of all modalities; the scheduled user updates the received global multimodal model using the local multimodal data to obtain a local multimodal model, and uploads the local multimodal model to the base station using the allocated uplink bandwidth; The server is directly deployed on the base station side or connected to the base station via an optical cable.

2. The wireless multimodal federated learning system for modality-heterogeneous users according to claim 1, characterized in that: The server decides the user scheduling vector in the nth communication round And the uplink communication bandwidth allocation result Where U is the total number of users, and the participating users in the nth communication round constitute the set The global multimodal model of the server is composed of M global unimodal sub-models, that is, The base station uses the wireless network to complete the communication between the user and the server, and collects the channel gain of user i in the nth communication round as U users constitute a set All modalities owned by the user constitute a collection The set of users with modality m is The modality set of user i is The number of modes is The local dataset is Among them D i is the size of the dataset, x i,m,j is the eigenvector of the sample mode, y i,j is the label of the sample; the scheduled user After updating locally, upload the local multimodal model The set of participating users with modal data m in the nth communication round 3. The wireless multimodal federated learning system for modality-heterogeneous users according to claim 2, characterized in that: During the local update process, user i uses the local multimodal dataset to calculate the multimodal loss function as follows: Where L(·) is the loss function; Indicates that the feature x of mode m is k,m,j Input to the global unimodal sub-model The output result is obtained, and then the output results of each modality owned by user i are averaged, and the output result of the multimodal model is obtained by using the decision-level fusion in the multimodal model; each unimodal sub-model in the multimodal model uses the unimodal features to complete the reasoning task of machine learning, and the sum of the user's unimodal loss function is expressed as: For user-owned modals The unimodal loss function is calculated as: Due to the heterogeneity of modalities, for the modalities that users are missing The unimodal loss function is defined as the global loss, and its gradient is defined as the global unimodal gradient, that is: Adding the two together gives the user's local loss function: H i (i n-1 )=F i (i n-1 )+G i (i n-1 ) In the user's local update process, the number of local training cycles in each communication round is set to 1, and the batch gradient descent method is used to obtain the local gradient: H i (i n-1 )=F i (i n-1 )+G i (i n-1 ) The local gradient is also composed of M local sub-gradients, namely: For user-owned modals The corresponding local sub-gradient is calculated as: For the user-missing modality, the local sub-gradient is defined as the global gradient Finally, the local update formula of user i is obtained:

4. The wireless multimodal federated learning system for modality-heterogeneous users according to claim 3, characterized in that: The global multimodal model on the server is obtained by aggregating the local multimodal models uploaded by the scheduled users. Its aggregate form is: The aggregation weight is defined as When no scheduled user has mode m, that is, At this time, the parameters of the single-modal sub-model of this modality are kept the same as at the end of the previous communication round, that is: The goal of the entire wireless multimodal federated learning system is to minimize the sum of the multimodal loss function and the unimodal loss function, which is expressed as:

5. The wireless multimodal federated learning system for modality-heterogeneous users according to claim 4, characterized in that: When user i performs local updates, computational delay and computational energy overhead are incurred. Assuming that all users have the same computing power, that is, the CPU frequency f and energy consumption coefficient α are the same; use β m represents the number of CPU cycles required to train a sample on modality m. According to the decision-level fusion method, user i in modality The number of CPU cycles required to train one sample is: The user's computing delay and computing energy cost are:

6. The wireless multimodal federated learning system for modality-heterogeneous users according to claim 5, characterized in that: After completing the local update, user i generates communication delay and communication energy overhead in the process of uploading the local multimodal model. The uplink communication of multiple users is realized through frequency division multiple access. According to Shannon's formula, in the nth communication round, user The uplink rate is: in, is the bandwidth allocated to user i, p is the uplink communication power, N0 is the power spectral density of white noise, is the channel gain; In an FDMA system, the total bandwidth allocated to participating users shall not exceed the total bandwidth of the base station used for wireless multi-mode federated services: According to their own modalities, different users upload single-modal sub-models of different modalities. Single-modal sub-models of the same modality have the same structure, so The data length corresponding to the single-mode sub-model of mode m is represented by the total data volume of the user uplink communication. Thus, it represents the uplink delay of user i in the nth communication round: Then the energy consumption of user i in the uplink communication process is given as:

7. The wireless multimodal federated learning system for modality-heterogeneous users according to claim 6, characterized in that: For each participating user in the nth communication round, within the specified maximum delay T max The local model is uploaded within 10 seconds, so the user The latency constraints need to be met: T i n,com +T i cmp ≤T max Assume that the energy allocated by the user to a single communication round of wireless multimodal federated learning is E add , thus the residual energy of user i in the nth communication round is: The residual energy of each communication round is accumulated and used in the high-cost communication round. For the entire training process, it is only necessary that the total residual energy of user i is not negative, that is, there is a constraint: Under the constraints of bandwidth, delay and energy, the decision on access users and bandwidth allocation is made with the loss function H(θ N ) is the objective function, and the performance optimization problem of wireless multimodal federated learning is expressed as: Where N→+∞ means that enough communication rounds are trained until the multimodal model converges.

8. The wireless multimodal federated learning system for modality-heterogeneous users according to claim 7, characterized in that: Based on the fact that the loss function H(θ) has the properties of γ smoothness and ρ-Lipschitz continuity, the global gradient of the single-modal sub-model has an upper bound And the difference between the local gradient and the global gradient of the unimodal sub-model has an upper bound Under the assumption that , the upper bound of the loss function is derived as: There is a form Where γ is the smoothness coefficient, ρ is the Lipschitz continuity coefficient, are the upper bound coefficient of the m-mode gradient and the m-mode gradient difference coefficient of user i, respectively. The ideal aggregation weight when all users are connected 9. The wireless multimodal federated learning system for modality-heterogeneous users according to claim 8, characterized in that: Subtract the constant term H(θ from the objective function of P1 0 ), it is split into the sum of the differences of a series of global loss functions In addition, for the long-term average energy constraint C5, we first construct a virtual queue about energy, and then give the virtual queue mean stability constraint C5′ equivalent to C5: After equivalent transformation of the objective function and C5, we obtain the standard structure that can be solved using the Lyapunov optimization method: stC1,C2,C3,C4,C5′ According to the Lyapunov optimization method, the Lyapunov drift function with penalty term is given: Substitute the upper bound of performance, and then amplify it through a series of inequalities, ignoring the constant term and a n Irrelevant terms, we get the upper bound for optimization and the transformed optimization problem: stC1,C2,C3,C4 Objective function J1(a n ,B n ) is the upper bound of the performance of multimodal federated learning, and the second is the energy cost of the current communication round. The Lyapunov penalty factor V>0 can balance the relationship between the two: when V→+∞, the optimization method tends to increase the energy cost to improve the performance of wireless multimodal federated learning, while when V→0, the optimization method strives to reduce the energy cost and ignores the performance of wireless multimodal federated learning. The specific setting of V is determined according to different energy cost and performance requirements.

10. The wireless multimodal federated learning system for modality-heterogeneous users according to claim 9, characterized in that: Due to a n is the combinatorial optimization variable, B n is a continuous optimization variable, which is written as the equivalent main problem P4 and decomposed into the combinatorial optimization subproblem P4.1 and the continuous optimization subproblem P4.2, respectively: stC1,C2,C3,C4. stC1,C3. stC2,C3,C4. in, represents the optimal solution of P4.2, and the objective function of P4.2 is: P4.2 becomes a convex problem through a series of transformations, and the optimal bandwidth is given by the KKT condition: The related variables satisfy the relationship: The remaining formulas involved in the numerical solution are as follows: Using the above relationship, the optimal bandwidth vector is obtained; P4.1 is solved by immune algorithm, assuming that the randomly generated primary antibody set is Where S is the total number of antibodies, Let g be the user access vector corresponding to the s-th antibody of generation 0, set the current generation g = 0, and construct the affinity function aff to measure the quality of the antibody according to the objective function in P4.1: in, ι>0 is the exponential coefficient for adjusting the dispersion of affinity. Currently, for infeasible solutions, the affinity function is set to 0; The antibody concentration function den of the immune algorithm measures the similarity between the antibody and other antibodies in the set. Its calculation formula is: The similarity function sim(·) is measured by comparing the Hamming distance dis(·) of two antibodies with the distance threshold Dis: The affinity of the antibody and the antibody concentration together determine the degree of motivation when the antibody is selected: Among them, ε1 and ε2 represent the weights corresponding to the affinity function and the antibody concentration function, respectively. The S / μ antibodies with the largest incentive in the selection antibody set constitute the set. Finally, each selected antibody is cloned into μ times of the original one, and these antibodies are randomly mutated to obtain the mutated set Select Collection The highest affinity antibodies, and randomly generated S / μ antibodies constitute the next generation antibody collection Then, the above operation is repeated for the new antibody set until the maximum number of iterations is reached, and the optimal

Citation Information

Patent Citations

  • Movie classification system based on multi-modal depth representation collaborative federated learning

    CN118468098A

  • Differential privacy defense mechanism performance evaluation method for cross-domain heterogeneous multi-modal federated learning

    CN118536154A

  • Hierarchical federal learning system and method in wireless mobility scene

    CN119623673A

  • Federated mining method and system for multimodal data based on multiple security policies

    US20250184362A1

  • Retransmission count and transmitting power allocation method for federated learning system, and gradient retransmission method

    WO2025050494A1