federated multi-modal large model direction preserving robust low-rank aggregation method to overcome client drift

CN122674079APending Publication Date: 2026-09-01GUANGXI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610782999.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0003]本发明的目的是为解决联邦多模态大模型微调中,因客观数据异质性(Non-IID)及算术平均规律导致的客户端漂移和“方向不稳定性”问题,而提供一种克服客户端漂移的联邦多模态大模型保向鲁棒低秩聚合方法

Benefits of technology

[0038] (1) Lightweight and efficient communication: Through the decoupled low-rank aggregation design on the server side, the complexity of uplink and downlink communication is strictly kept at the scale of low-rank space (i.e., Compared to full parameter fine-tuning of large multimodal models, the parameter transmission volume is compressed by orders of magnitude, greatly alleviating the communication bottleneck in distributed training of large multimodal models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122674079A_ABST
    Figure CN122674079A_ABST
Patent Text Reader

Abstract

This invention discloses a robust low-rank aggregation method for orientation preservation in federated multimodal large models that overcomes client drift, belonging to the fields of federated learning and privacy computing. The invention uses a pre-trained language model built into the multimodal large model as its backbone. Each client completes local LoRA training based on non-independent identically distributed data and uploads a low-rank matrix. The server generates a binary mask matrix with a uniform pruning rate to perform sparse pruning of the low-rank matrix. A task-aware diagonal scaling matrix and a cross-client normalization factor are autonomously constructed to compensate for scaling and normalize the low-rank matrix. The two types of low-rank matrices are then decoupled and weighted, and the global matrix is ​​distributed to the clients to update model weights. This invention solves the problems of client drift, low-rank orientation tampering, and long-tail knowledge forgetting in traditional federated aggregation under data heterogeneity, improving the orientation preservation and robustness of federated multimodal large model aggregation. It is suitable for privacy-constrained scenarios such as healthcare, finance, and edge IoT.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to federated learning, multimodal large language models, privacy computing, and distributed collaborative optimization, specifically a method for orientation-preserving robust low-rank aggregation of federated multimodal large models that overcomes client drift. Background Technology

[0002] In recent years, multimodal large language models have demonstrated outstanding performance in vertical fields such as image understanding and natural language processing. However, with the exponential growth of model parameters, their training and deployment face severe technical bottlenecks constrained by physical laws and objective hardware conditions: First, under the limitations of distributed computing and physical network transmission, updating all parameters of a large model with hundreds of billions of parameters requires enormous network bandwidth and GPU memory, and the latency of a single communication often far exceeds the tolerance limit of a normal business cycle. Second, in real-world scenarios such as healthcare, finance, or edge IoT, due to data privacy protection regulations, data cannot be uploaded centrally, necessitating the use of federated learning architectures for isolated local training. Combining federated learning with efficient parameter fine-tuning (such as LoRA) is a standard paradigm for alleviating communication bottlenecks. However, in the real physical world, the distribution of data collected by clients in different regions or at different levels is affected by the objective environment, exhibiting highly non-independent identically distributed (Non-IID) characteristics. Traditional federated aggregation algorithms (such as FedAvg) suffer from severe "client drift" when faced with this objectively existing data heterogeneity. Analysis from the perspective of Singular Value Decomposition (SVD) in low-rank spaces reveals that parameters representing long-tail knowledge in specific domains (such as rare disease features) in local models often have relatively small singular values. During global aggregation, due to the simple arithmetic averaging operation, clients with large datasets or feature-dominated data (whose head singular values ​​are extremely large) will forcibly distort and cover the low-rank subspace orientation of clients with smaller datasets, resulting in "directional instability." This "directional distortion," caused by both data distribution patterns and the fusion characteristics of linear algebraic addition, ultimately leads to catastrophic forgetting of long-tail-specific knowledge in the federated multimodal large model. Summary of the Invention

[0003] The purpose of this invention is to address the client drift and "directional instability" problems caused by objective data heterogeneity (Non-IID) and arithmetic mean regularity in the fine-tuning of federated multimodal large models. This invention provides a direction-preserving robust low-rank aggregation method for federated multimodal large models that overcomes client drift. This method adaptively protects the directional robustness of the low-rank matrices of each client during cloud aggregation, mitigating mutual interference of domain knowledge from different clients by reducing the difference between singular values ​​at the head and tail. This method requires no modification to the client's local training logic, is naturally adapted to low-bandwidth physical network environments, and possesses excellent communication efficiency and system deployment compatibility.

[0004] The technical solution to achieve the objective of this invention is:

[0005] A robust low-rank aggregation method for preserving orientation in a federated multimodal large model to overcome client drift is disclosed. The method involves the following entities: a client, a server, a multimodal large model, and a pre-trained language model and LoRA module deployed within the multimodal large model. The client is deployed at the data owner, possessing local data storage and computation capabilities, used for efficient parameter fine-tuning of the multimodal large model on locally highly non-independent and identically distributed (Non-IID) data. The server is deployed at a central node, responsible for performing amplitude-based sparsity pruning, complementary parameter scaling, and cross-client decoupled aggregation to protect the orientation robustness of the low-rank matrices of each client and overcome client drift. The pre-trained language model and LoRA module are deployed at both the client and server ends. The client trains and updates the low-rank matrix locally, while the server only aggregates the low-rank matrix. The method includes the following steps:

[0006] Assuming there are a total of The client participating in collaborative training, of which the first Each client is represented as ; in the In round-robin communication, this federated multimodal large model uses a built-in pre-trained language model as its backbone, targeting dimensions of... Freezing weights in pre-trained language models Each client trains a LoRA module based on its local non-independent and identically distributed data, and generates data with dimensions of [dimensions to be filled in]. and Low-rank update matrix and The size of the lower rank ;

[0007] During the server-side aggregation phase, the server sets a globally uniform sparsity pruning rate. And generate an element-wise binary mask matrix. and The binary mask matrix is ​​used to perform a masking operation on the low-rank matrix uploaded by the client to obtain the pruned sparse low-rank matrix. and The server is configured to perform federated complementarity parameter scaling, and the construction dimension is [dimension number missing]. Task-aware diagonal scaling matrix , of which The diagonal element of the rank is denoted as The scaled matrix obtained after compensation and scaling by the server is denoted as... The server uses a cross-client global normalization factor based on the intrinsic representation of parameters. The normalized matrix is ​​obtained. Finally, the global low-rank matrix generated by the server through decoupling and aggregation is denoted as follows: and The data is then distributed to each client; each client uses the received global low-rank matrix to fuse the model weights and obtain the complete weight increment. ;

[0008] 1) Local training and parameter upload:

[0009] Each client Perform standard LoRA training on local data to obtain the updated low-rank matrix. and The parameters of the low-rank matrix obtained from the training are then uploaded to the server.

[0010] 2) Server-side adaptive pruning:

[0011] After receiving the low-rank parameters uploaded by each client, the server performs sparsity pruning based on the parameter magnitude, retaining the first 1-p proportional elements with the largest absolute values ​​of the matrix elements, and setting the remaining matrix elements to zero, thus obtaining the pruned sparse low-rank matrix. and ;

[0012] 3) Federated complementary parameter scaling:

[0013] The server uses the matrix uploaded by the client. Based on the statistical characteristics, a dedicated diagonal scaling matrix is ​​constructed for each client. The server uses this diagonal scaling matrix to prune the corresponding client's matrix. Perform rank-dimension compensation scaling to obtain the scaled matrix. ;

[0014] 4) Heterogeneity-aware cross-client normalization and aggregation decoupling:

[0015] The server uses a diagonal scaling matrix. The constructed global normalization factor applies to the scaled matrix Normalization process is performed to obtain The server then prunes the low-rank matrices from each client. Compared with the normalized matrix Aggregation and decoupling are performed, and the global low-rank matrix is ​​obtained by weighted averaging. and normalized aggregation The final server will and Distributed to each client, each client according to... and Calculate the global weight increment It also updates the local multimodal large model, completing this round of communication iteration.

[0016] The specific process of server-side adaptive pruning in step 2) is as follows:

[0017] The server receives the set of low-rank parameters uploaded by all clients. Then, the server uses a globally uniform sparsity pruning rate. Generate the corresponding binary mask matrix element by element. and ;

[0018] The server updates the low-rank matrix uploaded by each client. and Retain the elements with the largest absolute values ​​within the matrix. The proportional element is set, and all remaining matrix elements are set to 0. The server calculates the pruned sparse low-rank matrix using the following Hadamard product operation. and , and It is calculated using the following formula:

[0019]

[0020]

[0021] in This represents the Hadamard product, which is an element-wise multiplication.

[0022] The specific process of scaling the federated complementary parameters in step 3) is as follows:

[0023] For each client The server construction dimension is And it belongs to the task-aware diagonal scaling matrix of the low-rank space The first diagonal scaling matrix Line number The diagonal elements of a column are defined as , of which Line number The diagonal element of the column corresponds to the first The scaling factor for each rank; the formula for calculating the scaling factor is:

[0024]

[0025] in Subsequently, the server uses this diagonal scaling matrix The low-rank matrix after pruning for the corresponding client Rank-dimension compensation scaling is performed, satisfying the following calculation formula:

[0026]

[0027] This yields the scaled low-rank matrix. .

[0028] The specific process of step 4), heterogeneity-aware cross-client normalization and aggregation decoupling, is as follows:

[0029] 4.1) Cross-client normalization: The server uses a normalization factor based on the intrinsic expressive power of the parameters.

[0030]

[0031] The server uses a diagonal matrix composed of normalization factors to scale the low-rank matrix obtained after the aforementioned scaling process. Perform a correction update to obtain a globally normalized low-rank matrix. ;

[0032] 4.2) Decoupling low-rank aggregation: The server decouples the low-rank matrices from each client after pruning. With the normalized low-rank matrix Perform aggregation and decoupling processing; where the low-rank matrix To perform the general feature projection function, the server uses either an arithmetic mean or a weighted average based on the amount of local data on the client side to perform aggregation calculations, resulting in a global low-rank matrix. :

[0033]

[0034] In the formula The aggregation weight coefficient is the one corresponding to the k-th client. When using arithmetic mean aggregation... =1 / K; when using data volume weighted aggregation The local data volume of the k-th client is the proportion of the total data volume of all clients. This is the low-rank matrix normalized by the server for each client. Perform global aggregation to obtain a global low-rank matrix. :

[0035]

[0036] 4.3) Model Deployment and Local Model Update: Finally, the server will aggregate the global low-rank matrix. and The data is distributed to each client. After receiving the global low-rank matrix, each client processes it according to the following formula. Calculate the model weight increments and update the pre-trained language model weights built into the local multimodal large model to complete the entire process of this round of federated collaborative training. This not only reduces the downlink communication complexity from Strictly maintain It also perfectly achieves directionally robust federated model fusion.

[0037] This technical solution achieves the following improvements:

[0038] (1) Lightweight and efficient communication: Through the decoupled low-rank aggregation design on the server side, the complexity of uplink and downlink communication is strictly kept at the scale of low-rank space (i.e., Compared to full parameter fine-tuning of large multimodal models, the parameter transmission volume is compressed by orders of magnitude, greatly alleviating the communication bottleneck in distributed training of large multimodal models.

[0039] (2) Privacy and security: Each client's private multimodal data (such as medical images, financial records, edge device logs, etc.) is always kept locally, and the low-rank matrix is ​​only shared between the server and the client. and This avoids the direct aggregation and transmission of sensitive data, meeting data compliance and strict privacy protection requirements.

[0040] (3) Robust aggregation and overcoming client drift: For heterogeneous client data with high non-independent identical distribution (Non-IID), adaptive pruning denoising and federated complementary parameter scaling implicitly and losslessly amplify the "long-tailed singular values" representing local domain-specific knowledge. This method effectively protects the directional stability of the low-rank matrices of each client, avoids the forced distortion of domain features of clients with smaller data volumes by clients with larger data volumes during aggregation, and completely alleviates the catastrophic forgetting and client drift problems that occur during the decoupling and aggregation process of global low-rank matrices.

[0041] (4) Dual improvement in global generalization and personalization performance: Compared with traditional federated aggregation baselines (such as FedAvg-LoRA), this technical solution retains more directional similarity and shows significant advantages in the generalization capability of global joint distribution and the local personalization performance of distribution to each client.

[0042] This technical solution overcomes the communication bottlenecks and challenges of heterogeneous, non-independent, and identically distributed data in federated scenarios for multimodal large models while protecting data privacy. Compared to traditional methods for fine-tuning federated large models, it not only significantly reduces communication overhead but also effectively preserves the client's long-tail-specific knowledge through orientation-protected robust low-rank aggregation. This solution provides an efficient, stable, and secure approach for the distributed collaborative deployment of multimodal large models in privacy-sensitive fields such as healthcare, finance, and the Internet of Things, promoting the widespread application of AI in edge computing and data silo scenarios. Attached Figure Description

[0043] Figure 1 This is a schematic diagram of the structural framework of an embodiment. Detailed Implementation

[0044] The present invention will be further described below with reference to the accompanying drawings and embodiments, but this is not intended to limit the scope of the invention.

[0045] Example:

[0046] This embodiment uses the "Smart Healthcare Multimodal Large-Scale Federated System" as an example. This system includes a server deployed at the headquarters of the National Health Commission, and multiple clients deployed at different levels of medical institutions (such as provincial tertiary hospitals, municipal specialized hospitals, and remote community clinics). Each hospital possesses multimodal private data containing medical record texts and CT / MRI images. Because tertiary hospitals have large amounts of data covering a wide range of diseases, while community clinics have smaller amounts of data but include rare long-tail cases from specific regions, the data distribution is extremely unbalanced (Non-IID).

[0047] Reference Figure 1 A robust low-rank aggregation method for orientation-preserving large federated multimodal models that overcomes client drift, the method comprising the following steps:

[0048] 1) Local training on non-independent and identically distributed data and uploading of low-rank parameters:

[0049] In the During round-robin communication, each hospital (client) Based on local private medical multimodal data, LoRA modules are trained on a locally pre-trained large-scale medical model. After training, each hospital will update the low-rank matrix. and (rank The parameters of the low-rank matrix obtained from training are encrypted and uploaded to the headquarters server via the medical private network. This reduces the amount of uplink communication data from the GB level to the MB level, breaking through the physical bandwidth limitations of the medical private network.

[0050] 2) Server-side adaptive pruning:

[0051] Due to variations in medical sensor equipment or overfitting from small local samples, the uploaded weights contain invalid noise. The headquarters server receives the low-rank matrix parameters uploaded by each hospital and then applies a globally uniform pruning rate. Generate a binary mask matrix and Perform the following masking operation, retaining the first mask with the largest absolute value. Scale parameter:

[0052]

[0053]

[0054] This operation is executed in parallel in the cloud, effectively "refining" the main medical focus of each hospital.

[0055] 3) Federated complementary parameter scaling:

[0056] The massive amounts of data from top-tier hospitals generate a powerful fusion "gravity," easily obscuring rare disease characteristics (extremely small tail singularities) uploaded by community clinics. Servers utilize a matrix that tends towards a uniform distribution. Statistical characteristics, constructing a diagonal scaling matrix for each hospital. The diagonal elements are calculated as follows:

[0057]

[0058] Subsequently, the low-rank matrix of the corresponding client was pruned. After compensating and scaling, we get: Since the basic nodes containing knowledge of rare diseases suffer a greater proportion of parameter loss during pruning, this formula adaptively calculates an amplification factor greater than 1, mathematically enhancing their tail singular values ​​without loss, thus endowing them with extremely strong resistance to torsion.

[0059] 4) Heterogeneity-aware cross-client normalization and decoupled aggregation:

[0060] To break the data dominance of top-tier hospitals, the server calculates a cross-client normalization factor based on the intrinsic expressive power of parameters:

[0061]

[0062] The low-rank matrix obtained by scaling each hospital as described above. Updated to a globally normalized version: .

[0063] Subsequently, decoupling and aggregation occur: a global low-rank matrix that undertakes the projection of general medical features. A global low-rank matrix carrying in-depth pathological knowledge .

[0064] Finally, the headquarters server sent a very small number of data packets to all hospitals. and Various hospitals utilize The weight update is complete. This example method successfully preserves the long-tail characteristics of primary healthcare while protecting data privacy, achieving robust global collaborative evolution.

[0065] To verify the effectiveness of this example in solving the client drift and orientation instability problems caused by the high degree of non-independent and identically distributed (Non-IID) data in the fine-tuning of federated multimodal large models, this embodiment conducted a systematic simulation experiment under the above-mentioned smart healthcare multimodal large model federated system, and demonstrated its technical effect through specific experimental data.

[0066] 1. Experiment and Dataset Configuration

[0067] This experiment simulates a scenario where multiple medical institutions (clients) collaboratively fine-tune a multimodal large model. To comprehensively evaluate the orientation-preserving robustness of this example for multimodal and heterogeneous data, the experiment introduces four standard benchmark datasets covering different multimodal tasks, corresponding to the feature distributions of local medical / general visual question answering and image-text understanding for different clients:

[0068] (1) Flickr dataset: used to evaluate cross-modal image and text retrieval and feature alignment capabilities;

[0069] (2) IconQA dataset: contains a large number of charts and abstract graphical question and answer, used to simulate clinically specific long-tail indicators and symbolic graphical reasoning;

[0070] (3) ScienceQA dataset: a multimodal science question answering dataset with high contextual dependence, used to test the generalization reasoning depth of the model;

[0071] (4)OCRVQA dataset: a text-intensive visual question answering dataset used to simulate the extraction of strong text features from medical documents, image reports, etc.

[0072] In this embodiment, the local data distribution of each client in the network exhibits a significant non-independent and identically distributed characteristic. The hidden layer dimension d of the fine-tuned base large model... in and d out For the standard dimensions of the large model, the rank of the LoRA module is set to r=8 or r=16.

[0073] 2. Verification and Comparative Analysis of Technical Effects

[0074] To definitively demonstrate the gains of each key technical step (server-side adaptive pruning, federated complementary parameter scaling, and heterogeneity-aware cross-client normalization) in overcoming client drift, the experiment selected several mainstream distributed large-model weight fusion / federated aggregation algorithms in the industry as comparison baselines, including:

[0075] (1) Zero-Shot: A large base model that has not been federated for fine-tuning;

[0076] (2) Multi-Task: The upper limit of the ideal state for centralized training;

[0077] (3) Task Arithmetic, DARE, Tie-merging, PCB-merging: Traditional parameter merging methods based on arithmetic average or simple symbol mask.

[0078] The accuracy (%) of each method on the above four multimodal datasets is shown in Table 1.

[0079] Table 1. Accuracy comparison of various aggregation methods on multimodal datasets (%)

[0080]

[0081] 3. Analysis of Technical Effects

[0082] Based on the experimental results in Table 1, the specific technical effects of each key step of the method in this example are as follows:

[0083] (1) The technical effect of adaptive sparsity pruning: Traditional methods such as Tie merging and PCB merging suffer from a decline in accuracy (only 34.46% and 37.01%, respectively) when faced with the highly heterogeneous IconQA dataset, due to their inability to effectively filter out small samples and hardware-specific noise. In this example, adaptive sparsity pruning in step 2) utilizes binary masks. and Accurately filtering out invalid low-dimensional noise with extremely small absolute values ​​successfully paves the way for subsequent aggregation and purification of core medical features.

[0084] (2) Technical effects of federated complementary parameter scaling and cross-client normalization: On IconQA and ScienceQA datasets, which contain a large amount of vertical domain-specific long-tail knowledge, traditional methods (such as task arithmetic) will use the feature hegemony of the data-rich to forcibly distort the low-rank subspace orientation of small nodes during simple addition fusion, resulting in orientation drift. However, the method in this example achieves an accuracy of 71.18% on IconQA and 80.48% on ScienceQA, even exceeding the upper limit of centralized multi-task training. This directly demonstrates that: the task-aware diagonal scaling matrix S constructed in step 3) k Combined with the intrinsic expression normalization factor in step 4), it can adaptively and losslessly amplify long-tail singular values, giving small nodes extremely strong anti-torsion ability, completely overcoming client drift and preventing catastrophic forgetting of long-tail knowledge.

[0085] (3) Synergistic effect of decoupling low-rank aggregation: by projecting the universal features onto the global low-rank matrix With the global low-rank matrix that carries knowledge of deep pathology Complete decoupling and aggregation are performed, and the final output is sent to the client. and This perfectly achieves directional robust fusion of low-rank matrices. Experimental data shows that this example strictly compresses the uplink and downlink physical communication complexity to a lightweight level of O(r × (d)). in + d out At the same time, it significantly improves the global generalization ability and local personalization performance of joint distribution, and realizes efficient and stable distributed large model co-evolution.

Claims

1. A robust low-rank aggregation method for orientation-preserving multimodal large models that overcomes client drift, characterized in that, The method involves the following entities: a client, a server, a multimodal large model, and a pre-trained language model and LoRA module deployed within the multimodal large model. The method includes the following steps: Assuming there are a total of The client participating in collaborative training, of which the first Each client is represented as ; in the In round-robin communication, this federated multimodal large model uses a built-in pre-trained language model as its backbone, targeting dimensions of... Freezing weights in pre-trained language models Each client trains a LoRA module based on its local non-independent and identically distributed data, and generates data with dimensions of [dimensions to be filled in]. and Low-rank update matrix and The size of the lower rank ; During the server-side aggregation phase, the server sets a globally uniform sparsity pruning rate. And generate an element-wise binary mask matrix. and The binary mask matrix is ​​used to perform a masking operation on the low-rank matrix uploaded by the client to obtain the pruned sparse low-rank matrix. and ; The server is configured to perform federated complementary parameter scaling, and the dimension is [missing information]. Task-aware diagonal scaling matrix , of which The diagonal element of the rank is denoted as The scaled matrix obtained after compensation and scaling by the server is denoted as... ; The server uses a cross-client global normalization factor based on the intrinsic representation of parameters. The normalized matrix is ​​obtained. ; Finally, the global low-rank matrices generated by the server through decoupling and aggregation are denoted as follows: and The data is then distributed to each client; each client uses the received global low-rank matrix to fuse the model weights and obtain the complete weight increment. ; 1) Local training and parameter upload: Each client Perform standard LoRA training on local data to obtain the updated low-rank matrix. and The parameters of the low-rank matrix obtained from the training are then uploaded to the server. 2) Server-side adaptive pruning: After receiving the low-rank parameters uploaded by each client, the server performs sparsity pruning based on the parameter magnitude, retaining the first 1-p proportional elements with the largest absolute values ​​of the matrix elements, and setting the remaining matrix elements to zero, thus obtaining the pruned sparse low-rank matrix. and ; 3) Federated complementary parameter scaling: The server uses the matrix uploaded by the client. Based on the statistical characteristics, a dedicated diagonal scaling matrix is ​​constructed for each client. The server uses this diagonal scaling matrix to prune the corresponding client's matrix. Perform rank-dimension compensation scaling to obtain the scaled matrix. ; 4) Heterogeneity-aware cross-client normalization and aggregation decoupling: The server uses a diagonal scaling matrix. The constructed global normalization factor applies to the scaled matrix Normalization process is performed to obtain The server then prunes the low-rank matrices from each client. Compared with the normalized matrix Aggregation and decoupling are performed, and the global low-rank matrix is ​​obtained by weighted averaging. and normalized aggregation The final server will and Distributed to each client, each client according to... and Calculate the global weight increment It also updates the local multimodal large model, completing this round of communication iteration.

2. The method for overcoming client drift in federated multimodal large model orientation-preserving robust low-rank aggregation according to claim 1, characterized in that, The specific process of server-side adaptive pruning in step 2) is as follows: The server receives the set of low-rank parameters uploaded by all clients. Then, the server uses a globally uniform sparsity pruning rate. Generate the corresponding binary mask matrix element by element. and ; The server updates the low-rank matrix uploaded by each client. and Retain the elements with the largest absolute values ​​within the matrix. The proportional element is set, and all remaining matrix elements are set to 0. The server calculates the pruned sparse low-rank matrix using the following Hadamard product operation. and , and It is calculated using the following formula: , , in This represents the Hadamard product, which is an element-wise multiplication.

3. The method for overcoming client drift in federated multimodal large model orientation-preserving robust low-rank aggregation according to claim 1, characterized in that, The specific process of scaling the federated complementary parameters in step 3) is as follows: For each client The server construction dimension is And it belongs to the task-aware diagonal scaling matrix of the low-rank space The first diagonal scaling matrix Line number The diagonal elements of a column are defined as , of which Line number The diagonal element of the column corresponds to the first Scaling factor for each rank; The scaling factor is calculated using the following formula: in Subsequently, the server uses this diagonal scaling matrix The low-rank matrix after pruning for the corresponding client Rank-dimension compensation scaling is performed, satisfying the following calculation formula: This yields the scaled low-rank matrix. .

4. The method for overcoming client drift in federated multimodal large model orientation-preserving robust low-rank aggregation according to claim 1, characterized in that, The specific process of step 4), heterogeneity-aware cross-client normalization and aggregation decoupling, is as follows: 4.1) Cross-client normalization: The server uses a normalization factor based on the intrinsic expressive power of the parameters. The server uses a diagonal matrix composed of normalization factors to scale the low-rank matrix obtained after the aforementioned scaling process. Perform a correction update to obtain a globally normalized low-rank matrix. ; 4.2) Decoupling low-rank aggregation: The server decouples the low-rank matrices from each client after pruning. With the normalized low-rank matrix Perform aggregation and decoupling processing; where the low-rank matrix To perform the general feature projection function, the server uses either an arithmetic mean or a weighted average based on the amount of local data on the client side to perform aggregation calculations, resulting in a global low-rank matrix. : In the formula The aggregation weight coefficient is the one corresponding to the k-th client. When using arithmetic mean aggregation... =1 / K; when using data volume weighted aggregation The local data volume of the k-th client is the proportion of the total data volume of all clients. This is the low-rank matrix normalized by the server for each client. Perform global aggregation to obtain a global low-rank matrix. : 4.3) Model Deployment and Local Model Update: Finally, the server will aggregate the global low-rank matrix. and The data is distributed to each client. After receiving the global low-rank matrix, each client processes it according to the following formula. Calculate the model weight increment, update the pre-trained language model weights built into the local multimodal large model, and complete the entire process of this round of federated collaborative training.