Federal large model adaptive low-rank fine tuning method and device, computer equipment and storage medium
By decomposing the low-rank adaptation matrix of a large language model into server-shared and client-private matrices and adopting a rank-aware aggregation strategy, the problems of low resource utilization and aggregation interference under heterogeneous and non-independent identically distributed data are solved, thus achieving efficient model training and improved generalization capabilities.
Patent Information
- Application Number
- CN202511055996.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-14
AI Technical Summary
In federated learning scenarios with heterogeneous resources and non-independent and identically distributed data, traditional federated learning methods suffer from low resource utilization, severe aggregation interference, and insufficient model generalization ability. In particular, during the fine-tuning of large language models, computational resources are not fully utilized, and communication overhead is high, resulting in low training efficiency.
The low-rank adaptation matrix of the large language model is decomposed into a server-shared matrix and a client-private matrix. The server-shared matrix is uniformly maintained by the server to represent a cross-domain general semantic representation, while the client-private matrix is updated independently by local data. The model is aggregated through a rank-aware aggregation strategy, and the actual effective rank is dynamically adjusted by combining an adaptive diagonal matrix to optimize resource allocation and reduce aggregation bias.
It improves resource utilization, reduces computational and storage redundancy, enhances model convergence and generalization capabilities, reduces communication overhead, adapts to devices with different resource environments, and improves model training efficiency and performance.
Smart Images

Figure CN120952106A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and more specifically to a method, apparatus, computer device, and storage medium for adaptive low-rank fine-tuning of federated large models. Background Technology
[0002] In key sectors such as finance and healthcare, with the deepening application of artificial intelligence technology, Large Language Models (LLMs) have demonstrated powerful capabilities in handling domain-specific tasks, leading to a growing demand for fine-tuning. However, data in these sectors is often scattered across multiple heterogeneous clients, such as financial institutions and medical facilities, and is subject to strict privacy regulations, preventing centralized training. This situation poses a significant challenge to the effective fine-tuning of large language models.
[0003] Traditional federated learning (FL), as a distributed machine learning framework, aims to train models using data distributed across various clients while protecting data privacy. However, applying traditional federated learning techniques to fine-tuning large language models has exposed several key technical bottlenecks, severely limiting the efficiency and effectiveness of model fine-tuning:
[0004] Insufficient resource heterogeneity adaptation: Existing federated learning methods typically require all clients participating in training to use the same fine-tuning modules and configurations, which leads to significant resource waste and performance imbalances in practical applications. Specifically, clients with strong computing power fail to fully utilize their computational resources when processing highly complex models (such as high-rank LoRA matrices); while clients with weaker computing power become bottlenecks in the overall training progress because they cannot effectively handle these highly complex models. For example, devices with weak computing power cannot process high-rank LoRA matrices, resulting in low overall training efficiency.
[0005] Aggregation interference problem: In traditional LoRA (Low-Rank Adaptation) aggregation, the A and B matrices are typically directly averaged. However, since matrix multiplication does not satisfy the distributive law, this simple averaging method leads to a significant deviation between the aggregation result and the ideal value. This aggregation interference problem is particularly prominent in non-independent and identically distributed (Non-IID) data scenarios, severely affecting the model's convergence and generalization ability.
[0006] High communication and computational overhead: Full parameter fine-tuning or unified high-rank LoRA update strategies require the transmission of a large number of parameters during implementation, which is an unbearable high-frequency communication burden for resource-constrained edge devices. High communication overhead not only limits training efficiency but may also lead to training interruptions or delays due to network bandwidth limitations, further affecting model performance.
[0007] Therefore, in federated learning scenarios with heterogeneous resources and non-independent and identically distributed data, how to efficiently achieve adaptive fine-tuning of large language models while effectively solving problems such as aggregation interference, low resource utilization, and insufficient generalization ability in traditional methods has become a key technical challenge that urgently needs to be addressed in the field of artificial intelligence. Summary of the Invention
[0008] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method, apparatus, computer device and storage medium for adaptive low-rank fine-tuning of federated large models.
[0009] To achieve the above objectives, the present invention adopts the following technical solution:
[0010] An adaptive low-rank fine-tuning method for federated large models is applied to federated learning systems containing servers and multiple heterogeneous clients. The method includes:
[0011] The low-rank adaptation matrix of the large language model is decomposed into a server-shared matrix and a client-private matrix. The server-shared matrix is maintained by the server and captures cross-domain general semantic representations, while the client-private matrix is independently updated by each client based on local data.
[0012] The client receives the shared matrix from the server and initializes the private matrix and the adaptive diagonal matrix;
[0013] The client dynamically trims low-contribution dimensions and adjusts the actual effective rank based on the standard deviation of each element in the adaptive diagonal matrix, and combines the shared matrix and the private matrix to obtain the updated private matrix;
[0014] The server receives the updated private matrices uploaded by each client, aggregates them using a rank-aware aggregation strategy, and fine-tunes the shared matrices using a public dataset to obtain a global model.
[0015] Deploy the global model to the server and each client.
[0016] This invention also provides a federated large model adaptive low-rank fine-tuning device, applied to a federated learning system comprising a server and multiple heterogeneous clients, the device comprising:
[0017] The decomposition unit is used to decompose the low-rank adaptation matrix of a large language model into a server-shared matrix and a client-private matrix. The server-shared matrix is maintained by the server and captures cross-domain general semantic representations, while the client-private matrix is independently updated by each client based on local data.
[0018] The initialization unit is used by the client to receive the shared matrix sent by the server and initialize the private matrix and the adaptive diagonal matrix.
[0019] The adjustment and combination unit is used by the client to dynamically prune low-contribution dimensions and adjust the actual effective rank based on the standard deviation of each element in the adaptive diagonal matrix, and combine the shared matrix and the private matrix to obtain the updated private matrix;
[0020] The aggregation fine-tuning unit is used by the server to receive the updated private matrices uploaded by each client, aggregate them using a rank-aware aggregation strategy, and fine-tune the shared matrix using a public dataset to obtain a global model.
[0021] The deployment unit is used to deploy the global model to the server and various clients.
[0022] The present invention also provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the above-described method.
[0023] The present invention also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0024] The advantages of this invention compared to existing technologies are as follows: By decomposing the low-rank adaptation matrix of a large language model into a server-shared matrix and a client-private matrix, optimized allocation of computing resources is achieved. The server-shared matrix is uniformly maintained by the server, capturing cross-domain general semantic representations and avoiding the repeated computation and storage of the same information on each client, thus saving computing resources. Simultaneously, the client-private matrix is independently updated by each client based on local data, allowing clients with stronger computing power to fully utilize their resources to process more complex model parts, while clients with weaker computing power can adapt to their resource limitations by simplifying the model structure. This resource adaptation mechanism improves resource utilization and solves the problem of insufficient heterogeneous resource adaptation in traditional methods. Furthermore, by introducing a rank-aware aggregation strategy, biases in the aggregation process are effectively reduced. This aggregation method considers the data volume and model complexity of different clients, making the aggregation result closer to the ideal value, thereby improving the model's convergence and stability, reducing aggregation interference, and enhancing the model's generalization ability.
[0025] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. Attached Figure Description
[0026] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a schematic diagram illustrating an application scenario of the adaptive low-rank fine-tuning method for federated large models provided in this embodiment of the invention.
[0028] Figure 2 A flowchart illustrating the adaptive low-rank fine-tuning method for a federated large model provided in an embodiment of the present invention;
[0029] Figure 3 A schematic block diagram of the adaptive low-rank fine-tuning device for a federated large model provided in an embodiment of the present invention;
[0030] Figure 4 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation
[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0032] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0033] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0034] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0035] Please see Figure 1 and Figure 2 , Figure 1 This is a schematic diagram illustrating an application scenario of the adaptive low-rank fine-tuning method for federated large models provided in this embodiment of the invention. Figure 2 This is a schematic flowchart illustrating the adaptive low-rank fine-tuning method for federated large models provided in this embodiment of the invention. This method is applied to a server that interacts with terminals. By decomposing the low-rank adaptation matrix of the large language model into a server-shared matrix and a client-private matrix, it optimizes the allocation of computing resources. The server-shared matrix is uniformly maintained by the server, capturing cross-domain common semantic representations and avoiding redundant computation and storage of the same information on each client, thus saving computing resources. Simultaneously, the client-private matrix is independently updated by each client based on local data. This allows clients with stronger computing power to fully utilize their resources to process more complex model parts, while clients with weaker computing power can simplify the model structure to adapt to their resource limitations. This resource adaptation mechanism improves resource utilization and solves the problem of insufficient resource heterogeneous adaptation in traditional methods. Furthermore, by introducing a rank-aware aggregation strategy, it effectively reduces bias in the aggregation process. This aggregation method considers the data volume and model complexity of different clients, making the aggregation result closer to the ideal value, thereby improving the model's convergence and stability, reducing aggregation interference, and enhancing the model's generalization ability.
[0036] Figure 2 This is a flowchart illustrating the adaptive low-rank fine-tuning method for federated large models provided in this embodiment of the invention. Figure 2 As shown, the method is applied to a federated learning system that includes a server and multiple heterogeneous clients, and includes the following steps S110 to S150.
[0037] S110. Decompose the low-rank adaptation matrix of the large language model into a server-shared matrix and a client-private matrix; wherein, the server-shared matrix is maintained by the server and captures cross-domain general semantic representations, and the client-private matrix is independently updated by each client based on local data;
[0038] Specifically, during the initialization phase of the federated learning system, the server decomposes the low-rank adaptation (LoRA) matrix of the large language model into two parts: a server-shared matrix A and a client-private matrix B. The shared matrix A is maintained uniformly by the server and is responsible for capturing cross-domain common semantic representations (such as basic language structures and common entity relationships), and remains frozen during training to avoid repeated transmission. The client-private matrix B is independently initialized by each client based on local data and is used to learn domain-specific knowledge (such as financial terminology and medical diagnostic logic).
[0039] In other words, by separating general and private knowledge, redundant computation and storage of shared components are avoided on each client, significantly reducing computational and storage redundancy. For example, financial institutions and medical institutions do not need to learn basic language models separately, but instead share the general semantic representation provided by the server, thus saving resources. Furthermore, the original data remains local to the client, with knowledge only indirectly transferred through updates to the private matrix B, complying with data privacy regulations.
[0040] In one embodiment, the decomposition of the low-rank adaptation matrix of the large language model into a server-shared matrix and a client-private matrix includes:
[0041] The server initializes a low-rank adaptive architecture for a large language model, which contains decomposable low-rank matrices.
[0042] Specifically, the server initializes a low-rank adaptation (LoRA) architecture for a large language model. The core of this architecture is a decomposable low-rank matrix. This matrix represents the incremental knowledge that the model needs to adjust during fine-tuning in a parameterized manner. Its rank is much smaller than the dimension of the original model parameter matrix, thereby significantly reducing computational and communication overhead.
[0043] The low-rank matrix is decomposed into two parts, one of which is designed as a server-shared matrix and the other is designed as a client-private matrix.
[0044] Specifically, during the initialization phase, the server explicitly decomposes the low-rank matrix into two sub-matrices:
[0045] Server-shared matrix (A): This matrix is maintained uniformly by the server and captures common semantic representations across domains (such as basic language structures and common entity relationships). It remains frozen during training to avoid redundant transmission and computation.
[0046] Client-side private matrix (B): Initialized independently by each client based on local data, used for learning domain-specific knowledge (such as financial terminology, medical diagnostic logic, etc.). The dimensions and rank of the private matrix can be dynamically adjusted according to client resources.
[0047] More specifically, assume the original low-rank matrix is M∈R d×d Its decomposition is M≈A·B, where A∈R d×r For a shared matrix, B∈R r×d This is a private matrix, where r << d represents the low-rank dimension. During initialization, the server can preset a base rank r based on the distribution of client resources, and then dynamically adjust it through an adaptive mechanism (e.g., reducing r for clients with weak computing power to reduce computational load).
[0048] In addition, domain-specific data (such as financial information and medical texts) is divided into several subsets based on the institution or device. Each client stores only its local data, ensuring that the original data does not leave its local machine. For example, a bank client stores financial transaction data, while a hospital client stores medical record data. During initialization, the server distributes a shared matrix A to all clients. The clients freeze A during local training and only update their private matrix B.
[0049] In other words, by separating the shared matrix A and the private matrix B, the redundant computation and storage of the common semantic representation on each client is avoided. For example, financial institutions and medical institutions do not need to learn the basic language model separately, but instead share A provided by the server, thus saving significant computing resources. Furthermore, the rank of the private matrix B can be dynamically adjusted according to the client's computing power (e.g., the rank is reduced from 64 to 32 for clients with weak computing power), enabling the model to adapt to devices with different resource environments and avoiding resource waste on clients with strong computing power and performance bottlenecks on clients with weak computing power. Additionally, the original data is always kept locally on the client, and knowledge is only indirectly transferred through updates to the private matrix B, complying with data privacy regulations in fields such as finance and healthcare. During the initialization phase, Laplace noise can be added to the initial values of either the shared matrix A or the private matrix B to further prevent data reconstruction attacks.
[0050] S120. The client receives the shared matrix sent by the server and initializes the private matrix and the adaptive diagonal matrix.
[0051] Specifically, at the start of each training round, the server initializes the shared matrix A and then distributes the frozen shared matrix A to all clients. Upon receiving it, the clients initialize their private matrix B (randomly initialized or based on the pre-trained model) and the adaptive diagonal matrix Λ (with diagonal elements initialized to 1 for dynamically adjusting feature importance).
[0052] More specifically, the server distributes a shared matrix A (capturing general semantics), while the client only updates its private matrix B (learning local domain knowledge). The parameter update formula is ΔW = B·Λ·A, where Λ is the client-adaptive diagonal matrix, dynamically adjusting the importance of each dimension. During local training, the shared matrix A is frozen, and only the private matrix B and the adaptive diagonal matrix Λ are fine-tuned, reducing computational load.
[0053] In other words, by pre-initializing the shared matrix, the client does not need to learn general semantics from scratch, accelerating model convergence. For example, a medical client can quickly focus on learning professional terminology after initialization, rather than basic language structure. The introduction of the adaptive diagonal matrix Λ provides the basis for subsequent dynamic pruning of low-contribution dimensions, enabling the model to automatically adjust its complexity based on data characteristics.
[0054] In one embodiment, the client receives a shared matrix from the server and initializes a private matrix and an adaptive diagonal matrix, including:
[0055] A communication connection is established between the client and the server, and the client receives the shared matrix sent by the server through this communication connection;
[0056] Specifically, the client and server establish a secure connection through encrypted communication protocols (such as TLS / SSL) to ensure the confidentiality and integrity of data (especially shared matrix A) during transmission. During the communication initialization phase, the server verifies the client's identity (e.g., certificate authentication) to prevent malicious nodes from accessing the network. For different network environments such as 5G and Wi-Fi, the client dynamically adjusts transmission parameters (e.g., chunk size, retransmission mechanisms) to ensure the reliable delivery of shared matrix A. For example, under low-bandwidth Wi-Fi, compressed transmission (e.g., sparse matrix coding) is used to reduce latency.
[0057] In addition, to adapt to the memory of edge devices, the server splits the shared matrix A into multiple fragments. The client downloads these fragments as needed and temporarily stores them in memory. After training is complete, only the necessary parts (such as the basic semantic layer) are retained. After receiving A, the client sets its parameters to be untrainable (i.e., frozen) to prevent accidental modification during local training and ensure the stability of the general semantic representation.
[0058] The client initializes a private matrix based on local data characteristics and resource conditions; at the same time, the client initializes an adaptive diagonal matrix.
[0059] Specifically, the client initializes the dimensions and rank of the private matrix B based on local data characteristics (such as the vocabulary distribution of financial texts and the entity types of medical data). For example, if a medical client detects high-frequency occurrences of disease names in local medical record data, it increases the initial rank of the disease-related dimensions in B. The client dynamically sets the initial rank of B based on its own computing power (such as GPU memory and CPU core count). Clients with strong computing power (such as bank servers) initialize a high-rank B (e.g., 64 dimensions), while weaker clients (such as mobile phones) initialize a low-rank B (e.g., 32 dimensions). Here, Λ is a diagonal matrix with the same number of columns as the private matrix B, and the diagonal elements λ... i Let represent the importance weight of the i-th feature. Initially, all λ i Set to 1, and dynamically adjust through training. To prevent excessively large values of Λ from causing training instability, the client applies a unit norm constraint (such as L2 normalization) to Λ to ensure that ∑λ i 2 =1.
[0060] In other words, by freezing the shared matrix A, the client only needs to update the private matrices B and Λ, reducing the number of parameters and computational load. For example, on IoT devices with limited computing power, the rank of B is reduced from 1024 dimensions in the full model to 64 dimensions, resulting in a 94% reduction in memory usage. Furthermore, the client initializes the rank of B based on local resources, avoiding resource waste for clients with high computing power and performance bottlenecks for clients with low computing power. Additionally, the client only receives the fragments of shared matrix A relevant to its local task, without needing to acquire data from other domains, reducing the risk of data leakage. For example, a medical client will not receive shared parameters from the financial sector.
[0061] S130. The client dynamically trims low-contribution dimensions and adjusts the actual effective rank based on the standard deviation of each element in the adaptive diagonal matrix, and combines the shared matrix and the private matrix to obtain the updated private matrix.
[0062] Specifically, during local training, the client periodically calculates the standard deviation of each element in the adaptive diagonal matrix Λ, prunes low-contribution dimensions with standard deviations below the threshold β, and updates the actual effective rank to the number of remaining dimensions. For example, a client with limited computing power in a health checkup center can automatically reduce the rank from 64 to 32 to reduce memory usage. Additionally, a unit norm constraint is introduced to ensure the stability of the row vectors in the private matrix B, preventing training divergence.
[0063] More specifically, the client calculates each element λ in Λ. i Standard deviation σ k , trim |λ i |<β·σ k For the low-contribution dimension, update the rank to r′. k =r k -#{i||λ i |<β·σ k For example, clients with weak computing power will automatically reduce the rank from 64 to 32 to decrease memory usage. Then, a unit norm constraint is introduced.
[0064]
[0065] This is to ensure the stability of the row vectors in matrix B and avoid training divergence. Where r... k The initial low rank set for client k; r′ k The actual effective rank (the number of dimensions involved in the calculation after dynamically adjusting) after pruning low-contribution dimensions for client k; λ i σ is the i-th element of the diagonal matrix Λ; β is the sparsity threshold (a hyperparameter that controls the pruning intensity and requires experimental tuning); k For client k, all λ i Standard deviation (a measure of λ) i (the degree of dispersion); #{...} represents the number of elements in the set.
[0066] In other words, the dynamic pruning mechanism enables the model to automatically adjust its complexity based on the client's computing power, avoiding resource waste on clients with high computing power and performance bottlenecks on clients with low computing power. For example, a bank client can maintain high rank to handle complex financial analysis, while a health checkup center client can achieve lightweight deployment by reducing rank. In addition, pruning low-contribution dimensions reduces invalid computation and accelerates the local training process.
[0067] In one embodiment, the client dynamically trims low-contribution dimensions and adjusts the actual effective rank based on the standard deviation of each element in the adaptive diagonal matrix, and combines the shared matrix and the private matrix to obtain an updated private matrix, including:
[0068] The client calculates the standard deviation of each diagonal element in the adaptive diagonal matrix;
[0069] Specifically, during the training process, the client counts each diagonal element λ in the adaptive diagonal matrix Λ by batch or epoch. i The numerical distribution of λ. For example, maintaining a sliding window of length T, recording the λ values in the most recent T updates. i The value of is calculated, and its standard deviation σ is calculated. i The baseline value μ can be set as the global mean (e.g., all λ values). i The average value (μ) or a domain-specific threshold (such as the weight benchmark for features related to disease diagnosis in a medical task). For example, in financial text classification, μ could be set as the average weight of transaction-related features.
[0070] The client uses a preset sparse threshold as a standard to identify low-contribution dimensions, and iterates through all diagonal elements of the adaptive diagonal matrix to identify dimensions with a standard deviation less than the sparse threshold multiplied by a certain benchmark value as low-contribution dimensions.
[0071] Specifically, a sparse threshold β (e.g., 0.1) is preset, and the identification criteria are dynamically adjusted in conjunction with a baseline value μ. The criteria for determining low contribution dimensions are:
[0072] σ i <β·μ;
[0073] The client iterates through all diagonal elements of Λ, marks the dimensions that satisfy the above conditions, and sorts them by σ. i Sort by size from smallest to largest, and prioritize trimming the dimension with the lowest standard deviation.
[0074] The client trims the identified low-contribution dimensions, that is, removes the rows and columns corresponding to the low-contribution dimensions from the private matrix and the adaptive diagonal matrix;
[0075] Specifically, the client synchronously removes rows and columns corresponding to low-contribution dimensions from the private matrix B and the adaptive diagonal matrix Λ. For example, if the 3rd dimension is marked as low contribution, the 3rd row and 3rd column of B, as well as the 3rd diagonal element of Λ, are deleted. After pruning, the client ensures that the proportion of remaining dimensions is not less than a preset sparsity rate (e.g., retaining at least 30% of the dimensions) to avoid over-pruning leading to insufficient model capacity.
[0076] After the pruning operation is completed, the client adjusts the actual effective rank of the private matrix based on the number of remaining dimensions;
[0077] Specifically, after the pruning operation, the remaining dimensions of the private matrix B become the new effective rank r′. The client feeds r′ back into the training process, and subsequent gradient updates are calculated only for the remaining dimensions. If model performance degrades after pruning (e.g., validation set accuracy drops below a threshold), the client can trigger a rank restoration operation to reactivate some low-contribution dimensions (e.g., by σ). i (Sorting restores the first 20% of dimensions).
[0078] The updated private matrix is obtained by combining the actual effective rank, the shared matrix, and the private matrix.
[0079] Specifically, the client combines the shared matrix A (frozen) and the pruned private matrix M′, and generates the updated incremental matrix B′, i.e., the updated private matrix, using the low-rank decomposition formula B′=A·Λ′·M′. Here, Λ′ is the pruned adaptive diagonal matrix. During training, the client only updates the parameters of M′ and Λ′, while A remains unchanged, ensuring the stability of the general semantic representation.
[0080] In other words, by pruning low-contribution dimensions, the number of parameters in the private matrix B is significantly reduced. For example, in medical text classification tasks, after pruning, the dimension of B is reduced from 128 to 64, resulting in a 50% reduction in memory usage and a 60% reduction in computational cost per forward propagation. Furthermore, the dynamic adjustment of the actual effective rank r′ allows the model to adapt to clients with different resource environments. Devices with weak computing power (such as mobile phones) can reduce the rank through pruning to avoid training interruptions due to computational overload. Additionally, the standard deviation of the adaptive diagonal matrix Λ reflects the fluctuations in feature importance. Pruning dimensions with low standard deviations (i.e., dimensions with stable weights) reduces noise interference, allowing the model to focus on key features. For example, in financial fraud detection, the accuracy of the model in identifying abnormal transaction patterns improves by 8% after pruning. Pruning low-contribution dimensions is equivalent to implicit regularization, reducing model complexity. Moreover, after pruning, the client only needs to upload the updated B′ and Λ′ (instead of the full matrix), and the communication volume is proportional to the number of dimensions. For example, when the dimension is reduced from 128 to 64, the communication volume is reduced by 50%, and the training speed is increased by 2 times in edge network environments.
[0081] S140. The server receives the updated private matrix uploaded by each client, aggregates it using a rank-aware aggregation strategy, and fine-tunes the shared matrix using a public dataset to obtain a global model.
[0082] Specifically, the server receives the private matrix (B' (multiplied by Λ), i.e., the updated private matrix) uploaded by the client, aligns each matrix to the maximum rank using zero-padding, and aggregates them according to data volume and rank weights. After aggregation, the global model is updated to ensure consistency with the client's update. Additionally, the server fine-tunes the shared matrix A using a small-scale public dataset (such as general text) to generate A', and merges it with the original A to obtain the global model, which is used to improve the model's generalization ability under non-independent and identically distributed data.
[0083] More specifically, the server receives the B′ matrix (multiplied by Λ) uploaded by the client, and aligns each client's B′ to the maximum rank r using zero-padding. max Aggregate by data volume and rank weight:
[0084]
[0085] Update the global model after aggregation:
[0086]
[0087] This ensures consistency with client updates. The server then fine-tunes matrix A using a small, public dataset to generate A. refinement And merge it with the original A to obtain the global model:
[0088] A′ global =α·A global +(1-α)·A refinement ;
[0089] To improve the model's generalization ability on non-independent and identically distributed data.
[0090] In other words, the rank-aware aggregation strategy considers the data volume and model complexity of different clients, avoiding the bias caused by traditional direct averaging, and is particularly effective with non-independent and identically distributed data. For example, when financial and medical data differ significantly, the aggregation result can more accurately reflect global features. Fine-tuning with common data enables the shared matrix A to capture a wider range of semantic features, improving the model's performance on unseen data.
[0091] In one embodiment, the server receives updated private matrices uploaded by each client, aggregates them using a rank-aware aggregation strategy, and fine-tunes the shared matrix using a public dataset to obtain a global model, including:
[0092] The server obtains or infers the actual effective rank information from the updated private matrix uploaded by each client;
[0093] Specifically, when uploading the updated private matrix B′, the client synchronously uploads its actual effective rank r′ (i.e., the pruned dimension number). For example, the client appends r′ to the matrix data via metadata fields (such as JSON headers). If the client does not explicitly transmit r′, the server can infer it from the matrix structure. For example, if the dimension of B′ is d×r′, r′ can be obtained directly from the column number; or by analyzing the distribution of singular values in the matrix (e.g., retaining the first r′ significant singular values). Additionally, the server maintains a rank dictionary, recording the historical rank information for each client, for reference in subsequent alignment and aggregation.
[0094] For private matrices whose actual effective rank is less than the maximum rank in the actual effective rank information, the server uses zero-padding to align them to the maximum rank in order to obtain the padded private matrix.
[0095] Specifically, the server counts the actual effective rank {r1′, r2′, ..., r} of all clients. n ′}, take the maximum value r max =max({r i r′}) is used as the alignment target. For example, if client A's r′ = 64 and client B's r′ = 128, then r′}) is the alignment target. max =128. For the actual effective rank r i ′ <r max Private matrix B i The server fills r to its right (column direction). max -r i The zero vector in column B generates the padded matrix. For example, if B i If the first column is 100×64, then after padding it will be 100×128. The zero columns of the padding column do not participate in subsequent calculations (such as gradient updates), and are only used for matrix dimension unification.
[0096] The server performs weighted aggregation on the filled private matrix based on the amount of data uploaded by each client and the actual effective rank, to obtain the aggregated private matrix.
[0097] Specifically, the server calculates the aggregation weight based on the amount of data uploaded by the client and the actual effective rank. The server then performs a weighted average on the padded matrix to generate the aggregated private matrix. After aggregation, the server can perform sparsification on the aggregated private matrix (e.g., retaining the first 90% of non-zero elements) to further reduce redundant calculations.
[0098] The server uses a public dataset to fine-tune the shared matrix to obtain the fine-tuned shared matrix;
[0099] Specifically, the server constructs a public dataset from trusted data sources (such as public medical image libraries and financial transaction records), ensuring that it is consistent with the client's task domain but does not overlap with the data. For example, the publicly available MIMIC-III dataset is used in the medical scenario, while historical transaction data before federated learning is used in the financial scenario. Furthermore, the server freezes the structure of the shared matrix A, updating only its parameters. A is optimized using gradient descent by minimizing a loss function (such as cross-entropy loss). To prevent overfitting, the server incorporates an L2 regularization term during fine-tuning to control the parameter size of A.
[0100] The server merges the fine-tuned shared matrix with the aggregated private matrix to generate a global model.
[0101] Specifically, the server generates a global model by combining the fine-tuned shared matrix A′ with the aggregated private matrix using a low-rank decomposition formula.
[0102] In other words, by using zero-padding to unify private matrices of different ranks to the maximum rank, the problem of rank discrepancies caused by client hardware limitations (such as computing power and memory) is solved. For example, in medical federated learning, mobile clients (rank 64) and server-level clients (rank 128) can collaborate seamlessly, avoiding aggregation failures due to rank mismatch. Furthermore, weighted aggregation combining data volume and actual effective rank ensures that model parameters from high-contribution clients (such as hospitals with large data volumes and high ranks) dominate the aggregation, improving the professionalism of the global model. Additionally, domain adaptation of the shared matrix A using a public dataset enables it to capture common semantic features (such as lesion patterns in medical images and transaction keywords in financial texts), reducing the impact of client data distribution differences.
[0103] S150. Deploy the global model to the server and each client.
[0104] Specifically, only lightweight components (shared matrix A, adaptive private matrix B') are stored, and parameters are further compressed through model pruning to adapt to resource-constrained environments. By deploying the complete model, complex tasks across domains (such as processing 1000+ user queries per second) can be supported. Each training round only transmits matrix differences (instead of all parameters), and combined with differential compression algorithms, bandwidth consumption is reduced.
[0105] In other words, the layered strategy enables the model to adapt to devices with different resource environments. For example, a full model can be deployed on a high-resource bank client to support high-concurrency financial question answering, while a lightweight component can be deployed on a low-resource medical examination center client to interpret medical examination reports. In addition, communication optimization reduces data transmission volume and supports low-latency real-time inference (such as when a user inputs "today's stock market trend", the local model combines financial terminology features and language understanding capabilities to quickly generate a response).
[0106] In one embodiment, deploying the global model to the server and each client includes:
[0107] The server stores the global model in a designated model storage area and configures the model services, including setting the model's input and output formats, API interfaces, and access permissions.
[0108] Specifically, after receiving the trained global model, the server stores it in a pre-defined model storage area. This storage area features high reliability, large capacity, and fast read / write capabilities to ensure secure storage and efficient access to the model data. Subsequently, the server configures the model service, including specifying the model's input and output formats (e.g., input is image data of a specific size, output is classification labels and corresponding probability values); setting up an API interface to provide a unified calling method for clients, facilitating interaction between clients and the server; and configuring access permissions, assigning different access permissions based on different client roles (e.g., regular users, administrators, etc.) to ensure the security of the model service and data privacy.
[0109] Each client sends a model download request to the server based on its own needs and the model service information provided by the server. After receiving the download request from the client, the server transmits the global model to the corresponding client.
[0110] Specifically, the server activates a listening mechanism, waiting in real time for model download requests from various clients. Upon receiving a download request, the server first verifies the request, checking if the client's access permissions meet the requirements. If the verification passes, the server determines the corresponding global model based on the request information and transmits the global model to the requesting client via a secure network transmission protocol (such as HTTPS).
[0111] After receiving the global model, the client adapts the global model according to its own hardware environment and resource limitations;
[0112] Specifically, each client analyzes the type of global model it needs to use based on its own business requirements and actual application scenarios. The client sends a model download request to the server through the communication interface, containing key information such as the client's identity and the required model identifier, so that the server can accurately identify and process it. Furthermore, after receiving the global model from the server, the client comprehensively assesses its own hardware environment and resource status, including CPU and GPU computing power, memory size, and storage space. Based on the assessment results, the client performs adaptation operations on the global model, such as quantizing the model to reduce the precision of model parameters and lower hardware resource requirements, or pruning the model to remove redundant neurons or connections to improve inference speed.
[0113] The client deploys the adapted global model on the local device and configures the corresponding inference environment and services.
[0114] Specifically, the client deploys the adapted global model on its local device, ensuring that the model files are correctly stored in the specified local location. Simultaneously, the client configures the appropriate inference environment according to the model's requirements, including installing necessary runtime libraries and framework versions. After configuration, the inference service is started, enabling the model to run normally on the local device and providing support for the client's business applications.
[0115] In other words, by storing the global model in a designated area and configuring a unified model service, the server can centrally manage model resources, facilitating model updates, maintenance, and monitoring. Simultaneously, unified API interfaces and access permission configurations make the use of the model service more standardized and secure, reducing errors and security risks caused by chaotic model management. Furthermore, configuring access permissions effectively controls access to the model service from different clients, preventing unauthorized access and data leaks. At the same time, using secure network transmission protocols to transmit model data ensures the confidentiality and integrity of the model during transmission, protecting the intellectual property rights of the model and the data privacy of users. In addition, each client can adapt and deploy the global model locally according to its specific needs and hardware environment, allowing the model to better adapt to different application scenarios and device conditions. This personalized deployment method improves the usability and performance of the model, providing users with a higher quality service. Furthermore, by deploying the adapted model on local devices, clients can perform model inference locally, reducing frequent requests to the server and mitigating the impact of network latency on business applications. This also reduces the server load, improving the stability and scalability of the entire system. In addition, by performing adaptation operations on the model, such as quantization and pruning, the client can improve the inference speed and efficiency of the model with limited local hardware resources, and achieve real-time or near real-time data processing and analysis to meet the real-time requirements of business applications.
[0116] In one embodiment, after deploying the global model to the server and each client, the method further includes:
[0117] The server establishes a global model update mechanism to retrain and fine-tune the model periodically or based on client feedback.
[0118] Specifically, the server presets a fixed time period, such as weekly, monthly, or quarterly. When this time period arrives, the server automatically triggers the global model update process. This periodic update method is suitable for situations where the business scenario is relatively stable and data changes have a certain periodicity. For example, in the field of financial risk control, monthly economic data and transaction patterns change relatively regularly; using monthly periodic model updates can promptly capture some long-term trend changes. In addition, during the process of each client using the global model for inference and business processing locally, various feedback information related to model performance will be collected. This feedback information includes, but is not limited to, indicators such as the model's prediction accuracy, false positive rate, and inference time, as well as special data samples and problem situations encountered in actual business scenarios. The client sends this feedback information to the server according to a preset format and communication protocol. The server summarizes and analyzes the received client feedback information, and when it finds that the model's performance indicators have dropped to a certain level or encountered many special problems, it triggers the global model update process. For example, in image recognition applications, if multiple clients report low accuracy in recognizing certain newly emerging object categories, the server will initiate a model update based on this feedback.
[0119] Once the update mechanism is triggered, the server begins collecting data for model retraining and fine-tuning. This data comes from a wide range of sources: on the one hand, it integrates special data samples and problem data from various clients; on the other hand, the server can obtain new business data from its own data storage system. After data collection, preprocessing operations are performed, including data cleaning to remove noise and erroneous data; data labeling to ensure correct labels for supervised learning tasks; and data normalization to unify the data to a uniform scale range to improve model training performance. Furthermore, if the business scenario changes significantly or model performance deteriorates severely, the server will choose to retrain the global model using a completely new dataset. During retraining, the server adjusts the model's hyperparameters, such as learning rate, batch size, and number of training epochs, to optimize the model's training effect. Appropriate optimization algorithms, such as stochastic gradient descent (SGD) and Adam, are used to update the model's parameters, enabling the model to better fit the new data distribution. Finally, when the business scenario changes only slightly, or when the goal is simply to solve some local problems, the server will fine-tune the existing global model. Fine-tuning typically involves training some or all layers of a pre-trained model with new data. Through fine-tuning, the model can be quickly adapted to new data features and business requirements without compromising its original performance. For example, in natural language processing, a language model pre-trained on a large corpus can be fine-tuned using specialized corpora for a specific domain (such as medicine or law) to improve its performance in that domain.
[0120] After retraining and fine-tuning the global model, the server validates the updated model. It evaluates the model using an independent test dataset, checking performance metrics such as accuracy, recall, and F1 score to ensure they meet expectations. If model performance fails to meet requirements, the server analyzes the reasons, adjusts training strategies or data, and retrains and fine-tunes until the model performance meets the standards. Once the validated updated global model is deployed, the server follows the previously described method to deploy the new model to the server and clients. This ensures that both the server and clients can use the latest version of the model for business processing, thereby improving the overall system performance and effectiveness.
[0121] In other words, by retraining and fine-tuning the model periodically or based on client feedback, the model can continuously adapt to new data distributions and business needs. As data accumulates and business scenarios change, the model can learn more features and patterns, thereby improving prediction accuracy, reducing false positives, and enhancing performance across various tasks. Furthermore, different clients may be in different business environments and data scenarios; the specific data and questions they provide can offer valuable information for model updates. Updating the model based on this feedback allows it to better adapt to various complex real-world application scenarios, enhancing the system's versatility and adaptability. Additionally, in a rapidly evolving technological and business environment, data and demands are constantly changing. Establishing a global model update mechanism enables the system to keep pace with these changes, maintaining technological advancement and business competitiveness. Through continuous model optimization, the system can provide higher-quality and more efficient services, attracting more users and customers, thus gaining a competitive edge in the market. For example, in the fintech field, timely updates to risk assessment models can better address market fluctuations and new financial risks, ensuring the stable operation and business expansion of financial institutions.
[0122] For example, in the current booming development of federated learning systems, this adaptive low-rank fine-tuning method for federated large models provides an efficient and secure solution for cross-domain collaboration. The following example, a federated large language model fine-tuning scenario involving institutions in the financial (bank A, securities B) and medical (hospital C, health checkup center D) domains, illustrates the specific application of this method in detail.
[0123] Data preparation stage:
[0124] Each institution in different sectors segments and stores its data locally based on its own business characteristics. Bank A possesses a large amount of financial information data, reflecting the dynamics of the financial market and investor sentiment; Securities Firm B focuses on securities trading data, including key information such as stock prices and trading volumes; Hospital C has accumulated rich medical record data, covering patients' symptoms, diagnoses, and treatment plans; and Health Checkup Center D stores various health checkup reports, providing detailed information on people's health status. To balance potential biases between data from different sectors, the server uses a common text dataset. Simultaneously, to ensure data consistency and security, the server standardizes the text token format, enabling effective processing of data from different sectors within the model. For sensitive medical record data, the server employs differential privacy protection technology, ensuring that the data provides valuable information for model training while protecting patient privacy.
[0125] Client training phase:
[0126] The server distributes the initialized general semantic shared matrix to clients in the financial and medical fields. This shared matrix, maintained by the server, aims to capture cross-domain general semantic representations, providing a basic knowledge framework for the model. Clients A (bank) and B (securities firm) in the financial field, and C (hospital) and D (health checkup center) in the medical field, initialize their private matrices based on local data. These private matrices are updated independently by each client and are used to learn domain-specific knowledge and features. For example, Bank A's private matrix learns key information and patterns in financial data, while Securities B's focuses on the patterns and trends in securities trading; Hospital C's private matrix extracts important features from medical records, and Health Checkup Center D's private matrix analyzes health indicators in health checkup reports. For Health Checkup Center D, with its limited data and computational resources, an adaptive low-rank fine-tuning method is used. It first initializes an adaptive diagonal matrix and then calculates the standard deviation of each diagonal element. Based on a preset sparsity threshold as a standard for identifying low-contribution dimensions, it iterates through all diagonal elements of the adaptive diagonal matrix, identifying dimensions with a standard deviation less than the sparsity threshold multiplied by a certain benchmark value as low-contribution dimensions. Next, the identified low-contribution dimensions are pruned, that is, the rows and columns corresponding to the low-contribution dimensions are removed from the private matrix and the adaptive diagonal matrix. After the pruning operation is completed, the actual effective rank of the private matrix is adjusted according to the number of remaining dimensions. In addition, the physical examination center D also uses norm constraints to prevent training divergence and ensure the stability of model training.
[0127] Server aggregation phase:
[0128] The server receives updated private matrices uploaded by each client. First, the server obtains or infers the actual effective rank information from these private matrices. For private matrices with an actual effective rank less than the maximum rank, the server aligns them to the maximum rank using zero-padding to obtain a padded private matrix. This ensures that all private matrices have the same dimension, facilitating subsequent aggregation operations. Then, the server performs weighted aggregation on the padded private matrices based on the amount of data uploaded by each client and their actual effective rank. Clients with larger data volumes and relatively higher actual effective rank have their uploaded private matrices given greater weight in the aggregation process. This weighted aggregation method allows for a more reasonable integration of the private information from each client, resulting in an aggregated private matrix. Next, the server fine-tunes the shared matrix using a public dataset. The public dataset contains a wide range of general textual information; by fine-tuning the shared matrix, the model's versatility can be further optimized, making it better adaptable to tasks in different domains. Finally, the server merges the fine-tuned shared matrix with the aggregated private matrices to generate a global model.
[0129] Deployment phase:
[0130] When deploying the global model, the resource availability and application requirements of different clients were fully considered. Bank A, with high resources, deployed a complete global model to support high-concurrency financial question-and-answer services, ensuring a fast and accurate response to user requests. Meanwhile, the low-resource medical examination center D deployed a lightweight component primarily for interpreting medical examination reports, meeting business needs while conserving computing resources. After receiving the global model, the client adapts it to its own hardware environment and resource limitations. For example, it adjusts model parameter settings and optimizes the inference process to ensure efficient operation on its local device. Then, the client deploys the adapted global model on its local device and configures the corresponding inference environment and services, preparing it for practical application.
[0131] Reasoning stage:
[0132] During inference, the client-side locally integrates general and domain-specific knowledge to generate responses. The shared matrix provides a cross-domain general semantic representation, while the private matrix contains domain-specific knowledge and features. By fusing the two, the model can combine general information with domain expertise to provide users with more accurate and comprehensive answers. For example, when a user asks a question related to financial investment, the model can not only use general semantics to understand the meaning of the question but also combine private knowledge in the financial domain to provide professional investment advice.
[0133] Model optimization phase:
[0134] During model operation, continuous optimization is performed. After running for a period of time, the medical examination center D detected a decline in inference performance, automatically triggering the rank adjustment mechanism. It recalculates the standard deviation of elements in the adaptive diagonal matrix, identifies and prunes low-contribution dimensions, and adjusts the actual effective rank of the private matrix to improve the model's inference efficiency. When bank A adds new data, there is no need to retrain the entire model; instead, only the relevant dimensions are incrementally fine-tuned. This approach allows for rapid adaptation to data changes, reduces training time and computational resource consumption, while maintaining model performance and stability.
[0135] Through the above specific applications in the financial and healthcare fields, the federated large model adaptive low-rank fine-tuning method fully demonstrates its advantages in cross-domain collaboration. It can effectively integrate data and knowledge from different fields to generate high-performance global models, while protecting data privacy and meeting the resource needs of different clients.
[0136] The aforementioned federated large model adaptive low-rank fine-tuning method optimizes the allocation of computing resources by decomposing the low-rank adaptation matrix of the large language model into a server-shared matrix and a client-private matrix. The server-shared matrix is uniformly maintained by the server, capturing cross-domain common semantic representations and avoiding the duplication of computation and storage of the same information on each client, thus saving computing resources. Meanwhile, the client-private matrix is independently updated by each client based on local data, allowing clients with stronger computing power to fully utilize their resources to process more complex parts of the model, while clients with weaker computing power can simplify the model structure to adapt to their resource limitations. This resource adaptation mechanism improves resource utilization and solves the problem of insufficient resource heterogeneous adaptation in traditional methods. In addition, by introducing a rank-aware aggregation strategy, the bias in the aggregation process is effectively reduced. This aggregation method considers the data volume and model complexity of different clients, making the aggregation result closer to the ideal value, thereby improving the convergence and stability of the model, reducing aggregation interference, and improving the model's generalization ability.
[0137] Figure 3 This is a schematic block diagram of a federated large model adaptive low-rank fine-tuning device 300 provided in an embodiment of the present invention. Figure 3 As shown, corresponding to the above-described adaptive low-rank fine-tuning method for federated large models, the present invention also provides a federated large model adaptive low-rank fine-tuning apparatus 300. This federated large model adaptive low-rank fine-tuning apparatus 300 includes a unit for performing the above-described adaptive low-rank fine-tuning method for federated large models, and the apparatus can be configured in a server. Specifically, please refer to... Figure 3 The federated large model adaptive low-rank fine-tuning device 300 is applied to a federated learning system comprising a server and multiple heterogeneous clients. The device includes:
[0138] Decomposition unit 301 is used to decompose the low-rank adaptation matrix of the large language model into a server-shared matrix and a client-private matrix; wherein, the server-shared matrix is maintained by the server and captures cross-domain general semantic representations, and the client-private matrix is independently updated by each client based on local data;
[0139] The initialization unit 302 is used for the client to receive the shared matrix sent by the server and initialize the private matrix and the adaptive diagonal matrix;
[0140] The adjustment and combination unit 303 is used by the client to dynamically prune low-contribution dimensions and adjust the actual effective rank based on the standard deviation of each element in the adaptive diagonal matrix, and combine the shared matrix and the private matrix to obtain the updated private matrix.
[0141] The aggregation fine-tuning unit 304 is used by the server to receive the updated private matrix uploaded by each client, aggregate it through a rank-aware aggregation strategy, and fine-tune the shared matrix using a public dataset to obtain a global model.
[0142] Deployment unit 305 is used to deploy the global model to the server and each client.
[0143] In one embodiment, the decomposition unit 301 includes:
[0144] The initialization module is used by the server to initialize a low-rank adaptive architecture for a large language model, which contains decomposable low-rank matrices.
[0145] The decomposition module is used to decompose a low-rank matrix into two parts, one of which is designed as a server-shared matrix and the other of which is designed as a client-private matrix.
[0146] In one embodiment, the receiving initialization unit 302 includes:
[0147] Establish a receiving module to establish a communication connection between the client and the server, and the client receives the shared matrix sent by the server through this communication connection;
[0148] The initialization module is used by the client to initialize a private matrix based on local data characteristics and resource conditions; at the same time, the client initializes an adaptive diagonal matrix.
[0149] In one embodiment, the adjustment coupling unit 303 includes:
[0150] The calculation module is used by the client to calculate the standard deviation of each diagonal element in the adaptive diagonal matrix;
[0151] The identification module is used by the client to identify low-contribution dimensions based on a preset sparse threshold, and to traverse all diagonal elements of the adaptive diagonal matrix to identify dimensions with a standard deviation less than the sparse threshold multiplied by a certain benchmark value as low-contribution dimensions.
[0152] The trimming module is used by the client to trim the identified low-contribution dimensions, that is, to remove the rows and columns corresponding to the low-contribution dimensions from the private matrix and the adaptive diagonal matrix.
[0153] The adjustment module is used to adjust the actual effective rank of the private matrix based on the number of remaining dimensions after the pruning operation is completed.
[0154] The combination module is used to combine the actual effective rank, the shared matrix, and the private matrix to obtain the updated private matrix.
[0155] In one embodiment, the aggregation fine-tuning unit 304 includes:
[0156] The acquisition module is used by the server to obtain or infer the actual effective rank information from the updated private matrix uploaded by each client;
[0157] The padding module is used to align private matrices with the maximum rank to the maximum rank using zero padding, for which the actual effective rank in the actual effective rank information is less than the maximum rank, so as to obtain the padded private matrix.
[0158] The aggregation module is used by the server to perform weighted aggregation on the filled private matrix based on the amount of data uploaded by each client and the actual effective rank, so as to obtain the aggregated private matrix.
[0159] The fine-tuning module is used by the server to fine-tune the shared matrix using a public dataset to obtain the fine-tuned shared matrix;
[0160] The fusion module is used by the server to merge the fine-tuned shared matrix with the aggregated private matrix to generate a global model.
[0161] In one embodiment, the deployment unit 305 includes:
[0162] The storage configuration module is used by the server to store the global model in a specified model storage area and configure model services, including setting the model's input and output formats, API interfaces, and access permissions.
[0163] The transmission module is used by each client to send a model download request to the server according to its own needs and the model service information provided by the server. After receiving the download request from the client, the server transmits the global model to the corresponding client.
[0164] The adaptation module is used by the client to adapt the global model according to its own hardware environment and resource limitations after receiving the global model;
[0165] The deployment configuration module is used by the client to deploy the adapted global model on the local device and configure the corresponding inference environment and services.
[0166] In one embodiment, the device further includes:
[0167] Establishment units are used to establish a global model update mechanism on the server, and to retrain and fine-tune the model periodically or based on client feedback.
[0168] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned federated large model adaptive low-rank fine-tuning device 300 and each unit can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.
[0169] The aforementioned federal large model adaptive low-rank fine-tuning device 300 can be implemented as a computer program, which can, for example... Figure 4 It runs on the computer device shown.
[0170] Please see Figure 4 , Figure 4 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a server, wherein the server can be a standalone server or a server cluster composed of multiple servers.
[0171] See Figure 4 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.
[0172] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform a federated large model adaptive low-rank fine-tuning method.
[0173] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0174] The internal memory 504 provides an environment for the execution of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a federated large model adaptive low-rank fine-tuning method.
[0175] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0176] The processor 502 is used to run a computer program 5032 stored in the memory to perform the following steps:
[0177] The low-rank adaptation matrix of the large language model is decomposed into a server-shared matrix and a client-private matrix. The server-shared matrix is maintained by the server and captures cross-domain general semantic representations, while the client-private matrix is independently updated by each client based on local data. The client receives the shared matrix from the server and initializes its private matrix and adaptive diagonal matrix. The client dynamically prunes low-contribution dimensions and adjusts the actual effective rank based on the standard deviation of each element in the adaptive diagonal matrix, and combines the shared matrix and private matrix to obtain the updated private matrix. The server receives the updated private matrices uploaded by each client, aggregates them using a rank-aware aggregation strategy, and fine-tunes the shared matrix using a public dataset to obtain the global model. The global model is then deployed to the server and each client.
[0178] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0179] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0180] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein when executed by a processor, the computer program causes the processor to perform the following steps:
[0181] The low-rank adaptation matrix of the large language model is decomposed into a server-shared matrix and a client-private matrix. The server-shared matrix is maintained by the server and captures cross-domain general semantic representations, while the client-private matrix is independently updated by each client based on local data. The client receives the shared matrix from the server and initializes its private matrix and adaptive diagonal matrix. The client dynamically prunes low-contribution dimensions and adjusts the actual effective rank based on the standard deviation of each element in the adaptive diagonal matrix, and combines the shared matrix and private matrix to obtain the updated private matrix. The server receives the updated private matrices uploaded by each client, aggregates them using a rank-aware aggregation strategy, and fine-tunes the shared matrix using a public dataset to obtain the global model. The global model is then deployed to the server and each client.
[0182] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0183] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0184] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0185] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0186] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0187] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An adaptive low-rank fine-tuning method for federated large models, characterized in that: The method, applied to a federated learning system comprising a server and multiple heterogeneous clients, includes: The low-rank adaptation matrix of the large language model is decomposed into a server-shared matrix and a client-private matrix. The server-shared matrix is maintained by the server and captures cross-domain general semantic representations, while the client-private matrix is independently updated by each client based on local data. The client receives the shared matrix from the server and initializes the private matrix and the adaptive diagonal matrix; The client dynamically trims low-contribution dimensions and adjusts the actual effective rank based on the standard deviation of each element in the adaptive diagonal matrix, and combines the shared matrix and the private matrix to obtain the updated private matrix; The server receives the updated private matrices uploaded by each client, aggregates them using a rank-aware aggregation strategy, and fine-tunes the shared matrices using a public dataset to obtain a global model. Deploy the global model to the server and each client.
2. The adaptive low-rank fine-tuning method for federated large models according to claim 1, characterized in that, The process of decomposing the low-rank adaptation matrix of the large language model into a server-shared matrix and a client-private matrix includes: The server initializes a low-rank adaptive architecture for a large language model, which contains decomposable low-rank matrices. The low-rank matrix is decomposed into two parts, one of which is designed as a server-shared matrix and the other is designed as a client-private matrix.
3. The adaptive low-rank fine-tuning method for the federated large model according to claim 1, characterized in that, The client receives the shared matrix from the server and initializes the private matrix and the adaptive diagonal matrix, including: A communication connection is established between the client and the server, and the client receives the shared matrix sent by the server through this communication connection; The client initializes a private matrix based on local data characteristics and resource conditions; at the same time, the client initializes an adaptive diagonal matrix.
4. The adaptive low-rank fine-tuning method for federated large models according to claim 1, characterized in that, The client dynamically trims low-contribution dimensions and adjusts the actual effective rank based on the standard deviation of each element in the adaptive diagonal matrix, and combines the shared matrix and the private matrix to obtain an updated private matrix, including: The client calculates the standard deviation of each diagonal element in the adaptive diagonal matrix; The client uses a preset sparse threshold as a standard to identify low-contribution dimensions, and iterates through all diagonal elements of the adaptive diagonal matrix to identify dimensions with a standard deviation less than the sparse threshold multiplied by a certain benchmark value as low-contribution dimensions. The client trims the identified low-contribution dimensions, that is, removes the rows and columns corresponding to the low-contribution dimensions from the private matrix and the adaptive diagonal matrix; After the pruning operation is completed, the client adjusts the actual effective rank of the private matrix based on the number of remaining dimensions; The updated private matrix is obtained by combining the actual effective rank, the shared matrix, and the private matrix.
5. The adaptive low-rank fine-tuning method for federated large models according to claim 1, characterized in that, The server receives updated private matrices uploaded by each client, aggregates them using a rank-aware aggregation strategy, and fine-tunes the shared matrix using a public dataset to obtain a global model, including: The server obtains or infers the actual effective rank information from the updated private matrix uploaded by each client; For private matrices whose actual effective rank is less than the maximum rank in the actual effective rank information, the server uses zero-padding to align them to the maximum rank in order to obtain the padded private matrix. The server performs weighted aggregation on the filled private matrix based on the amount of data uploaded by each client and the actual effective rank, to obtain the aggregated private matrix. The server uses a public dataset to fine-tune the shared matrix to obtain the fine-tuned shared matrix; The server merges the fine-tuned shared matrix with the aggregated private matrix to generate a global model.
6. The adaptive low-rank fine-tuning method for federated large models according to claim 1, characterized in that, The deployment of the global model to the server and each client includes: The server stores the global model in a designated model storage area and configures the model services, including setting the model's input and output formats, API interfaces, and access permissions. Each client sends a model download request to the server based on its own needs and the model service information provided by the server. After receiving the download request from the client, the server transmits the global model to the corresponding client. After receiving the global model, the client adapts the global model according to its own hardware environment and resource limitations; The client deploys the adapted global model on the local device and configures the corresponding inference environment and services.
7. The adaptive low-rank fine-tuning method for federated large models according to claim 1, characterized in that, After deploying the global model to the server and each client, the process also includes: The server establishes a global model update mechanism to retrain and fine-tune the model periodically or based on client feedback.
8. A federated large model adaptive low-rank fine-tuning device, characterized in that, An apparatus for use in a federated learning system comprising a server and multiple heterogeneous clients, the apparatus comprising: The decomposition unit is used to decompose the low-rank adaptation matrix of a large language model into a server-shared matrix and a client-private matrix. The server-shared matrix is maintained by the server and captures cross-domain general semantic representations, while the client-private matrix is independently updated by each client based on local data. The initialization unit is used by the client to receive the shared matrix sent by the server and initialize the private matrix and the adaptive diagonal matrix. The adjustment and combination unit is used by the client to dynamically prune low-contribution dimensions and adjust the actual effective rank based on the standard deviation of each element in the adaptive diagonal matrix, and combine the shared matrix and the private matrix to obtain the updated private matrix; The aggregation fine-tuning unit is used by the server to receive the updated private matrices uploaded by each client, aggregate them using a rank-aware aggregation strategy, and fine-tune the shared matrix using a public dataset to obtain a global model. The deployment unit is used to deploy the global model to the server and various clients.
9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.
Citation Information
Cited By
Multi-mode knee osteoarthritis auxiliary diagnosis method based on resource adaptive federal learning
CN121687460A
Multi-modal osteoarthritis auxiliary diagnosis method based on federated learning with resource adaptation
CN121687460B
Heterogeneous computing power network service zero-interruption endogenous security dynamic immunization method and system
CN121770909A