Big model training system and method based on federal learning and electronic equipment

By introducing orthogonal frequency division multiple access network communication and low-rank adapter rank selection technology in federated learning systems, large-scale model training problems under low data transmission volume and low resource occupation are solved, and resource consumption and hardware requirements are reduced.

CN119940477AInactive Publication Date: 2025-05-06BEIJING UNIV OF CIVIL ENG & ARCHITECTURE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411996184.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

It is difficult for the prior art to realize large-scale model training under low data transmission and low resource occupation conditions. Especially for small companies or individuals, the resource consumption and hardware requirements for developing large models are high.

Method used

By introducing orthogonal frequency division multiple access network communication into the federated learning system, the client downloads the initial global model from the server and freezes it, trains the matrix of the low-rank adapter through local sample data, performs rank selection and uploads to the server, and the server updates the global matrix and issues it to the client until the preset round is reached.

Benefits of technology

Large-scale model training with low data transmission volume and low resource occupation is realized, reducing the resource consumption and hardware requirements of large-scale model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940477A_ABST
    Figure CN119940477A_ABST
Patent Text Reader

Abstract

The invention provides a federated learning-based large model training system and method and electronic equipment, the system comprises a server and a plurality of clients communicating through an orthogonal frequency division multiple access network, and each client downloads an initial global model from the server and freezes the initial global model; each client trains a first to-be-trained matrix and a second to-be-trained matrix of the low-rank adapter through the local sample data to obtain a first gradient matrix and a second gradient matrix; each client performs rank selection on the first gradient matrix and the second gradient matrix to determine a first rank selection matrix and a second rank selection matrix; the server determines a first global matrix and a second global matrix of the current round based on a first rank selection matrix and a second rank selection matrix which are received in the last round and uploaded by all clients; and each client adjusts model parameters of the initial global model based on the first global matrix and the second global matrix which are finally received, and the model parameters serve as model parameters of the target large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and more specifically, to a large model training system, method and electronic device based on federated learning. Background Art

[0002] Federated learning, also known as federated machine learning, can solve the privacy problem when jointly training models and is widely used in most industries. Large companies or research institutions have sufficient resources to develop large models, but for small companies or individuals, it is almost impossible to develop their own large models. Therefore, a large model training architecture with low cost, low resource consumption, and low transmission volume is an urgent problem to be solved. Summary of the invention

[0003] The purpose of the embodiments of the present application is to provide a large model training system, method and electronic device based on federated learning, which realizes large-scale model training based on federated learning with low data transmission volume and low resource occupancy.

[0004] In a first aspect, the present invention provides a large model training system based on federated learning, the system comprising a server and multiple clients communicating via an orthogonal frequency division multiple access network, wherein:

[0005] Each client downloads the initial global model from the server and freezes it;

[0006] Each client trains the first to-be-trained matrix and the second to-be-trained matrix of the low-rank adapter using local sample data to obtain a first gradient matrix and a second gradient matrix;

[0007] Each client performs ranking on the first gradient matrix and the second gradient matrix to determine the first ranked matrix and the second ranked matrix and upload them to the server;

[0008] The server determines the first global matrix and the second global matrix of this round based on the first rank selection matrix and the second rank selection matrix uploaded by all clients received in the previous round, and sends them to each client, so that each client updates the first matrix to be trained and the second matrix to be trained in the next round;

[0009] When the preset round is reached, each client adjusts the model parameters of the initial global model based on the last received first global matrix and second global matrix as the model parameters of the target large model.

[0010] In an optional implementation, each client obtains the first gradient matrix by the following method: and the second gradient matrix

[0011]

[0012] Among them, the first gradient matrix is the loss function f relative to the first global matrix A t-1 The gradient matrix, the second gradient matrix is the loss function f relative to the second global matrix B t-1 The gradient matrix of , t is the current round, is the local sample data set of client k, d is the data sample, For client k A sub-sample data set is formed by randomly selecting ξ data samples from .

[0013] In an optional implementation, each client determines the first rank selection matrix by the following method: and the second rank selection matrix

[0014]

[0015] in, for The rth column of for The rth column of t,R is the corresponding ranking index.

[0016] In an optional implementation, the server determines the first global matrix A of round t by the following method: t and the second global matrix B t :

[0017]

[0018] Where η is the learning rate, s t,k is the transmission index, q t,k is the transmission interruption probability of client k, P is the probability, D k for The number of samples.

[0019] In an optional implementation, each client adjusts the parameters of the initial global model in the following manner to generate a target large model:

[0020] W=W0+B t A t T ;

[0021] Among them, W is the model parameter of the target large model, and W0 is the model parameter of the initial global model.

[0022] In the second aspect, the present invention provides a large model training method based on federated learning, which is suitable for a server and multiple clients communicating through an orthogonal frequency division multiple access network, wherein each client downloads an initial global model from the server and freezes it; each client trains a first to-be-trained matrix and a second to-be-trained matrix of a low-rank adapter through local sample data to obtain a first gradient matrix and a second gradient matrix; each client ranks the first gradient matrix and the second gradient matrix to determine a first rank-selected matrix and a second rank-selected matrix and uploads them to the server; the server determines the first global matrix and the second global matrix of this round based on the first rank-selected matrices and the second rank-selected matrices uploaded by all clients received in the previous round, and sends them to each client, so that each client updates the first to-be-trained matrix and the second to-be-trained matrix of the next round; when the preset round is reached, each client adjusts the model parameters of the initial global model based on the first global matrix and the second global matrix received last, as the model parameters of the target large model.

[0023] In an optional implementation, each client obtains the first gradient matrix by the following method: and the second gradient matrix

[0024]

[0025] Among them, the first gradient matrix is the loss function f relative to the first global matrix A t-1 The gradient matrix, the second gradient matrix is the loss function f relative to the second global matrix B t-1 The gradient matrix of , t is the current round, is the local sample data set of client k, d is the data sample, For client k A sub-sample data set is formed by randomly selecting ξ data samples from .

[0026] In an optional implementation, each client determines the first rank selection matrix by the following method: and the second rank selection matrix

[0027]

[0028] in, for The rth column of for The rth column of t,R is the corresponding ranking index.

[0029] In a third aspect, the present invention provides an electronic device comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate through the bus, and the processor executes the machine-readable instructions to perform the steps of the large model training method based on federated learning as described in any of the aforementioned embodiments.

[0030] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of a large model training method based on federated learning as described in any of the aforementioned embodiments are executed.

[0031] The present application provides a large model training system, method and electronic device based on federated learning, the system includes a server and multiple clients communicating through an orthogonal frequency division multiple access network, wherein each client downloads an initial global model from the server and freezes it; each client trains the first to-be-trained matrix and the second to-be-trained matrix of the low-rank adapter through local sample data to obtain the first gradient matrix and the second gradient matrix; each client ranks the first gradient matrix and the second gradient matrix to determine the first rank selection matrix and the second rank selection matrix and uploads them to the server; the server determines the first global matrix and the second global matrix of this round based on the first rank selection matrix and the second rank selection matrix uploaded by all clients received in the previous round and sends them to each client, so that each client updates the first to-be-trained matrix and the second to-be-trained matrix of the next round; when the preset round is reached, each client adjusts the model parameters of the initial global model based on the last received first global matrix and the second global matrix as the model parameters of the target large model. Through the combination of federated learning and low-rank adaptive fine-tuning model, large-scale model training based on federated learning can be realized under the conditions of low data transmission volume and low resource occupation, reducing the resource consumption and hardware requirements of large-scale model training. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0033] Figure 1 A schematic diagram of the structure of a large model training system based on federated learning provided in an embodiment of the present application;

[0034] Figure 2 A flowchart of the steps of large model training provided in an embodiment of the present application;

[0035] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0036] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0037] Figure 1 A structural diagram of a large model training system based on federated learning provided in an embodiment of the present application. Figure 2 A flowchart of the steps of large model training provided in an embodiment of the present application.

[0038] like Figure 1 As shown, in one embodiment of the present application, a large model training system based on federated learning is provided, and the system includes a server and multiple clients communicating through an orthogonal frequency division multiple access (OFDMA) network.

[0039] The server here is equipped with one antenna, and the server's resources can be divided into N resource blocks RB, denoted by N = {1, ..., N}. Multiple clients can be wireless devices, denoted by K = {1, ..., K}, each client is equipped with one antenna and a local sample data set, and the local sample data set is denoted by D k ={d k,l}, contains l data samples.

[0040] When training a large model, each client will be assigned a corresponding RB for communication and data transmission with the server.

[0041] Before model training begins, a target client can initiate a large model training request to the server. The server responds to the large model training request sent by the target client and broadcasts it to other clients. After other clients respond, the server can determine all clients participating in this large model training.

[0042] like Figure 2 As shown, the large model can be trained by the following steps:

[0043] S1. Each client downloads the initial global model from the server and freezes it.

[0044] The initial global model is an untrained model, such as ChatGPT. The parameters of the initial global model can be expressed as

[0045] S2. Each client trains the first to-be-trained matrix and the second to-be-trained matrix of the low-rank adapter using local sample data to obtain a first gradient matrix and a second gradient matrix.

[0046] Each client is configured with a low-dimensional low-rank adapter. The low-rank adapter includes the first matrix to be trained and the second matrix to be trained Where dimensions Q and L satisfy min{Q,L}>>R.

[0047] Each client can train the low-rank adapter based on local sample data and minimize the loss function Update the first matrix to be trained and the second matrix to be trained:

[0048]

[0049] Where, f(ΔW t-1 ; W0,d k,d ) is the loss function of a single data sample. Here, the loss function of each client can be cross entropy, MSE, etc., and each client uses the same loss function.

[0050] The client can obtain the first gradient matrix in the following way and the second gradient matrix

[0051]

[0052] Among them, the first gradient matrix is the loss function f relative to the first global matrix A t-1 The gradient matrix, the second gradient matrix is the loss function f relative to the second global matrix B t-1 The gradient matrix of , t is the current round, is the local sample data set of client k, d is the data sample, For client k A sub-sample data set is formed by randomly selecting ξ data samples from is the data sample d in the tth round of client k relative to the first global matrix A t-1 The gradient matrix of is the data sample d in round t relative to the second global matrix B t-1 The gradient matrix of .

[0053] S3. Each client performs ranking on the first gradient matrix and the second gradient matrix to determine the first ranked matrix and the second ranked matrix and upload them to the server.

[0054] Each client can determine the first rank selection matrix by the following method and the second rank selection matrix

[0055]

[0056] in, for The rth column of for The rth column of t,r is the corresponding ranking index. t,r is a value of 0 or 1. In this way, the number of bits of data sent by client k to the server in the tth round of parameter fine-tuning can be:

[0057]

[0058] V max =R(Q+L);

[0059] Among them, V max is the total number of model parameters, and M is the number of bits used to store a single model parameter.

[0060] That is, it can be seen that by controlling the ranking index, the amount of data transmission between the client and the server can be reduced.

[0061] S4. The server determines the first global matrix and the second global matrix of this round based on the first rank selection matrix and the second rank selection matrix uploaded by all clients received in the previous round.

[0062] The server can determine the first global matrix A of round t in the following way: t and the second global matrix B t :

[0063]

[0064]

[0065] Where η is the learning rate, s t,k is the transmission index, q t,k is the transmission interruption probability of client k, P is the probability, D k for The number of samples.

[0066] S5. The server sends the first global matrix and the second global matrix to each client, so that each client updates the first matrix to be trained and the second matrix to be trained for the next round.

[0067] S6. When the preset round is reached, each client adjusts the model parameters of the initial global model based on the first global matrix and the second global matrix received last, and uses them as the model parameters of the target large model.

[0068] Here, each client can adjust the parameters of the initial global model in the following way to generate the target large model:

[0069] W=W0+B t A t T ;

[0070] Among them, W is the model parameter of the target large model, and W0 is the model parameter of the initial global model.

[0071] That is, the client can be based on the first global matrix A sent by the server in the last round t and the second global matrix B t , and the model parameters of the initial global model to determine the final model parameters of the target large model.

[0072] The present application provides a large model training system based on federated learning, which trains model parameters through a federated learning mechanism and a fine-tuning model. The amount of data transmitted between the client and the server is reduced, the communication delay is low, and the hardware requirements for the client are not high, while reducing the possibility of data transmission interruption.

[0073] In one embodiment of the present application, the transmission interruption probability of each client may be determined in the following manner:

[0074] Let θ k and are the CPU frequency and CPU cycles required for a sample of client k, respectively. Then the local fine-tuning delay time (computation delay) of client k ) can be derived from the following formula:

[0075]

[0076] The transmission energy consumed by client k for local fine-tuning is:

[0077]

[0078] where ∈ k is the energy consumption coefficient of client k.

[0079] Considering the inherent CSI (channel state information) estimation error and feedback delay, it is impossible for the server to obtain perfect CSI in reality. t,k,n It is represented as the resource block allocation indicator, where δ t,k,n =1 means allocating resource block n of the server to client k.

[0080] The actual CSI from client k to server on resource block n can be expressed as:

[0081]

[0082] in represents the estimated CSI, represents the estimation error of CSI. Therefore, the channel capacity of client k can be expressed as:

[0083]

[0084] Among them, N0, B and P max denote the power spectral density, the bandwidth of a resource block RB, and the transmit power of client k respectively. Let τ max For the maximum delay budget of each fine-tuning round, the transmission interruption probability q of client k is t,k It can be expressed as:

[0085]

[0086]

[0087] The transmission energy consumption is is the transmission delay, C t,k is the channel capacity of client k in the tth round of fine-tuning, R t,k is the actual transmission channel capacity of client k in the tth round of fine-tuning, V t is the number of bits of data sent from the client to the server in the tth round of parameter fine-tuning, Represents the estimated CSI, N0, B and P max denote the power spectral density, the bandwidth of a resource block RB, and the transmit power of client k, respectively, max is the maximum latency budget for each fine-tuning round, To calculate the delay, is the transmission delay, δ t,k,n An indicator is allocated for a resource block.

[0088] In one embodiment of the present application, the allocation scheme of the resource blocks of the server can be determined by:

[0089] The average F-norm of the global gradient is bounded by the optimality gap and is given by:

[0090]

[0091] Where ΔW * =B * (A * ) T Is the best adapter, is a Lipschitz continuous system with a positive modulus L, F(ΔWt ) is the global loss function, ‖·‖ F represents the F-norm operator, σ 2 is the variance of the stochastic gradient.

[0092] The above parameters meet:

[0093]

[0094] The squared 2-norm of any gradient is bounded, that is, for a constant G 2 ,have

[0095] This formula shows that the dominant factor causing the optimality gap is mainly transmission interruption Sum rank selection The resulting fine-tuning error.

[0096] In addition, the more gradient matrix columns are transmitted, the more likely transmission interruption will occur.

[0097] To improve the performance of the Fed-LoRA model framework, the problem can be posed by minimizing the optimality gap as follows:

[0098]

[0099] Resource block allocation constraints:

[0100]

[0101] CPU frequency constraints:

[0102]

[0103] The delay and energy consumption constraints for each round are:

[0104]

[0105] Among them, τ max and E max represents the maximum delay and energy budget for each round of fine-tuning.

[0106] It can be seen that longer transmission delays reduce the interrupt probability. The CPU frequency can be expressed as It can be concluded that:

[0107]

[0108] It can be seen that along with Therefore, the optimal CPU frequency satisfies And can be obtained by binary search, where satisfy and

[0109] In fixed and In the case of , the original problem can be simplified to the RB allocation problem, expressed as:

[0110]

[0111] Since client k and resource block N are regarded as two disjoint sets, the RB allocation problem can be expressed as The maximum matching of the bipartite graph can be optimized and solved by the Hungarian algorithm.

[0112] In addition, considering the low dimension of R, the optimal number of selected columns can be obtained by binary search. RB allocation and number of selected columns Perform alternating optimization until the RB allocation scheme {δ t,k,n} and rank selection number constant.

[0113] In a specific embodiment of the present application, a large model training system based on federated learning includes a wireless network consisting of a server and K = 8 clients, and the maximum number of columns of the low-rank matrix in the client, the number of bits required to store a parameter, the transmission power, the maximum CPU frequency and the delay budget are R = 16, M = 16, P = 20dBm, θ max =1GHz and τ max =4s. The number of resource blocks, the bandwidth of each resource block, and the AWGN power spectrum density are set to N=8, B=2MHz, and N0=-174dBm / Hz, respectively.

[0114] The server can randomly allocate resource blocks to each client. When generating the rank selection matrix, the client can select all columns, that is,

[0115] Based on the same inventive concept, a large model training method based on federated learning is also provided in an embodiment of the present application, which is suitable for a server and multiple clients communicating through an orthogonal frequency division multiple access network, wherein each client downloads an initial global model from the server and freezes it; each client trains the first to-be-trained matrix and the second to-be-trained matrix of the low-rank adapter through local sample data to obtain the first gradient matrix and the second gradient matrix; each client ranks the first gradient matrix and the second gradient matrix to determine the first rank-selected matrix and the second rank-selected matrix and uploads them to the server; the server determines the first global matrix and the second global matrix of this round based on the first rank-selected matrices and the second rank-selected matrices uploaded by all clients received in the previous round, and sends them to each client, so that each client updates the first to-be-trained matrix and the second to-be-trained matrix of the next round; when the preset round is reached, each client adjusts the model parameters of the initial global model based on the first global matrix and the second global matrix received last as the model parameters of the target large model.

[0116] See also Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 3 As shown in , the electronic device 300 includes a processor 310 , a memory 320 and a bus 330 .

[0117] The memory 320 stores machine-readable instructions executable by the processor 310. When the electronic device 300 is running, the processor 310 communicates with the memory 320 through the bus 330. When the machine-readable instructions are executed by the processor 310, the steps of a large model training method based on federated learning in the above method embodiment can be executed. The specific implementation method can be found in the method embodiment, which will not be repeated here.

[0118] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can execute the steps of a large model training method based on federated learning in the above method embodiment. The specific implementation method can be found in the method embodiment, which will not be repeated here.

[0119] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0120] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0121] In addition, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0122] Furthermore, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0123] It should be noted that if the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can essentially be embodied in the form of a software product, or the part that contributes to the prior art or the part of the technical solution. The computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM) random access memory (RAM), disk or optical disk, and other media that can store program codes.

[0124] In this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0125] The above description is only an embodiment of the present application and is not intended to limit the protection scope of the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A large model training system based on federated learning, characterized in that: The system includes a server and multiple clients communicating via an orthogonal frequency division multiple access network, wherein: Each client downloads the initial global model from the server and freezes it; Each client trains the first to-be-trained matrix and the second to-be-trained matrix of the low-rank adapter using local sample data to obtain a first gradient matrix and a second gradient matrix; Each client performs ranking on the first gradient matrix and the second gradient matrix to determine the first ranked matrix and the second ranked matrix and upload them to the server; The server determines the first global matrix and the second global matrix of this round based on the first rank selection matrix and the second rank selection matrix uploaded by all clients received in the previous round, and sends them to each client, so that each client updates the first matrix to be trained and the second matrix to be trained in the next round; When the preset round is reached, each client adjusts the model parameters of the initial global model based on the last received first global matrix and second global matrix as the model parameters of the target large model.

2. The system according to claim 1, characterized in that Each client obtains the first gradient matrix in the following way and the second gradient matrix Among them, the first gradient matrix is the loss function f relative to the first global matrix A t-1 The gradient matrix, the second gradient matrix is the loss function f relative to the second global matrix B t-1 The gradient matrix of , t is the current round, is the local sample data set of client k, d is the data sample, For client k A sub-sample data set is formed by randomly selecting ξ data samples from .

3. The system according to claim 1, characterized in that Each client determines the first rank selection matrix by the following method and the second rank selection matrix in, for The rth column of for The rth column of t,R is the corresponding ranking index.

4. The system according to claim 1, characterized in that The server determines the first global matrix A of round t by the following method: t and the second global matrix B t : Where η is the learning rate, s t,k is the transmission index, q t,k is the transmission interruption probability of client k, P is the probability, D k for The number of samples.

5. The system according to claim 1, characterized in that Each client adjusts the parameters of the initial global model in the following way to generate the target large model: W=W0+B t A t T ; Among them, W is the model parameter of the target large model, and W0 is the model parameter of the initial global model.

6. A large model training method based on federated learning, characterized in that: Applicable to a server and multiple clients communicating via an OFDMA network, where: Each client downloads the initial global model from the server and freezes it; Each client trains the first to-be-trained matrix and the second to-be-trained matrix of the low-rank adapter using local sample data to obtain a first gradient matrix and a second gradient matrix; Each client performs ranking on the first gradient matrix and the second gradient matrix to determine the first ranked matrix and the second ranked matrix and upload them to the server; The server determines the first global matrix and the second global matrix of this round based on the first rank selection matrix and the second rank selection matrix uploaded by all clients received in the previous round, and sends them to each client, so that each client updates the first matrix to be trained and the second matrix to be trained in the next round; When the preset round is reached, each client adjusts the model parameters of the initial global model based on the last received first global matrix and second global matrix as the model parameters of the target large model.

7. The method according to claim 6, characterized in that Each client obtains the first gradient matrix in the following way and the second gradient matrix Among them, the first gradient matrix is the loss function f relative to the first global matrix A t-1 The gradient matrix, the second gradient matrix is the loss function f relative to the second global matrix B t-1 The gradient matrix of , t is the current round, is the local sample data set of client k, d is the data sample, For client k A sub-sample data set is formed by randomly selecting ξ data samples from .

8. The method according to claim 6, characterized in that Each client determines the first rank selection matrix by the following method and the second rank selection matrix in, for The rth column of for The rth column of t,R is the corresponding ranking index.

9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory through the bus, and the processor executes the machine-readable instructions to perform the steps of the large model training method based on federated learning as described in any one of claims 6 to 8.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the large model training method based on federated learning as described in any one of claims 6 to 8.

Citation Information

Patent Citations

  • Multi-agent air-ground network resource allocation method based on federated learning

    CN116546462A

  • Federal learning resource optimization design method based on statistical channel state information

    CN116633462A

  • Federal intelligent language translation method based on low-rank adaptation

    CN118446232A

  • Parameter efficient big language fine-tuning federal learning framework

    CN118504526A

  • Industrial large model training method and system based on efficient fine tuning and federated learning

    CN118982074A