Heterogeneous federated learning-based network traffic large model construction method

Through the heterogeneous federated learning framework and LoRA module optimization, the privacy, resource and heterogeneity issues of large language models in network traffic analysis are solved, and efficient and secure network traffic classification is achieved.

CN120811731APending Publication Date: 2025-10-17SUZHOU UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511131001.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional large language models have data privacy and sample imbalance issues in network traffic analysis, high computing resource consumption, and low training efficiency in heterogeneous federated learning environments, making it difficult to meet real-time and security requirements.

Method used

A heterogeneous federated learning framework is adopted. By decoupling the teacher model into a simulator and adapter, the LoRA module is used for parameter optimization and fine-tuning, and the weighted stacking algorithm is combined to aggregate the client model to achieve efficient training and privacy protection.

Benefits of technology

It significantly improves the accuracy of network traffic classification, reduces training costs and resource requirements, improves training stability and efficiency, meets data security requirements, and is suitable for network traffic analysis in heterogeneous environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120811731A_ABST
    Figure CN120811731A_ABST
Patent Text Reader

Abstract

The invention discloses a network traffic large model construction method based on heterogeneous federated learning, which comprises the following steps: analyzing, cleaning and formatting network traffic data, and constructing question and answer pairs; performing semantic equivalence reappearing on the public data set by using a large language model to generate an enhanced data set; performing architecture optimization on the generative basic model to replace an original output layer with a customizable network classification head; decoupling the teacher model into a simulator and an adapter; initializing a LoRA module for a simulator and an adapter of the student model; the client adaptively adjusts the LoRA rank according to local resources; the client only finely adjusts parameters of the LoRA module of the adapter and freezes other parameters; the server adopts a weighted stacking algorithm to aggregate the LoRA modules uploaded by the client; the method has good universality and expandability.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of artificial intelligence and network technology, and particularly relates to a network traffic large model construction method based on heterogeneous federated learning. BACKGROUND

[0002] To solve the generalization bottleneck of traditional machine learning models, researchers have begun to explore the application of large language models (LLM) with stronger generality and representation ability for network traffic analysis. The basic principle of this idea is to use the powerful learning ability of large models to autonomously extract deeper and more essential semantic features from raw traffic data, in order to obtain better generalization performance. Although this idea has great potential in theory, it faces a series of more severe challenges in application, which together constitute the current technical difficulties that need to be broken through. Specifically, these challenges and their reasons are as follows: First, there are fundamental restrictions of data privacy and sample imbalance. Network traffic data often contains user privacy, business secrets and other highly sensitive information, which is strictly limited by data security regulations (such as GDPR), resulting in the inability to centralize data from all parties for model training. At the same time, malicious traffic samples are extremely sparse in real network environments, resulting in a serious class imbalance problem in training data, which affects the model training effect. Second, there is a contradiction between performance and cost of large models. The large number of parameters and complex structure of large models result in the consumption of tremendous computing resources and time in the training process, and the inference speed is slow, which conflicts with the real-time requirements of low delay and high timeliness of network traffic analysis business. Finally, the introduction of the federated learning framework has produced new technical difficulties. Although federated learning provides a path to solve the problem of data privacy, it introduces new difficulties when combined with large models. The core reasons are as follows: First, the hardware devices of different clients (data holders) have great differences in computing power, storage space and network bandwidth (i.e. device heterogeneity), which will seriously interfere with the synchronization and convergence efficiency of distributed training; second, the high intensity of large models on client local computing resources exceeds the carrying capacity of most ordinary devices, making the training process difficult to proceed efficiently and stably, and even unable to complete. SUMMARY

[0003] The purpose of the present application is to provide a network traffic large model construction method based on heterogeneous federated learning, which solves the problems of "high cost and low efficiency" when large models are applied to network traffic analysis, solves the problems of "data islands" and privacy protection in the model training process, and solves the "multiple bottleneck" problem of large model training in the heterogeneous federated learning environment.

[0004] Technical scheme: The network traffic large model construction method based on heterogeneous federated learning provided by the present application comprises the following steps:

[0005] (1) Analyze, clean and format network traffic data to construct structured question-answer pairs containing instructions and expected outputs;

[0006] (2) Use large language models to perform semantic equivalent restatement on public datasets to generate enhanced datasets;

[0007] (3) Perform architecture optimization on generative base models, replacing the original output layer with a customizable network classification head consisting of a fully connected layer and a Softmax function;

[0008] (4) Decouple the teacher model into an emulator and an adapter, where the emulator extracts general semantic features and the adapter performs classification decisions;

[0009] (5) Migrate teacher model knowledge to student model through two-stage knowledge distillation: general ability migration: minimize the feature difference loss between student model emulator and teacher model emulator; specific task migration: align the adapter decision boundary through KL divergence;

[0010] (6) Initialize LoRA modules for the emulator and adapter of the student model;

[0011] (7) The client adjusts the LoRA rank adaptively according to local resources;

[0012] (8) The client only fine-tunes the adapter LoRA module parameters, freezing other parameters;

[0013] (9) The server uses a weighted stacking algorithm to aggregate the LoRA modules uploaded by the clients.

[0014] Further, in step (1), the specific process of data formatting includes: using tshark tool to extract frame layer, Ethernet layer, IP layer, TCP / UDP layer field; and performing load data truncation processing.

[0015] Further, in step (3), the network classification head is constructed as follows: extract the high-dimensional feature vector Z output by the generative large language model; map Z to logits through a fully connected layer, and convert logits to probability distribution using the Softmax function; specifically as follows:

[0016] logits=W·Z+b

[0017] Where W is a matrix of dimension Cxd, and b is a bias vector of dimension C. W and b are training parameters of the classification head.

[0018] Y pred =Softmax(logits)

[0019] Further, in step (5), the feature difference loss formula is as follows:

[0020]

[0021] wherein, represents the average value of all sample losses in a batch when using batch training, represents the feature output by the student model simulator, represents the feature output by the teacher model simulator.

[0022] Further, in step (5), the adapter decision boundary is aligned by KL divergence, and the loss function is:

[0023]

[0024] wherein, represents the original logits output of the adapter of the student model.

[0025] Further, in step (6), the update process satisfies:

[0026] ΔW=BA

[0027] h=Wx+ΔWx

[0028] wherein, A∈R r×k and B∈R d×r is a low-rank decomposition matrix, r is the rank of the LoRA module, and r

[0029] Further, in step (7), the formula is as follows:

[0030]

[0031] wherein, r min and r max respectively define the minimum and maximum boundaries of the global LoRA rank, to constrain the range of the calculation result; α, β, γ are three hyperparameters, respectively used to weigh the influence weight of the three heterogeneous factors of data volume, data complexity and hardware resources in the rank calculation model; and is the normalized feature of the original feature of each client.

[0032] Further, in step (8), the details are as follows: the simulator original weight is frozen during training by the client the adapter original weight the simulator LoRA module the updated parameter is the LoRA module of the adapter

[0033] Further, in step (9), the weighted stacked aggregation algorithm includes: setting the adapter LoRA module uploaded by the client s as where m and n are the dimensions of the original weight matrix, r s is the local LoRA rank of the client s, then the global LoRA module of the next round is is generated as follows:

[0034] Weighted and vertically stacked matrix A, stack each client matrix by its corresponding scaling factor p s Then stack all the results along the vertical direction.

[0035]

[0036] Horizontally stacked matrix B, stack all the clients matrix is directly stacked along the horizontal direction, and no weighting is performed here

[0037]

[0038] After stacking, the dimensions of the new global LoRA module become:

[0039]

[0040] The electronic device described in the application comprises a processor and a memory, wherein the memory stores a computer program, and when the program is executed by the processor, the method steps of any of the above are realized.

[0041] Beneficial effects: Compared with the prior art, the present application has the following remarkable advantages: the present application realizes high-precision network traffic classification, and the performance exceeds that of existing methods: the present application combines the powerful representation capability of a large model with an architecture optimized for classification tasks, significantly improving the recognition accuracy. Compared with the most advanced network traffic classification model based on machine learning, the accuracy is improved by 1.5%-36.36% on the USTC-TFC-2016 dataset, by 4.92%-40% on the ISCX-Botnet-2014 dataset, and by 16.79%-75.12% on the CSTNET-2023 dataset. The training and application costs of the model are significantly reduced, and the efficiency is greatly improved: the present application realizes precise training of a small number of parameters by adopting the "emulator-adapter" separation architecture and the parameter efficient fine-tuning (LoRA) technology. Compared with the traditional full-parameter fine-tuning method, the present application reduces the required training data volume by about 90% with only about 1%-3% performance loss; compared with existing large model federated learning frameworks such as FedOT and FedIT, the required training parameter volume can be reduced by up to 88.74%. This greatly reduces the requirements for client hardware resources (computing power, memory) and overall training overhead. The present application enhances the training stability and convergence speed in a heterogeneous federated environment: by stacking LoRA modules for aggregation, the convergence speed of the present method is improved by about 40%-60% compared with traditional federated learning frameworks, and the performance fluctuation during training is smaller, effectively avoiding model performance degradation and significantly improving the robustness of the global model. The present application follows the core principles of federated learning, and the original network traffic data containing user privacy and commercial secrets of each client remains local during the entire collaborative training process and does not need to be uploaded to the central server. Only lightweight model parameters (adapter LoRA modules) that do not contain raw information are exchanged in an encrypted channel, effectively eliminating the risk of data leakage and meeting strict data security and compliance requirements. The technical framework proposed by the present application is decoupled from the specific large language model underlying architecture and has good portability. It can be flexibly applied to various mainstream generative pre-training models on the market, helping users quickly and cost-effectively transform them into specialized models suitable for specific downstream tasks such as network traffic analysis, and has strong ease of use and technical compatibility. BRIEF DESCRIPTION OF DRAWINGS

[0042] Figure 1 is a flowchart of the present application;

[0043] Figure 2 is an optimization comparison chart of the present application. DETAILED DESCRIPTION

[0044] The technical solutions of the present application will be further described below in conjunction with the drawings.

[0045] As Figure 1 shown, the embodiment of the present application provides a network traffic large model construction method based on heterogeneous federated learning, comprising the following steps:

[0046] S1, data analysis, cleaning and formatting. Define tshark extraction field list, specifically including:

[0047] Frame layer: frame.time, frame.len, frame.protocols..

[0048] Ethernet layer: eth.dst, eth.src, eth.type...

[0049] IP layer: ip.src, ip.dst, ip.proto, ip.ttl...

[0050] TCP layer: tcp.srcport, tcp.flags, tcp.payload..

[0051] UDP layer: udp.srcport, udp.length...

[0052] Construct tshark command, extract specified fields, and parse and format each row of data, including escaping special characters (such as \n converted to truncated characters to avoid processing long payloads, ignoring fields with no value), each row corresponding to a packet.

[0053] Adjust the data processed after parsing to the format of supervised instruction fine-tuning, taking the ustc-tfc dataset as an example as follows:

[0054] {

[0055] Instruction: Given the following traffic data <packet>that contains protocol fields, traffic features, and payloads. Please conduct the ENCRYPTED MALWARE DETECTION TASK to determine which application category the encrypted beign or malicious traffic belongs to. The categories include 'BitTorrent, FTP, Facetime, Gmail, MySQL, outlook, SMB, Skype, Weibo,

[0056] World of Warcraft, Cridex, Geodo, Htbot, Miuref, Neris, Nsis-ay, Shifu, Tinba, Virut, Zeus'.\n <packet>:Data packet after S12 parsing

[0057] Output: corresponding label (such as BitTorrent)

[0058] }

[0059] S2. After data parsing, 100 items are extracted from the existing public datasets of different categories (such as USTC TFC2016, ISCX BOTNET 2014, etc.) to construct the public dataset D p . After getting D p Afterwards, we use the currently powerful DeepSeek language model to construct more new sample pairs. Each group (I, O) is constructed into 3 new sample pairs, namely:

[0060] {(I1,O),(I2,O),(I3,O)}

[0061] The new public dataset D is obtained by the above method. public

[0062] S3. Optimize the architecture of the original generative model. Here, we optimize two models by replacing their dedicated network classification heads to adapt to the network traffic classification task. The two models are chatglm2-6B and chatglm2-2B. The optimized models are set as the teacher model and student model respectively for subsequent model decoupling and distillation compression. The specific principle can be compared with the original generative model. Figure 2 The optimized model reasoning is divided into two stages. The first stage is feature extraction. base In the model inference process, the input data X is mapped to a high-dimensional feature space to generate a feature vector Z. The specific process is as follows:

[0063] Z=M base (X)

[0064] The second stage is classification prediction. Through the extracted high-dimensional features, the new classification head H cls Accepts the feature vector Z and converts it into classification probabilities as follows:

[0065] Y pred =H cls (Z)

[0066] A standard classification head usually consists of one or more fully connected layers (Linear Layer) and a Softmax activation function. Assume that network traffic has C categories (for example: HTTP, FTP, DNS, P2P, etc.) and the dimension of the feature vector Z is d. The internal calculation of the classification head can be decomposed into:

[0067] logits = W · Z + b

[0068] W is a matrix of dimension C x d, b is a bias vector of dimension C, and W and b are trainable parameters of the classification head.

[0069] The logits are converted to a probability distribution using the Softmax function. The probability value for each class is between 0 and 1, and the sum of the probabilities for all classes is 1.

[0070] Y pred = Softmax(logits)

[0071] For the i-th class, the formula for calculating the predicted probability is:

[0072]

[0073] S4, the server will determine whether it is the first round of training, if so, the teacher model will be decoupled. First, select the optimized teacher model chatglm2-6b for model decoupling, which can be represented as M T represents the entire optimized teacher model, ε represents the simulator of the teacher model, represents the adapter of the teacher model. We set the first 2 / 3 of the model as ε, and the remaining 1 / 3 as

[0074] S5, the optimized student model chatglm2-2b is also decoupled, which can be represented as M S represents the entire optimized student model, ε represents the simulator of the student model, represents the adapter of the student model. General ability transfer, in order to let the student model simulator have the powerful general feature extraction ability of the teacher model simulator, construct the loss function:

[0075]

[0076] When using batch training, take the average of all sample losses in a batch, represents the feature output by the student model simulator, represents the feature output by the teacher model simulator.

[0077] Specific task transfer ability learning, in order to let the student model adapter better align the teacher model adapter on the decision boundary, establish the adapter optimization formula specific process as follows:

[0078] First, compute the softened probability distribution of the teacher model's adapter for input sample x:

[0079]

[0080] where, represents the logits output of the teacher model's adapter, and τ is the temperature coefficient.

[0081] Similarly, compute the softened probability distribution of the student model's adapter for input sample x:

[0082]

[0083] where, represents the logits output of the student model's adapter.

[0084] Construct the loss function for the decision boundary learning of the adapter:

[0085]

[0086] τ 2 to compensate for the impact of temperature scaling on gradients.

[0087] The final loss function for the entire distillation process is:

[0088] L total = δL base + εL KD + ωL CE

[0089] L CE is the cross-entropy loss between the student model and the true label. Here we set δ to 0.7, ε to 0.2, and ω to 0.1, so that the student model's simulator can maintain as close to the performance of the teacher model as possible, focusing on the extraction of general features.

[0090] S6、After the distillation and compression of the model, initialize the corresponding LoRA module for ε * and . The goal is to efficiently fine-tune the model with a small number of trainable parameters. This process involves low-rank decomposition and updating of the weight matrices in specific layers of the model. Suppose there are a series of weight matrices that need to be adapted in the model ε * and . Take the weight matrix W0∈R d×k of any fully connected layer or attention module as an example, its update process during model fine-tuning is as follows:

[0091] For the original weight matrix W0, its update amount △W is represented as the product of two low-rank matrices:

[0092] AW = BA

[0093] h = Wx + AWx

[0094] where A ∈ R r×k and B ∈ R d×r is a low-rank decomposition matrix, r is the rank of the LoRA module, and r « min(d, k), and h is the output of the layer

[0095] For any target weight matrix in the adapter its forward propagation process is modified as:

[0096]

[0097] where x is the input vector, is the output vector. and is a low-rank matrix injected for the adapter, is its rank, and At initialization, is usually randomly generated by a Gaussian distribution, is initialized as a zero matrix to ensure that the training starts before

[0098] Similarly for any target weight matrix in the simulator ε * its forward propagation process is modified as:

[0099]

[0100] where, and is a low-rank matrix injected for the simulator, is its rank.

[0101] After initialization or at the beginning of a new round of federated learning, the server Server distributes the simulator ε * , the adapter corresponding LoRA module (the first three are sent only in the first round of federated learning) and to the local client for fine-tuning training. Assuming that at the beginning of the tth iteration, the server sends the model components to the client as follows:

[0102]

[0103] S7, when the local client accepts the model, it will start the adaptive low-rank adaptation mechanism for efficient model training in the heterogeneous federated learning scenario. First, the local client resources will be standardized:​​

[0104]

[0105] Using D s , E s and C s to represent the three key heterogeneous features specific to the s-th client: local data volume, data distribution complexity (e.g., information entropy), and hardware computing power. To standardize these features, we first need to determine their global ranges. Therefore, we define D max , E max , C max , and D min , E min , C min , which represent the global maximum and minimum values of the above three indicators across all client sets. Based on these extreme values, we normalize the original features of each client to obtain and as inputs for subsequent calculations.

[0106] Next, we use the above quantitative indicators to adaptively adjust the low-rank mechanism. The core goal is to enable clients with high-quality data and strong computing resources to perform more complex local training, while resource-constrained clients perform lightweight updates, thereby achieving the best balance between training efficiency and model performance on a global level. The specific formula is as follows:

[0107]

[0108] In the above formula, r min and r max define the minimum and maximum boundaries of the global LoRA rank to constrain the range of the calculation results. In addition, α, β, γ are three adjustable hyperparameters that are used to weigh the influence of data volume, data complexity, and hardware resources in the rank calculation model. Here we set α = 0.34, β = 0.33, γ = 0.33 to balance the three factors and maintain optimal stability.

[0109] S8, when the local client receives the model training, to achieve the core goal of the present application, i.e., only fine-tuning the adapter part most relevant to the specific task, the client freezes the original weights of the simulator Adapter original weights Simulator LoRA module Therefore, during the entire local training process, the only trainable parameter is the LoRA module of the adapter

[0110] S82, the training objective of client s is to minimize the loss function L s on the local dataset D s . Let the model parameter set be Θ, where the trainable part is The optimization process of local training can be represented as:

[0111]

[0112] where η is the learning rate, is the gradient of the loss function with respect to the current model parameters. According to the aforementioned freezing strategy, the gradient is only calculated and updated for , i.e.,

[0113]

[0114] After several epochs of local training, client s obtains a set of LoRA parameters of the adapter optimized by the local data, denoted as

[0115] S9, when the local training is completed, the client s uploads the training results to the server. In order to greatly reduce the communication bandwidth requirement and protect the data privacy, the client only uploads the adapter LoRA module parameters with extremely small parameter quantity obtained by training:

[0116]

[0117] After the server collects the LoRA module updates from all clients, it performs aggregation operations to form the global model of the next round. The server aggregates all the LoRA modules uploaded by the clients in a weighted stacking manner. For each client s, its uploaded LoRA weight The server will allocate a scaling factor p s to each client according to the proportion of the local data volume |D s | in the total data volume:

[0118]

[0119] where S is the total number of clients participating in the current round of training, and

[0120] Then the server will perform asymmetric weighted sum and stacking operations on all the collected local LoRA modules . Let the adapter LoRA module uploaded by client s be and where m and n are the dimensions of the original weight matrix, and r s is the local LoRA rank of client s, which can be different. The global LoRA module of the next round is The matrix A is weighted and vertically stacked, each client

[0121] The matrix A is weighted and vertically stacked, each client The matrix is multiplied by its corresponding scaling factor p s Then all the results are stacked along the vertical direction.

[0122]

[0123] The matrix B is horizontally stacked, all clients' results are The matrix is directly stacked along the horizontal direction, no weighting here

[0124]

[0125] Where The symbol represents the stacking operation.

[0126] After stacking, the dimension of the new global LoRA module becomes:

[0127]

[0128] Importantly, this construction guarantees that the aggregated global update Mathematically, it is exactly equal to the weighted sum of individual local updates:

[0129]

[0130] This method ingeniously applies the scaling factor only to one of the matrices, thus avoiding the noise in the aggregation of the traditional method, achieving accurate and noise-free aggregation. After aggregation, the server distributes the new global LoRA module with increased dimension to all participating clients. After receiving, the clients update their local models using the module and start the next round of local training based on it; return to step S6. This loop is repeated until the model converges.< / packet> < / packet>

Claims

1. A method for constructing a large network traffic model based on heterogeneous federated learning, characterized in that: The following steps are involved: (1) Parse, clean, and format network traffic data to construct structured question-answer pairs containing instructions and expected outputs; (2) Using large language models to perform semantically equivalent restatements of public datasets to generate enhanced datasets; (3) The architecture of the generative base model is optimized to replace the original output layer with a customizable network classification head, which consists of a fully connected layer and a softmax function; (4) Decoupling the teacher model into a simulator and an adapter, where the simulator extracts common semantic features and the adapter performs classification decisions; (5) Transferring teacher model knowledge to the student model through two-stage knowledge distillation: General capability transfer: minimizing the feature difference loss between the student model simulator and the teacher model simulator; Task-specific transfer: aligning adapter decision boundaries via KL divergence; (6) Initialize the LoRA module for the simulator and adapter of the student model; (7) The client adaptively adjusts the LoRA rank based on local resources; (8) The client only fine-tunes the adapter LoRA module parameters and freezes other parameters; (9) The server uses a weighted stacking algorithm to aggregate the LoRA modules uploaded by the client.

2. A method for constructing a large network traffic model based on heterogeneous federated learning according to claim 1, characterized in that: In step (1), the specific process of data formatting includes: using the tshark tool to extract the frame layer, Ethernet layer, IP layer, and TCP / UDP layer fields; and truncating the payload data.

3. The method for constructing a large network traffic model based on heterogeneous federated learning according to claim 1 is characterized in that: In step (3), the network classification head is constructed by extracting the high-dimensional feature vector Z output by the generative large language model; mapping Z to logits through a fully connected layer, and converting logits into a probability distribution using the Softmax function; specifically as follows: logits=W·Z+b Where W is a matrix of dimension C×d, · is a bias vector of dimension C, and W and b are the training parameters of the classification head. Y pred =Softmax(logits)。 4. The method for constructing a large network traffic model based on heterogeneous federated learning according to claim 1 is characterized in that: In step (5), the feature difference loss formula is as follows: in, Indicates that when batch training is used, the average loss of all samples in a batch is taken. represents the characteristics of the student model simulator output, Represents the features of the teacher model simulator output.

5. The method for constructing a large network traffic model based on heterogeneous federated learning according to claim 1 is characterized in that: In step (5), the adapter decision boundary is aligned by KL divergence, and the loss function is: in, Represents the raw logits output of the adapter from the student model.

6. The method for constructing a large network traffic model based on heterogeneous federated learning according to claim 1 is characterized in that: In step (6), the update process satisfies: ΔW=BA h=Wx+ΔWx Where A∈R r×k and B∈R d×r It is the low-rank decomposition matrix, which is the rank of the LoRA module, and r<<min(d,k), and h is the output of the current layer.

7. The method for constructing a large network traffic model based on heterogeneous federated learning according to claim 1 is characterized in that: In step (7), the formula is as follows: Among them, r min and r max The minimum and maximum bounds of the global LoRA rank are defined to constrain the range of the calculation results; α, β, and γ are three hyperparameters used to weigh the influence of three heterogeneous factors, namely data volume, data complexity, and hardware resources, in the rank calculation model. and The normalized features of the original features of each client.

8. The method for constructing a large network traffic model based on heterogeneous federated learning according to claim 1, characterized in that: In step (8), the details are as follows: The client freezes the original weights of the simulator during training Adapter original weight Emulator LoRA module Update parameter is the adapter's LoRA module 9. The method for constructing a large network traffic model based on heterogeneous federated learning according to claim 1, characterized in that: In step (9), the weighted stacking aggregation algorithm includes: assuming that the adapter LoRA module uploaded by client s and Where m and n are the dimensions of the original weight matrix, r s is the local LoRA rank of client s, then the global LoRA module of the next round Generate as follows: The weighted matrix A is stacked vertically, and each client The matrix is ​​multiplied by its corresponding scaling factor p s , and then stack all the results vertically. Matrix B is stacked horizontally, with all clients The matrices are stacked directly in the horizontal direction, and no weighting is performed here. After stacking, the dimensions of the new global LoRA module become:

10. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program, and when the program is executed by the processor, the method steps according to any one of claims 1 to 9 are implemented.

Citation Information

Cited By

  • Model training method and device based on federated learning, medium and equipment

    CN122452693A

  • Model training method and device based on federated learning, medium and equipment

    CN122452693B