Large model federal splitting privacy protection method based on homomorphic encryption

By splitting the large language model into client and server sub-models and using compressed sensing and homomorphic encryption technologies, the problems of communication and resource limitations in federated learning are solved, and efficient training and privacy protection of large language models in specific fields are achieved.

CN120768520APending Publication Date: 2025-10-10KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510755506.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Directly training or fine-tuning large language models (LLMs) in the federated learning paradigm incurs high communication overhead and poses great challenges to the storage and computing resources of edge devices, while also posing the risk of sensitive information leakage.

Method used

A large-model federated splitting privacy protection method based on homomorphic encryption is adopted to split the large language model into a client sub-model and a server sub-model. The activation value is compressed using compressed sensing technology, and the activation value is encrypted and transmitted using the homomorphic encryption algorithm CKKS, reducing communication and computing overhead and utilizing the computing power of the server side to fine-tune the model.

Benefits of technology

It effectively reduces the storage and computing resource requirements of edge clients, improves model training speed and privacy protection capabilities, solves the problems of resource limitations and privacy leakage, and realizes efficient training and privacy protection of large language models in specific fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120768520A_ABST
    Figure CN120768520A_ABST
Patent Text Reader

Abstract

The invention relates to a big model federal splitting privacy protection method based on homomorphic encryption, and belongs to the technical field of big language models and privacy protection. The method comprises the following steps: splitting a large language model into a client side sub-model and a server side sub-model; after the client finishes each round of local training, performing compression processing on an intermediate activation value generated by each client sub-model through a compressed sensing technology; then encrypting and transmitting the compressed activation value by using a homomorphic encryption algorithm, thereby reducing the calculation overhead of the homomorphic encryption algorithm and the communication overhead between the client and the server, and providing privacy protection capability for transmitted sensitive data; the server terminal model is updated by using the gradient of the client, and each client uses local private data to finely adjust the server terminal model, so that the original data of the client is not out of the domain and is available and invisible, and the data privacy security is protected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a large model federated splitting privacy protection method based on homomorphic encryption, belonging to the field of large language models and privacy protection technology. Background Art

[0002] With the booming development of artificial intelligence, data and models have become key factors in promoting technological progress. Massive data contains rich information and knowledge, which has enabled large language models (LLMs) to achieve remarkable results in key areas such as natural language understanding, multimodal reasoning, and complex decision-making. Their performance gains depend on the supply quality of multimodal, cross-domain, heterogeneous data resources. It is expected that available public data sources will be exhausted around 2026. In order to break through the performance bottleneck of general models in vertical fields, the industry has turned to exploring collaborative training of private data, but this model carries the risk of sensitive information leakage. Federated learning allows cross-departmental model training without leaving the local data. It can use distributed data sources to provide large models with large amounts of data, alleviating the problem of data source exhaustion for large models.

[0003] However, due to the large parameter size of LLMs, directly training or fine-tuning them within the federated learning paradigm incurs high communication overhead and poses significant challenges to the storage and computing resources of edge devices. Furthermore, during the distributed training and fine-tuning of large models based on federated learning (FL), although servers cannot directly access training samples, attacks on FL have demonstrated that FL still faces threats such as inference attacks. These issues pose significant challenges to improving the performance of LLMs in specific domains and hinder the development and application of large models.

[0004] As an emerging distributed training paradigm, split learning (SL) shows great potential for training models on resource-constrained devices and can overcome the shortcomings of FL. SL divides the model between the client and a central server through model partitioning. The client retains the input layer (such as word embeddings) and output layer, while the server hosts most of the intermediate layers. This design effectively reduces the computational burden on the client by offloading intensive computation to the server, exchanging only activation values / activation gradients rather than raw data or complete model parameters. SL also incorporates the advantages of FL parallel learning to further improve training efficiency. In addition to model partitioning, SF also features periodic aggregation of server-side and client-side sub-models to achieve model synchronization after multiple rounds of training, which conforms to the design principles of FL. Summary of the Invention

[0005] The purpose of the present application is to provide a homomorphic encryption-based large model federated split privacy protection method, aiming to solve the technical problems of high communication overhead and storage and computing resource limitations of edge clients in the direct training or fine-tuning of LLMs in the federated learning paradigm.

[0006] To achieve the above-mentioned purpose, the present application provides a homomorphic encryption-based large model federated split privacy protection method, which uses compression sensing technology to compress the transmitted activation values, and uses homomorphic encryption algorithm CKKS to encrypt the transmitted activation values, thereby reducing the computational overhead of the homomorphic encryption algorithm and the communication overhead between the client and the server; by splitting the model into a client sub-model and a server sub-model, the large language model can be efficiently run in a distributed environment, fully utilizing the data resources of each participant, improving the training speed and inference ability of the model, and providing ideas for the privacy protection requirements of LLMs in specific fields under split fine-tuning, including the following steps:

[0007] Step 1: the server splits the large language model into a client sub-model and a server sub-model, and distributes the client sub-model to each client;

[0008] Step 2: after the splitting is completed, each client uses local private data to train the client sub-model, and after each round of training is completed, the weights generated by the client sub-model are compressed by two-dimensional discrete cosine transform, compressing the high-dimensional model parameters into one-dimensional parameters;

[0009] Step 3: after compression, the client performs privacy protection on the one-dimensional parameters compressed by the homomorphic encryption algorithm, and after each round of training is completed, the client and the server encrypt the weights transmitted between them, and sends the ciphertext weights to the server, and the server merges and updates the server sub-model after receiving the ciphertext weights from each client, thereby updating the client and completing the adjustment of the large language model using the local data of the client.

[0010] The client sub-model includes the first N layers of the large language model, and the server sub-model includes the last M layers of the large language model.

[0011] The principle of splitting the large language model is to migrate the intensive computation of the large language model to the server side to reduce the demand for storage and computing resources of edge clients. At the same time, considering the data transmission and communication efficiency, avoid excessive data transmission between the client and the server.

[0012] The step 2 specifically includes the following steps:

[0013] Step 2.1: define the variables in the training process, the number of clients is K, is the global initial model gradient, Represents a collection of trainable LoRa adapters for server-side pre-trained models, where represents the decomposition matrix of the k-th LoRA adapter, The total number of trainable LoRa adapters for server-side pre-trained models, is the local model weight of client i during the tth round of communication, and the perception basis matrix is , the sparse orthogonal basis matrix is ​​Ψ, and the compression rate is r;

[0014] Step 2.2: Based on the defined variables, after receiving the first N layers of the large language model, each client performs local training using local private data to obtain the weights generated by the client sub-model. ;

[0015] Step 2.3: For an n×n two-dimensional matrix , two-dimensional discrete cosine transform coefficients for:

[0016]

[0017] Among them, G(u) and G(v) are normalization coefficients, and the expressions are:

[0018]

[0019] The client will output the weight of Converted into a sparse vector s, the expression is:

[0020]

[0021] in, is a sparse orthogonal basis matrix with n rows and n columns;

[0022] Step 2.4: Compressed into one-dimensional parameters :

[0023]

[0024] in, It is a perceptual basis matrix with m rows and n columns, and the compression ratio is expressed as r=m / n.

[0025] Based on step 2 above, high-dimensional vector parameters can be compressed into one-dimensional vectors, greatly reducing the computational and communication burden of subsequent homomorphic encryption, making client-side fine-tuning of the server model more efficient.

[0026] The step 3 specifically includes the following steps:

[0027] Step 3.1: The key generation center generates the public key pk, private key sk, and relinearization key rlk of the homomorphic encryption algorithm CKKS (Cheon-Kim-Kim-Song), and distributes the private key to each client;

[0028] in, , CKKS.Enc() is the encryption algorithm, Indicates the ciphertext obtained after encryption;

[0029] in, , CKKS.Dec() is the decryption algorithm, and p is the plaintext obtained by decryption;

[0030] in, is homomorphic addition, is homomorphic multiplication;

[0031] Step 3.2: Based on the obtained one-dimensional parameters , use CKKS.Enc() algorithm to encrypt one-dimensional parameters get:

[0032]

[0033] The ciphertext weight obtained after encryption and client data true labels Upload to the server via wired or wireless channels;

[0034] Step 3.3: The server receives the ciphertext weights of all clients , using its own model parameters and the LoRA adapter of the server ti round Perform forward propagation and calculate the predicted value :

[0035]

[0036] in, Represents a given model parameter and trainable LoRA adapter set Input data The mapping relationship between the predicted value and the predicted value and the true label Calculate the loss function L:

[0037]

[0038] Where D is the model batch size, is the loss function, and then the server performs backpropagation Generate weights for server terminal models ;

[0039] Step 3.4: The server will Send it to each client, and the client will receive it. Then, use CKKS.Dec() to decrypt the plaintext information. :

[0040]

[0041] Step 3.5: Each client gets the plaintext Finally, the complete weight value is reconstructed using the two-dimensional inverse discrete cosine transform:

[0042]

[0043] in, is the two-dimensional inverse discrete cosine transform matrix, expressed as:

[0044]

[0045] Based on the full weight value Update the client sub-model weights and conduct the next round of training until the model converges.

[0046] The beneficial effects of the present invention are:

[0047] (1) This paper proposes a large-model federated splitting privacy protection method based on homomorphic encryption. By splitting LLMs into client-side sub-models and server-side sub-models, intensive computation is migrated to the server side. The server side with powerful computing power hosts most of the middle layers, and the client side only has the first N layers of Transformers of the large language model, effectively solving the storage and computing resource limitations of edge clients. At the same time, the distributed parallel training mechanism is used to efficiently use private data to fine-tune the server-side model.

[0048] (2) During the large-scale model task training process, the present invention compresses the intermediate activation values ​​generated by the edge client sub-model through compressed sensing technology, compresses the high-dimensional model parameters into a one-dimensional vector, and avoids the reduction of training efficiency due to insufficient client resources or communication delays by reducing the data that needs to be transmitted between the client and the server;

[0049] (3) This invention uses the homomorphic encryption algorithm CKKS to implement the inference attack faced by malicious attackers in the split fine-tuning of large language models. The activation value after client compression is encrypted and transmitted through the CKKS encryption algorithm, and the server-side sub-model is updated through homomorphic calculation. It protects the private data of edge clients safely and efficiently, greatly improving the privacy protection capability in split fine-tuning.

[0050] (4) In the process of large-scale model task training, especially when facing privacy-sensitive data, the present invention can solve the problems of data privacy and data source scarcity through federated split learning while avoiding the problem of insufficient client resources, providing new ideas for further integration and innovation of large language model split fine-tuning and privacy protection. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a schematic flow chart of the present invention;

[0052] Figure 2 It is a schematic diagram of the model splitting and client fine-tuning server model of the present invention. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0054] The diagrams provided in the following examples and the settings of specific parameter values ​​in the model are mainly for illustrating the basic concept of the present invention and for simulation verification of the present invention. In specific application environments, appropriate adjustments can be made according to actual scenarios and needs.

[0055] Example 1: A privacy protection method for large-scale model federated splitting based on homomorphic encryption, such as Figure 1 Shown, including:

[0056] Step 1: The server splits the large language model into a client sub-model and a server sub-model, and sends the client sub-model to each client.

[0057] The client sub-model includes the first N layers of Transformers of the large language model, and the server sub-model includes the last M layers of Transformers of the large language model.

[0058] Specifically, in this embodiment, the server first splits the LLMs into a client sub-model and a server sub-model. The client sub-model contains the first N layers, and the server sub-model contains the last M layers. The client sub-model is then distributed to each client. This embodiment uses the GPT2-s split as an example. The client sub-model contains the first three Transformer layers of GPT2-s, and the server sub-model contains the last nine Transformer layers of GPT2-s. Three edge clients are used to participate in LLMs model fine-tuning.

[0059] Step 2: After the split is completed, each client uses local private data to train the client sub-model. After each round of training, the weights generated by the client sub-model are compressed through a two-dimensional discrete cosine transform to compress the high-dimensional model parameters into one-dimensional parameters.

[0060] Step 2.1: Define the variables in the training process, the number of clients is K, is the global initial model gradient, Represents a collection of trainable LoRa adapters for server-side pre-trained models, where represents the decomposition matrix of the k-th LoRA adapter, The total number of trainable LoRa adapters for server-side pre-trained models, is the local model weight of client i during the tth round of communication, and the perception basis matrix is , the sparse orthogonal basis matrix is ​​Ψ, and the compression rate is r;

[0061] Step 2.2: Based on the defined variables, after receiving the first N layers of the large language model, each client performs local training using local private data to obtain the weights generated by the client sub-model. ;

[0062] Step 2.3: For an n×n two-dimensional matrix , the two-dimensional discrete cosine transform coefficient Ψ(u,v) is:

[0063]

[0064] Among them, G(u) and G(v) are normalization coefficients, and the expressions are:

[0065]

[0066] The client will output the weight of Converted into a sparse vector s, the expression is:

[0067]

[0068] in, is a sparse orthogonal basis matrix with n rows and n columns;

[0069] Step 2.4: Compressed into one-dimensional parameters :

[0070]

[0071] in, It is a perceptual basis matrix with m rows and n columns, and the compression ratio is expressed as r=m / n.

[0072] Specifically, in this embodiment, client 1 uses local private data to train the client sub-model and generate weight values ; Use the sparse orthogonal basis matrix Ψ to convert the weight value Convert it into a sparse vector s, and then use the compressed sensing algorithm to convert the sparse weight value Perform compression processing to obtain the compressed one-dimensional parameters .

[0073] Furthermore, set the weight value , after compression, output one-dimensional parameters .

[0074] Step 3: After compression is completed, the client uses a homomorphic encryption algorithm to protect the privacy of the compressed one-dimensional parameters. After each round of training, each client encrypts the weights transmitted between the client and the server and sends the ciphertext weights to the server. After receiving the ciphertext weights from each client, the server merges and updates the server-side sub-model, thereby updating the client and completing the adjustment of the large language model using the client's local data.

[0075] Step 3.1: The key generation center generates the public key pk, private key sk, and relinearization key rlk of the homomorphic encryption algorithm CKKS, and distributes the private key to each client;

[0076] in, , CKKS.Enc() is the encryption algorithm, Indicates the ciphertext obtained after encryption;

[0077] in, , CKKS.Dec() is the decryption algorithm, and p is the plaintext obtained by decryption;

[0078] in, is homomorphic addition, is homomorphic multiplication;

[0079] Step 3.2: Based on the obtained one-dimensional parameters , use CKKS.Enc() algorithm to encrypt one-dimensional parameters get:

[0080]

[0081] The ciphertext weight obtained after encryption and client data true labels Upload to the server via wired or wireless channels;

[0082] Step 3.3: The server receives the ciphertext weights of all clients , using its own model parameters and the LoRA adapter of the server ti round Perform forward propagation and calculate the predicted value :

[0083]

[0084] in, Represents a given model parameter and trainable LoRA adapter set Input data The mapping relationship between the predicted value and the predicted value and the true label Calculate the loss function L:

[0085] , )

[0086] Where D is the model batch size, is the loss function, and then the server performs backpropagation Generate weights for server terminal models ;

[0087] Step 3.4: The server will Send it to each client, and the client will receive it. Then, use CKKS.Dec() to decrypt the plaintext information. :

[0088]

[0089] Step 3.5: Each client gets the plaintext Finally, the complete weight value is reconstructed using the two-dimensional inverse discrete cosine transform:

[0090]

[0091] in, is the two-dimensional inverse discrete cosine transform matrix, expressed as:

[0092]

[0093] Based on the full weight value Update the client sub-model weights and conduct the next round of training until the model converges.

[0094] Specifically, in this embodiment, the client 1 uses the public key pk generated by the key generation center and the encryption algorithm CKKS.Enc() to convert the compressed one-dimensional parameter Encrypt and get the ciphertext:

[0095]

[0096] will be sent to the server through a wired or wireless channel, and the encrypted

[0097] .

[0098] Further, the remaining client 2 and client 3 are sequentially executed step 2, step 3, the server receives After, the CKKS.HomAdd() algorithm is executed to obtain the aggregation result :

[0099]

[0100] Further, the aggregation result is used to update the service end sub-model, and .

[0101] Further, the is sent to the client 1, client 2 and client 3;

[0102] Among them, the weight of the client 2 is:

[0103]

[0104] Among them, the weight of the client 3 is:

[0105]

[0106] .

[0107] Further, each client receives , decrypts by the private key sk, and then reconstructs the complete weight value by using the two-dimensional discrete cosine transform. The client performs the next round of training based on the complete weight value until the model converges.

[0108] Specifically, in the embodiment, the complete weight value is:

[0109] .

[0110] The specific embodiments of the application are described in detail above in combination with the drawings, but the application is not limited to the above-mentioned embodiments, and various changes can be made within the knowledge of those skilled in the art without departing from the purpose of the application.​​

Claims

1. A privacy protection method for large-scale model federated splitting based on homomorphic encryption, characterized by: The following steps are involved: Step 1: The server splits the large language model into a client sub-model and a server sub-model, and sends the client sub-model to each client; Step 2: After the split is completed, each client uses local private data to train the client sub-model. After each round of training, the weights generated by the client sub-model are compressed using a two-dimensional discrete cosine transform to compress the high-dimensional model parameters into one-dimensional parameters. Step 3: After compression is completed, the client uses a homomorphic encryption algorithm to protect the privacy of the compressed one-dimensional parameters. After each round of training, each client encrypts the weights transmitted between the client and the server and sends the ciphertext weights to the server. After receiving the ciphertext weights from each client, the server merges and updates the server-side sub-model, thereby updating the client and completing the adjustment of the large language model using the client's local data.

2. A large-model federated splitting privacy protection method based on homomorphic encryption according to claim 1, characterized in that: The client sub-model includes the first N layers of Transformers of the large language model, and the server sub-model includes the last M layers of Transformers of the large language model.

3. According to the privacy protection method for large-scale model federated splitting based on homomorphic encryption in claim 1, it is characterized in that: The step 2 specifically includes the following steps: Step 2.1: Define the variables in the training process, the number of clients is K, is the global initial model gradient, Represents a collection of trainable LoRa adapters for server-side pre-trained models, where represents the decomposition matrix of the k-th LoRA adapter, The total number of trainable LoRa adapters for server-side pre-trained models, is the local model weight of client i during the tth round of communication, and the perception basis matrix is , the sparse orthogonal basis matrix is ​​Ψ, and the compression rate is r; Step 2.2: Based on the defined variables, after receiving the first N layers of the large language model, each client performs local training using local private data to obtain the weights generated by the client sub-model. ; Step 2.3: For an n×n two-dimensional matrix , two-dimensional discrete cosine transform coefficients for: ; Among them, G(u) and G(v) are normalization coefficients, and the expressions are: ; The client will output the weight of Converted into a sparse vector s, the expression is: ; in, is a sparse orthogonal basis matrix with n rows and n columns; Step 2.4: Compressed into one-dimensional parameters : ; in, It is a perceptual basis matrix with m rows and n columns, and the compression ratio is expressed as r=m / n.

4. A large-model federated splitting privacy protection method based on homomorphic encryption according to claim 1, characterized in that: The step 3 specifically includes the following steps: Step 3.1: The key generation center generates the public key pk, private key sk, and relinearization key rlk of the homomorphic encryption algorithm CKKS, and distributes the private key to each client; in, is the encryption algorithm, Indicates the ciphertext obtained after encryption; in, is the decryption algorithm, and p is the plaintext obtained by decryption; in, is homomorphic addition, is homomorphic multiplication; Step 3.2: Based on the obtained one-dimensional parameters ,use Algorithm encryption one-dimensional parameters get: ; The ciphertext weight obtained after encryption and the true label of client data Upload to the server via wired or wireless channels; Step 3.3: The server receives the ciphertext weights of all clients , using its own model parameters and server Wheel LoRA Adapter Perform forward propagation and calculate the predicted value : ; in, Represents a given model parameter and trainable LoRA adapter set Input data The mapping relationship between the predicted value and the predicted value and the true label Calculate the loss function L: ; Where D is the model batch size, is the loss function, and then the server performs backpropagation Generate weights for server terminal models ; Step 3.4: The server will Send it to each client, and the client will receive it. Afterwards, use Decryption operation obtains plaintext information : ; Step 3.5: Each client gets the plaintext Finally, the complete weight value is reconstructed using the two-dimensional inverse discrete cosine transform: ; in, is the two-dimensional inverse discrete cosine transform matrix, expressed as: ; Based on the full weight value Update the client sub-model weights and conduct the next round of training until the model converges.

Citation Information

Cited By

  • Large model-oriented split privacy protection training method

    CN121561979A

  • Federal splitting large model training method and system based on Stackelberg game

    CN121562807A