Large model federal splitting trusted data space construction method based on rotation part secret sharing
By using the large-scale language model federated split trusted data space construction method that rotates part secretly shared in large-scale language model training, the model splitting strategy and client selection are optimized, and the problem of low data privacy protection and computing resource utilization efficiency faced by large-scale language models during training is solved, and federated learning with high performance and high privacy protection is achieved.
Patent Information
- Application Number
- CN202510250779.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-27
AI Technical Summary
Large language models face the problems of data privacy protection and low efficiency in utilization of computing resources during training, especially when integrating private data.
A large-model federated split trusted data space construction method based on rotational part secret sharing is adopted, and federated learning with high performance and high privacy protection is achieved by optimizing model splitting strategies and client selection. This method splits the big model into multiple model partitions, each model partition is trained by a group of client devices, and data privacy is protected by rotating partial secret sharing technology.
Effectively allocating the client's computing resources improves the speed of model training and inference efficiency, enhances data privacy protection capabilities, and solves the problems of low computing resource utilization and difficulty in data privacy protection faced by large language models during training.
Smart Images

Figure CN120218233A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for constructing a large model federated split trusted data space based on rotating partial secret sharing, belonging to the technical fields of federated learning and large language model splitting. Background Art
[0002] In the current booming development of artificial intelligence, data and models have become the key elements driving technological progress. Massive amounts of data contain rich information and knowledge, and deep learning models, especially large language models (LLMs), with their powerful learning and reasoning capabilities, can extract deep-seated laws and patterns from the data. However, with the continuous expansion of model scale and the increasing awareness of data privacy protection, traditional centralized data processing and model training methods face many challenges.
[0003] The "Action Plan for the Development of Trusted Data Spaces (2024 - 2028)" emphasizes the importance of trusted data spaces as data circulation and utilization infrastructure, and clearly points out to carry out research on core technologies of trusted data spaces such as privacy computing and high-performance encrypted computing, and explore the integrated innovation of large models and trusted data spaces. This indicates that ensuring the trustworthiness and security of data has become the key direction for industry development during the process of data processing and utilization. However, the training of large models depends on large-scale, diverse, and high-quality public datasets, and is facing the problem that the data sources will be exhausted around 2026. To improve their performance in specific fields, it is necessary to integrate data from private entities (such as hospitals and banks). These highly sensitive data bring privacy challenges and hinder the application of large models.
[0004] Federated learning is an emerging distributed machine learning method. It allows multiple participants to jointly train a global model without sharing the original data. Each participant only trains and updates the model locally, and then encrypts and sends the model parameters or gradient information to the server for aggregation, thereby achieving collaborative learning of the model. This method greatly reduces the risk of data leakage because the original data always remains local and is not transmitted over the network. Even if data leakage occurs during network communication, attackers cannot obtain the complete dataset. Therefore, the combination of federated learning and large models, that is, federated large models, can solve the privacy leakage problem faced by current large models. Summary of the Invention
[0005] The object of the present invention is to provide a method for constructing a large model federated split trusted data space based on rotating partial secret sharing. On the basis of protecting data privacy, by optimizing the model splitting strategy and enhancing data credibility, high-performance and high-privacy protection federated learning is achieved. Through reasonable model splitting, large language models can operate efficiently in a distributed environment, making full use of the computing resources of each participating party and improving the training speed and inference efficiency of the model. In the process of model splitting, the present invention pays attention to maintaining the interpretability of the model. By reasonably dividing the hierarchical structure and functional modules of the model, the role and contribution of each sub-module become clearer.
[0006] To achieve the above object, the present invention provides a method for constructing a large model federated split trusted data space based on rotating partial secret sharing, including the following steps:
[0007] Step 1: Split the large model into multiple model partitions, and each model partition is trained by a group of client devices. The clients are divided into several groups, and there are a certain number of clients in each client group. The client group waits for the allocation of model partitions as a candidate client group;
[0008] Step 2: After determining the number of model partitions and candidate client groups, solve an integer optimization problem to select an optimal group of client devices for the allocation of model partitions. All clients within each client group are responsible for the training task of the same model partition;
[0009] Step 3: During the task training process, data privacy protection is carried out by the method of rotating partial secret sharing. Each client divides its local model into multiple random parameter blocks, and selects a parameter block through a dynamic indexing mechanism. The client divides the selected parameter block into two secret shares, and one of the secret shares is sent to the next client within the same client group for allocation for secure aggregation, and the other secret share is securely aggregated with the secret share sent by the previous client.
[0010] The specific steps of Step 1 include the following steps:
[0011] Step 1.1: Perform model partitioning according to the computational graph structure of the model, divide the model into multiple sub-parts, and each sub-part corresponds to a model partition;
[0012] Step 1.2: Adjust the number of model partitions according to the number of client groups participating in the task training, so that the number of model partitions is equal to the number of client groups selected in each communication round;
[0013] Step 1.3: Combine data flow and communication efficiency to balance the computational load of each model partition and avoid excessive data transmission between clients.
[0014] Step 2 specifically includes the following steps:
[0015] Step 2.1: Define problem variables, including the total number of model partitions Z and the total number of candidate client groups N; use a binary variable α zn to indicate whether the model partition z is assigned to the client group n;
[0016] Step 2.2: Define the objective function A zn , representing the fitness between the model partition and the client:
[0017]
[0018] where D zn represents the transmission time of the model partition on the client, and T zn represents the average training time of the model partition on the client group, and a and b are weight hyperparameters;
[0019] Step 2.3: Define the constraints. Each model partition can only be assigned to one client group At the same time, all clients within a client group can only be used to train the same model partition If the GPU memory of a client in the group is not sufficient to train the model partition, no assignment is made: α zn = 0, where is the GPU memory of client n in the client group i , and |w Pz | is the number of parameters of the model partition z;
[0020] After completing the definition of the problem variables, objective function, and constraints, the integer optimization problem is expressed as: That is, maximize the sum of all fitnesses, representing faster transmission duration and training time;
[0021] Step 2.5: Since the constraint matrix is totally unimodular, convert the integer optimization problem into a linear programming problem and use a standard linear programming solver for solution to make the solution process more efficient;
[0022] When solving the linear programming problem, α zn takes values between [0, 1]. Since the constraint matrix is totally unimodular, the optimal solution of the linear programming still satisfies α zn ∈ {0, 1};
[0023] Step 2.7: In the optimal solution, α zn = 1 if and only if the assignment of the model partition z to the client group n is the optimal assignment that maximizes the total fitness, otherwise α zn = 0. Therefore, αzn When α = 1, the model partition z is assigned to the client group n. zn When α = 0, no assignment is made. Through the above steps, an optimal set of client devices can be effectively selected for model training, while ensuring the reasonable allocation and utilization of resources;
[0024] The above steps consider the computing resources and network bandwidth of the clients to ensure that the selected clients can effectively process the assigned model partitions. This allocation method enables different parts of the model to be processed in parallel on different clients, thereby improving the training efficiency.
[0025] Step 3 specifically includes the following steps:
[0026] Step 3.1: Define the client group to be assigned as a cell, and each cell is a set of clients; each cell is randomly divided into K j subgroups, that is, a smaller set of clients. The number of clients in the cell is an integer multiple of the subgroups;
[0027] Step 3.2: Let be the global initial model, is the local model of client n i at the t-th round of communication, is the local model of the j-th cell at the t-th round of communication, is the local model under the k-th parameter block, is the aggregation secret on client n i ;
[0028] Step 3.3: After receiving the model partition, the client performs local training to obtain the local model
[0029] Step 3.4: Each client n i cuts the weighted local model into K j parameter blocks: where θ i is the aggregation weight of the local model i of client n . The number of parameter blocks is equal to the number of subgroups in this cell. In different communication rounds, the client generates secret shares for different parameter blocks, and the dynamic block index is: φk = [(k + t) mod K j + 1, indicating that in each round t, for each subgroup k to which each client n i belongs, the shared parameter block is determined. k is the index of the subgroup, t is the current communication round, and K j is the total number of parameter blocks. The mod operation is used when the sum of k + t exceeds K jWhen the result loops back to 1, the index value loops among K j parameter blocks;
[0030] Step 3.5: After determining the parameter block to be encrypted the client n i splits into and and forwards to the next client in the same subgroup, where and represent two random secret parameter blocks derived from ;
[0031] Step 3.6: When the latter client receives the secret share from the former client, it aggregates the received secret and the local secret into to obtain where represents one of the two secret parameter blocks formed by splitting the selected parameter block by the former client n l ;
[0032] Step 3.7: Sum the aggregated secrets of the k-th subgroup to and splice all the subgroup parameter blocks as the complete cell model Then broadcast to the clients in the client group for the next round of local training.
[0033] The beneficial effects of the present invention are:
[0034] (1) A method for constructing a large model federated split trusted data space based on rotated partial secret sharing proposed by the present invention can effectively allocate the computing resources of clients by optimizing the model partitioning and client selection strategies, avoiding performance bottlenecks caused by excessive computing load of a single client. At the same time, the distributed training method makes full use of the computing resources of each participant, improving the overall training efficiency;
[0035] (2) Through a reasonable model splitting method, the present invention enables a large language model to share computing tasks among multiple clients. This splitting method not only improves the computing efficiency but also enhances the interpretability of the model. Each client only needs to focus on its allocated model partition, thus simplifying the training process and reducing the requirements for the computing resources of each client;
[0036] (3) The present invention adopts an integer optimization method to reasonably select the client devices participating in the training, fully considering the computing power, memory size, and network bandwidth of each client, thereby ensuring the optimal allocation of resources and avoiding the reduction of training efficiency caused by insufficient client resources or communication delays;
[0037] (4) During the training process of large-scale model tasks, especially when facing privacy-sensitive data, the present invention can solve the problems of data privacy and scarce data sources through the method of federated learning; in addition, the present invention adopts a method of rotating partial secret sharing during the local training process, enabling clients to exchange encrypted information through a partial secret sharing mechanism without directly transmitting model data between each other, thereby reducing communication overhead and enhancing privacy protection capabilities. At the same time, it also provides new ideas for the further combination and innovation of large models and trusted data spaces. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a schematic flow diagram of the present invention;
[0039] Figure 2 is a schematic diagram of splitting the model of the present invention and allocating clients. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] In order to make the use, technical solutions, and advantages of the present invention clearer and easier to understand, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0041] In the following examples, the diagrams provided and the setting of specific parameter values in the model are mainly for illustrating the basic concept of the present invention and conducting simulation verification on the present invention. In a specific application environment, appropriate adjustments can be made according to the actual scenario and requirements.
[0042] Embodiment 1: The embodiment of the present invention provides a method for constructing a large model federated split trusted data space based on rotating partial secret sharing, as Figure 1 shown, including:
[0043] Step 1: Split the large model into multiple model partitions, each model partition is trained by a group of client devices, divide the clients into several groups, and there are several clients in each client group. The client group serves as a candidate client group waiting for the allocation of model partitions.
[0044] Specifically, 40 clients are selected, and every 4 clients are grouped into one group, that is, there are 10 candidate client groups. Then, according to the computational graph structure of the model, the model is partitioned. Here, it is set that the model is divided into 4 model partitions, and the computational load of each model partition is kept balanced.
[0045] During the splitting process at this time, each model partition can be regarded as an independent computational task model and split according to its computational graph structure, including multiple layers or multiple functional modules. Each model partition will be assigned to a group of client devices for training in each communication round. The number of split model partitions corresponds to the number of selected client groups in that round.
[0046] Step 2: After determining the number of model partitions and candidate client groups, solve an integer optimization problem to select an optimal set of client devices for the assignment of model partitions. All clients within each client group are responsible for the training task of the same model partition.
[0047] Specifically, the goal of the integer optimization problem is to select a set of clients and determine the model partitions each client is responsible for, so as to maximize the training efficiency of the overall system. At this time, the total number of model partitions is 4, and the total number of candidate client groups is 10. Use a binary variable α zn to represent whether to assign model partition z to client group n; define a fitness A zn to represent the matching degree between the model partition and the client.
[0048] The optimization problem is:
[0049] The constraint conditions are: Each model partition can only be assigned to one client group All clients within one client group can only be used to train the same model partition If the GPU memory of the clients within the group is not sufficient to train the model partition, no assignment is made: α zn = 0; that is no assignment is made, and the value range of the binary variable: α zn ∈{0,1},
[0050] Furthermore, in the optimal solution, the value of α zn depends on whether the (z,n) pair is part of the optimal assignment. If it is optimal to assign model partition z to client group n, then α zn = 1, otherwise α zn = 0;
[0051] Now there are 4 model partitions z = 1, 2, 3, 4, and 10 candidate client groups n = 1, 2,..., 10. Assume that at this time for the A of model partition 1 1n , there is maxA 1n = A 11 = 0.9, (α 11 = 1), for the A of model partition 2 2n , there is maxA 2n = A22 = 0.8, (α 22 = 1), for A in model partition 3 3n , there is maxA 3n = A 33 = 0.7, (α 33 = 1), for A in model partition 4 4n , there is maxA 4n = A 44 = 0.9, (α 44 = 1), the sum of the result optimization problems is 3.3, that is, model partition 1 is assigned to client group 1, model partition 2 is assigned to client group 2, model partition 3 is assigned to client group 3, and model partition 4 is assigned to client group 4. Here, all α 11 , α 22 , α 33 , α 44 except are 0. zn
[0052] Step 3: During the task training process, data privacy protection is carried out by rotating some secret sharing. Each client divides its local model into multiple random parameter blocks and selects a parameter block through a dynamic indexing mechanism. The client divides the selected parameter block into two secret shares, one of which is sent to the next client within the same client group for assignment for secure aggregation, and the other secret share is securely aggregated with the secret share sent by the previous client.
[0053] Specifically, the situation after assigning model partition 1 to client group 1 is described here. At this time, client group 1 is divided into 2 subgroups, each subgroup has 2 clients, and these 4 clients all use their local data to train the received model partition to obtain local models After that, clients n1, n2 in subgroup 1 respectively divide the weighted local model into 2 parameter blocks and Suppose the parameter block selected through dynamic indexing is and Then is divided into and is divided into and Then and are aggregated into and are aggregated into The two secrets in subgroup 1 are aggregated into a subgroup secret: Among them
[0054] The subgroup 2 in client group 1 also operates according to the above steps, then aggregates the two subgroup secrets into a new cell model, and broadcasts the new cell model to the 4 clients in this cell for the next round of training. The operation steps of the remaining client groups are the same as those of client group 1.
[0055] Furthermore, during the task training process, each client will load its local data and train the model partition assigned to it; when each client is training, it will perform encrypted communication and aggregation through the rotated partial secret sharing technology without uploading the local data to other clients or the central server.
[0056] For the dynamic index part: If the sum of k + t is less than K j , then (k + t) mod K j is equal to k + t. If the sum of k + t is greater than or equal to K j , then (k + t) mod K j is equal to k + t - K j , where k is the index of the subgroup, t is the current communication round, and K j is the total number of parameter blocks. The mod operation is used to make the result loop back to 1 when the sum of k + t exceeds K j , so that the index value loops among the K j parameter blocks.
[0057] Furthermore, if a cell has 3 subgroups (K j = 3), in the first round (t = 1):
[0058] The dynamic block index φ of the first subgroup (k = 1) k = [(1 + 1) mod 3] + 1 = 3;
[0059] The dynamic block index φ of the second subgroup (k = 2) k = [(2 + 1) mod 3] + 1 = 1;
[0060] The dynamic block index φ of the third subgroup (k = 3) k = [(3 + 1) mod 3] + 1 = 2;
[0061] During the rotated partial secret sharing process, the clients generate and transmit secrets in parallel. For client n i it does not need to wait for the encryption result of the previous client n l in the secret sharing chain to be generated before it can operate.
[0062] During the task training process, by rotating some of the secret sharing, data privacy can be protected from model inference attacks or membership inference attacks, while reducing the communication overhead of large model training.
[0063] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.
Claims
1. A method for constructing a large-model federated split trusted data space based on rotating partial secret sharing, characterized in that: The following steps are involved: Step 1: Split the large model into multiple model partitions, each model partition is trained by a group of client devices, and the clients are divided into several groups, each of which has a certain number of clients. The client groups serve as candidate client groups waiting for the allocation of model partitions; Step 2: After determining the number of model partitions and candidate client groups, an optimal set of client devices is selected for model partition allocation by solving an integer optimization problem. All clients in each client group are responsible for the training task of the same model partition. Step 3: During the task training process, data privacy protection is performed by rotating partial secret sharing. Each client divides its local model into multiple random parameter blocks and selects a parameter block through a dynamic indexing mechanism. The client divides the selected parameter block into two secret shares, one of which is sent to the next client in the same client group for security aggregation, and the other secret share is securely aggregated with the secret share sent by the previous client.
2. According to the method of claim 1, a large-model federated split trusted data space construction method based on rotating partial secret sharing is characterized in that: The step 1 specifically comprises the following steps: Step 1.1: Partition the model according to the model's computational graph structure, dividing the model into multiple sub-parts, each sub-part corresponding to a model partition; Step 1.2: Adjust the number of model partitions according to the number of client groups participating in task training, so that the number of model partitions is equal to the number of client groups selected in each communication round; Step 1.3: Combine data flow and communication efficiency to balance the computational load of each model partition.
3. According to the method of claim 1, the method is characterized in that: The step 2 specifically includes the following steps: Step 2.1: Define the problem variables, including the total number of model partitions Z, the total number of candidate client groups N; use a binary variable α zn Indicates whether to assign model partition z to client group n; Step 2.2: Define the objective function A zn , which represents the fitness between the model partition and the customer: Among them, D zn represents the transmission time of the model partition on the client, T zn represents the average training time of the model partition on the client group, a and b are weight hyperparameters; Step 2.3: Define the constraint that each model partition can only be assigned to one client group At the same time, all clients in a client group can only be used to train the same model partition If the GPU memory of any client in the group is insufficient to train the model partition, no allocation will be made: α zn =0, in Is client n in the client group i GPU memory, |w Pz | is the number of parameters of the model partition z; Step 2.4: After completing the definition of the problem variables, objective function and constraints, the integer optimization problem is expressed as: Step 2.5: Convert the integer optimization problem into a linear programming problem and solve it using a standard linear programming solver. Step 2.6: When solving a linear programming problem, α zn Takes a value between [0,1]; Step 2.7: In the optimal solution, α zn =1 means that assigning model partition z to client group n is the optimal assignment to maximize the total fitness, α zn =0 when no allocation is made.
4. According to the method of claim 1, the method is characterized in that: The step 3 specifically includes the following steps: Step 3.1: Define the client group to be assigned as cells, each cell is a client set; each cell is randomly divided into K j subgroups, the number of clients in a cell is an integer multiple of the subgroups; Step 3.2: Order is the global initial model, Is client n i The local model at the tth round of communication, is the local model of the jth cell in the tth round of communication, Is a local model The kth parameter block under Is client n i Aggregate secrets on; Step 3.3: After receiving the model partition, the client performs local training to obtain the local model Step 3.4: Each client n i The weighted local model Cut to K j parameter blocks: where θ i For client n i Local Model The aggregation weight of the parameter block is equal to the number of subgroups in this cell. In different communication rounds, the client generates secret shares for different parameter blocks. The dynamic block index is: φk=[(k+t)modK j ]+1, which means that in each round t, for each client n i The subgroup k to which it belongs determines the shared parameter block, where k is the index of the subgroup, t is the current communication round, and K j is the total number of parameter blocks. The mod operation is used when the sum of k+t exceeds K j When the result loops back to 1, the index value is K j Cycle through parameter blocks; Step 3.5: Determine the parameter block to be encrypted After that, client n i Will Split into and and will forwarded to the next client in the same subgroup, where and Representing two origins A random secret parameter block; Step 3.6: When the latter client receives the secret share from the previous client, it will receive the secret and local secrets Aggregate to get in Represents the previous client n l One of the two secret parameter blocks formed by splitting the selected parameter block; Step 3.7: Sum the aggregate secrets of the kth subgroup to And concatenate all subgroup parameter blocks As a complete cell model Then Broadcast to the clients in the client group for the next round of local training.