Anti-collusion privacy protection data training method and data outsourcing training method
By combining fully homomorphic encryption and random mask matrices, the collusion attack problem in privacy-preserving model training is solved, data security and model performance are improved, communication overhead is reduced, and efficient data outsourcing and model training are promoted.
Patent Information
- Application Number
- CN202511027493.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-09-30
AI Technical Summary
Many existing machine learning solutions have problems with incomplete privacy protection and collusion attacks in privacy-preserving model training, especially when integrating data from multiple parties, it is difficult to ensure data security and model performance.
A fully homomorphic encryption algorithm is used to generate public parameters. The client generates a client key pair and deploys a private key. Data is encrypted using a random mask matrix. The server or cloud service aggregates and processes the ciphertext. The client updates the final model parameters and designs a distributed key management mechanism to avoid collusion attacks.
It effectively protects the security of client data, reduces communication overhead, reduces the risk of sensitive data leakage, enhances the credibility of the system, and achieves efficient model training and data outsourcing.
Smart Images

Figure CN120729612A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data security technology, and in particular relates to an anti-collusion privacy protection data training method and a data outsourcing training method. Background Art
[0002] With the rapid development of technology and the widespread application of big data, the "Model as a Service" (MaaS) model based on intelligent data analysis has become a key driver of innovation and development. Leveraging massive amounts of data to train high-precision models not only provides more efficient and accurate decision support, but also offers users a personalized and convenient experience, significantly improving service efficiency and quality. However, reliance on a single source of data often exposes limitations that hinder model performance. Leveraging data from multiple sources for training offers a promising solution for improving model performance.
[0003] Due to the sensitivity of data, data sharing during the model training phase is prohibited. To better facilitate the integration of multi-party data, various existing machine learning solutions can be applied to privacy-preserving model training. However, these existing machine learning solutions still suffer from issues such as incomplete privacy protection and collusion attacks. Summary of the Invention
[0004] In order to solve the above problems existing in the prior art, the present invention provides an anti-collusion privacy protection data training method and a data outsourcing training method.
[0005] The technical problem to be solved by the present invention is achieved through the following technical solutions: In a first aspect, the present invention provides a privacy-preserving data training method for anti-collusion, which is applied to a data processing system including a server and multiple clients. The method includes: The server initializes the parameters of the model to be trained and generates public parameters based on a fully homomorphic encryption algorithm; Each client generates a client key pair based on the public parameters and deploys the client private key in the client key pair to the client enclave on the server; and generates a random mask matrix; The server generates a target public key based on the public key parameters and public parameters in each client key pair; Each client uses the target public key to encrypt the local training data to obtain the initial ciphertext vector; and then encrypts the initial ciphertext vector again based on the random mask matrix to obtain the client ciphertext vector; The server aggregates the ciphertext vectors of each client to obtain an aggregated ciphertext vector; each client enclave processes the aggregated ciphertext vector based on the client private key to obtain an intermediate vector; and updates the parameters of the to-be-trained model based on the intermediate vector.
[0006] In a second aspect, the present invention provides a privacy-preserving data outsourcing training method that is resistant to collusion and is applied to a target data processing system, the target data processing system including a model owner, a cloud server, and multiple clients. The method includes: The model owner initializes the parameters of the model to be trained and deploys the initial model parameters of the model to be trained in the model owner's enclave on the cloud server; public parameters are generated based on a fully homomorphic encryption algorithm; Each client generates a client key pair based on the public parameters, and deploys the client private key in the client key pair to the client enclave on the cloud server; and generates a random mask matrix; The model owner generates the target public key based on the public key parameters and public parameters in each client key pair; Each client uses the target public key to encrypt the local training data to obtain the initial ciphertext vector; and then encrypts the initial ciphertext vector again based on the random mask matrix to obtain the client ciphertext vector; The cloud server aggregates the ciphertext vectors of each client to obtain an aggregated ciphertext vector. The model owner enclave determines the ciphertext pair based on the aggregated ciphertext vector and the initial model parameters. Each client enclave processes the ciphertext pair based on the client's private key to obtain an intermediate vector. The parameters of the to-be-trained model are updated based on the intermediate vector to obtain the target model parameters. The target model parameters are sent to the model owner.
[0007] The present invention provides a collusion-resistant privacy-preserving data training method and a data outsourcing training method. A server initializes parameters of a model to be trained and generates public parameters based on a fully homomorphic encryption algorithm. Each client generates a client key pair based on the public parameters and deploys the client private key in the client key pair in a client enclave on the server. A random mask matrix is generated. The server generates a target public key based on the public key parameters and public parameters in each client key pair. Each client encrypts local training data using the target public key to obtain an initial ciphertext vector. The initial ciphertext vector is then re-encrypted based on the random mask matrix to obtain a client ciphertext vector. The server aggregates the client ciphertext vectors to obtain an aggregated ciphertext vector. Each client enclave processes the aggregated ciphertext vector based on the client private key to obtain an intermediate vector. Parameters of the model to be trained are updated based on the intermediate vector.
[0008] This invention protects client data by designing a random mask matrix, which can avoid possible collusion attacks and ensure data security. By designing a distributed key management mechanism, the need for a trusted third party is eliminated. Even an honest but curious trusted execution environment (TEE) (such as a client enclave) cannot steal aggregated data or model parameters, thus enhancing the credibility of the system. By performing homomorphic double encryption on client data and using the encrypted aggregated data for model training, the confidentiality of each client's sensitive information and data label distribution is protected.
[0009] The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 1 is a schematic structural diagram of a data processing system for a privacy-preserving data training method for anti-collusion provided by an embodiment of the present invention; Figure 2 1 is a flow chart of a privacy-preserving data training method for resisting collusion provided by an embodiment of the present invention; Figure 3 1 is a schematic structural diagram of a target data processing system for a privacy-preserving data outsourcing training method for anti-collusion provided by an embodiment of the present invention; Figure 4 This is a flowchart of an anti-collusion privacy protection data outsourcing training method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0011] The present invention will be further described in detail below with reference to specific examples, but the embodiments of the present invention are not limited thereto.
[0012] The embodiment of the present invention provides a privacy protection data training method for anti-collusion, which is applied to a data processing system, referring to Figure 1 , the data processing system includes a server and multiple clients, see Figure 2 , the method comprises the following steps: S10. The server initializes the parameters of the model to be trained and generates public parameters based on the fully homomorphic encryption algorithm.
[0013] For example, the client is the data holder with the same characteristics required for model training, and hopes to provide data for the model training process without leaking the data, and participate in the encryption system initialization and data security outsourcing process.
[0014] The server is a model holder and trainer with powerful computing resources and Intel SGX-enabled machines. It hopes to train the model without leaking the aggregated model data, obtain high-precision models, and participate in the encryption system initialization, data security outsourcing, and micro-batch model training process.
[0015] The server first publishes the training tasks and incentive mechanism to attract clients to join the training process. Assume that the server holds the initial model parameters of the model to be trained for ,in, Indicates the number of model training times, Indicates the number of data dimensions. The server expands each parameter into an N-dimensional column vector with equal elements.
[0016] Afterwards, the server establishes a multi-key CKKS cryptosystem (MK-CKKS) that supports packaged fully homomorphic encryption and can support calculations between multiple ciphertexts. This cryptosystem supports the following operations: : Given an integer and a safety parameter , generate public parameters , public parameters Random polynomials can be encapsulated in , the maximum power of the polynomial and coefficient modulus , error distribution and key distribution .
[0017] S20. Each client generates a client key pair based on the public parameters, and deploys the client private key in the client key pair in the client enclave on the server; and generates a random mask matrix.
[0018] For example, each client starts from the public parameters Generate the client key pair in ,in, Represents the client private key, and Polynomials in Represents the public key parameters, assuming the client is recorded as , client Request enclave service from the server and obtain the client enclave deployed on the server, which is recorded as ; Client The client private key Send and deploy to the corresponding client enclave Among them, the subscript Indicates the client ID.
[0019] Optionally, in step S20, generating a random mask matrix may specifically include: S201. Each client generates a security seed shared among all clients.
[0020] S202: Every two clients generate a shared random mask matrix based on a security seed and a pseudo-random generator.
[0021] For example, the client can use the Diffie-Hellman algorithm to generate a secure seed shared among all clients. ; Any two clients, such as and , you can use safe seeds and a pseudo-random generator Generate a shared random mask matrix, which can include and Among them, the subscript Representation and subscript Different client numbers, Indicates the data dimension index.
[0022] S30. The server generates a target public key according to the public key parameters and public parameters in each client key pair.
[0023] Optionally, in step S30, the server generates a target public key according to the public key parameters and public parameters in each client key pair, which may specifically include: The server combines the public key parameters in each client key pair and the random polynomial contained in the public parameters to obtain the target public key.
[0024] For example, the CKKS encryption system established by the server also supports the following operations: : The encryption system merges the public key parameters of all client key pairs and public parameters The random polynomial contained in As the target public key, denoted as , and publish it; among them, Indicates the total number of clients.
[0025] S40. Each client uses the target public key to encrypt the local training data to obtain an initial ciphertext vector; and based on the random mask matrix, the initial ciphertext vector is secondary encrypted to obtain a client ciphertext vector.
[0026] Optionally, in step S40, each client uses the target public key to encrypt the local training data to obtain an initial ciphertext vector, which may specifically include: S401. Each client normalizes the local training data to obtain an initial data matrix.
[0027] For example, local training data refers to the data stored locally on each client that can be used to train the model to be trained. The local training data owned by the client can be expressed as ,in The size of , , , Indicates the number of data dimensions.
[0028] Optionally, in step S401, each client performs normalization processing on the local training data to obtain an initial data matrix, which may specifically include: S4011. Each client generates an initial maximum value vector and an initial minimum value vector of each client based on the maximum value and minimum value in the local training data.
[0029] For example, each client calculates the maximum value of each dimension in the local training data and minimum value , and generate the initial maximum value vector corresponding to each client , and generate the initial minimum vector corresponding to each client , sent to the server. Among them, represents local training data, Represents the data dimension index, Indicates the number of data dimensions, Indicates the client ID.
[0030] S4012. The server receives the initial maximum vectors and initial minimum vectors of all clients, and determines the global maximum vector and global minimum vector therefrom; adds random perturbations to the global maximum vector and global minimum vector to obtain the perturbed maximum vector and perturbated minimum vector.
[0031] For example, the server receives the initial maximum value vector and the initial minimum value vector of all clients and calculates the global maximum value (denoted as ) vector and the global minimum (denoted as ) vector, and add random perturbations to the global maximum vector and the global minimum vector to obtain the perturbation maximum vector and the perturbation minimum vector, and send them back to each client. Among them, random perturbations (i.e. random vectors) can be achieved by generating small random numbers S4013. Each client normalizes the local training data based on the maximum perturbation vector and the minimum perturbation vector to obtain an initial data matrix.
[0032] For example, each client Based on the normalized local training data of the maximum perturbation vector and the minimum perturbation vector, the client The corresponding matrix , which can be specifically expressed as:
[0033] ; Among them, the subscript Indicates the The data and size of each client, subscript Indicates the number of data dimensions, , the superscript 0 is used as the marker of the initial matrix, Both indicate client numbers.
[0034] Afterwards, each client In the corresponding matrix Then, the matrix The row vector of Multiply each element by , construct the initial data matrix , and then the matrix The row vector of Multiply each element in the local training dataset by , and insert at the beginning of the row vector Get the matrix ,in,
[0035] ,
[0036] ; Among them, the superscript Indicates use Middle Matrix constructed from column vectors, subscript Representation matrix Column vector index, subscript Representation matrix column vector indexing, S402: Encrypt the initial data matrix using the target public key to obtain an initial ciphertext vector.
[0037] For example, each client For the initial data matrix and Each column vector in the matrix and Use the target public key Encrypt to get the initial ciphertext vector. The initial ciphertext vector can include the corresponding initial data matrix The first initial ciphertext vector, and the corresponding initial data matrix The second initial ciphertext vector.
[0038] Optionally, in step S40, the initial ciphertext vector is encrypted twice based on the random mask matrix to obtain the client ciphertext vector, which may specifically include: S403 : For each client, accumulate the random mask matrix shared between the first client whose client number is smaller than the current client and the current client to obtain a first accumulated mask matrix.
[0039] For example, each client is referred to as the local client, and the local client number of the client is assumed to be , the client numbers of other clients are recorded as , determine the client number of other clients With this client number The size of Less than The client is the first client, and the random mask matrix shared by the first client and this client is and Accumulate and get the first accumulation mask matrix, For example, it can be recorded as .
[0040] S404: Accumulate the random mask matrix corresponding to the second client whose client number is greater than that of the current client and the random mask matrix corresponding to the current client to obtain a second accumulated mask matrix.
[0041] For example, Greater than The client is used as the second client, and the random mask matrix shared by the second client and this client is and Accumulate and get the second accumulation mask matrix, For example, it can be recorded as .
[0042] S405 : Calculate the difference between the first cumulative mask matrix and the second cumulative mask matrix to obtain a target mask matrix.
[0043] For example, the target mask matrix can be obtained by subtracting the first cumulative mask matrix from the second cumulative mask matrix, which are respectively expressed as ,as well as .
[0044] S406: Use the target mask matrix to perform secondary encryption on the initial ciphertext vector to obtain the client ciphertext vector.
[0045] For example, each client can use the target mask matrix Column vector in The column vector of the first initial ciphertext vector in the initial ciphertext vector (that is, the column vector after encryption with the target public key) ) are added to form the target mask matrix Column vector in and the column vector of the second initial ciphertext vector in the initial ciphertext vector (that is, the column vector after encryption with the target public key ) are added together to obtain the client ciphertext vector corresponding to each client and Afterwards, each client will send the client ciphertext vector and Sent to the server.
[0046] S50: The server aggregates the ciphertext vectors of each client to obtain an aggregated ciphertext vector; each client enclave processes the aggregated ciphertext vector based on the client private key to obtain an intermediate vector; and updates the parameters of the to-be-trained model based on the intermediate vector.
[0047] For example, the server sends the client ciphertext vector of each client to and Perform aggregation to obtain the aggregated ciphertext vector, which can be recorded as and , thus eliminating all random masks.
[0048] Optionally, in step S50, each client enclave on the server processes the client ciphertext vector based on the client private key to obtain an intermediate vector, which may specifically include: S501. The server performs inner product calculation on the aggregated ciphertext vector and the initial model parameters of the model to be trained to obtain inner product data.
[0049] For example, the server calculates the inner product of the aggregated ciphertext vector and the initial model parameters based on the homomorphic property of the ciphertext to obtain the inner product data. The specific calculation formula can be expressed as: ; in, represents the inner product data, Indicates the In the training round The model parameter state of a mini-batch training sample set, middle Indicates the index of the current model training round number, represents the index of the mini-batch training sample set in each round of training, Represents the corresponding feature dimension vector, Indicates the number of data dimensions. Here, the training data used in each round of training is all mini-batch training sample sets of all clients.
[0050] S502: Perform logistic regression calculation on the inner product data to obtain a ciphertext pair.
[0051] For example, the server performs a logistic regression calculation on the inner product data based on the homomorphic property of the ciphertext to obtain a ciphertext pair, which can be expressed as: ; in, represents a ciphertext pair, Represents the corresponding feature dimension The CKKS ciphertext in , Represents the corresponding feature dimension The CKKS ciphertext in , S503. Each client enclave processes the ciphertext pair based on the client private key to obtain an intermediate vector.
[0052] For example, each client is deployed in a client enclave on the server Use ciphertext pairs and each client's private key The intermediate vector corresponding to each client is calculated and can be expressed as: ; in, Indicates the intermediate vector corresponding to each client.
[0053] In this embodiment, by designing an effective encoding strategy and polynomial multiplication strategy under a single cloud, the computational overhead caused by the fast Fourier transform in the inner product calculation is avoided, the inner product multiplication of homomorphic encryption is accelerated, and the communication and computational overhead is significantly reduced.
[0054] Optionally, in step S50, updating the parameters of the to-be-trained model based on the intermediate vector may specifically include: S504. Each client enclave divides the intermediate vectors into multiple batches of training sample sets, and calculates the sum of the intermediate vectors contained in each batch of training sample sets to obtain sub-training data for each client in each batch. S505: The server aggregates the sub-training data of each client in each batch to obtain target training data for each batch; S506 , updating the parameters of the to-be-trained model based on the target training data of each batch.
[0055] For example, each client enclave Use a shared random number generator to generate multiple batches of training sample sets from intermediate vectors. The training sample set is a mini-batch training sample set that can contain intermediate vectors, and the sum of the intermediate vectors contained in the training sample set of each batch can be expressed as , get the sub-training data of each client in each batch and return it to the server. The number of samples in each training sample set can also be returned here The server receives data from each client enclave. The sub-training data of each batch returned is aggregated to obtain the target training data of each batch, which can be expressed as: ; in, Represents the target training data for each batch.
[0056] Here, we can also set the number of samples in each training sample set. Aggregate to get the total number of samples of target training data .
[0057] Based on the target training data of each batch and the total number of samples in each batch, the model to be trained is trained in batches, and the mini-batch stochastic gradient descent (MSGD) technique is used to update the parameters of the model to be trained.
[0058] It should be noted that after the to-be-trained model is updated using a single batch of target training data, the updated model parameters are obtained. Then, step S501 is returned to perform inner product calculation on the aggregated ciphertext vector and the updated model parameters to obtain new inner product data, and then a new intermediate vector is obtained. The parameters of the to-be-trained model are updated based on the new intermediate vector. The above process is iterated until the training is completed to obtain the target model parameters.
[0059] The present invention provides a collusion-resistant privacy-preserving data training method, which is applied to a data processing system. The data processing system includes a server and multiple clients. The server initializes parameters of a model to be trained and generates public parameters based on a fully homomorphic encryption algorithm. Each client generates a client key pair based on the public parameters and deploys the client private key in the client key pair in a client enclave on the server. A random mask matrix is generated. The server generates a target public key based on the public key parameters and public parameters in each client key pair. Each client encrypts local training data using the target public key to obtain an initial ciphertext vector. The initial ciphertext vector is then re-encrypted based on the random mask matrix to obtain a client ciphertext vector. The server aggregates the client ciphertext vectors to obtain an aggregated ciphertext vector. Each client enclave processes the aggregated ciphertext vector based on the client private key to obtain an intermediate vector. Parameters of the model to be trained are updated based on the intermediate vector.
[0060] This invention protects client data by designing a random mask matrix, which can avoid possible collusion attacks and ensure data security. By designing a distributed key management mechanism, the need for a trusted third party is eliminated. Even an honest but curious trusted execution environment (TEE) (such as a client enclave) cannot steal aggregated data or model parameters, thus enhancing the credibility of the system. By performing homomorphic double encryption on client data and using the encrypted aggregated data for model training, the confidentiality of each client's sensitive information and data label distribution is protected.
[0061] In summary, the present invention achieves the purpose of resisting collusion attacks, reduces communication overhead and alleviates the risk of sensitive data leakage, thereby achieving efficient data outsourcing and further promoting model training efficiency.
[0062] Corresponding to the above-mentioned anti-collusion privacy protection data training method, in order to expand the application scenario of the present invention, the embodiment of the present invention also provides an anti-collusion privacy protection data outsourcing training method, which is applied to the target data processing system, referring to Figure 3 ,The target data processing system includes the model owner, the cloud server, and multiple clients, e.g. Figure 4 As shown, the method may include: S110. The model owner initializes the parameters of the model to be trained and deploys the initial model parameters of the model to be trained in the model owner's enclave on the cloud server; and generates public parameters based on a fully homomorphic encryption algorithm.
[0063] For example, this embodiment is suitable for outsourcing model training while ensuring the security of training data and model parameters. This embodiment introduces the role of the model owner, who, lacking computing resources, is responsible for publishing training tasks and public parameters to clients. It also introduces a cloud service provider (CSP), who assumes the training responsibilities previously performed by the server. The CSP does not need to access the training data or model itself to train the model, thus providing privacy protection for the model parameters.
[0064] The client is the data holder with the same characteristics as those required for model training. They hope to provide data for model training without leaking the data, and participate in the encryption system initialization and data security outsourcing process. Model owners are model holders and training task publishers who lack computing resources. They hope to train models in outsourced scenarios without leaking model parameters or aggregated model data, obtain high-precision models, and participate in the initialization of encryption systems, data security outsourcing, and micro-batch model training processes. CSP is a model trainer with powerful computing capabilities and support for Intel SGX. It can provide computing support when the computing tasks are large, and participate in the encryption system initialization, data security outsourcing and micro-batch model training process.
[0065] It should be noted that the model owner requests the enclave service from the CSP and obtains the model owner enclave deployed by the CSP. After the remote attestation verification is passed, the model owner outsources the initial model parameters to the model owner enclave through a secure channel. .
[0066] S120. Each client generates a client key pair based on the public parameters, and deploys the client private key in the client key pair in the client enclave on the cloud server; and generates a random mask matrix.
[0067] Optionally, generate a random mask matrix, including: Each client generates a secure seed shared among all clients; Every two clients generate a shared random mask matrix based on a secure seed and a pseudo-random generator. S130: The model owner generates a target public key based on the public key parameters and public parameters in each client key pair.
[0068] Optionally, the model owner generates a target public key based on the public key parameters and public parameters in each client key pair, including: the model owner combines the public key parameters in each client key pair and the random polynomial contained in the public parameters to obtain the target public key.
[0069] S140. Each client encrypts the local training data using the target public key to obtain an initial ciphertext vector; and performs secondary encryption on the initial ciphertext vector based on a random mask matrix to obtain a client ciphertext vector.
[0070] Optionally, each client uses the target public key to encrypt local training data to obtain an initial ciphertext vector, including: each client normalizes the local training data to obtain an initial data matrix; and uses the target public key to encrypt the initial data matrix to obtain an initial ciphertext vector.
[0071] Optionally, each client normalizes the local training data to obtain an initial data matrix, including: each client generates an initial maximum value vector and an initial minimum value vector of each client based on the maximum value and minimum value in the local training data; the server receives the initial maximum value vectors and initial minimum value vectors of all clients, and determines the global maximum value vector and the global minimum value vector therefrom; random perturbations are added to the global maximum value vector and the global minimum value vector to obtain a perturbation maximum value vector and a perturbation minimum value vector; each client normalizes the local training data based on the perturbation maximum value vector and the perturbation minimum value vector to obtain an initial data matrix. Optionally, based on the random mask matrix, the initial ciphertext vector is encrypted twice to obtain a client ciphertext vector, including: for each client, accumulating a random mask matrix shared between a first client whose client number is smaller than the current client and the current client to obtain a first accumulated mask matrix; accumulating a random mask matrix shared between a second client whose client number is larger than the current client and the current client to obtain a second accumulated mask matrix; calculating the difference between the first accumulated mask matrix and the second accumulated mask matrix to obtain a target mask matrix; and using the target mask matrix to encrypt the initial ciphertext vector twice to obtain a client ciphertext vector.
[0072] S150: The cloud server aggregates the ciphertext vectors of each client to obtain an aggregated ciphertext vector. The model owner enclave determines a ciphertext pair based on the aggregated ciphertext vector and the initial model parameters. Each client enclave processes the ciphertext pair based on the client private key to obtain an intermediate vector. The parameters of the to-be-trained model are updated based on the intermediate vector to obtain target model parameters. The target model parameters are sent to the model owner.
[0073] Optionally, the model owner enclave determines a ciphertext pair based on the aggregated ciphertext vector and the initial model parameters, including: the model owner enclave performs an inner product calculation on the aggregated ciphertext vector and the initial model parameters of the model to be trained to obtain inner product data; and performs a logistic regression calculation on the inner product data to obtain a ciphertext pair.
[0074] For example, after obtaining the ciphertext pair Afterwards, the ciphertext can be Re-encrypt, in the ciphertext A random vector is added to prevent linear attacks.
[0075] Optionally, updating the parameters of the to-be-trained model based on the intermediate vectors includes: each client enclave divides the intermediate vectors into multiple batches of training sample sets, and calculates the sum of the intermediate vectors contained in the training sample sets of each batch to obtain sub-training data of each client in each batch; the cloud server aggregates the sub-training data of each client in each batch to obtain target training data for each batch; and updating the parameters of the to-be-trained model based on the target training data of each batch.
[0076] Illustratively, after completing the parameter update of the model to be trained in this embodiment, the CSP returns the trained target model parameters to the model owner through a secure channel.
[0077] For the specific implementation process of each step in this embodiment, please refer to the specific implementation process of the corresponding steps in the anti-collusion privacy protection data training method provided in the first aspect, which will not be repeated here.
[0078] The present invention provides a collusion-resistant privacy-preserving data outsourcing training method, which is applied to a data processing system. The data processing system includes a server and multiple clients. The server initializes parameters of a model to be trained and generates public parameters based on a fully homomorphic encryption algorithm. Each client generates a client key pair based on the public parameters and deploys the client private key in the client key pair in a client enclave on the server. A random mask matrix is generated. The server generates a target public key based on the public key parameters and public parameters in each client key pair. Each client encrypts local training data using the target public key to obtain an initial ciphertext vector. The initial ciphertext vector is then re-encrypted based on the random mask matrix to obtain a client ciphertext vector. The server aggregates the client ciphertext vectors to obtain an aggregated ciphertext vector. Each client enclave processes the aggregated ciphertext vector based on the client private key to obtain an intermediate vector. Parameters of the model to be trained are updated based on the intermediate vector.
[0079] In summary, the present invention achieves the purpose of resisting collusion attacks, reduces communication overhead and mitigates the risk of sensitive data leakage, thereby realizing efficient data and model training outsourcing, and further promoting training efficiency.
[0080] It should be noted that the terms "first," "second," and the like are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in sequences other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention.
[0081] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification.
[0082] Although the present invention is described herein in conjunction with various embodiments, in the process of implementing the claimed invention, those skilled in the art can understand and implement other variations of the disclosed embodiments by viewing the drawings and the disclosed content. In the description of the present invention, the word "comprising" does not exclude other components or steps, "one" or "a" does not exclude multiple situations, and "multiple" means two or more, unless otherwise clearly and specifically defined. In addition, certain measures are recorded in different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0083] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.
Claims
1. A privacy-preserving data training method that is resistant to collusion, characterized in that: Applied to a data processing system, the data processing system includes a server and multiple clients, and the method includes: The server initializes parameters of the model to be trained and generates public parameters based on a fully homomorphic encryption algorithm; Each of the clients generates a client key pair based on the public parameters, and deploys the client private key in the client key pair in a client enclave on the server; and generates a random mask matrix; The server generates a target public key according to the public key parameters in each of the client key pairs and the public parameters; Each client encrypts the local training data using the target public key to obtain an initial ciphertext vector; and performs secondary encryption on the initial ciphertext vector based on the random mask matrix to obtain a client ciphertext vector; The server aggregates the ciphertext vectors of each client to obtain an aggregated ciphertext vector; each client enclave processes the aggregated ciphertext vector based on the client private key to obtain an intermediate vector; and updates the parameters of the model to be trained based on the intermediate vector.
2. The anti-collusion privacy protection data training method according to claim 1, characterized in that: Each of the client enclaves on the server processes the client ciphertext vector based on the client private key to obtain an intermediate vector, including: The server performs an inner product calculation on the aggregated ciphertext vector and the initial model parameters of the to-be-trained model to obtain inner product data; Performing a logistic regression calculation on the inner product data to obtain a ciphertext pair; Each of the client enclaves processes the ciphertext pair based on the client private key to obtain an intermediate vector.
3. The anti-collusion privacy protection data training method according to claim 2, characterized in that: Each of the clients encrypts the local training data using the target public key to obtain an initial ciphertext vector, including: Each client performs normalization processing on the local training data to obtain an initial data matrix; The initial data matrix is encrypted using the target public key to obtain an initial ciphertext vector.
4. The anti-collusion privacy protection data training method according to claim 3, characterized in that: The generating of the random mask matrix comprises: Each of the clients generates a security seed shared among all the clients; Every two clients generate a shared random mask matrix based on the security seed and a pseudo-random generator.
5. The anti-collusion privacy protection data training method according to claim 4, characterized in that: The method further comprises: performing secondary encryption on the initial ciphertext vector based on the random mask matrix to obtain a client ciphertext vector; For each of the clients, accumulating the random mask matrix shared between the first client whose client number is smaller than the client's and the client to obtain a first accumulated mask matrix; Accumulate the random mask matrix shared between the second client whose client number is greater than that of the current client and the current client to obtain a second accumulated mask matrix; Calculating a difference between the first cumulative mask matrix and the second cumulative mask matrix to obtain a target mask matrix; The initial ciphertext vector is encrypted twice using the target mask matrix to obtain a client ciphertext vector.
6. The anti-collusion privacy protection data training method according to claim 1, characterized in that: The server generates a target public key according to the public key parameters in each of the client key pairs and the public parameters, including: The server combines the public key parameters in each of the client key pairs and the random polynomial included in the public parameters to obtain a target public key.
7. The anti-collusion privacy protection data training method according to claim 3, characterized in that: Each of the clients normalizes the local training data to obtain an initial data matrix, including: Each of the clients generates an initial maximum value vector and an initial minimum value vector of each client based on the maximum value and the minimum value of each dimension in the local training data; The server receives the initial maximum value vectors and initial minimum value vectors of all clients, and determines a global maximum value vector and a global minimum value vector therefrom; adds random perturbations to the global maximum value vector and the global minimum value vector to obtain a perturbation maximum value vector and a perturbation minimum value vector; Each of the clients performs normalization processing on the local training data based on the maximum disturbance vector and the minimum disturbance vector to obtain an initial data matrix.
8. The anti-collusion privacy protection data training method according to claim 2, characterized in that: The updating of parameters of the to-be-trained model based on the intermediate vector includes: Each of the client enclaves divides the intermediate vectors into multiple batches of training sample sets, and calculates the sum of the intermediate vectors contained in each batch of training sample sets to obtain sub-training data of each of the clients in each batch; The server aggregates the sub-training data of each client in each batch to obtain target training data for each batch; Parameters of the to-be-trained model are updated based on the target training data of each batch.
9. A privacy-preserving data outsourcing training method that is resistant to collusion, characterized in that: Applied to a target data processing system comprising a model owner, a cloud server, and multiple clients, the method comprises: The model owner initializes the parameters of the model to be trained and deploys the initial model parameters of the model to be trained in the model owner enclave on the cloud server; generates public parameters based on a fully homomorphic encryption algorithm; Each of the clients generates a client key pair based on the public parameters, and deploys the client private key in the client key pair in a client enclave on the cloud server; and generates a random mask matrix; The model owner generates a target public key according to the public key parameters in each of the client key pairs and the public parameters; Each client encrypts the local training data using the target public key to obtain an initial ciphertext vector; and performs secondary encryption on the initial ciphertext vector based on the random mask matrix to obtain a client ciphertext vector; The cloud server aggregates the ciphertext vectors of each client to obtain an aggregated ciphertext vector; the model owner enclave determines a ciphertext pair based on the aggregated ciphertext vector and the initial model parameters; each client enclave processes the ciphertext pair based on the client private key to obtain an intermediate vector; based on the intermediate vector, the parameters of the model to be trained are updated to obtain target model parameters; and the target model parameters are sent to the model owner.