Privacy protection-based linear regression method for horizontally dividing data
Through encryption technology, the original symmetric matrix and product vector in horizontal federated learning can be solved, and the problem that model updates may leak data privacy is realized, and the secure transmission and operation of data in the federated learning process is realized.
Patent Information
- Application Number
- CN202510350150.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-24
AI Technical Summary
In horizontal federated learning, although only model updates are shared rather than raw data, an attacker may infer information about training data through analyzing model updates, threatening the privacy of user data.
A linear regression method is proposed for horizontally dividing data and based on privacy protection. The original symmetry matrix and product vector are encrypted into ciphertext through encryption technology, and encryption operations are performed on the server side to ensure the security of data during transmission and operation.
It effectively protects the privacy of private data, prevents attackers from inferring original data information through model updates, and ensures the security of data during federated learning.
Smart Images

Figure CN120217320A_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, the field of machine learning technology, and particularly relates to a method and apparatus for horizontally partitioning data and performing linear regression based on privacy protection. Background Art
[0002] As the public and policymakers become increasingly aware of the importance of privacy, the demand for privacy-preserving machine learning in data practices is also on the rise. Access to data is subject to increasing scrutiny, and research on privacy-respecting tools such as horizontal federated learning is becoming increasingly active. Ideally, horizontal federated learning can enable cooperation among data stakeholders while protecting the privacy of individuals and institutions.
[0003] In horizontal federated learning, even though horizontal federated learning only shares model updates rather than raw data, attackers may still infer information about the training data by analyzing these updates, and this inference attack threatens the privacy of user data. Summary of the Invention
[0004] The following is an overview of the subject matter described in detail in this document. This overview is not intended to limit the scope of protection of the claims.
[0005] Embodiments of this application provide a method for horizontally partitioning data and performing linear regression based on privacy protection, which ensures the security of private data during data transmission and operation.
[0006] To achieve the above object, a first aspect of the embodiments of the present application proposes a linear regression method for horizontally partitioning data and based on privacy protection, including: each of at least two data owners respectively uses a number of its own private data samples to obtain ciphertexts related to their respective original symmetric matrices and ciphertexts of original product vectors, where the original symmetric matrix and the original product vector are respectively the symmetric matrix and the product vector corresponding to the number of the data samples; each of the data owners sends the ciphertexts related to their respective original symmetric matrices and the ciphertexts of the original product vectors to a first server, and the first server, after receiving the ciphertexts related to the original symmetric matrices and the ciphertexts of the original product vectors sent by each of at least two data owners, cooperates with a second server to obtain target linear regression coefficients. The linear regression method based on privacy protection includes: the first server receives the ciphertexts related to the original symmetric matrices and the ciphertexts of the original product vectors sent by each of at least two data owners, and obtains a processed result of a pseudo-sample symmetric matrix and a processed result of a pseudo-sample product vector, where the processed result of the pseudo-sample symmetric matrix and the processed result of the pseudo-sample product vector are respectively the results obtained by processing a pseudo-sample symmetric matrix and a pseudo-sample product vector, and the pseudo-sample symmetric matrix and the pseudo-sample product vector are respectively the symmetric matrix and the product vector corresponding to a number of pseudo-data samples; the first server uses the ciphertexts related to the original symmetric matrices and the processed result of the pseudo-sample symmetric matrix to obtain the sum of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix, and uses the ciphertexts of the original product vectors and the processed result of the pseudo-sample product vector to obtain the sum of the ciphertexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector; the first server sends the sum of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector to the second server; the second server receives and decrypts respectively the sum of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector to obtain the sum of the plaintexts of the original symmetric matrices of all data owners including the influence of the pseudo-sample symmetric matrix and the sum of the plaintexts of the original product vectors of all data owners including the influence of the pseudo-sample product vector;The second server obtains the pseudo-sample encrypted linear regression coefficients including the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector by using the sum of the original symmetric matrix plaintexts of all data owners including the influence of the pseudo-sample symmetric matrix and the sum of the original product vector plaintexts of all data owners including the influence of the pseudo-sample product vector, and sends the pseudo-sample encrypted linear regression coefficients to the first server. Alternatively, the second server performs regularization processing on the sum of the original symmetric matrix plaintexts of all data owners including the influence of the pseudo-sample symmetric matrix, and obtains the pseudo-sample encrypted linear regression coefficients including the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector according to the sum of the original product vector plaintexts of all data owners including the influence of the pseudo-sample product vector and the sum of the original symmetric matrix plaintexts of all data owners after regularization processing, and sends the pseudo-sample encrypted linear regression coefficients to the first server. The first server receives the pseudo-sample encrypted linear regression coefficients including the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector, and at least partially eliminates the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector from the pseudo-sample encrypted linear regression coefficients to obtain the target linear regression coefficients.
[0007] In some embodiments, each of the several private data samples of its own includes an original input feature column vector and a corresponding original label; each of the several pseudo-data samples includes a pseudo-input feature column vector and a corresponding pseudo-label; the original symmetric matrix of any one of the data owners is equal to the sum of the products of each of the original input feature column vectors of the data owner and its own transpose, and the original product vector of any one of the data owners is equal to the sum of the products of each of the original labels of the data owner and the corresponding original input feature column vectors; the pseudo-sample symmetric matrix is the sum of the products of the several pseudo-input feature column vectors included in the several pseudo-data samples and their own transposes, and the pseudo-sample product vector is the sum of the products of the several pseudo-input feature column vectors included in the several pseudo-data samples and the corresponding pseudo-labels; the ciphertext related to the original symmetric matrix and the ciphertext of the original product vector are obtained by the data owner using an encryption technique that supports addition calculation on ciphertext to encrypt the plaintext related to the original symmetric matrix and the original product vector respectively, and the encryption technique that supports addition calculation on ciphertext satisfies: the plaintext is respectively encrypted by the encryption technique that supports addition calculation on ciphertext to obtain ciphertext, the sum of the ciphertext is obtained through the arithmetic addition operation between the ciphertexts, and the sum of the ciphertexts corresponds to the arithmetic sum of multiple plaintexts after decryption; the first server uses the ciphertext related to the original symmetric matrix and the processing result of the pseudo-sample symmetric matrix to obtain the sum of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix, specifically by using the technique of addition calculation on ciphertext included in the encryption technique that supports addition calculation on ciphertext to accumulate the ciphertexts of the original symmetric matrices of multiple data owners and the processing result of the pseudo-sample symmetric matrix; the specific process of the first server using the ciphertext of the original product vector and the processing result of the pseudo-sample product vector to obtain the sum of the ciphertexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector is to use the technique of addition calculation on ciphertext included in the encryption technique that supports addition calculation on ciphertext to accumulate the ciphertexts of the original product vectors of multiple data owners and the processing result of the pseudo-sample product vector.
[0008] In some embodiments, the first server receives the encrypted linear regression coefficients of the pseudo-samples including the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector, and at least partially eliminates the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector from the encrypted linear regression coefficients of the pseudo-samples to obtain the target linear regression coefficients, including: the first server encrypts the pseudo-input matrix composed of a plurality of pseudo-input feature column vectors included in the plurality of pseudo-data samples to obtain the ciphertext of the pseudo-input matrix, and the first server sends the ciphertext of the pseudo-input matrix to the second server; the second server calculates the inverse matrix of the sum of the original symmetric matrix plaintexts of all data owners including the influence of the pseudo-sample symmetric matrix, that is, the first inverse matrix, and the second server obtains the product of the ciphertext of the pseudo-input matrix and the first inverse matrix based on the encryption method of the pseudo-input matrix by the first server to obtain the product of the ciphertext of the pseudo-input matrix and the first inverse matrix, and the second server sends the product of the ciphertext of the pseudo-input matrix and the first inverse matrix to the first server; the first server decrypts the product of the ciphertext of the pseudo-input matrix and the first inverse matrix based on its own encryption method of the pseudo-input matrix to obtain the product of the plaintext of the pseudo-input matrix and the first inverse matrix, and the first server uses the product of the plaintext of the pseudo-input matrix and the first inverse matrix to at least partially eliminate the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector from the encrypted linear regression coefficients of the pseudo-samples including the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector to obtain the target linear regression coefficients.
[0009] In some embodiments, each of at least two of the data owners uses the same encryption technology to encrypt the intermediate variables related to a plurality of its own private data samples respectively to obtain the ciphertext related to the original symmetric matrix and the ciphertext of the original product vector of each; the sum of the encrypted texts of the original symmetric matrices of the multiple data owners including the influence of the pseudo-sample symmetric matrix obtained by the first server is the sum of the encrypted texts of the original symmetric matrices of all the data owners including the influence of the pseudo-sample symmetric matrix, and the sum of the encrypted texts of the original product vectors of the multiple data owners including the influence of the pseudo-sample product vector obtained by the first server is the sum of the encrypted texts of the original product vectors of all the data owners including the influence of the pseudo-sample product vector.
[0010] In some embodiments, the at least two data owners respectively use at least two different encryption techniques to encrypt the intermediate variables related to several private data samples of their own, obtaining the ciphertexts related to their respective original symmetric matrices and the ciphertexts of the original product vectors; the sum of the ciphertexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix obtained by the first server is the result obtained by adding the influence of the pseudo-sample symmetric matrix to the sum of the ciphertexts of the original symmetric matrices of multiple data owners using the same encryption technique; the sum of the ciphertexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector obtained by the first server is the result obtained by adding the influence of the pseudo-sample product vector to the sum of the ciphertexts of the original product vectors of multiple data owners using the same encryption technique; the first server sends the sum of the ciphertexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector to the second server; the second server receives the sum of the ciphertexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector, which are respectively related to different encryption methods, and uses different decryption techniques corresponding to various encryption techniques to decrypt the sum of the ciphertexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector corresponding to at least two different encryption techniques, so as to obtain the sum of the plaintexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix and the sum of the plaintexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector; the second server accumulates the sum of the plaintexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix to obtain the sum of the plaintexts of the original symmetric matrices of all data owners containing the influence of the pseudo-sample symmetric matrix, and accumulates the sum of the plaintexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector to obtain the sum of the plaintexts of the original product vectors of all data owners containing the influence of the pseudo-sample product vector.
[0011] In some embodiments, the at least two data owners use at least two different encryption techniques including a first encryption technique to encrypt intermediate variables related to a plurality of their respective private data samples, obtaining ciphertexts related to their respective original symmetric matrices and ciphertexts of original product vectors; the first server further encrypts the ciphertext of the original symmetric matrix encrypted with other encryption techniques or the corresponding accumulated result using the first encryption technique, and further encrypts the ciphertext of the original product vector encrypted with the other encryption techniques or the corresponding accumulated result, where the other encryption technique is an encryption technique other than the first encryption technique among the at least two different encryption techniques; the first server then uses the result of the further encryption to obtain the sum of the ciphertexts of the original symmetric matrices of all data owners encrypted with the first encryption technique and including the influence of the pseudo-sample symmetric matrix, and obtains the sum of the ciphertexts of the original product vectors of all data owners encrypted with the first encryption technique and including the influence of the pseudo-sample product vector; the first server sends the sum of the ciphertexts of the original symmetric matrices of all data owners encrypted with the first encryption technique and including the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of all data owners encrypted with the first encryption technique and including the influence of the pseudo-sample product vector to the second server; the second server receives and decrypts the sum of the ciphertexts of the original symmetric matrices of all data owners encrypted with the first encryption technique and including the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of all data owners encrypted with the first encryption technique and including the influence of the pseudo-sample product vector respectively using the decryption technique corresponding to the first encryption technique, and then further decrypts using the decryption technique corresponding to the other encryption technique to obtain the sum of the plaintexts of the original symmetric matrices of all data owners including the influence of the pseudo-sample symmetric matrix and the sum of the plaintexts of the original product vectors of all data owners including the influence of the pseudo-sample product vector.
[0012] In some embodiments, each of at least two of the data owners encrypts using the same encryption technique, which is a homomorphic encryption technique that supports addition operations on ciphertexts and encrypts using a first key, where the first key is the public key used in the homomorphic encryption technique; the ciphertext related to the original symmetric matrix is the ciphertext of the original symmetric matrix; the processed result of the pseudo-sample symmetric matrix is the ciphertext obtained by homomorphically encrypting the pseudo-sample symmetric matrix or the negative value of the pseudo-sample symmetric matrix using the first key, and the processed result of the pseudo-sample product vector is the ciphertext obtained by homomorphically encrypting the pseudo-sample product vector or the negative value of the pseudo-sample product vector using the first key; the sum of the ciphertexts of the original symmetric matrices of all data owners affected by the pseudo-sample symmetric matrix is the result obtained by using the addition operation on ciphertexts included in the homomorphic encryption technique to accumulate the processed result of the pseudo-sample symmetric matrix and the ciphertexts of the original symmetric matrices of all data owners, and performing optional regularization processing; the sum of the ciphertexts of the original product vectors of all data owners affected by the pseudo-sample product vector is the result obtained by using the addition operation on ciphertexts included in the homomorphic encryption technique to accumulate the processed result of the pseudo-sample product vector and the ciphertexts of the original product vectors of all data owners; the second server decrypts the sum of the ciphertexts of the original symmetric matrices of all data owners affected by the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of all data owners affected by the pseudo-sample product vector, using the decryption technique corresponding to the homomorphic encryption technique and decrypting using a second key, where the second key is the private key used in the homomorphic encryption technique.
[0013] In some embodiments, each of at least two of the data owners encrypts using the same encryption technique, which is to encrypt by masking the plaintext; the ciphertext of the original product vector of any one of the data owners is the sum of the original product vector of the data owner and the masked product vector of the data owner; the ciphertext related to the original symmetric matrix satisfies that the ciphertext of the original symmetric matrix can be calculated using the ciphertext related to the original symmetric matrix, and the ciphertext of the original symmetric matrix of any one of the data owners is the sum of the original symmetric matrix of the data owner and the masked symmetric matrix of the data owner; the processed result of the pseudo-sample symmetric matrix is the pseudo-sample symmetric matrix or the negative value of the pseudo-sample symmetric matrix, and the processed result of the pseudo-sample product vector is the pseudo-sample product vector or the negative value of the pseudo-sample product vector; the sum of the ciphertexts of the original symmetric matrices of all data owners affected by the pseudo-sample symmetric matrix is the result obtained by accumulating the processed result of the pseudo-sample symmetric matrix and the ciphertexts of the original symmetric matrices of all data owners, and performing an optional regularization process; the sum of the ciphertexts of the original product vectors of all data owners affected by the pseudo-sample product vector is the result obtained by accumulating the processed result of the pseudo-sample product vector and the ciphertexts of the original product vectors of all data owners; the second server decrypts the sum of the ciphertexts of the original symmetric matrices of all data owners affected by the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of all data owners affected by the pseudo-sample product vector, which is that the second server subtracts the sum of the masked symmetric matrices of all data owners from the sum of the ciphertexts of the original symmetric matrices of all data owners affected by the pseudo-sample symmetric matrix, and subtracts the sum of the masked product vectors of all data owners from the sum of the ciphertexts of the original product vectors of all data owners affected by the pseudo-sample product vector, to obtain the sum of the plaintexts of the original symmetric matrices of all data owners affected by the pseudo-sample symmetric matrix and the sum of the plaintexts of the original product vectors of all data owners affected by the pseudo-sample product vector.
[0014] In some embodiments, the masked symmetric matrix and the masked product vector of any one of the data owners are respectively the symmetric matrix and the product vector corresponding to a plurality of masked data samples of the data owner, wherein each of the plurality of masked data samples of any one of the data owners includes a masked input feature column vector and a corresponding masked label; the masked symmetric matrix of any one of the data owners is the cumulative result of the product of each masked input feature column vector and its transpose among the plurality of masked data samples of the data owner, and the masked product vector of any one of the data owners is the cumulative result of the product of each masked input feature column vector and the corresponding masked label among the plurality of masked data samples of the data owner; the ciphertext related to the original symmetric matrix of any one of the data owners is the ciphertext of the original symmetric matrix or the product of a splicing matrix and any orthogonal matrix, wherein the splicing matrix is composed of each original input feature column vector and each masked input feature column vector of the data owner.
[0015] To achieve the above object, a second aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the linear regression method according to any one of claims 1 to 9.
[0016] The embodiments of the present application at least include the following beneficial effects: When the data owner encrypts the private data it owns and sends the ciphertext related to the original symmetric matrix and the ciphertext of the original product vector to the first server, since the first server does not know the method by which the data owner encrypts the private data, the first server cannot know the real situation of the private data; in addition, after receiving the ciphertext related to the original symmetric matrix and the ciphertext of the original product vector, the first server also performs an encryption process on the ciphertext related to the original symmetric matrix and the ciphertext of the original product vector again through the pseudo-sample encryption method known only to itself. Therefore, after receiving the sum of the ciphertexts of the original symmetric matrices of the multiple data owners including the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of the multiple data owners including the influence of the pseudo-sample product vector, the second server also cannot crack the real situation of the private data. In summary, except for the data owner holding its own private data, neither the first server nor the second server can steal the private data based on their own conditions, thus ensuring the security of the private data.
[0017] Other features and advantages of the present application will be described in the following specification, and some of them will become obvious from the specification, or be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the specification, claims, and drawings. Brief Description of the Drawings
[0018] The drawings are used to provide a further understanding of the technical solutions of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solutions of the present application, and do not constitute a limitation to the technical solutions of the present application.
[0019] Figure 1 It is an optional flowchart for the method for horizontally dividing data and performing linear regression based on privacy protection provided by an embodiment of the present application;
[0020] Figure 2 It is an optional schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed Description of the Embodiments
[0021] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0022] In the description of the present application, the meaning of "a number of" is one or more, the meaning of "a plurality of" is two or more, and understandings such as "greater than", "less than", "exceeding", etc. do not include the present number, and understandings such as "above", "below", "within", etc. include the present number.
[0023] It should be noted that although functional module division is performed in the system schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the system or the flowchart. Terms such as "first", "second", etc. in the specification, claims or the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0024] It should be noted that in each specific embodiment of the present application, when it comes to performing relevant processing based on data related to the characteristics of the target object, such as the attribute information of the target object or the set of attribute information, the permission or consent of the target object will be obtained first. Moreover, the collection, use and processing of these data will comply with relevant laws, regulations and standards. Among them, the target object can be a user. In addition, when an embodiment of the present application needs to obtain the attribute information of the target object, it will obtain the separate permission or separate consent of the target object through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary data related to the target object for the normal operation of the embodiment of the present application will be obtained.
[0025] To facilitate the understanding of the technical solutions provided by the embodiments of the present application, some key terms used in the embodiments of the present application will be explained here:
[0026] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0027] Linear Regression is a type of regression analysis that uses a least-squares function called the linear regression equation to model the relationship between one or more independent variables and a dependent variable. This function is a linear combination of one or more model parameters called regression coefficients. The case with only one independent variable is called simple linear regression, and the case with more than one independent variable is called multiple linear regression. The above independent variables are also called input features. The above dependent variable is also called a label, which refers to the target variable or the variable to be predicted, representing the result or value that the regression model is expected to estimate based on the given input features. An example of linear regression is as follows: Suppose we want to build a regression model to predict house prices based on various features such as house size, number of bedrooms, location, etc.; then, the label will be the actual price of the house, and the input features include the size of the house, the number of bedrooms, the location, and any other relevant factors. Federated Learning is an emerging fundamental technology in artificial intelligence. Its design goal is to carry out efficient machine learning among multiple participants or computing nodes while ensuring information security during big data exchange, protecting terminal data and personal data privacy, and ensuring legal compliance. Federated Learning can perform large-scale training on the devices that generate data, and these sensitive data remain with the data owners, being collected and trained locally. After local training, the central training coordinator obtains the training contributions of each node by acquiring the updates of the distributed model without accessing the actual sensitive data.
[0028] Homomorphic Encryption (HE) is a widely used cryptographic tool, which is an encryption algorithm that satisfies the property of homomorphic operations on ciphertexts. That is, after the data is encrypted by homomorphic encryption, specific calculations are performed on the ciphertexts, and the plaintext obtained after the corresponding homomorphic decryption of the ciphertext calculation result is equivalent to directly performing the same calculation on the plaintext data, realizing the "computable but invisible" of the data. Therefore, it is widely used in scenarios such as privacy-protected cloud service computing, outsourced computing, and horizontal federated learning, and is a direction of emerging privacy technologies. Homomorphic encryption usually uses a pair of keys, including a public key and a private key. The public key is used to encrypt the plaintext data, that is, the plaintext data is encrypted with the public key to obtain the corresponding ciphertext data, and the public key is usually held by multiple nodes that need to perform encryption operations; the private key is used to decrypt the ciphertext data, that is, the ciphertext data is decrypted with the private key to obtain the corresponding plaintext data, and the private key is usually held by one or a few trusted nodes that need to perform decryption operations. Homomorphic encryption technologies include additive homomorphic encryption technology that supports addition calculations on ciphertexts, multiplicative homomorphic encryption technology that supports multiplication calculations on ciphertexts, and fully homomorphic encryption technology that satisfies both additive and multiplicative homomorphic properties and thus supports addition and multiplication calculations on ciphertexts. Generally, the more types of calculation methods that the homomorphic encryption technology needs to support, the higher the implementation complexity of the encryption technology. For example, the implementation complexity of the above-mentioned fully homomorphic encryption technology is usually higher than that of the above-mentioned additive homomorphic encryption technology and also higher than that of the above-mentioned multiplicative homomorphic encryption technology.
[0029] Based on this, the embodiments of the present application provide a method and device for horizontally partitioning data and performing linear regression based on privacy protection, which can ensure the security of data transmission and operation in a horizontal federated learning platform.
[0030] The method and device for horizontally partitioning data and performing linear regression based on privacy protection provided by the embodiments of the present application are specifically described through the following embodiments. First, the method for horizontally partitioning data and performing linear regression based on privacy protection in the embodiments of the present application is described.
[0031] The linear regression method for horizontally partitioning data and based on privacy protection provided by the embodiments of the present application relates to the field of computer technology. The linear regression method for horizontally partitioning data and based on privacy protection provided by the embodiments of the present application can be applied to a terminal, can also be applied to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the linear regression method for horizontally partitioning data and based on privacy protection, etc., but is not limited to the above forms.
[0032] The present application can be used in many general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0033] It should be noted that in each specific embodiment of the present application, when it comes to performing relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain sensitive personal information of users, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.
[0034] The following further elaborates on the embodiments of the present application in conjunction with the accompanying drawings. As Figure 1 shown, Figure 1FIG. 0 is an alternative process schematic diagram of the linear regression method for horizontally partitioning data and based on privacy protection provided by an embodiment of the present application. The linear regression method for horizontally partitioning data and based on privacy protection is applied to a horizontal federated learning platform. A data owner is communicatively connected to the horizontal federated learning platform. The horizontal federated learning platform includes a first server and a second server. The first server is communicatively connected to the second server. The first server performs a machine learning task based on a federated model. The second server is used to coordinate encryption services. The linear regression method for horizontally partitioning data and based on privacy protection can be executed by a server, or can also be executed by a terminal, or can also be executed by a server in cooperation with a terminal. The linear regression method for horizontally partitioning data and based on privacy protection includes, but is not limited to, the following steps S110 to S140:
[0035] Step S110, the first server receives the ciphertext related to the original symmetric matrix and the ciphertext of the original product vector sent by each of at least two data owners, and obtains the processing results of the pseudo-sample symmetric matrix and the pseudo-sample product vector. Among them, the processing results of the pseudo-sample symmetric matrix and the pseudo-sample product vector are the results obtained by processing the pseudo-sample symmetric matrix and the pseudo-sample product vector respectively, and the pseudo-sample symmetric matrix and the pseudo-sample product vector are the symmetric matrix and the product vector corresponding to a plurality of pseudo-data samples respectively; the first server uses the ciphertext related to the original symmetric matrix and the processing result of the pseudo-sample symmetric matrix to obtain the sum of the original symmetric matrix ciphertexts of multiple data owners including the influence of the pseudo-sample symmetric matrix, and uses the ciphertext of the original product vector and the processing result of the pseudo-sample product vector to obtain the sum of the original product vector ciphertexts of multiple data owners including the influence of the pseudo-sample product vector; the first server sends the sum of the original symmetric matrix ciphertexts of multiple data owners including the influence of the pseudo-sample symmetric matrix and the sum of the original product vector ciphertexts of multiple data owners including the influence of the pseudo-sample product vector to the second server;
[0036] Step S120, the second server receives and decrypts the sum of the original symmetric matrix ciphertexts of multiple data owners including the influence of the pseudo-sample symmetric matrix and the sum of the original product vector ciphertexts of multiple data owners including the influence of the pseudo-sample product vector respectively, so as to obtain the sum of the original symmetric matrix plaintexts of all data owners including the influence of the pseudo-sample symmetric matrix and the sum of the original product vector plaintexts of all data owners including the influence of the pseudo-sample product vector;
[0037] Step S130: The second server obtains the pseudo-sample encrypted linear regression coefficients affected by the pseudo-sample symmetric matrix and the pseudo-sample product vector by using the sum of the original symmetric matrix plaintexts of all data owners containing the influence of the pseudo-sample symmetric matrix and the sum of the original product vector plaintexts of all data owners containing the influence of the pseudo-sample product vector, and sends the pseudo-sample encrypted linear regression coefficients to the first server. Alternatively, the second server performs regularization processing on the sum of the original symmetric matrix plaintexts of all data owners containing the influence of the pseudo-sample symmetric matrix, and obtains the pseudo-sample encrypted linear regression coefficients affected by the pseudo-sample symmetric matrix and the pseudo-sample product vector according to the sum of the original product vector plaintexts of all data owners containing the influence of the pseudo-sample product vector and the sum of the original symmetric matrix plaintexts of all data owners after regularization processing, and sends the pseudo-sample encrypted linear regression coefficients to the first server;
[0038] Step S140: The first server receives the pseudo-sample encrypted linear regression coefficients affected by the pseudo-sample symmetric matrix and the pseudo-sample product vector, and at least partially eliminates the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector from the pseudo-sample encrypted linear regression coefficients to obtain the target linear regression coefficients.
[0039] It can be understood that horizontal federated learning refers to the case where there is less user overlap and more user feature overlap among the data owners participating in the joint modeling. By splitting the private data of each data owner according to the user dimension and taking out the part of the private data where the user features are the same but the users are different for training. It faces the problem of joint learning of different sample subsets in the same feature space, which is the key difference between it and vertical-horizontal federated learning and federated transfer learning. It is also called horizontal federated learning divided by samples and can be applied to scenarios where the private data of each data owner participating in horizontal federated learning has the same feature space and different sample spaces. The IEEE has released the standard document IEEE Standard 3652.1-2020: IEEE Guide for Architectural Framework and Application of Federated Machine Learning, which gives standardized definitions and descriptions of horizontal federated learning, vertical-horizontal federated learning, and federated transfer learning. According to the above standard, horizontal federated learning follows the original framework setting of horizontal federated learning and faces the data scenario of sample joint, that is, the feature spaces of the private data of each data owner are consistent, and the relationship between the private data sets of each data owner and the whole samples is a "horizontal division" relationship.
[0040] The essence of horizontal federated learning is the union of samples, which is applicable to scenarios where data owners have the same business type but different customer reach (i.e., more feature overlaps and fewer user overlaps). It is often applied to data with different samples but similar features. In traditional machine learning engine modeling, usually, the data required for model training is gathered in a single data center and then the model is trained. In horizontal federated learning, it can be regarded as distributed model training based on samples. All data is distributed to different machines, and each machine uses its local private data to collaboratively train the model while protecting the privacy of local private data. Horizontal federated learning can be considered as distributed model training based on samples. The system scale of horizontal federated learning can vary from a few to millions of data owners. There are two typical scenarios: cross-device horizontal federated learning and cross-silo horizontal federated learning. The former targets a relatively large-scale and heterogeneous client group; the latter targets scenarios composed of a few but resource-rich nodes (corresponding to large organizations such as e-commerce enterprises and banks), and each node contains a relatively large amount of private data.
[0041] A typical example is that each mobile phone is used by an individual, forming a collection of private data. There are different mobile phones, and basically, the features obtained by each mobile phone are the same, but the private data is different. On the premise that the data privacy on each mobile phone is protected, horizontal federated learning can aggregate the private data samples on these mobile phones to train the machine learning engine model. For example, the Gboard system of Google Input can form a federation of many devices installed with Gboard, integrate multi-party data to build horizontal federated learning, so as to improve the accuracy of the input word prediction task of the input method for users in different industries and with different input habits; and the above-mentioned many devices installed with Gboard also include mobile phones.
[0042] Two enterprises with the same business but located in different regions have user groups from their respective regions, and the intersection between them is very small. However, their businesses are very similar, so the user features recorded are the same. At this time, horizontal federated learning can be used to build a joint model.
[0043] In the financial scenario, horizontal federated learning is applicable to joint modeling among financial institutions, that is, the business scenarios among data owners are similar, the feature spaces of private data are the same, and the intersection of user groups is small. For example, two bank institutions in different regions have a very small intersection of user groups, but their businesses are very similar, so the feature spaces of private data are the same. However, due to some specific business scenarios, such as micro and small enterprise credit, etc., the modeling samples available to each data owner are relatively few, so it is difficult to build models using traditional machine learning engine algorithms separately. In this case, horizontal federated learning can be used to jointly use the sample data among multiple different institutions to expand the sample space for model training, so as to build a more accurate and better generalizable model, including a risk control model. The features of micro, small and medium-sized enterprises often used in horizontal federated learning include multi-source information such as enterprise operation data, tax data, industrial and commercial data, and payment data.
[0044] In intelligent medical technology, private data such as patients' diseases, pathological reports, and test results are scattered in multiple medical institutions. Each medical institution has data samples of different patients. Therefore, multiple medical institutions can also use horizontal federated learning technology to jointly train a machine learning engine model. Securities dealers can also use horizontal federated learning technology: two branches of a securities firm in different regions, or two local securities firms in different regions, serve different customer groups, but the financial attributes of the customers are roughly similar. In this scenario, the features of customers often used in horizontal federated learning usually include information such as the types of securities purchased, the purchase amount, and the holding time, as well as customer information such as occupation, gender, age, annual income, and family situation.
[0045] The embodiments of this application are based on a common federated learning system model. The embodiments of this application include two servers that do not collude to steal user privacy data. One can be called a Machine-Learning Engine, and the other can be called a Crypto Service Provider (CSP). In the following text, the Machine-Learning Engine is called the first server, and the Crypto Service Provider is called the second server. In addition, in the subsequent detailed description of multiple embodiments of the present invention, the embodiments of the present invention also include M data owners, where M is greater than or equal to 2, and the forms of private data owned by each data owner are as follows.
[0046] The private data owned by the k-th data owner among the M data owners is represented as a set of original data samples where k is any positive integer greater than or equal to 1 and less than or equal to M, that is, 1 ≤ k ≤ M. The above set of original data samples D k contains n k -n k-1a set of n original data samples k -n k-1 is greater than or equal to 1; the set D of the above-mentioned original data samples k The i-th original data sample included is denoted as where i is greater than or equal to n k-1 +1 and less than or equal to n k that is, n k-1 +1 ≤ i ≤ n k . The above-mentioned i-th original data sample is In represents the original input feature column vector included in the i-th original data sample, and y i represents the corresponding private original label, which can be called the original label included in the i-th original data sample; the above-mentioned private original input feature column vector includes δ terms, and δ is greater than or equal to 1; δ is also called the number of attributes of the original input feature column vector or the number of input attributes in linear regression. The embodiments of the present application are directed to common scenarios where δ is greater than or equal to 2.
[0047] For the sake of concise expression, the above-mentioned n k -n k-1 original input feature column vectors constitute an original input feature matrix with n k -n k-1 rows and δ columns The above-mentioned n k -n k-1 original labels y i (i = n k-1 +1, n k-1 +2, …, n k ) constitute an original label column vector containing n k -n k-1 terms The above-mentioned original input feature matrix X k includes n k -n k-1 rows, and the i-th row is the transpose of the original input feature column vector In the present application document, the matrix X is called the original input feature matrix of the k-th data owner, and is simply referred to as the original input matrix; the above-mentioned column vector k includes n -n k terms, and the i-th term is the original label y k-1 (i = n i +1, n k-1 +2, …, n k-1 ) In the present application document, the column vector k is called the original label column vector of the k-th data owner, and is simply referred to as the original label vector. In the present application document, the matrix X Abbreviated as the original label column vector of the k-th data owner.
[0048] In the scenario of credit assessment for small, medium and micro enterprises, the above-mentioned private input feature column vectors can include multi-source information such as a certain enterprise's business data, tax data, industrial and commercial data, payment data, etc., and the corresponding y i can be data related to the credit level of the enterprise. Note that the private input feature column vectors can also be represented in other forms. For example, in the form of a private input feature row vector. In this case, only the description of the embodiments needs to be changed accordingly, and such changes are well-known.
[0049] In the first and second embodiments, all data owners use the same encryption method. In the first embodiment, all data owners use the homomorphic encryption method. In the second embodiment, all data owners use the masked encryption method.
[0050] First, the first embodiment of the present application is described. The main difference between the first embodiment and the second embodiment described later is that in the first embodiment, each data owner uses the homomorphic encryption method in the prior art to encrypt its own private original data sample, and the second server decrypts the intermediate ciphertext obtained in the calculation process based on the decryption method corresponding to the homomorphic encryption method used by the data owner.
[0051] Specifically, Figure 1 before step S110 in, each data owner encrypts its own private data, including but not limited to the following steps S210 to S220:
[0052] Step S210, each of at least two data owners uses its own private original feature column vector and original label to determine its own original symmetric matrix and original product vector. Among them, the original symmetric matrix of any data owner is equal to the sum of the products of each original feature column vector of the data owner and its own transpose, and the original product vector of any data owner is equal to the sum of the products of each original label and the corresponding original feature column vector.
[0053] When represented in matrix form, the above-mentioned original symmetric matrix is a symmetric matrix obtained by multiplying the transpose of the original input matrix by the original input matrix itself, and the above-mentioned original product vector is the product of the transpose of the original input matrix and the original label column vector. Note that the above-mentioned original symmetric matrix and original product vector are intermediate results calculated from the original data samples private to the corresponding data owners. Before sending the intermediate results calculated from the private original data samples to other nodes, any data owner usually needs to encrypt the intermediate results to protect their private data; therefore, the above-mentioned original symmetric matrix is also called the symmetric matrix to be encrypted, and the above-mentioned original product vector is also called the product vector to be encrypted.
[0054] The implementation of this step will be described in detail below. The k-th data owner calculates the original symmetric matrix of the k-th data owner based on the private data it owns. and the original product vector of the k-th data owner The above-mentioned original symmetric matrix A k is the cumulative result of the products between the input feature column vectors and their transposes in each data sample owned by the k-th data owner. Specifically, it is for each of the n k -n k-1 data samples, the input feature column vector is obtained and its transpose to calculate the product and then the cumulative result of adding up the n k -n k-1 products is obtained. Here, n k-1 +1 ≤ i ≤ n k , that is, the above-mentioned original symmetric matrix A k is the cumulative result of the products between each input feature column vector of the k-th data owner and its transpose; when represented in matrix form, the above-mentioned original symmetric matrix is a symmetric matrix obtained by multiplying the transpose of the original input matrix X k of the k-th data owner by itself, and can also be called the symmetric result of the original input matrix of the k-th data owner, where represents one of the feature column vectors of a certain data owner, and X k represents the combined matrix formed by the transposes of each feature column vector of the k-th data owner as rows, that is, each row of X k is the transpose of each feature column vector. On the other hand, the above-mentioned original product vector is the cumulative result of the products between the input feature column vectors and the corresponding private labels in each data sample owned by the k-th data owner, and is for n k -n k-1Each of the data samples contains an input feature column vector obtained the product with the corresponding private label y i product Then for n k -n k-1 pieces The accumulated result obtained by accumulation, where n k-1 +1 ≤ i ≤ n k , that is, the above column vector is the accumulated result of the product of each input feature column vector and the corresponding label of the k-th data owner; when represented in matrix form, the above original product vector is the transpose of the original input matrix x of the k-th data owner k transpose and the original label column vector product, which can also be called the product of the original input matrix and the label of the k-th data owner.
[0055] Step S220, each of at least two data owners uses the additive homomorphic encryption technology that supports addition calculation on ciphertexts to encrypt their respective original symmetric matrices and original product vectors respectively, obtaining the ciphertexts of their respective original symmetric matrices and the ciphertexts of the original product vectors; each data owner sends the ciphertexts of their respective original symmetric matrices and the ciphertexts of the original product vectors to the first server.
[0056] Each of the M data owners has to execute this step. The implementation method for the k-th data owner among the M data owners is given below, where 1 ≤ k ≤ M. The k-th data owner obtains the first key, and then uses the additive homomorphic encryption technology that supports addition calculation on ciphertexts, and uses the first key to perform homomorphic encryption on its own original symmetric matrix A k and the original product vector respectively, obtaining the ciphertext A' of the original symmetric matrix k = Enc key1 (A k ) and the ciphertext of the original product vector Then send the ciphertext A' of the original symmetric matrix k and the ciphertext of the original product vector to the first server. The above key1 represents the first key, Enc key1(Δ) represents homomorphic encryption processing of the plaintext Δ therein based on the first key; and the above-mentioned first key is usually sent by the second server to each of the M data owners, and is usually the public key used in homomorphic encryption technology. The above-mentioned additive homomorphic encryption technology that supports addition calculation on ciphertexts not only uses the first key usually called the public key, but also uses the second key usually called the private key. Usually, the second server has the second key and uses the second key to decrypt the received ciphertext, and the details will be introduced below. The above-mentioned additive homomorphic encryption technology that supports addition calculation on ciphertexts can be the Pallier homomorphic encryption method.
[0057] The implementation manner of this step will be described in detail below. The ciphertext A' of the above-mentioned original symmetric matrix k and the ciphertext of the original product vector respectively satisfy A' k = Enc key1 (A k ) and where key1 represents the first key, and Enc key1 (A k ) represents homomorphic encryption processing of A k based on the first key. A' k represents the ciphertext of the original symmetric matrix obtained by homomorphic encryption of the original symmetric matrix A k of the k-th data owner. Correspondingly, A' k can be called the ciphertext of the original symmetric matrix of the k-th data owner; on the other hand, represents homomorphic encryption processing of based on the first key, represents the ciphertext obtained by homomorphic encryption of the original product vector of the k-th data owner. Correspondingly, can be called the ciphertext of the original product vector of the k-th data owner.
[0058] In the step S110 including but not limited to steps S210 to S220, each of at least two data owners uses a number of private data samples to obtain the ciphertext related to the original symmetric matrix and the ciphertext of the original product vector, and sends the above two ciphertexts to the first server; the ciphertext related to the original symmetric matrix and the ciphertext of the original product vector are the ciphertexts of the original symmetric matrix and the ciphertext of the original product vector obtained by encrypting the original symmetric matrix and the original product vector respectively using an encryption technology that supports addition calculation on ciphertexts.
[0059] It can be understood that before sending the intermediate result calculated from the original data sample owned by itself to the first server, the data owner can ensure that the first server cannot know the true situation of the intermediate result by performing homomorphic encryption processing on the intermediate result based on the first key owned by each; in addition, the data owner has the intermediate result calculated from the private data, the first server obtains the ciphertext of the intermediate result after homomorphic encryption, and the second key used to decrypt the homomorphic encrypted ciphertext is stored in the second server. Therefore, except for the data owner who holds the intermediate result calculated from its own private data, the first server cannot know the above intermediate result, thus ensuring the security of the private data of the data owner. In addition, as will be described below, the first server further encrypts the ciphertext of the intermediate result after homomorphic encryption before sending it to the second server; correspondingly, after the second server decrypts the received ciphertext with the second key, it still cannot know the above intermediate result because of the above further encryption performed by the first server. As can be seen from the above, neither the first server nor the second server can know the above intermediate result calculated from the private data of the data owner through the information received by themselves, thus protecting the privacy of the data owner. In the subsequent steps of this embodiment, the first server receives the ciphertext of the original symmetric matrix and the ciphertext of the original product vector sent by each of at least two data owners, and cooperates with the second server to obtain the linear regression coefficient. The subsequent steps of this embodiment include step S110, step 120, step S130 and step S140, which will be described below.
[0060] Figure 1In method step S110, the first server receives the ciphertexts of the original symmetric matrices and the ciphertexts of the original product vectors sent by each of at least two data owners, processes the ciphertexts of the original symmetric matrices sent by at least two data owners to obtain the sum of the ciphertexts of the original symmetric matrices of multiple data owners, and processes the ciphertexts of the original product vectors sent by at least two data owners to obtain the sum of the ciphertexts of the original product vectors of multiple data owners; on the other hand, the first server obtains the symmetric matrix and the product vector corresponding to the pseudo-data sample, which are respectively called the pseudo-sample symmetric matrix and the pseudo-sample product vector; then, the first server adds the processing result of the pseudo-sample symmetric matrix to the sum of the ciphertexts of the original symmetric matrices of multiple data owners to obtain the sum of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix. In addition, the first server adds the processing result of the pseudo-sample product vector to the sum of the ciphertexts of the original product vectors of multiple data owners to obtain the sum of the ciphertexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector; finally, the first server sends the sum of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector to the second server; step S110 includes but is not limited to the following steps S310 to S340:
[0061] In step S310, the first server receives the ciphertexts of the original symmetric matrices and the ciphertexts of the original product vectors sent by each of at least two data owners; the first server performs an accumulation process and an optional regularization process on the ciphertexts of the original symmetric matrices sent by at least two data owners to obtain the sum of the ciphertexts of the original symmetric matrices of multiple data owners, which is simply called the first ciphertext sum of the symmetric matrix; the first server also performs an accumulation process on the ciphertexts of the original product vectors sent by at least two data owners to obtain the sum of the ciphertexts of the original product vectors of multiple data owners, which is simply called the first ciphertext sum of the product vector.
[0062] Specifically, the first server receives the ciphertexts of the original symmetric matrices, i.e., ciphertext A', sent by each of M data owners k , and the ciphertexts of the original product vectors, i.e., ciphertext where k = 1, 2,..., M; then, the first server uses the technology of performing addition calculation on ciphertexts included in the homomorphic encryption technology to perform accumulation and optional regularization processing on the ciphertexts of the original symmetric matrices of the M data owners received to obtain the sum of the ciphertexts of the original symmetric matrices of the M data owners, i.e., ciphertext And perform accumulation on the ciphertexts of the original product vectors of the M data owners received to obtain the sum of the ciphertexts of the original product vectors of the M data owners, i.e., ciphertext The ciphertext sum of the original symmetric matrices of multiple data owners (referred to as the first ciphertext sum of symmetric matrices for short), that is, the ciphertext sum A′ of the original symmetric matrices of the above M data owners, is the ciphertext Enc key1 (λI) and the ciphertext A′ of the original symmetric matrices of M data owners k (k = 1, 2... M) obtained by cumulative processing, where λ in the regularization term λI is a real number greater than or equal to zero, and I is the identity matrix; the above regularization process is an optional step, because when λ in the regularization term λI is equal to zero, it is equivalent to not performing the regularization process, but only performing cumulative processing on the ciphertext of the symmetricized result of the original input matrices of the M data owners received to obtain the ciphertext sum of the original symmetric matrices of the M data owners Specifically, the above regularization process can also be performed in other steps. For details, please refer to the subsequent description. On the other hand, the ciphertext sum of the original product vectors of multiple data owners (referred to as the first ciphertext sum of product vectors for short), that is, the ciphertext sum of the original product vectors of the above M data owners is the ciphertext of the original product vectors of M data owners The cumulative result of
[0063] Step S320, the first server uses a number of pseudo-data samples to obtain the corresponding symmetric matrix and product vector of the pseudo-data samples, which are respectively called the pseudo-sample symmetric matrix and the pseudo-sample product vector
[0064] Specifically, for further encryption, the first server obtains τ pseudo-data samples, where τ is greater than or equal to 1 and less than δ. The first server uses the pseudo-input matrix E and the pseudo-label column vector included in the pseudo-data samples Calculate the symmetric matrix obtained by multiplying the transpose of the pseudo-input matrix E by itself, that is, the pseudo-sample symmetric matrix φ = E T E, and the transpose E of the pseudo-input matrix E T And the product of the pseudo-label column vector , that is, the pseudo-sample product vector The values of each item in the above τ pseudo-data samples are usually taken as random numbers
[0065] The implementation manner of this step is described in detail below. As mentioned above, the above δ is the number of items included in the input feature column vector in the data samples owned by the data owner , which is also called the number of attributes of the input feature column vector or the number of input attributes in linear regression. The above τ pseudo-data samples can be represented as a pseudo-data sample set The i-th pseudo-data sample included in the set S can be represented as where i is greater than or equal to 1 and less than or equal to τ, that is, 1 ≤ i ≤ τ. The above-mentioned i-th pseudo-data sample is In represents the pseudo-input feature column vector included in the i-th pseudo-data sample, f i represents the corresponding pseudo-label, which can be called the pseudo-label included in the i-th pseudo-data sample; the above-mentioned pseudo-input feature column vector also includes a δ term, just like the private input feature column vector mentioned above In this application document, for the convenience of expression, the above-mentioned τ pseudo-input feature column vectors constitute a pseudo-input feature matrix with τ rows and δ columns The above-mentioned τ pseudo-labels f i (i = 1, 2,..., τ) constitute a pseudo-label column vector containing τ terms The above-mentioned pseudo-input feature matrix E includes τ rows, and the i-th row is the transpose of the pseudo-input feature column vector of In this application document, the pseudo-input feature matrix is simply referred to as the pseudo-input matrix; the above-mentioned pseudo-label column vector includes τ terms, and the i-th term is the pseudo-label f i (i = 1, 2,..., τ).
[0066] The above-mentioned pseudo-sample symmetric matrix φ is the cumulative result of the product of the pseudo-input feature column vector in the pseudo-data sample and its transpose, and is for each pseudo-input feature column vector included in each of the τ pseudo-data samples obtained and its transpose of the product and then the cumulative result obtained by accumulating the τ is Here 1 ≤ i ≤ τ; when expressed in matrix form, the above-mentioned φ = E T E is a symmetric matrix obtained by multiplying the transpose of the pseudo-input matrix E by itself, simply referred to as the result of symmetrizing the pseudo-input matrix, and also simply referred to as the pseudo-sample symmetric matrix. On the other hand, the above-mentioned pseudo-sample product vector is the cumulative result of the product of the pseudo-input feature column vector and the pseudo-label in each pseudo-data sample, and is for each pseudo-input feature column vector included in each of the τ pseudo-data samples obtained and the corresponding pseudo-label f i of the product and then the cumulative result obtained by accumulating the τ is Here 1 ≤ i ≤ τ; when expressed in matrix form, the above-mentioned is the transpose of the pseudo-input matrix E and the pseudo-label column vector The product, abbreviated as the product of the pseudo-input matrix and the pseudo-label, is also simply referred to as the pseudo-sample product vector.
[0067] Step S330, the first server adds the sum of the first ciphertexts of the symmetric matrix to the processing result of the pseudo-sample symmetric matrix to obtain the sum of the original ciphertexts of the symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix; the first server adds the sum of the first ciphertexts of the product vector to the processing result of the pseudo-sample product vector to obtain the sum of the original ciphertexts of the product vectors of multiple data owners including the influence of the pseudo-sample product vector. The first server sends the sum of the original ciphertexts of the symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix and the sum of the original ciphertexts of the product vectors of multiple data owners including the influence of the pseudo-sample product vector to the second server.
[0068] The processing result of the pseudo-sample symmetric matrix is the ciphertext obtained by homomorphic encryption of the plaintext of the pseudo-sample symmetric matrix (i.e., the above pseudo-sample symmetric matrix φ = E T E), or the ciphertext obtained by homomorphic encryption after taking the negative sign of the plaintext of the pseudo-sample symmetric matrix; the processing result of the pseudo-sample product vector is the ciphertext obtained by homomorphic encryption of the plaintext of the pseudo-sample product vector (i.e., the above pseudo-sample product vector ) or the ciphertext obtained by homomorphic encryption after taking the negative sign of the plaintext of the pseudo-sample product vector. To obtain the above processing results of the pseudo-sample symmetric matrix and the pseudo-sample product vector, the first server obtains the first key usually sent by the second server, and uses the additive homomorphic encryption technology that supports addition calculation on ciphertexts to perform homomorphic encryption using the first key; as mentioned above, the above first key is usually the public key used in homomorphic encryption technology. The above processing result of the pseudo-sample symmetric matrix is to perform homomorphic encryption on the pseudo-sample symmetric matrix φ to obtain the ciphertext φ' = Enc key1 (φ), or the ciphertext φ' = Enc key1 (-φ) obtained by taking the negative sign of the pseudo-sample symmetric matrix φ and then performing homomorphic encryption; the above processing result of the pseudo-sample product vector is the ciphertext of the pseudo-sample product vector obtained by performing homomorphic encryption on the pseudo-sample product vector or the ciphertext of the pseudo-sample product vector obtained by taking the negative sign of the pseudo-sample product vector and then performing homomorphic encryption The ciphertext of the pseudo-sample product vector obtained by taking the negative sign and then performing homomorphic encryption
[0069] The first server uses the technology of performing addition calculations on multiple ciphertexts in homomorphic encryption to add the sum of the first ciphertexts of the symmetric matrix and the processing result of the pseudo-sample symmetric matrix to obtain the sum of the original symmetric matrix ciphertexts of multiple data owners including the influence of the pseudo-sample symmetric matrix, and adds the sum of the first ciphertexts of the product vector and the processing result of the pseudo-sample product vector to obtain the sum of the original product vector ciphertexts of multiple data owners including the influence of the pseudo-sample product vector. Specifically, the first server adds the sum of the first ciphertexts of the symmetric matrix A' and the processing result of the pseudo-sample symmetric matrix φ' to obtain the sum of the original symmetric matrix ciphertexts of multiple data owners including the influence of the pseudo-sample symmetric matrix C' = A' + φ'; and adds the sum of the first ciphertexts of the product vector and the processing result of the pseudo-sample product vector to obtain the sum of the original product vector ciphertexts of multiple data owners including the influence of the pseudo-sample product vector Then, the first server sends the sum C' of the original symmetric matrix ciphertexts of multiple data owners including the influence of the pseudo-sample symmetric matrix and the sum of the original product vector ciphertexts of multiple data owners including the influence of the pseudo-sample product vector to the second server.
[0070] This step further encrypts the ciphertext A' and using the ciphertext related to the pseudo-data sample. The processing result of the pseudo-sample symmetric matrix can be φ' = Enc key1 (φ) or φ' = Enc key1 (-φ), and the processing result of the pseudo-sample product vector can be or Correspondingly, this step can include the following four implementation manners:
[0071] Implementation manner 1: The sum C' of the original symmetric matrix ciphertexts of multiple data owners including the influence of the pseudo-sample symmetric matrix = A' + φ' = Enc key1 (A) + Enc key1 (φ), and the sum
[0072] of the original product vector ciphertexts of multiple data owners including the influence of the pseudo-sample product vector key1 (A) + Enc key1 (-φ), and the sum
[0073] Implementation manner 3: The sum C' of the original symmetric matrix ciphertexts of multiple data owners including the influence of the pseudo-sample symmetric matrix = A' + φ' = Enckey1 (A) + Enc key1 (φ), and the sum of the original ciphertexts of the product vectors of multiple data owners including the influence of the pseudo-sample product vector
[0074] Embodiment 4: The sum of the original ciphertexts of the symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix C' = A' + φ' = Enc key1 (A) + Enc key1 (-φ), and the sum of the original ciphertexts of the product vectors of multiple data owners including the influence of the pseudo-sample product vector
[0075] Taking Embodiment 1 as an example below, the implementation of this step will be described in detail. As described in Step S220, key1 used in this step represents the first key, and Enc key1 (Δ) represents performing homomorphic encryption processing on the plaintext Δ therein based on the first key. The above φ' = Enc key1 (φ) represents the ciphertext obtained by performing homomorphic encryption on the pseudo-sample symmetric matrix φ based on the first key, and the above represents the ciphertext obtained by performing homomorphic encryption on the pseudo-sample product vector based on the first key.
[0076] The sum of the original ciphertexts of the symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix C' = A' + φ' represents the ciphertext obtained by performing an addition calculation on the ciphertexts A' and φ' using the technique of adding multiple ciphertexts in the homomorphic encryption technology, and the sum of the original ciphertexts of the product vectors of multiple data owners including the influence of the pseudo-sample product vector represents using the technique of adding multiple ciphertexts in the homomorphic encryption technology to perform an addition calculation on the ciphertext and the ciphertext to obtain the ciphertext.
[0077] Both the above-mentioned step S310 and step S330 are operations of adding ciphertexts. The first server first accumulates the ciphertexts of the original symmetric matrices sent by at least two data owners and then adds the processing result of the pseudo-sample symmetric matrix, and first accumulates the ciphertexts of the original product vectors sent by at least two data owners and then adds the processing result of the pseudo-sample product vector. In fact, the order of adding the above ciphertexts can be adjusted arbitrarily; for example, the first server first accumulates the ciphertexts of the original symmetric matrices sent by one or any number of data owners, then adds the processing result of the pseudo-sample symmetric matrix, and finally accumulates the ciphertexts of the original symmetric matrices sent by the remaining several data owners; on the other hand, the first server first accumulates the ciphertexts of the original product vectors sent by one or any number of data owners, then adds the processing result of the pseudo-sample product vector, and finally accumulates the ciphertexts of the original product vectors sent by the remaining several data owners.
[0078] It should be noted that in the first embodiment, since each data owner uses the homomorphic encryption method to encrypt their respective original symmetric matrices respectively, therefore, in step S330, the first server adds the sum of the first ciphertexts of the symmetric matrices and the processing result of the pseudo-sample symmetric matrix to obtain the sum of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix, which includes the ciphertexts related to the original symmetric matrices of all data owners, that is, the sum of the first ciphertexts of the symmetric matrices in the first embodiment. Therefore, in step S330, the sum of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix is equivalent to the sum of the ciphertexts of the original symmetric matrices of all data owners including the influence of the pseudo-sample symmetric matrix; similarly, since each data owner uses the homomorphic encryption method to encrypt their respective original product vectors respectively, therefore, the sum of the ciphertexts of the original product vectors of multiple data owners including the influence of the pseudo-sample symmetric matrix in step S330 is equivalent to the sum of the ciphertexts of the original product vectors of all data owners including the influence of the pseudo-sample symmetric matrix.
[0079] Step S340, the first server encrypts the pseudo-input matrix E composed of several pseudo-input feature column vectors to obtain the ciphertext of the pseudo-input matrix Then send the ciphertext of the pseudo-input matrix To the second server; the above-mentioned several pseudo-input feature column vectors are part of the corresponding several pseudo-data samples, as described above. The encryption of the pseudo-input matrix E can be implemented in different ways, and two implementation methods are given below:
[0080] Implementation method 1: The first server generates or obtains an invertible matrix Θ of τ rows and τ columns, and this matrix Θ is called a multiplication encryption mask matrix, and as described above, the above-mentioned τ is greater than or equal to 2; the first server multiplies the invertible matrix Θ by the pseudo-input matrix E to obtain the multiplication-encrypted pseudo-input matrix As the ciphertext of the pseudo-input matrix.
[0081] Embodiment 2: The first server uses a homomorphic encryption method that satisfies additive homomorphism and supports multiplication of ciphertext and plaintext, and encrypts the pseudo-input matrix E with a third key to obtain the ciphertext of the pseudo-input matrix That is And the third key, i.e., key3, is usually the public key in the homomorphic encryption that satisfies additive homomorphism and supports multiplication of ciphertext and plaintext. The above homomorphic encryption method that satisfies additive homomorphism and supports multiplication of ciphertext and plaintext not only uses the third key, which is usually called the public key, but also uses the fourth key, which is usually called the private key. And the above fourth key is used to decrypt the received ciphertext. The details will be introduced below. The above homomorphic encryption method that satisfies additive homomorphism and supports multiplication of ciphertext and plaintext can be the Pallier homomorphic encryption method.
[0082] Figure 1 In step S130 of, the second server receives the sum of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector sent by the first server, and then decrypts the above two results respectively to obtain the first plaintext sum of the symmetric matrix including the influence of the pseudo-sample symmetric matrix and the first plaintext sum of the product vector including the influence of the pseudo-sample product vector; thereafter, the second server uses the above two decrypted results to calculate the pseudo-sample encrypted linear regression coefficients including the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector, and sends the above pseudo-sample encrypted linear regression coefficients to the first server. Step S130 includes but is not limited to the following steps S410 to S430:
[0083] Step S410, the second server receives the sum of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector sent by the first server, and then decrypts the above two ciphertext sums respectively to obtain the sum of the plaintexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix and the sum of the plaintexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector.
[0084] Specifically, the second server receives the sum C′ of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix sent by the first server, and the sum of the ciphertexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector The second server uses the second key to perform homomorphic decryption on the above C′ to obtain the sum C of the plaintexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix, and for the above Perform homomorphic decryption to obtain the sum of the plaintexts of the original product vectors of multiple data owners that includes the influence of the product vectors of the pseudo-samples The above-mentioned second key is usually the private key used in the homomorphic encryption technology that supports addition calculation on ciphertexts, and the private key is used to decrypt the received ciphertexts. The sum C of the plaintexts of the original symmetric matrices of multiple data owners that includes the influence of the pseudo-sample symmetric matrices is only encrypted by the pseudo-sample symmetric matrices related to the pseudo-data samples, and there is no more homomorphic encryption; on the other hand, the sum of the plaintexts of the original product vectors of multiple data owners that includes the influence of the product vectors of the pseudo-samples is only encrypted by the pseudo-sample product vectors related to the pseudo-data samples, and there is no more homomorphic encryption.
[0085] To explain the effect of this step, it is necessary to introduce the technical principle related to the addition calculation of ciphertexts in the well-known homomorphic encryption technology, that is, the result obtained by performing addition calculation on multiple ciphertexts is equivalent to the result obtained by performing homomorphic encryption after performing addition calculation on the corresponding multiple plaintexts. Correspondingly, the sum of the ciphertexts of the original symmetric matrices of multiple data owners obtained in step S310 is equal to the ciphertext obtained after homomorphic encryption of the sum of the plaintexts of the original symmetric matrices of multiple data owners, and the sum of the ciphertexts of the original product vectors of multiple data owners obtained in step S310 is equal to the ciphertext obtained after homomorphic encryption of the sum of the plaintexts of the original product vectors of multiple data owners. The sum of the ciphertexts of the original symmetric matrices of the above-mentioned multiple data owners is the sum of the ciphertexts of the original symmetric matrices of the regularized M data owners, that is The sum of the plaintexts of the original symmetric matrices of the above-mentioned multiple data owners is the sum of the plaintexts of the original symmetric matrices of the regularized M data owners, that is The sum of the ciphertexts of the original symmetric matrices of the above-mentioned multiple data owners is equal to the ciphertext obtained after homomorphic encryption of the sum of the plaintexts of the original symmetric matrices of multiple data owners, and can be expressed as On the other hand, the sum of the ciphertexts of the original product vectors of the above-mentioned multiple data owners is the sum of the ciphertexts of the original product vectors of the M data owners, that is The sum of the plaintexts of the original product vectors of the above-mentioned multiple data owners is the sum of the plaintexts of the original product vectors of the M data owners, that is The sum of the original product vectors of the above-mentioned multiple data owners is equal to the ciphertext obtained after homomorphic encryption of the sum of the plaintexts of the original product vectors of multiple data owners, and can be expressed as
[0086] This step includes an optional regularization process. The optional regularization process introduced in step S330 can be moved to this step. That is, when the optional regularization process is not adopted in step S330, the sum C of the plaintexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix does not include the regularization term λI. Correspondingly, the regularization process can be performed in this step, that is, adding the regularization term λI. More generally, the regularization process can also be performed in step S330 and in this step at the same time. The final effect of the regularization process is the superposition of the effects of the regularization processes in the two steps. The related implementation methods are easy to obtain and will not be elaborated here.
[0087] As mentioned above, the result obtained by performing an addition calculation on multiple ciphertexts in the homomorphic encryption technology is equivalent to the result obtained by performing a homomorphic encryption after performing an addition calculation on the corresponding multiple plaintexts. Correspondingly, the sum of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix obtained in step S330 is equal to the ciphertext obtained after performing a homomorphic encryption on the sum of the plaintexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix. The sum of the ciphertexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector obtained in step S330 is equal to the ciphertext obtained after performing a homomorphic encryption on the sum of the plaintexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector. The sum of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix is the ciphertext C′ = A′ + φ′, that is, the ciphertext obtained by adding the first sum of symmetric matrix ciphertexts A′ and the processing result φ′ of the pseudo-sample symmetric matrix. The sum of the plaintexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix is C = A + φ, that is, the plaintext obtained by adding the first sum of symmetric matrix plaintexts A and the pseudo-sample symmetric matrix φ. The sum of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix is equal to the ciphertext obtained after performing a homomorphic encryption on the sum of the plaintexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix, which can be expressed as C′ = Enc key1 (C) = Enc key1 (A + φ). On the other hand, the sum of the ciphertexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector is the ciphertext That is, the first sum of product vector ciphertexts Plus the processing result of the pseudo-sample product vector The obtained ciphertext; the sum of the plaintexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector is the plaintext That is, the first sum of product vector plaintexts Plus the pseudo-sample product vector The obtained plaintext; the sum of the original product vector ciphertexts of multiple data owners including the influence of the pseudo-sample product vector is equal to the ciphertext obtained after the sum of the original product vector plaintexts of multiple data owners including the influence of the pseudo-sample product vector is homomorphically encrypted, and can be expressed as
[0088] The implementation manner of this step, that is, step S410, is described in detail below. The sum C of the original symmetric matrix plaintexts of multiple data owners including the influence of the pseudo-sample symmetric matrix is dec key2 (C′), and the sum of the original product vector plaintexts of multiple data owners including the influence of the pseudo-sample product vector where the second key key2 is the private key used in the additive homomorphic encryption technology that supports addition calculation on ciphertexts, and dec key2 (Δ) represents decrypting the ciphertext Δ therein based on the second key. As described above, only the second server needs to decrypt with the second key. Therefore, in this embodiment, usually only the second server holds the second key to achieve a better confidentiality effect. The above C = dec key2 (C′) can be derived from C′ = Enc key1 (C) = Enc key1 (A + φ) introduced earlier. By dec key2 (C′) = dec key2 (Enc key1 (C)) = dec key2 (Enc key1 (A + φ)) = C = A + φ; and the above can be derived from the introduced earlier, through
[0089] From the above description, it can be known that the sum C of the original symmetric matrix plaintexts of multiple data owners including the influence of the pseudo-sample symmetric matrix satisfies C = A + φ, that is, it is the result of the sum A of the first symmetric matrix plaintexts plus the pseudo-sample symmetric matrix φ; and the sum of the original product vector plaintexts of multiple data owners including the influence of the pseudo-sample product vector satisfies that is, it is the result of the sum of the first product vector plaintexts plus the pseudo-sample product vector .
[0090] It can be easily seen that although the second server decrypts the ciphertext of the homomorphic encryption, it does not know the results A and calculated from the private data of the M data owners because there is an influence of the pseudo-sample symmetric matrix φ in the matrix C decrypted by the second server, and the vector There are pseudo-sample product vectors in it influence.
[0091] Step S420: The second server calculates the pseudo-sample encrypted linear regression coefficients including the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector using the sum of the plaintexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix and the sum of the plaintexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector, and sends the above pseudo-sample encrypted linear regression coefficients to the first server.
[0092] Specifically, the second server uses the sum C of the plaintexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix obtained in the previous step and the sum First, calculate the inverse matrix G = C of the sum C of the plaintexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix -1 , and then use the above inverse matrix G = C abbreviated as the first inverse matrix -1 and the sum of the plaintexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector to calculate the pseudo-sample encrypted linear regression coefficient vector including the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector Then send the above pseudo-sample encrypted linear regression coefficient vector to the first server. The above pseudo-sample encrypted linear regression coefficient vector includes the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector, which can also be called including the influence of the foregoing τ pseudo-data samples.
[0093] The implementation manner of this step is described in detail below. From the above it is easy to see that includes the influence of the pseudo-sample symmetric matrix φ and the pseudo-sample product vector obtained from the pseudo-data samples, and the target linear regression coefficient vector to be finally obtained. Therefore, in the subsequent steps, it is necessary to at least partially eliminate the influence of the pseudo-data samples in the above pseudo-sample encrypted linear regression coefficient vector ; and theoretically, it is to completely eliminate the influence of the pseudo-sample symmetric matrix φ and the pseudo-sample product vector obtained from the pseudo-data samples.
[0094] Step S430: The second server receives the encrypted pseudo-input matrix ciphertext and then uses the first inverse matrix, that is, the matrix G = C -1 , to obtain the product of the encrypted pseudo-input matrix ciphertext and the first inverse matrix Then, send the above product Ξ to the first server. The above product Ξ is simply referred to as the product of the pseudo-input matrix ciphertext and the first inverse matrix.
[0095] For the two encryption methods given in step S340, there are corresponding two implementation manners in this step, as follows:
[0096] Implementation manner 1: When the first encryption method is used in step S340, that is, multiplication encryption is performed using the multiplication encryption mask matrix Θ, then the product of the pseudo-input matrix ciphertext and the first inverse matrix is the product of the pseudo-input matrix ciphertext and the first inverse matrix, that is
[0097] Implementation manner 2: When the second encryption method is used in step S340, that is, a homomorphic encryption method that satisfies additive homomorphism and supports multiplication of ciphertext and plaintext, then the second server uses the techniques of multiplication of ciphertext and plaintext and addition of ciphertext supported by this homomorphic encryption method to calculate the product of the pseudo-input matrix ciphertext and the first inverse matrix; specifically, the second server calculates the pseudo-input matrix ciphertext and the product of the plaintext G of the first inverse matrix, that is Here, the calculation includes the multiplication operation of ciphertext and plaintext in homomorphic encryption and the addition operation of ciphertext in homomorphic encryption, and uses the techniques of addition of ciphertext and multiplication of ciphertext and plaintext included in the homomorphic encryption method that satisfies additive homomorphism and supports multiplication of ciphertext and plaintext.
[0098] An embodiment of the present application Figure 1 In the method step S140, the first server receives the pseudo-sample encrypted linear regression coefficients including the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector, and eliminates the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector at least partially to obtain the target linear regression coefficients. Step S140 includes but is not limited to the following steps S510 to S530:
[0099] Step S510, after the first server receives the product Ξ of the pseudo-input matrix ciphertext and the first inverse matrix sent by the second server, it performs corresponding decryption operations to obtain the product of the pseudo-input matrix plaintext and the first inverse matrix, that is, the product of the pseudo-input matrix plaintext and the inverse matrix of the sum of the original product vector plaintexts of multiple data owners including the influence of the pseudo-sample symmetric matrix.
[0100] Specifically, the first server obtains the matrix Ψ that satisfies Ψ = EC -1 That is, the product of the pseudo-input matrix plaintext and the first inverse matrix. Here, the matrix Ψ is the inverse matrix C of the sum of the original symmetric matrix plaintexts of multiple data owners including the influence of the pseudo-sample symmetric matrix and the pseudo-input matrix plaintext E -1The product. For the two encryption methods given in step S340, there are also two corresponding implementation manners in this step, as follows:
[0101] Implementation manner 1: When the first server encrypts the pseudo-input matrix, i.e., matrix E, by using the multiplication encryption mask matrix Θ in step S340, the first server calculates the inverse matrix Θ -1 of the multiplication encryption mask matrix Θ and the product of the ciphertext of the pseudo-input matrix and the first inverse matrix Ξ to obtain the product Ψ of the plaintext of the pseudo-input matrix and the first inverse matrix, i.e., calculate Ψ = Θ -1 Ξ.
[0102] Implementation manner 2: When the first server uses a homomorphic encryption method that satisfies additive homomorphism and supports multiplication of ciphertext and plaintext to homomorphically encrypt the pseudo-input matrix E by using the third key in step S340, the first server uses the corresponding homomorphic decryption method to homomorphically decrypt the product of the ciphertext of the pseudo-input matrix and the first inverse matrix, i.e., matrix Ξ, by using the fourth key key4 to obtain the product Ψ of the plaintext of the pseudo-input matrix and the first inverse matrix, i.e., matrix Ψ = dec key4 (Ξ), where the fourth key key4 is usually the private key used for decryption in a homomorphic encryption method that satisfies additive homomorphism and supports multiplication of ciphertext and plaintext.
[0103] Step S520: The first server receives the pseudo-sample encrypted linear regression coefficients containing the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector, and eliminates the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector in at least part of them to obtain the intermediate linear regression coefficients.
[0104] Specifically, the first server receives the vector of pseudo-sample encrypted linear regression coefficients containing the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector obtained from the pseudo-data samples sent by the second server and then uses the product Ψ of the plaintext of the pseudo-input matrix and the first inverse matrix obtained in the previous step, as well as the pseudo-input matrix E and the pseudo-label column vector obtained by the first server in step S320 to calculate the intermediate linear regression coefficients The above calculated theoretically eliminates the influence of the pseudo-data samples in the vector of pseudo-sample encrypted linear regression coefficients containing the influence of the pseudo-data samples, thereby preliminarily obtaining the target linear regression coefficient vector to be sought Actually, considering the influence of various factors including calculation errors, the above calculated at least partially eliminates the influence of the pseudo-data samples in the vector of pseudo-sample encrypted linear regression coefficients containing the influence of the pseudo-data samples
[0105] When step S330 adopts implementation mode 1, 2, 3 or 4 described in this step, then this step also correspondingly adopts one of the four implementation modes of this step. The four implementation modes of this step are implementation mode 1, 2, 3 and 4 described below:
[0106] When the method in implementation mode 1 is used in step S330, the calculation formula for the intermediate linear regression coefficient is Or
[0107] When the method in implementation mode 2 is used in step S330, the calculation formula for the intermediate linear regression coefficient is Or
[0108] When the method of matrix multiplication in implementation mode 3 is used in step S330, the calculation formula for the intermediate linear regression coefficient is Or
[0109] When the method of matrix multiplication in implementation mode 4 is used in step S330, the calculation formula for the intermediate linear regression coefficient is Or
[0110] In step S530, usually the first server performs some processing related to the above homomorphic encryption and homomorphic decryption on the intermediate linear regression coefficient obtained in the previous step After that, the target linear regression coefficient vector to be obtained is finally obtained The above-mentioned Satisfies The processing related to homomorphic encryption and homomorphic decryption is well-known prior art. In some cases, the intermediate linear regression coefficient obtained in the previous step Is equal to the target linear regression coefficient vector The summary of the confidentiality effect of this embodiment is as follows: From the above implementation process, it can be known that neither the first server nor the second server knows the private data of the data owner, including the symmetric matrix first plaintext sum A and the product vector first plaintext sum calculated from the private data Thus achieving the effect of protecting data privacy.
[0111] In the first embodiment, the data owner uses the homomorphic encryption technology that supports addition calculation on ciphertexts to encrypt the private data. The data owner can also use other encryption technologies to encrypt their private data as long as the encryption technology also supports addition calculation on ciphertexts. The encryption technology that supports addition calculation on ciphertexts can generally be described as: the result of adding multiple ciphertexts obtained by encrypting with this encryption technology is equal to the result obtained by encrypting the sum of the corresponding multiple plaintexts with this encryption technology; generally, the encryption technology that supports addition calculation on ciphertexts satisfies that according to the result of adding multiple ciphertexts obtained by encrypting with this encryption technology, the result obtained by encrypting the sum of the corresponding multiple plaintexts with this encryption technology can be obtained.
[0112] The following gives the second embodiment for the scenario where the data owner uses a mask to encrypt their private data. In one embodiment of the present application, Figure 1 In step S110, each of at least two data owners uses a number of private data samples to obtain the ciphertext related to the original symmetric matrix and the ciphertext of the original product vector, and sends the above two ciphertexts to the first server.
[0113] Regarding the differences between the second embodiment and the first embodiment in the above text, the second embodiment modifies the method steps in the first embodiment. The differences between the second embodiment and the first embodiment include that each data owner in the second embodiment uses the mask encryption method to encrypt the private data of the grid respectively. First, the second embodiment modifies steps S210 to S220 in the first embodiment to steps S610 to S620, that is, in the second embodiment, before step S110, it includes but is not limited to the following steps S610 to S620:
[0114] Step S610, each of at least two data owners sends the mask input matrix U k and the mask label column vector contained in the mask data sample to the second server, or sends the mask symmetric matrix and the mask product vector to the second server.
[0115] The masked data samples, masked input matrices, and masked label column vectors used by the above-mentioned respective data owners are actually respectively similar to the pseudo-data samples, pseudo-input matrices, and pseudo-label column vectors used by the first server in the first embodiment. Here, different names are used to better distinguish them from the pseudo-data samples, pseudo-input matrices, and pseudo-label column vectors used by the first server in the first embodiment and this embodiment. The following is a detailed introduction to this step. Each of the M data owners needs to execute this step. The following gives the specific implementation for the k-th data owner among the M data owners, where M is an integer greater than or equal to 2, and 1 ≤ k ≤ M. The k-th data owner obtains ρ k masked data samples as masks for encryption, where ρ k is greater than or equal to 1. The ρ k masked data samples can be represented as a masked data sample set This set R k contains the i-th masked data sample which can be represented as where 1 ≤ i ≤ ρ k . The above i-th masked data sample, i.e., in, represents the masked input feature column vector contained in the i-th masked data sample, and v i represents corresponding masked label, which can be called the masked label contained in the i-th masked data sample; the above masked input feature column vector contains δ terms. The above ρ k masked input feature column vectors are represented as a matrix of ρ k rows and δ columns The above ρ k pseudo-labels v i (i = 1, 2,..., τ) are represented as a column vector containing ρ k terms The above matrix U k includes ρ k rows, and the i-th row is the transpose of the pseudo-input feature column vector In this application document, the matrix U is called the masked input feature matrix of the k-th data owner, and is simply referred to as the masked input matrix; the above column vector k includes ρ terms, and the i-th term is the masked label v k (i = 1, 2,..., ρ i ), in this application document, the column vector k is called the masked label column vector of the k-th data owner.
[0116] The masked symmetric matrix Λ of the k-th data owner mentioned above k is the cumulative result of the product of each masked input feature column vector in the masked data sample of the k-th data owner and its transpose, and is for ρ k each of the ρ obtained the product of its transpose and then the cumulative result obtained by accumulating the ρ Here, 1 ≤ i ≤ ρ k ones is the cumulative result, that is Here, 1 ≤ i ≤ ρ k ; When represented in matrix form, the above-mentioned is the symmetric matrix obtained by multiplying the transpose of the masked input matrix U k by itself, simply referred to as the masked symmetric matrix. On the other hand, the masked product vector of the k-th data owner mentioned above is the cumulative result of the product of each masked input feature column vector in the masked data sample of the k-th data owner and the corresponding masked label, and is for ρ k each of the ρ obtained the product with the corresponding masked label v i and then the cumulative result obtained by accumulating the ρ Here, 1 ≤ i ≤ ρ k ones is the cumulative result, that is Here, 1 ≤ i ≤ ρ k ; When represented in matrix form, the above-mentioned is the product of the transpose of the masked input matrix U k and the masked label column vector , simply referred to as the masked product vector.
[0117] This step can also have other implementation manners. One other implementation manner is that the second server sends the masked input matrix U k contained in the masked data sample and the masked label column vector to the k-th data owner, or sends the masked symmetric matrix and the masked product vector to the k-th data owner. A simpler implementation manner includes that the second server and the k-th data owner save the same codebook containing multiple matrices and vectors, and any one of the two parties sends the serial numbers of a matrix and a vector in the codebook to the other party to determine the masked input matrix U k and the masked label column vector or the masked symmetric matrix and the masked product vector
[0118] Step S620: Each of at least two data owners encrypts to obtain the ciphertext related to its own original symmetric matrix and the ciphertext of the original product vector based on the technology of encrypting its own private data using a mask; each data owner sends the ciphertext related to its own original symmetric matrix and the ciphertext of the original product vector to the first server. Each of the M data owners needs to execute this step. The implementation manner is given below for the k-th data owner among the M data owners, where 1 ≤ k ≤ M.
[0119] The ciphertext related to the above original symmetric matrix obtained by the k-th data owner can have various forms, as long as it satisfies that the original symmetric matrix can be calculated from the ciphertext related to the original symmetric matrix. The forms of the ciphertext related to the original symmetric matrix include the following two ciphertext forms:
[0120] Ciphertext form 1): First, obtain the concatenated matrix of the k-th data owner's original input matrix and the k-th data owner's mask input matrix Then, obtain the above concatenated matrix Γ k and the product Π k with any orthogonal matrix Θ k = Γ k Θ k . The product of the above concatenated matrix and the orthogonal matrix is used as the ciphertext related to the original symmetric matrix. It is easy to see that the above concatenated matrix Γ k is composed of each original input feature column vector and each mask input feature column vector of the k-th data owner; when represented in matrix form, the above concatenated matrix Γ k is composed of the k-th data owner's original input matrix X k and the mask input matrix U k .
[0121] Ciphertext form 2): Obtain the sum of the original symmetric matrix A k and the mask symmetric matrix , that is, matrix A' k = A k + Λ k , as the ciphertext related to the original symmetric matrix. It is easy to see
[0122] The above ciphertext form 2 can be called the ciphertext of the original symmetric matrix, that is, the ciphertext of the original symmetric matrix is the sum of the original symmetric matrix and the mask symmetric matrix.
[0123] On the other hand, the ciphertext of the above-mentioned original product vector obtained by the k-th data owner is the sum of the k-th data owner's original product vector and the k-th data owner's masked product vector , that is, the column vector It is easy to see that
[0124] It can be understood that before sending the intermediate result calculated from the original data sample owned by itself to the first server, the data owner encrypts the intermediate result based on the technology of encrypting its own private data using a mask, which can ensure that the first server cannot know the true situation of the intermediate result; in addition, the data owner has the intermediate result calculated from the private data, and the first server obtains the ciphertext of the intermediate result after mask encryption. The mask-related information (i.e., the mask input matrix and the mask label column vector, or the mask symmetric matrix and the mask product vector) used to decrypt the ciphertext of the mask encryption is stored in the second server. Therefore, except for the data owner who holds the intermediate result calculated from its own private data, the first server cannot know the above intermediate result, thus ensuring the security of the data owner's private data. In addition, as will be described below, the first server further encrypts the ciphertext of the intermediate result encrypted by the mask before sending it to the second server; correspondingly, after decrypting the received ciphertext with the mask-related information, the second server still cannot know the above intermediate result because of the above further encryption performed by the first server. As can be seen from the above, neither the first server nor the second server can know the above intermediate result calculated from the private data of the data owner through the information they receive, thus protecting the privacy of the data owner. In the subsequent steps of this embodiment, the first server receives the ciphertext related to the original symmetric matrix and the ciphertext of the original product vector sent by each of at least two data owners, and cooperates with the second server to obtain the linear regression coefficient. The subsequent steps of this embodiment include step 120, step S130, and step S140, which will be described below.
[0125] Figure 1In method step S110 of [the above], the first server receives the ciphertexts related to the original symmetric matrices and the ciphertexts of the original product vectors sent by each of at least two data owners, processes the ciphertexts related to the original symmetric matrices sent by the at least two data owners to obtain the sum of the ciphertexts of the original symmetric matrices of the multiple data owners, and processes the ciphertexts of the original product vectors sent by the at least two data owners to obtain the sum of the ciphertexts of the original product vectors of the multiple data owners; on the other hand, the first server obtains the symmetric matrix and the product vector corresponding to the pseudo-data sample, which are respectively referred to as the pseudo-sample symmetric matrix and the pseudo-sample product vector; then, the first server adds the processing result of the pseudo-sample symmetric matrix to the sum of the ciphertexts of the original symmetric matrices of the multiple data owners to obtain the first ciphertext sum of the symmetric matrix including the influence of the pseudo-sample symmetric matrix. In addition, the first server adds the processing result of the pseudo-sample product vector to the sum of the ciphertexts of the original product vectors of the multiple data owners to obtain the first ciphertext sum of the product vector including the influence of the pseudo-sample product vector; finally, the first server sends the sum of the ciphertexts of the original symmetric matrices of the multiple data owners including the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of the multiple data owners including the influence of the pseudo-sample product vector to the second server; for the characteristics of Embodiment 2 relative to Embodiment 1, Embodiment 2 modifies steps S310 to S340 in Embodiment 1 to obtain steps S710 to S740 in Embodiment 2. That is, in Embodiment 2, step S110 includes but is not limited to the following steps S710 to S740:
[0126] Step S710, the first server receives the ciphertexts related to the original symmetric matrices and the ciphertexts of the original product vectors sent by each of at least two data owners; the first server performs an accumulation process and an optional regularization process on the ciphertexts related to the original symmetric matrices sent by the at least two data owners to obtain the sum of the ciphertexts of the original symmetric matrices of the multiple data owners, which is simply referred to as the first ciphertext sum of the symmetric matrix; the first server also performs an accumulation process on the ciphertexts of the original product vectors sent by the at least two data owners to obtain the sum of the ciphertexts of the original product vectors of the multiple data owners, which is simply referred to as the first ciphertext sum of the product vector.
[0127] In step S620, the ciphertexts related to the original symmetric matrices sent by the k-th data owner (1 ≤ k ≤ M) have two different ciphertext forms. Correspondingly, the method for processing the ciphertexts related to the original symmetric matrices in this step also includes the following two preprocessing steps to obtain the preprocessing result:
[0128] The first preprocessing step of the ciphertexts related to the original symmetric matrix) In step S620, the k-th data owner adopts ciphertext form 1, that is Then the first server calculates the preprocessing result of the ciphertexts related to the original symmetric matrix as It is easy to see A' k satisfies The second preprocessing step of the ciphertext related to the original symmetric matrix) The k-th data owner adopts ciphertext form 2 in step S620, that is Then the preprocessing step of the first server does not require any operation, and directly uses the above ciphertext form 2 as the preprocessing result of the ciphertext related to the original symmetric matrix.
[0129] As mentioned above, the above ciphertext form 2 can be called the ciphertext of the original symmetric matrix. Then it is easy to see that the two preprocessing results obtained by the above two preprocessing steps are both equal to the ciphertext of the original symmetric matrix.
[0130] After the above preprocessing step, the first server performs an accumulation and optional regularization process on the preprocessing results of the ciphertexts related to the original symmetric matrices of M data owners, that is, the ciphertexts of the original symmetric matrices, to obtain the sum of the ciphertexts of the original symmetric matrices of multiple data owners (simply referred to as the first ciphertext sum of the symmetric matrix). The above regularization process can also be performed in other steps. For details, please refer to the subsequent description. On the other hand, the ciphertext of the original product vector received by the first server is the ciphertext of the original product vector Thus, the first server accumulates the ciphertexts of the original product vectors of M data owners received, that is, the ciphertext of the product of the original input matrix and the original label column vector to obtain the sum of the ciphertexts of the original product vectors of multiple data owners (simply referred to as the first ciphertext sum of the product vector).
[0131] Step S720, the first server uses a number of pseudo-data samples to obtain the corresponding symmetric matrix and product vector of the pseudo-data samples, which are respectively called the pseudo-sample symmetric matrix and the pseudo-sample product vector.
[0132] This step remains unchanged compared with the corresponding step in Embodiment 1. For relevant details, please refer to the corresponding step in Embodiment 1.
[0133] Step S730, the first server adds the first ciphertext sum of the symmetric matrix to the processing result of the pseudo-sample symmetric matrix to obtain the sum of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix; the first server adds the first ciphertext sum of the product vector to the processing result of the pseudo-sample product vector to obtain the sum of the ciphertexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector. The first server sends the sum of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector to the second server.
[0134] The processed result φ' of the pseudo-sample symmetric matrix is the pseudo-sample symmetric matrix (i.e., the above-mentioned pseudo-sample symmetric matrix φ = E T E), that is, φ' = φ, or it is the result of taking the negative sign of the pseudo-sample symmetric matrix, that is, φ' = -φ; the processed result of the pseudo-sample product vector is the pseudo-sample product vector (i.e., the above-mentioned pseudo-sample product vector ), that is or it is the result of taking the negative sign of the pseudo-sample product vector, that is
[0135] The first server adds the sum of the first ciphertexts of the symmetric matrix to the processed result of the pseudo-sample symmetric matrix to obtain the sum of the original symmetric matrix ciphertexts of multiple data owners including the influence of the pseudo-sample symmetric matrix, and adds the sum of the first ciphertexts of the product vector to the processed result of the pseudo-sample product vector to obtain the sum of the original product vector ciphertexts of multiple data owners including the influence of the pseudo-sample product vector. Specifically, the first server adds the sum of the first ciphertexts of the symmetric matrix A' to the processed result φ' of the pseudo-sample symmetric matrix to obtain the sum of the original symmetric matrix ciphertexts of multiple data owners including the influence of the pseudo-sample symmetric matrix C' = A' + φ'; and adds the sum of the first ciphertexts of the product vector to the processed result of the pseudo-sample product vector to obtain the sum of the original product vector ciphertexts of multiple data owners including the influence of the pseudo-sample product vector Then, the first server sends the sum of the original symmetric matrix ciphertexts of multiple data owners including the influence of the pseudo-sample symmetric matrix C' and the sum of the original product vector ciphertexts of multiple data owners including the influence of the pseudo-sample product vector to the second server.
[0136] This step uses intermediate variables related to the pseudo-data sample, that is, the processed result of the pseudo-sample symmetric matrix and the processed result of the pseudo-sample product vector, to further encrypt the ciphertexts A' and . The processed result of the pseudo-sample symmetric matrix can be φ' = φ or φ' = -φ, and the processed result of the pseudo-sample product vector can be or Correspondingly, this step can include the following four implementation manners:
[0137] Implementation manner 1: The sum of the original symmetric matrix ciphertexts of multiple data owners including the influence of the pseudo-sample symmetric matrix C' = A' + φ' = A' + φ, and the sum of the original product vector ciphertexts of multiple data owners including the influence of the pseudo-sample product vector
[0138] Embodiment 2: The sum C' of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix is C' = A' + φ' = A' - φ, and the sum of the ciphertexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector
[0139] Embodiment 3: The sum C' of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix is C' = A' + φ' = A' + φ, and the sum of the ciphertexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector
[0140] Embodiment 4: The sum C' of the ciphertexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix is C' = A' + φ' = A' - φ, and the sum of the ciphertexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vector
[0141] Both the above step S710 and step S730 are operations of adding ciphertexts. It is that the first server first accumulates the ciphertexts of the original symmetric matrices sent by at least two data owners and then adds the processing result of the pseudo-sample symmetric matrix, and first accumulates the ciphertexts of the original product vectors sent by at least two data owners and then adds the processing result of the pseudo-sample product vector. In fact, the order of adding the above ciphertexts can be adjusted arbitrarily; for example, the first server first accumulates the ciphertexts of the original symmetric matrices sent by one or any number of data owners, then adds the processing result of the pseudo-sample symmetric matrix, and finally accumulates the ciphertexts of the original symmetric matrices sent by the remaining several data owners; on the other hand, the first server first accumulates the ciphertexts of the original product vectors sent by one or any number of data owners, then adds the processing result of the pseudo-sample product vector, and finally accumulates the ciphertexts of the original product vectors sent by the remaining several data owners.
[0142] It can be understood that in the second embodiment, since all data owners encrypt their respective private data using masked encryption, therefore, the sum of all the first ciphertexts of the symmetric matrices received by the first server is obtained by masked encryption, and the sum of all the first ciphertexts of the product vectors is also obtained by masked encryption. Therefore, when summing the ciphertexts of the original symmetric matrices based on different encryption methods, the obtained result contains the influence of the private data of all data owners, and when summing the original product vectors based on the encryption method, the obtained result contains the influence of the private data of all data owners. Therefore, in the second embodiment, the sum of the original symmetric matrix ciphertexts of multiple data owners containing the influence of the pseudo-sample symmetric matrix obtained by the first server summation is equivalent to the sum of the original symmetric matrix ciphertexts of all data owners containing the influence of the pseudo-sample symmetric matrix, and the sum of the original product vector ciphertexts of multiple data owners containing the influence of the pseudo-sample symmetric matrix obtained by the first server summation is equivalent to the sum of the original product vector ciphertexts of all data owners containing the influence of the pseudo-sample symmetric matrix.
[0143] Step S740, the first server encrypts the pseudo-input matrix E composed of several pseudo-input feature column vectors to obtain the ciphertext of the pseudo-input matrix Then the ciphertext of the pseudo-input matrix is sent to the second server; the above-mentioned several pseudo-input feature column vectors are part of the corresponding several pseudo-data samples, as described above.
[0144] This step remains unchanged compared to the corresponding step in the first embodiment. For relevant details, please refer to the corresponding step in the first embodiment.
[0145] Figure 1 In step S130 of [description of relevant content], the second server receives the sum of the original symmetric matrix ciphertexts of multiple data owners containing the influence of the pseudo-sample symmetric matrix and the sum of the original product vector ciphertexts of multiple data owners containing the influence of the pseudo-sample product vector sent by the first server, and then decrypts the above two results respectively to obtain the sum of the original symmetric matrix plaintexts of multiple data owners containing the influence of the pseudo-sample symmetric matrix and the sum of the original product vector plaintexts of multiple data owners containing the influence of the pseudo-sample product vector; thereafter, the second server uses the above two decrypted results to calculate the pseudo-sample encrypted linear regression coefficients containing the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector, and sends the above pseudo-sample encrypted linear regression coefficients to the first server. For the characteristics of the second embodiment relative to the first embodiment, the second embodiment modifies steps S410 to S430 in the first embodiment to obtain steps S810 to S830 in the second embodiment. That is, in the second embodiment, step S130 includes but is not limited to the following steps S810 to S830:
[0146] Step S810: The second server receives the sum of the original symmetric matrix ciphertexts of multiple data owners containing the influence of the pseudo-sample symmetric matrix and the sum of the original product vector ciphertexts of multiple data owners containing the influence of the pseudo-sample product vector sent by the first server, and then decrypts the above two ciphertext sums respectively to obtain the sum of the original symmetric matrix plaintexts of multiple data owners containing the influence of the pseudo-sample symmetric matrix and the sum of the original product vector plaintexts of multiple data owners containing the influence of the pseudo-sample product vector.
[0147] To decrypt the above two ciphertext sums, the second server receives the masked input matrix U sent by at least two data owners respectively k and the masked label column vector or receives the masked symmetric matrix sent by at least two data owners respectively and the masked product vector and then uses them to decrypt the sum of the original symmetric matrix ciphertexts of multiple data owners containing the influence of the pseudo-sample symmetric matrix and the sum of the original product vector ciphertexts of multiple data owners containing the influence of the pseudo-sample product vector respectively. When step S610 uses other implementation manners, so that it is not the k-th data owner who sends the masked input matrix U k and the masked label column vector or the masked symmetric matrix and the masked product vector then this step also needs to use the corresponding implementation manner, and it is no longer the case that the second server receives the masked input matrix U k and the masked label column vector or the masked symmetric matrix and the masked product vector The corresponding implementation manner used in this step is obvious to those skilled in the art and will not be elaborated here.
[0148] Specifically, the second server decrypts the sum C' of the original symmetric matrix ciphertexts of multiple data owners containing the influence of the pseudo-sample symmetric matrix by calculating that is, subtracting the sum of the masked symmetric matrices of multiple data owners from the sum of the original symmetric matrix ciphertexts of multiple data owners containing the influence of the pseudo-sample symmetric matrix
[0149] And decrypting the sum of the original product vector ciphertexts of multiple data owners containing the influence of the pseudo-sample product vector is by calculating that is, subtracting the sum of the masked product vectors of multiple data owners from the sum of the original product vector ciphertexts of multiple data owners containing the influence of the pseudo-sample product vector It can be easily seen that the sum C of the plaintexts of the original symmetric matrices of multiple data owners calculated to include the influence of the pseudo-sample symmetric matrix satisfies And the sum of the plaintexts of the original product vectors of multiple data owners calculated to include the influence of the pseudo-sample product vectors Satisfies
[0150] This step includes an optional regularization process. The optional regularization process introduced in step S710 can be moved to this step. That is, when the optional regularization process is not adopted in step S710, the sum C of the plaintexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix does not include the regularization term λI. Correspondingly, the regularization process can be performed in this step, that is, adding the regularization term λI. More generally, the regularization process can also be performed in step S710 and at the same time in this step. Finally, the effect of the regularization process is the superposition of the effects of the regularization processes in the two steps. The relevant implementation methods are easily obtained and will not be elaborated here.
[0151] Step S820: The second server uses the sum of the plaintexts of the original symmetric matrices of multiple data owners including the influence of the pseudo-sample symmetric matrix and the sum of the plaintexts of the original product vectors of multiple data owners including the influence of the pseudo-sample product vectors to calculate the pseudo-sample encrypted linear regression coefficients including the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vectors, and sends the above pseudo-sample encrypted linear regression coefficients to the first server.
[0152] This step remains unchanged compared to the corresponding step in Embodiment 1. For relevant details, please refer to the corresponding step in Embodiment 1.
[0153] Step S830: The second server receives the ciphertext of the pseudo-input matrix sent by the first server Then use the first inverse matrix, that is, matrix G = C -1 To obtain the product of the ciphertext of the pseudo-input matrix And the first inverse matrix Then send the above product Ξ to the first server. The above product Ξ is simply referred to as the product of the ciphertext of the pseudo-input matrix and the first inverse matrix.
[0154] This step remains unchanged compared to the corresponding step in Embodiment 1. For relevant details, please refer to the corresponding step in Embodiment 1.
[0155] Embodiment 2 of this application Figure 1In the method step S140, the first server receives the pseudo-sample encrypted linear regression coefficients containing the influences of the pseudo-sample symmetric matrix and the pseudo-sample product vector, and at least partially eliminates the influences of the pseudo-sample symmetric matrix and the pseudo-sample product vector to obtain the target linear regression coefficients. In view of the characteristics of Embodiment 2 relative to Embodiment 1, Embodiment 2 modifies steps S510 to S530 in Embodiment 1 to obtain steps S910 to S920 in Embodiment 2. That is, in Embodiment 2, step S140 includes but is not limited to the following steps S910 to S920:
[0156] Step S910, after the first server receives the product Ξ of the encrypted pseudo-input matrix and the first inverse matrix sent by the second server, it performs corresponding decryption operations to obtain the product of the plaintext pseudo-input matrix and the first inverse matrix, that is, the product of the plaintext pseudo-input matrix and the inverse matrix of the sum of the original symmetric matrix ciphertexts of multiple data owners containing the influence of the pseudo-sample symmetric matrix.
[0157] This step remains unchanged compared to the corresponding step in Embodiment 1. For relevant details, please refer to the corresponding step in Embodiment 1.
[0158] Step S920, the first server receives the pseudo-sample encrypted linear regression coefficients containing the influences of the pseudo-sample symmetric matrix and the pseudo-sample product vector, and at least partially eliminates the influences of the pseudo-sample symmetric matrix and the pseudo-sample product vector to obtain the desired target linear regression coefficient vector The above Meet Note that since each data owner in Embodiment 2 only uses mask encryption and does not use homomorphic encryption, generally the first server no longer needs to perform the processing related to homomorphic encryption and homomorphic decryption described in step S530 of Embodiment 1.
[0159] In the above Embodiment 1 and Embodiment 2, each data owner uses the same encryption technology that supports addition calculation on ciphertexts to obtain the ciphertexts related to their respective original symmetric matrices and the ciphertexts of the original product vectors. Correspondingly, the sum of the original symmetric matrix ciphertexts of multiple data owners containing the influence of the pseudo-sample symmetric matrix obtained by the first server is the sum of the original symmetric matrix ciphertexts of all data owners containing the influence of the pseudo-sample symmetric matrix, and the sum of the original product vector ciphertexts of multiple data owners containing the influence of the pseudo-sample product vector obtained by the first server is the sum of the original product vector ciphertexts of all data owners containing the influence of the pseudo-sample product vector. In Embodiment 1, homomorphic encryption is used, and in Embodiment 2, the technology of adding a mask to the plaintext to obtain the ciphertext is used for encryption. In fact, each data owner can also use other encryption technologies that support addition calculation on ciphertexts other than the above homomorphic encryption and mask encryption.
[0160] Each data owner can also use different encryption techniques that support addition operations on ciphertexts to obtain ciphertexts related to their respective original symmetric matrices and ciphertexts of the original product vectors. The following briefly describes an example of each data owner using different encryption techniques that support addition operations on ciphertexts. Assume that some data owners use homomorphic encryption, while all the remaining data owners use masked encryption. Correspondingly, some data owners use the method of Embodiment 1, while all the remaining data owners use the method of Embodiment 2. Only a slight modification can be made to step S410 in Embodiment 1 or step S810 in Embodiment 2, and several other related steps can be slightly modified.
[0161] First, Embodiment 3 is briefly introduced. In Embodiment 3, at least two data owners respectively encrypt intermediate variables related to a number of their own private data samples using at least two different encryption techniques to obtain ciphertexts related to their respective original symmetric matrices and ciphertexts of the original product vectors. The at least two different encryption techniques include the homomorphic encryption technique and the encryption technique of obtaining ciphertexts by masking plaintexts used in Embodiment 1 and Embodiment 2 respectively; the sum of the ciphertexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix obtained by the first server is the sum of at least two ciphertexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix corresponding to at least two different encryption techniques respectively, and the sum of the ciphertexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector obtained by the first server is the sum of at least two ciphertexts of the original product vectors of data owners containing the influence of the pseudo-sample product vector corresponding to at least two different encryption techniques respectively; the first server sends the sum of the ciphertexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector to the second server;
[0162] The second server receives and decrypts the sums of the ciphertexts of the original symmetric matrices of multiple data owners corresponding to at least two different encryption techniques, each containing the influence of the pseudo-sample symmetric matrix, and the sums of the ciphertexts of the original product vectors of multiple data owners corresponding to at least two different encryption techniques, each containing the influence of the pseudo-sample product vector, using different decryption techniques corresponding to different encryption techniques, to obtain the sums of the plaintexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix and the sums of the plaintexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector. Then, the sum of the sums of the plaintexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix is accumulated to obtain the sum of the plaintexts of the original symmetric matrices of all data owners containing the influence of the pseudo-sample symmetric matrix, and the sum of the sums of the plaintexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector is accumulated to obtain the sum of the plaintexts of the original product vectors of all data owners containing the influence of the pseudo-sample product vector.
[0163] The difference between Embodiment 3 and Embodiment 1 and Embodiment 2 is that each data owner can independently choose to use homomorphic encryption, masked encryption, or other third encryption techniques that support addition calculation on ciphertexts to encrypt their own private data respectively, obtaining the ciphertexts of the original symmetric matrix and the ciphertexts of the original product vector corresponding to each encryption method respectively. After receiving the ciphertexts of the original symmetric matrix and the ciphertexts of the original product vector corresponding to each encryption method, the first server sums the ciphertext of the original symmetric matrix corresponding to one of the encryption methods with the pseudo-sample symmetric matrix, obtaining the sums of the ciphertexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix and the sums of the ciphertexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector corresponding to at least two encryption methods respectively. The first server sends the sums of the ciphertexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix and the sums of the ciphertexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector to the second server. The second server decrypts the sums of the ciphertexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix respectively based on the encryption method, obtaining the sum of the plaintexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix. The second server then decrypts the sums of the ciphertexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector respectively based on the encryption method, obtaining the sum of the plaintexts of the original product vectors of multiple data owners.
[0164] Embodiment 3 includes the following steps S1010, S1020, S1030, and S1040, and the functions of these steps are respectively similar to Figure 1Steps S110 to S140 therein will be briefly introduced below.
[0165] Step S1010: At least two data owners use at least two different encryption techniques to encrypt the intermediate variables related to a number of their respective private data samples, respectively, to obtain the ciphertexts related to their respective original symmetric matrices and the ciphertexts of the original product vectors. Each of the at least two data owners uses one encryption technique to encrypt the intermediate variables related to a number of their respective private data samples to obtain the ciphertexts related to the original symmetric matrix and the ciphertexts of the original product vector of this data owner; each data owner sends the ciphertexts of their respective original symmetric matrices and the ciphertexts of the original product vectors to the first server. The at least two different encryption techniques include the homomorphic encryption technique and the encryption technique of obtaining ciphertexts by masking plaintexts used in Embodiment 1 and Embodiment 2 respectively, and may also include other encryption techniques that support addition calculation on ciphertexts. The above encryption technique of obtaining ciphertexts by masking plaintexts is hereinafter briefly referred to as the masking encryption technique.
[0166] Step S1020 includes the following operations:
[0167] The first server receives the ciphertexts related to the original symmetric matrices and the ciphertexts of the original product vectors sent by each of the at least two data owners; the first server accumulates the ciphertexts related to the original symmetric matrices and the ciphertexts of the original product vectors obtained by using the same encryption technique respectively to obtain the sum of the ciphertexts of the original symmetric matrices of multiple data owners corresponding to this encryption technique and the sum of the ciphertexts of the original product vectors of multiple data owners corresponding to this encryption technique, which are respectively briefly referred to as the first ciphertext sum of the symmetric matrix corresponding to this encryption technique and the first ciphertext sum of the product vector corresponding to this encryption technique. That is to say, the first server obtains: the first ciphertext sum of the symmetric matrix corresponding to the homomorphic encryption technique, the first ciphertext sum of the product vector corresponding to the homomorphic encryption technique, the first ciphertext sum of the symmetric matrix corresponding to the masking encryption technique, the first ciphertext sum of the product vector corresponding to the masking encryption technique, the first ciphertext sum of the symmetric matrix corresponding to other encryption techniques that support addition calculation on ciphertexts, and the first ciphertext sum of the product vector corresponding to other encryption techniques that support addition calculation on ciphertexts.
[0168] On the other hand, the first server obtains the symmetric matrix and the product vector corresponding to the pseudo-data sample, which are respectively called the pseudo-sample symmetric matrix and the pseudo-sample product vector. Then, the first server respectively uses the methods introduced in the corresponding steps of the first and second embodiments to obtain the processing results of the pseudo-sample symmetric matrix and the pseudo-sample product vector corresponding to the homomorphic encryption technology, and the processing results of the pseudo-sample symmetric matrix and the pseudo-sample product vector corresponding to the mask encryption technology; in addition, the first server also obtains the processing results of the pseudo-sample symmetric matrix and the pseudo-sample product vector corresponding to other encryption technologies that support addition calculation on ciphertexts, and they satisfy that they can be added to the ciphertexts obtained by the other encryption technologies.
[0169] Finally, the first server adds the first ciphertext sum of the symmetric matrix corresponding to the homomorphic encryption technology to the processing result of the pseudo-sample symmetric matrix corresponding to the homomorphic encryption technology to obtain the sum of the original symmetric matrix ciphertexts of multiple data owners including the influence of the pseudo-sample symmetric matrix corresponding to the homomorphic encryption technology, adds the first ciphertext sum of the product vector corresponding to the homomorphic encryption technology to the processing result of the pseudo-sample product vector corresponding to the homomorphic encryption technology to obtain the sum of the original product vector ciphertexts of multiple data owners including the influence of the pseudo-sample product vector corresponding to the homomorphic encryption technology, adds the first ciphertext sum of the symmetric matrix corresponding to the mask encryption technology to the processing result of the pseudo-sample symmetric matrix corresponding to the mask encryption technology to obtain the sum of the original symmetric matrix ciphertexts of multiple data owners including the influence of the pseudo-sample symmetric matrix corresponding to the mask encryption technology, adds the first ciphertext sum of the product vector corresponding to the mask encryption technology to the processing result of the pseudo-sample product vector corresponding to the mask encryption technology to obtain the sum of the original product vector ciphertexts of multiple data owners including the influence of the pseudo-sample product vector corresponding to the mask encryption technology, adds the first ciphertext sum of the symmetric matrix corresponding to the other encryption technology to the processing result of the pseudo-sample symmetric matrix corresponding to the other encryption technology to obtain the sum of the original symmetric matrix ciphertexts of multiple data owners including the influence of the pseudo-sample symmetric matrix corresponding to the other encryption technology, adds the first ciphertext sum of the product vector corresponding to the other encryption technology to the processing result of the pseudo-sample product vector corresponding to the other encryption technology to obtain the sum of the original product vector ciphertexts of multiple data owners including the influence of the pseudo-sample product vector corresponding to the other encryption technology, and then sends the above-mentioned multiple ciphertext sums to the second server.
[0170] It is easy to see that the sum of the ciphertexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix obtained by the first server is the sum of the ciphertexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix corresponding to at least two different encryption techniques respectively, and the sum of the ciphertexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector obtained by the first server is the sum of the ciphertexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector corresponding to at least two different encryption techniques respectively; the first server sends the sum of the ciphertexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector to the second server.
[0171] Step S1030: The second server receives the sum of the ciphertexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix corresponding to at least two different encryption techniques respectively, and the sum of the ciphertexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector corresponding to at least two different encryption techniques respectively, sent by the first server; the second server uses the decryption techniques corresponding to the at least two different encryption techniques respectively to decrypt the above ciphertext sums to obtain the sum of the plaintexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix and the sum of the plaintexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector; then, the second server accumulates the sum of the plaintexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample symmetric matrix to obtain the sum of the plaintexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample symmetric matrix corresponding to all data owners, and accumulates the sum of the plaintexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector to obtain the sum of the plaintexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector corresponding to all data owners; the second server uses the sum of the plaintexts of the original symmetric matrices of multiple data owners containing the influence of the pseudo-sample symmetric matrix and the sum of the plaintexts of the original product vectors of multiple data owners containing the influence of the pseudo-sample product vector corresponding to all data owners to calculate the pseudo-sample encrypted linear regression coefficients containing the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector, and sends the above pseudo-sample encrypted linear regression coefficients to the first server. More details of this step are described below.
[0172] The second server receives the sum of the ciphertexts of the original symmetric matrices of multiple data owners corresponding to the homomorphic encryption technology, which contains the influence of the pseudo-sample symmetric matrix, the sum of the ciphertexts of the original symmetric matrices of multiple data owners corresponding to the masking encryption technology, which contains the influence of the pseudo-sample symmetric matrix, the sum of the ciphertexts of the original symmetric matrices of multiple data owners corresponding to other encryption technologies, which contains the influence of the pseudo-sample symmetric matrix, the sum of the ciphertexts of the original product vectors of multiple data owners corresponding to the homomorphic encryption technology, which contains the influence of the pseudo-sample product vector, the sum of the ciphertexts of the original product vectors of multiple data owners corresponding to the masking encryption technology, which contains the influence of the pseudo-sample product vector, and the sum of the ciphertexts of the original product vectors of multiple data owners corresponding to other encryption technologies, which contains the influence of the pseudo-sample product vector.
[0173] The second server uses the homomorphic decryption technology corresponding to the homomorphic encryption technology to decrypt the sum of the ciphertexts of the original symmetric matrices of multiple data owners corresponding to the homomorphic encryption technology, which contains the influence of the pseudo-sample symmetric matrix, and the sum of the ciphertexts of the original product vectors of multiple data owners corresponding to the homomorphic encryption technology, which contains the influence of the pseudo-sample product vector, to obtain the sum of the plaintexts of the original symmetric matrices of multiple data owners corresponding to the homomorphic encryption technology, which contains the influence of the pseudo-sample symmetric matrix, and the sum of the plaintexts of the original product vectors of multiple data owners corresponding to the homomorphic encryption technology, which contains the influence of the pseudo-sample product vector; the second server uses the decryption technology corresponding to the masking encryption technology to decrypt the sum of the ciphertexts of the original symmetric matrices of multiple data owners corresponding to the masking encryption technology, which contains the influence of the pseudo-sample symmetric matrix, and the sum of the ciphertexts of the original product vectors of multiple data owners corresponding to the masking encryption technology, which contains the influence of the pseudo-sample product vector, to obtain the sum of the plaintexts of the original symmetric matrices of multiple data owners corresponding to the masking encryption technology, which contains the influence of the pseudo-sample symmetric matrix, and the sum of the plaintexts of the original product vectors of multiple data owners corresponding to the masking encryption technology, which contains the influence of the pseudo-sample product vector; the second server uses the decryption technology corresponding to other encryption technologies to decrypt the sum of the ciphertexts of the original symmetric matrices of multiple data owners corresponding to other encryption technologies, which contains the influence of the pseudo-sample symmetric matrix, and the sum of the ciphertexts of the original product vectors of multiple data owners corresponding to other encryption technologies, which contains the influence of the pseudo-sample product vector, to obtain the sum of the plaintexts of the original symmetric matrices of multiple data owners corresponding to other encryption technologies, which contains the influence of the pseudo-sample symmetric matrix, and the sum of the plaintexts of the original product vectors of multiple data owners corresponding to other encryption technologies, which contains the influence of the pseudo-sample product vector.
[0174] Subsequently, the second server accumulates the sums of the plaintexts of the original symmetric matrices of multiple data owners corresponding to the above homomorphic encryption technology that include the influence of the pseudo-sample symmetric matrix, the sums of the plaintexts of the original symmetric matrices of multiple data owners corresponding to the masking encryption technology that include the influence of the pseudo-sample symmetric matrix, and the sums of the plaintexts of the original symmetric matrices of multiple data owners corresponding to other encryption technologies that include the influence of the pseudo-sample symmetric matrix, to obtain the sum of the plaintexts of the original symmetric matrices of all data owners corresponding to all data owners that include the influence of the pseudo-sample symmetric matrix; on the other hand, the second server accumulates the sums of the plaintexts of the original product vectors of multiple data owners corresponding to the above homomorphic encryption technology that include the influence of the pseudo-sample product vector, the sums of the plaintexts of the original product vectors of multiple data owners corresponding to the masking encryption technology that include the influence of the pseudo-sample product vector, and the sums of the plaintexts of the original product vectors of multiple data owners corresponding to other encryption technologies that include the influence of the pseudo-sample product vector, to obtain the sum of the plaintexts of the original product vectors of all data owners corresponding to all data owners that include the influence of the pseudo-sample product vector.
[0175] Finally, the second server uses the sum of the plaintexts of the original symmetric matrices of all data owners corresponding to all data owners that include the influence of the pseudo-sample symmetric matrix and the sum of the plaintexts of the original product vectors of all data owners corresponding to all data owners that include the influence of the pseudo-sample product vector to calculate the pseudo-sample encrypted linear regression coefficient that includes the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector, and sends the above pseudo-sample encrypted linear regression coefficient to the first server, where the pseudo-sample encrypted linear regression coefficient is equal to the column vector obtained by multiplying the inverse matrix of the sum of the plaintexts of the original symmetric matrices of all data owners corresponding to all data owners that include the influence of the pseudo-sample symmetric matrix by the sum of the plaintexts of the original product vectors of all data owners corresponding to all data owners that include the influence of the pseudo-sample product vector.
[0176] In step S1040, the first server receives the pseudo-sample encrypted linear regression coefficient that includes the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector, and eliminates at least part of the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector therein to obtain the target linear regression coefficient. This step basically follows the corresponding steps in Embodiment 1 and Embodiment 2, but it should be noted that: the sum of the plaintexts of the original symmetric matrices of all data owners corresponding to all data owners that include the influence of the pseudo-sample symmetric matrix includes the influence of three pseudo-sample symmetric matrices, and the sum of the plaintexts of the original product vectors of all data owners corresponding to all data owners that include the influence of the pseudo-sample product vector includes the influence of three pseudo-sample product vectors, and corresponding modifications are made to the relevant steps. In addition, it may be necessary to perform some corresponding operations according to the requirements of the other encryption technologies that support addition calculation on ciphertexts, which are well known to those skilled in the art and will not be elaborated here.
[0177] Finally, Example 4 is briefly introduced. In Example 4, at least two data owners respectively use three different encryption techniques including the first encryption technique to encrypt the intermediate variables related to several private data samples of their own, obtaining the ciphertexts related to their respective original symmetric matrices and the ciphertexts of the original product vectors. Among them, the three different encryption techniques include the homomorphic encryption technique used in Example 1, which is called the first encryption technique, the encryption technique that masks the plaintext to obtain the ciphertext, simply referred to as the mask encryption technique, used in Example 2, and a third encryption technique that supports addition calculation on the ciphertext; both the mask encryption technique and the third encryption technique satisfy that the ciphertexts encrypted by the mask encryption technique and the third encryption technique can be further encrypted by the first encryption technique, that is, the homomorphic encryption technique, to obtain the ciphertexts further encrypted by the first encryption technique; the first server further encrypts the ciphertexts of the original symmetric matrices encrypted by the mask encryption technique and the third encryption technique or the corresponding accumulated results, and further encrypts the ciphertexts of the original product vectors encrypted by the mask encryption technique and the third encryption technique or the corresponding accumulated results; the first server then uses the above further encrypted results to obtain the sum of the ciphertexts of the original symmetric matrices of all data owners encrypted by the first encryption technique and including the influence of the pseudo-sample symmetric matrix, and obtains the sum of the ciphertexts of the original product vectors of all data owners encrypted by the first encryption technique and including the influence of the pseudo-sample product vector; the first server sends the sum of the ciphertexts of the original symmetric matrices of all data owners encrypted by the first encryption technique and including the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of all data owners encrypted by the first encryption technique and including the influence of the pseudo-sample product vector to the second server;
[0178] The second server receives and decrypts the sum of the ciphertexts of the original symmetric matrices of all data owners encrypted by the first encryption technique and including the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of all data owners encrypted by the first encryption technique and including the influence of the pseudo-sample product vector respectively by using the decryption technique corresponding to the first encryption technique, and then further decrypts them by using the two decryption techniques corresponding to the mask encryption technique and the third encryption technique respectively to obtain the sum of the plaintexts of the original symmetric matrices of all data owners including the influence of the pseudo-sample symmetric matrix and the sum of the plaintexts of the original product vectors of all data owners including the influence of the pseudo-sample product vector.
[0179] Example 4 includes the following steps S1110, step S1120, step S1130, and step S1140. The functions of these steps are respectively similar to steps S110, step S120, step S130, and step S140 in Example 1 and Example 2. A brief introduction to these steps is given below.
[0180] Step S1110: At least two data owners use the above-mentioned at least three different encryption techniques, namely the homomorphic encryption technique used in Embodiment 1, the masked encryption technique used in Embodiment 2, and the third encryption technique that supports addition calculation on ciphertexts, to encrypt the intermediate variables related to a number of their respective private data samples, obtaining the ciphertexts related to their respective original symmetric matrices and the ciphertexts of the original product vectors. Each of the at least two data owners uses one encryption technique to encrypt the intermediate variables related to a number of its own private data samples, obtaining the ciphertexts related to the original symmetric matrix and the ciphertexts of the original product vector of this data owner; each data owner sends the ciphertexts of its own original symmetric matrix and the ciphertexts of the original product vector to the first server.
[0181] Step S1120 includes the following operations: The first server receives the ciphertexts related to the original symmetric matrices and the ciphertexts of the original product vectors sent by each of the at least two data owners; the first server respectively accumulates the ciphertexts related to the original symmetric matrices and the ciphertexts of the original product vectors obtained by using the same encryption technique to obtain the sum of the ciphertexts of the original symmetric matrices of multiple data owners corresponding to this encryption technique and the sum of the ciphertexts of the original product vectors of multiple data owners corresponding to this encryption technique, which are respectively abbreviated as the first ciphertext sum of the symmetric matrix corresponding to this encryption technique and the first ciphertext sum of the product vector corresponding to this encryption technique. That is to say, the first server obtains: the first ciphertext sum of the symmetric matrix corresponding to the homomorphic encryption technique, the first ciphertext sum of the product vector corresponding to the homomorphic encryption technique, the first ciphertext sum of the symmetric matrix corresponding to the masked encryption technique, the first ciphertext sum of the product vector corresponding to the masked encryption technique, the first ciphertext sum of the symmetric matrix corresponding to the third encryption technique, and the first ciphertext sum of the product vector corresponding to the third encryption technique. On the other hand, the first server obtains the symmetric matrix and the product vector corresponding to the pseudo data sample, which are respectively called the pseudo sample symmetric matrix and the pseudo sample product vector. Then, the first server respectively uses the methods introduced in the corresponding steps of Embodiment 1 to obtain the processing results of the pseudo sample symmetric matrix and the pseudo sample product vector corresponding to the homomorphic encryption technique.
[0182] Subsequently, the first server uses the homomorphic encryption technology and the first key described in the first embodiment to further encrypt the sum of the first ciphertexts of the symmetric matrix corresponding to the masked encryption technology, the sum of the first ciphertexts of the product vector corresponding to the masked encryption technology, the sum of the first ciphertexts of the symmetric matrix corresponding to the third encryption technology, and the sum of the first ciphertexts of the product vector corresponding to the third encryption technology, respectively, so as to obtain the results of the sum of the first ciphertexts of the symmetric matrix corresponding to the masked encryption technology after further homomorphic encryption, the results of the sum of the first ciphertexts of the product vector corresponding to the masked encryption technology after further homomorphic encryption, the results of the sum of the first ciphertexts of the symmetric matrix corresponding to the third encryption technology after further homomorphic encryption, and the results of the sum of the first ciphertexts of the product vector corresponding to the third encryption technology after further homomorphic encryption.
[0183] Finally, the first server accumulates the sum of the first ciphertexts of the symmetric matrix corresponding to the homomorphic encryption technology, the processing result of the pseudo-sample symmetric matrix corresponding to the homomorphic encryption technology, the results of the sum of the first ciphertexts of the symmetric matrix corresponding to the masked encryption technology after further homomorphic encryption, and the results of the sum of the first ciphertexts of the symmetric matrix corresponding to the third encryption technology after further homomorphic encryption, to obtain the sum of the original symmetric matrix ciphertexts of all data owners containing the influence of the pseudo-sample symmetric matrix further encrypted by the homomorphic encryption technology; on the other hand, the first server accumulates the sum of the first ciphertexts of the product vector corresponding to the homomorphic encryption technology, the processing result of the pseudo-sample product vector corresponding to the homomorphic encryption technology, the results of the sum of the first ciphertexts of the product vector corresponding to the masked encryption technology after further homomorphic encryption, and the results of the sum of the first ciphertexts of the product vector corresponding to the third encryption technology after further homomorphic encryption, to obtain the sum of the original product vector ciphertexts of all data owners containing the influence of the pseudo-sample product vector further encrypted by the homomorphic encryption technology. The first server sends the sum of the original symmetric matrix ciphertexts of all data owners containing the influence of the pseudo-sample symmetric matrix further encrypted by the homomorphic encryption technology, and the sum of the original product vector ciphertexts of all data owners containing the influence of the pseudo-sample product vector further encrypted by the homomorphic encryption technology, to the second server.
[0184] In the first implementation manner of the above step S1120, for other encryption technologies other than the first encryption technology, that is, the masked encryption technology and the third encryption technology, the first server first accumulates the ciphertexts related to the original symmetric matrix and the ciphertexts of the original product vector obtained by using the first encryption technology, the masked encryption technology, and the third encryption technology, respectively, and then further encrypts the accumulated results with the first encryption technology; that is to say, the first server further encrypts the accumulated results corresponding to the ciphertexts of the original symmetric matrix encrypted by the masked encryption technology and the third encryption technology with the first encryption technology, and further encrypts the accumulated results corresponding to the ciphertexts of the original product vector encrypted by the masked encryption technology and the third encryption technology.
[0185] The above-mentioned method of first accumulating the encrypted ciphertext and then further encrypting it can be modified to first further encrypt and then accumulate the encrypted ciphertext, so as to obtain another implementation manner of step S1120. In this other implementation manner, the first server first further encrypts the ciphertext related to the original symmetric matrix obtained by encrypting with the mask encryption technology and the third encryption technology, and the ciphertext of the original product vector, and then accumulates the results of the above further encryption. That is to say, the first server uses the homomorphic encryption technology and the first key described in the first embodiment to further encrypt the ciphertext of the original symmetric matrix obtained by encrypting with the mask encryption technology and the third encryption technology, and the ciphertext of the original product vector obtained by encrypting with the mask encryption technology and the third encryption technology, and then accumulates the results of the above further encryption, so as to obtain the results of further homomorphic encryption of the sum of the first ciphertexts of the symmetric matrix corresponding to the mask encryption technology used in the first implementation manner of step S120, the results of further homomorphic encryption of the sum of the first ciphertexts of the product vector corresponding to the mask encryption technology, the results of further homomorphic encryption of the sum of the first ciphertexts of the symmetric matrix corresponding to the third encryption technology, and the results of further homomorphic encryption of the sum of the first ciphertexts of the product vector corresponding to the third encryption technology; thereafter, the first server then follows the subsequent pseudo-sample encryption operations using the above results of further homomorphic encryption in the first implementation manner of step S120.
[0186] Step S1130, the second server receives the sum of the ciphertexts of the original symmetric matrices of all data owners including the influence of the pseudo-sample symmetric matrix further encrypted by the homomorphic encryption technology sent by the first server, and the sum of the ciphertexts of the original product vectors of all data owners including the influence of the pseudo-sample product vector further encrypted by the homomorphic encryption technology; the second server first uses the homomorphic decryption technology corresponding to the above homomorphic encryption technology and the second key described in the first embodiment to decrypt the above two sums of ciphertexts respectively, and then further decrypts them using the two decryption technologies corresponding to the mask encryption technology and the third encryption technology respectively, so as to obtain the sum of the plaintexts of the original symmetric matrices of all data owners including the influence of the pseudo-sample symmetric matrix and the sum of the plaintexts of the original product vectors of all data owners including the influence of the pseudo-sample product vector, which can also be respectively referred to as the sum of the first plaintexts of the symmetric matrix corresponding to all data owners including the influence of the pseudo-sample symmetric matrix, and the sum of the first plaintexts of the product vector corresponding to all data owners including the influence of the pseudo-sample product vector; the second server uses the sum of the first plaintexts of the symmetric matrix corresponding to all data owners including the influence of the pseudo-sample symmetric matrix and the sum of the first plaintexts of the product vector corresponding to all data owners including the influence of the pseudo-sample product vector to calculate the pseudo-sample encrypted linear regression coefficients including the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector, and sends the above pseudo-sample encrypted linear regression coefficients to the first server. Some details of this step are introduced as follows.
[0187] The second server first uses the homomorphic decryption technology corresponding to the above homomorphic encryption technology, and uses the second key described in the first embodiment to decrypt the sum of the original symmetric matrix ciphertexts of all data owners further encrypted by the homomorphic encryption technology and containing the influence of the pseudo-sample symmetric matrix, and the sum of the original product vector ciphertexts of all data owners further encrypted by the homomorphic encryption technology and containing the influence of the pseudo-sample product vector, respectively, to obtain the sum of the original symmetric matrix ciphertexts of all data owners containing the influence of the pseudo-sample symmetric matrix after homomorphic decryption, and the sum of the original product vector ciphertexts of all data owners containing the influence of the pseudo-sample product vector after homomorphic decryption. The second server then uses the decryption technology corresponding to the mask encryption technology and the decryption technology corresponding to the third encryption technology to decrypt the sum of the original symmetric matrix ciphertexts of all data owners containing the influence of the pseudo-sample symmetric matrix after the above homomorphic decryption and the sum of the original product vector ciphertexts of all data owners containing the influence of the pseudo-sample product vector after homomorphic decryption, respectively, to obtain the sum of the original symmetric matrix plaintexts of all data owners containing the influence of the pseudo-sample symmetric matrix (i.e., the first plaintext sum of the symmetric matrix containing the influence of the pseudo-sample symmetric matrix corresponding to all data owners) and the sum of the original product vector plaintexts of all data owners containing the influence of the pseudo-sample product vector (i.e., the first plaintext sum of the product vector containing the influence of the pseudo-sample product vector corresponding to all data owners), where the decryption technology corresponding to the mask encryption technology is the decryption technology introduced in step S810 of the second embodiment. In addition, the encrypted linear regression coefficient of the pseudo-sample is equal to the column vector obtained by multiplying the inverse matrix of the first plaintext sum of the symmetric matrix containing the influence of the pseudo-sample symmetric matrix corresponding to all data owners by the first plaintext sum of the product vector containing the influence of the pseudo-sample product vector corresponding to all data owners.
[0188] In step S1140, the first server receives the encrypted linear regression coefficient of the pseudo-sample containing the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector, and eliminates the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector in at least part of it to obtain the target linear regression coefficient. This step follows the corresponding steps in the first and second embodiments, but it should be noted that: it may be necessary to perform some corresponding operations according to the requirements of the third encryption technology, which are well known to those skilled in the art and will not be elaborated here.
[0189] The linear regression method provided by the embodiments of the present application can be applied to a variety of horizontal federated learning scenarios. The following takes the financial scenario as an example for detailed description.
[0190] Data owners include the Second Bank in the First Region and the Fourth Bank in the Third Region. The private data of the Second Bank includes user information and deposit and loan data in the First Region, and the private data of the Fourth Bank includes user information and deposit and loan data in the Third Region. The user information is private input features, and the deposit and loan information is private labels. It can be understood that the private data of the Second Bank and the Fourth Bank includes private input features that are also user information and private labels that are also deposit and loan data. The horizontal federated learning carried out by the Second Bank and the Fourth Bank is horizontal federated learning. In the embodiments of the present invention, a cooperative and restrictive relationship is formed between the First Server and the Second Server, thereby ensuring the security of real user information and deposit and loan information.
[0191] In addition, referring to Figure 2 , Figure 2 schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0192] A processor 201, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0193] A memory 202, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 202 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 202 and are called by the processor 201 to execute the method for horizontally dividing data and linear regression based on privacy protection in the embodiments of the present application. For example, execute the method steps S110 to S140 described above in Figure 1 ;
[0194] An input / output interface 203, which is used to implement information input and output;
[0195] A communication interface 204, which is used to implement communication and interaction between this device and other devices, and can implement communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0196] The bus 205 transmits information between various components of the device (such as the processor 201, the memory 202, the input / output interface 203, and the communication interface 204).
[0197] Among them, the processor 201, the memory 202, the input / output interface 203, and the communication interface 204 achieve communication connections with each other inside the device through the bus 205.
[0198] The embodiment of the present application also provides a storage medium. The storage medium is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned method for horizontally dividing data and performing linear regression based on privacy protection. For example, execute the Figure 1 method steps S110 to step S140 described above.
[0199] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0200] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0201] Those skilled in the art can understand that Figure 1 the technical solutions shown in do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figure, or combine certain steps, or different steps.
[0202] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0203] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.
[0204] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0205] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist simultaneously. Here, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one)" or a similar expression thereof refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0206] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the above-mentioned unit division is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of systems or units can be in electrical, mechanical, or other forms.
[0207] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0208] In addition, the functional units in various embodiments of the present application may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0209] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0210] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.
Claims
1. A linear regression method for horizontally partitioning data based on privacy protection, characterized in that: include: Each of at least two data owners uses several private data samples to obtain the ciphertext related to their original symmetric matrix and the ciphertext of the original product vector, wherein the original symmetric matrix and the original product vector are the symmetric matrix and the product vector corresponding to the several data samples, respectively; each of the data owners sends the ciphertext related to their original symmetric matrix and the ciphertext of the original product vector to the first server, and the first server receives the ciphertext related to the original symmetric matrix and the ciphertext of the original product vector sent by each of the at least two data owners, and cooperates with the second server to obtain the target linear regression coefficient. The linear regression method based on privacy protection includes: The first server receives the ciphertext related to the original symmetric matrix and the ciphertext of the original product vector sent by each of at least two of the data owners, and obtains the pseudo-sample symmetric matrix processing result and the pseudo-sample product vector processing result, wherein the pseudo-sample symmetric matrix processing result and the pseudo-sample product vector processing result are the results obtained by processing the pseudo-sample symmetric matrix and the pseudo-sample product vector, respectively, and the pseudo-sample symmetric matrix and the pseudo-sample product vector are the symmetric matrix and the product vector corresponding to a number of pseudo-data samples, respectively; the first server uses the ciphertext related to the original symmetric matrix and the pseudo-sample symmetric matrix processing result to obtain the sum of the ciphertexts of the original symmetric matrices of multiple data owners affected by the pseudo-sample symmetric matrix, and uses the ciphertext of the original product vector and the pseudo-sample product vector processing result to obtain the sum of the ciphertexts of the original product vectors of multiple data owners affected by the pseudo-sample product vector; the first server sends the sum of the ciphertexts of the original symmetric matrices of the multiple data owners affected by the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of the multiple data owners affected by the pseudo-sample product vector to the second server; The second server receives and respectively decrypts the sum of the ciphertexts of the multiple data owners' original symmetric matrices affected by the pseudo-sample symmetric matrix and the sum of the ciphertexts of the multiple data owners' original product vectors affected by the pseudo-sample product vector, so as to obtain the sum of the plaintexts of all data owners' original symmetric matrices affected by the pseudo-sample symmetric matrix and the sum of the plaintexts of all data owners' original product vectors affected by the pseudo-sample product vector; The second server obtains the pseudo-sample encrypted linear regression coefficient including the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector by using the sum of the plaintexts of the original symmetric matrices of all data owners including the influence of the pseudo-sample symmetric matrix and the sum of the plaintexts of the original product vectors of all data owners including the influence of the pseudo-sample product vector, and sends the pseudo-sample encrypted linear regression coefficient to the first server, or the second server performs regularization processing on the sum of the plaintexts of the original symmetric matrices of all data owners including the influence of the pseudo-sample symmetric matrix, obtains the pseudo-sample encrypted linear regression coefficient including the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector according to the sum of the plaintexts of the original product vectors of all data owners including the influence of the pseudo-sample product vector and the sum of the plaintexts of the original symmetric matrices of all data owners after regularization, and sends the pseudo-sample encrypted linear regression coefficient to the first server; The first server receives the pseudo-sample encrypted linear regression coefficient including the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector, at least partially eliminates the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector from the pseudo-sample encrypted linear regression coefficient, and obtains the target linear regression coefficient.
2. The linear regression method according to claim 1, characterized in that Each of the several private data samples contains an original input feature column vector and a corresponding original label; each of the several pseudo data samples contains a pseudo input feature column vector and a corresponding pseudo label; the original symmetric matrix of any of the data owners is equal to the sum of the products of the original input feature column vectors of the data owner and their own transposition, and the original product vector of any of the data owners is equal to the sum of the products of the original labels of the data owner and the corresponding original input feature column vector; the pseudo sample symmetric matrix is the sum of the products of the pseudo input feature column vectors contained in the several pseudo data samples and their own transposition. The pseudo sample product vector is the sum of the products of several pseudo input feature column vectors and corresponding pseudo labels contained in the several pseudo data samples; the ciphertext associated with the original symmetric matrix and the ciphertext of the original product vector are obtained by the data owner using an encryption technology that supports addition calculation of ciphertexts to encrypt the plaintext associated with the original symmetric matrix and the original product vector respectively, and the encryption technology that supports addition calculation of ciphertexts satisfies: the plaintexts are respectively obtained by the encryption technology that supports addition calculation of ciphertexts to obtain the ciphertexts, the sum of the ciphertexts is obtained by arithmetic addition operation between the ciphertexts, and the sum of the ciphertexts corresponds to the arithmetic sum of multiple plaintexts after decryption; The first server uses the ciphertext related to the original symmetric matrix and the processing result of the pseudo-sample symmetric matrix to obtain the sum of the ciphertexts of the original symmetric matrices of multiple data owners affected by the pseudo-sample symmetric matrix, specifically using the technology of performing addition calculation on ciphertext included in the encryption technology supporting addition calculation on ciphertext, to accumulate the ciphertexts of the original symmetric matrices of multiple data owners and the processing result of the pseudo-sample symmetric matrix; The specific process by which the first server uses the ciphertext of the original product vector and the processing result of the pseudo-sample product vector to obtain the sum of the ciphertexts of the original product vectors of the multiple data owners affected by the pseudo-sample product vector is to utilize the technology for performing addition calculations on ciphertexts included in the encryption technology that supports addition calculations on ciphertexts, and accumulate the ciphertexts of the original product vectors of multiple data owners and the processing results of the pseudo-sample product vectors.
3. The linear regression method according to claim 2, characterized in that: The first server receives a pseudo-sample encrypted linear regression coefficient including the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector, and at least partially eliminates the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector from the pseudo-sample encrypted linear regression coefficient to obtain a target linear regression coefficient, including: The first server encrypts a pseudo input matrix composed of a plurality of pseudo input feature column vectors included in the plurality of pseudo data samples to obtain a pseudo input matrix ciphertext, and the first server sends the pseudo input matrix ciphertext to the second server; The second server calculates the inverse matrix of the sum of the plaintexts of the original symmetric matrices of all data owners affected by the pseudo sample symmetric matrix, that is, the first inverse matrix. The second server obtains the product of the pseudo input matrix ciphertext and the first inverse matrix based on the encryption method of the pseudo input matrix by the first server, and obtains the product of the pseudo input matrix ciphertext and the first inverse matrix. The second server sends the product of the pseudo input matrix ciphertext and the first inverse matrix to the first server. The first server decrypts the product of the pseudo-input matrix ciphertext and the first inverse matrix based on its own encryption method for the pseudo-input matrix to obtain the product of the pseudo-input matrix plaintext and the first inverse matrix. The first server uses the product of the pseudo-input matrix plaintext and the first inverse matrix to at least partially eliminate the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector in the pseudo-sample encrypted linear regression coefficient that includes the influence of the pseudo-sample symmetric matrix and the pseudo-sample product vector to obtain the target linear regression coefficient.
4. The linear regression method according to claim 1, characterized in that: Each of at least two of the data owners uses the same encryption technology to encrypt the intermediate variables related to several private data samples owned by itself, and obtains the ciphertext related to their respective original symmetric matrices and the ciphertext of the original product vector; the sum of the ciphertexts of the original symmetric matrices of the multiple data owners affected by the pseudo-sample symmetric matrix obtained by the first server is the sum of the ciphertexts of the original symmetric matrices of all the data owners affected by the pseudo-sample symmetric matrix, and the sum of the ciphertexts of the original product vectors of the multiple data owners affected by the pseudo-sample product vector obtained by the first server is the sum of the ciphertexts of the original product vectors of all the data owners affected by the pseudo-sample product vector.
5. The linear regression method according to claim 1, characterized in that: The at least two data owners respectively use at least two different encryption technologies to encrypt intermediate variables related to several private data samples owned by themselves, to obtain ciphertexts related to their respective original symmetric matrices and ciphertexts of the original product vectors; The sum of the ciphertexts of the original symmetric matrices of the data owners, any one of which includes the influence of the pseudo-sample symmetric matrix, obtained by the first server is the sum of the ciphertexts of the original symmetric matrices of the multiple data owners using the same encryption technology plus the influence of the pseudo-sample symmetric matrix; The sum of the ciphertexts of the original product vectors of the data owners, any one of which includes the influence of the pseudo sample product vector, obtained by the first server is the sum of the ciphertexts of the original product vectors of the multiple data owners using the same encryption technology plus the influence of the pseudo sample product vector; The first server sends the sum of the ciphertexts of the at least two original symmetric matrices of the data owners affected by the pseudo sample symmetric matrices and the sum of the ciphertexts of the at least two original product vectors of the data owners affected by the pseudo sample product vectors to the second server; The second server receives the sum of the ciphertexts of at least two original symmetric matrices of the data owners including the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the at least two original product vectors of the data owners including the influence of the pseudo-sample product vector, respectively, which are associated with different encryption methods, and uses different decryption techniques corresponding to various encryption techniques to respectively decrypt the sum of the ciphertexts of the at least two original symmetric matrices of the data owners including the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the at least two original product vectors of the data owners including the influence of the pseudo-sample product vector, respectively, corresponding to the at least two different encryption techniques, to obtain the sum of the plaintexts of the at least two original symmetric matrices of the data owners including the influence of the pseudo-sample symmetric matrix and the sum of the plaintexts of the at least two original product vectors of the data owners including the influence of the pseudo-sample product vector; The second server accumulates the sum of the plaintexts of the original symmetric matrices of the at least two data owners including the influence of the pseudo-sample symmetric matrix to obtain the sum of the plaintexts of the original symmetric matrices of all data owners including the influence of the pseudo-sample symmetric matrix, and accumulates the sum of the plaintexts of the original product vectors of the at least two data owners including the influence of the pseudo-sample product vector to obtain the sum of the plaintexts of the original product vectors of all data owners including the influence of the pseudo-sample product vector.
6. The linear regression method according to claim 1, characterized in that: The at least two data owners use at least two different encryption technologies including the first encryption technology to encrypt intermediate variables related to several private data samples of their own private ownership, and obtain ciphertexts related to their respective original symmetric matrices and ciphertexts of original product vectors; the first server uses the first encryption technology to further encrypt the ciphertexts of the original symmetric matrices encrypted by other encryption technologies or the corresponding accumulated results, and further encrypts the ciphertexts of the original product vectors encrypted by the other encryption technologies or the corresponding accumulated results, wherein the other encryption technology is an encryption technology other than the first encryption technology among the at least two different encryption technologies. ; The first server further utilizes the result of the further encryption to obtain the sum of the ciphertexts of the original symmetric matrices of all data owners including the influence of the pseudo-sample symmetric matrix encrypted by the first encryption technology, and obtains the sum of the ciphertexts of the original product vectors of all data owners including the influence of the pseudo-sample product vector encrypted by the first encryption technology; The first server sends the sum of the ciphertexts of the original symmetric matrices of all data owners including the influence of the pseudo-sample symmetric matrix encrypted by the first encryption technology and the sum of the ciphertexts of the original product vectors of all data owners including the influence of the pseudo-sample product vector encrypted by the first encryption technology to the second server; The second server receives and uses the decryption technology corresponding to the first encryption technology to decrypt the sum of the ciphertexts of the original symmetric matrices of all data owners including the influence of the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of all data owners including the influence of the pseudo-sample product vector encrypted by the first encryption technology, and then uses the decryption technology corresponding to the other encryption technologies for further decryption to obtain the sum of the plaintexts of the original symmetric matrices of all data owners including the influence of the pseudo-sample symmetric matrix and the sum of the plaintexts of the original product vectors of all data owners including the influence of the pseudo-sample product vector.
7. The linear regression method according to claim 4, characterized in that: Each of at least two of the data owners uses the same encryption technology for encryption, which is an additive homomorphic encryption technology that supports addition calculations on ciphertexts and uses a first key for encryption, wherein the first key is a public key used in the additive homomorphic encryption technology; the ciphertext associated with the original symmetric matrix is the ciphertext of the original symmetric matrix; the pseudo-sample symmetric matrix processing result is the ciphertext obtained by homomorphically encrypting the pseudo-sample symmetric matrix or the negative value of the pseudo-sample symmetric matrix using the first key, and the pseudo-sample product vector processing result is the pseudo-sample product vector or the negative value of the pseudo-sample product vector using the first key. The obtained ciphertext; the sum of the ciphertexts of the original symmetric matrices of all data owners affected by the pseudo-sample symmetric matrix is obtained by using the technology of performing addition calculation on ciphertexts included in the additive homomorphic encryption technology, accumulating the processing result of the pseudo-sample symmetric matrix and the ciphertexts of the original symmetric matrices of all the data owners, and performing optional regularization processing; the sum of the ciphertexts of the original product vectors of all data owners affected by the pseudo-sample product vector is obtained by using the technology of performing addition calculation on ciphertexts included in the additive homomorphic encryption technology, accumulating the processing result of the pseudo-sample product vector and the ciphertexts of the original product vectors of all the data owners; The second server decrypts the sum of the ciphertexts of the original symmetric matrices of all data owners affected by the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of all data owners affected by the pseudo-sample product vector, by utilizing the decryption technology corresponding to the additive homomorphic encryption technology and using a second key for decryption, wherein the second key is the private key used in the additive homomorphic encryption technology.
8. The linear regression method according to claim 4, characterized in that: Each of at least two of the data owners uses the same encryption technology for encryption, which is to encrypt by using the technology of masking the plaintext; the ciphertext of the original product vector of any of the data owners is the sum of the original product vector of the data owner and the mask product vector of the data owner; the ciphertext associated with the original symmetric matrix satisfies that the ciphertext of the original symmetric matrix can be calculated using the ciphertext associated with the original symmetric matrix, and the ciphertext of the original symmetric matrix of any of the data owners is the sum of the original symmetric matrix of the data owner and the mask symmetric matrix of the data owner; The pseudo-sample symmetric matrix processing result is the pseudo-sample symmetric matrix or the negative value of the pseudo-sample symmetric matrix, and the pseudo-sample product vector processing result is the pseudo-sample product vector or the negative value of the pseudo-sample product vector; the sum of the ciphertexts of the original symmetric matrices of all data owners affected by the pseudo-sample symmetric matrix is the result obtained by accumulating the pseudo-sample symmetric matrix processing result and the ciphertexts of the original symmetric matrices of all the data owners, and performing optional regularization processing; the sum of the ciphertexts of the original product vectors of all data owners affected by the pseudo-sample product vector is the result obtained by accumulating the pseudo-sample product vector processing result and the ciphertexts of the original product vectors of all the data owners; The second server decrypts the sum of the ciphertexts of the original symmetric matrices of all data owners affected by the pseudo-sample symmetric matrix and the sum of the ciphertexts of the original product vectors of all data owners affected by the pseudo-sample product vector, respectively. The second server subtracts the sum of the masked symmetric matrices of all data owners from the sum of the ciphertexts of the original symmetric matrices of all data owners affected by the pseudo-sample symmetric matrix, and subtracts the sum of the masked product vectors of all data owners from the sum of the ciphertexts of the original product vectors of all data owners affected by the pseudo-sample product vector, respectively, to obtain the sum of the plaintexts of the original symmetric matrices of all data owners affected by the pseudo-sample symmetric matrix and the sum of the plaintexts of the original product vectors of all data owners affected by the pseudo-sample product vector.
9. The linear regression method according to claim 8, characterized in that: The masked symmetric matrix and the masked product vector of any of the data owners are respectively the symmetric matrix and the product vector corresponding to a number of masked data samples of the data owner, wherein each of the several masked data samples of any of the data owners comprises a masked input feature column vector and a corresponding masked label; the masked symmetric matrix of any of the data owners is the cumulative result of the product of each of the masked input feature column vectors in the several masked data samples of the data owner and their transpose, and the masked product vector of any of the data owners is the cumulative result of the product of each of the masked input feature column vectors in the several masked data samples of the data owner and the corresponding masked label; The ciphertext associated with the original symmetric matrix of any of the data owners is the ciphertext of the original symmetric matrix or the product of a spliced matrix and an arbitrary orthogonal matrix, wherein the spliced matrix is composed of each of the original input feature column vectors and each of the masked input feature column vectors of the data owner.
10. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the linear regression method according to any one of claims 1 to 9 when executing the computer program.
Citation Information
Patent Citations
Privacy protection-based linear regression method and system, storage medium and equipment
CN117171498A
Linear regression method based on privacy protection
CN118313481A
Method and apparatus for carrying out multi-party joint dimension reduction processing on private data
WO2021190424A1
AU2020101237A4