A data regression processing method, terminal and device

CN115495709BActive Publication Date: 2026-09-25WUYI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210998415.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-19
Publication Date
2026-09-25
Estimated Expiration
2042-08-19

AI Technical Summary

Benefits of technology

[0059]本申请实施例通过至少两个用户终端利用自己持有的掩码加密自己持有的私有数据样本,得到加密矩阵和加密乘积;至少两个用户终端把所述加密矩阵和所述加密乘积发送或等效的发送至第一计算节点;至少两个用户终端把所述掩码发送或等效的发送至第二计算节点;基于所述加密矩阵,以及所述掩码,通过所述第一计算节点和所述第二计算节点合作求得线性回归系数。降低了加密的复杂度和数据传输时的数据量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115495709B_ABST
    Figure CN115495709B_ABST
Patent Text Reader

Abstract

The application relates to a privacy protection-based linear regression method, which comprises the following steps: at least two user terminals encrypt private data samples held by the user terminals by using masks held by the user terminals, to obtain an encryption matrix and an encryption product; the at least two user terminals send or equivalently send the encryption matrix and the encryption product to a first computing node; the at least two user terminals send or equivalently send the masks to a second computing node; and linear regression coefficients are obtained by cooperation of the first computing node and the second computing node based on the encryption matrix and the masks. The technical scheme of the application reduces the encryption complexity and the data amount during data transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer science, and more particularly to a privacy-preserving linear regression method. Background Technology

[0002] Linear regression is a statistical analysis method that uses regression analysis from mathematical statistics to determine the interdependencies between variables. As the simplest machine learning model, it has wide applications in finance, economics, and medicine. Linear regression uses a least-squares function called the linear regression equation to model the relationship between one or more independent variables and a dependent variable. This function is a linear combination of one or more model parameters called regression coefficients. The case with only one independent variable is called univariate linear regression, and the case with more than one independent variable is called multiple linear regression. Linear regression, also known as linear regression fitting, is a method of predicting one dependent variable (e.g., children's height) based on one or more independent variables (e.g., father and mother's height).

[0003] Linear regression can be used to fit a predictive model to a dataset of observed values ​​x and y. Once such a model is established, for a new independent variable x, without a corresponding dependent variable y, the fitted model can predict the value of the dependent variable y. Linear regression learns a vector of regression coefficients. Make in, It is a column vector including the d terms, (·) T This represents the transpose of a matrix. Assume there are n training samples, where the i-th (i = 1, 2, ..., n) training sample is a column vector containing d items. And the corresponding output variable, i.e., the dependent variable y. i It can be represented as a column vector. The i-th training sample includes d independent variables, i.e., d features, in column vector form. The j-th term (j=1,2,…,d)x j,i It is the value of the j-th independent variable, i.e., the feature, while the regression coefficient vector The j-th term β j It is the regression coefficient corresponding to the j-th independent variable, i.e., the feature.

[0004] Put y i (i = 1, 2, ..., n) can be written as a column vector Bundle Write as a matrix Therefore, by definition and A=X T Let X + λI, where λ is the regularization coefficient. Then, the regression coefficient vector... It can be achieved by solving linear equations This is a well-known existing technology.

[0005] Horizontal federated learning, also known as sample-based federated learning, involves participants possessing different data samples. Through collaboration, they improve the performance of the training model. These participants may be data-holding enterprises, user clients, or mobile communication terminals. There can be at least two user terminals and up to n, where n is a positive integer greater than or equal to 2. A key objective of horizontal federated learning is to protect the data privacy of each participant's private data, ensuring that each participant's private data samples are not leaked. Note that this embodiment is a typical example of the present invention, targeting a scenario where all user terminals are mobile communication terminals. In this scenario, typically each mobile terminal has only one data sample, while the amount of data from mobile terminals can reach hundreds of millions.

[0006] Federated learning is a distributed machine learning technique. Its core idea is to train models in a distributed manner across multiple data sources that have local data. Without exchanging local individual or sample data, it constructs a global model based on virtual fused data by exchanging model parameters or intermediate results. This achieves a balance between data privacy protection and data sharing computation, namely, a new application paradigm of "data is available but not visible" and "the model moves while the data does not move".

[0007] Horizontal federated learning, also known as sample-partitioned federated learning or example-partitioned federated learning, can be applied to scenarios where the datasets of the various participants in federated learning share the same feature space but have different sample spaces, similar to horizontally partitioning data in a tabular view. In fact, the term "horizontal" comes from the term "horizontal partition." "Horizontal partitioning" is widely used in traditional scenarios where database records are displayed in tabular form, such as when records in a table are horizontally divided into different groups according to rows, and each row contains complete data features. Summary of the Invention

[0008] In view of this, this application proposes a privacy-preserving linear regression method, which can reduce the complexity of encryption and the amount of data transmitted while ensuring privacy protection during data sharing in scenarios involving multiple user terminals.

[0009] On the one hand, embodiments of this application propose a privacy-preserving linear regression method, the method comprising:

[0010] At least two user terminals use their own masks to encrypt their own private data samples, obtaining an encryption matrix and an encryption product;

[0011] At least two user terminals send the encryption matrix and the encryption product, or an equivalent transmission, to the first computing node;

[0012] At least two user terminals send the mask or an equivalent to the second computing node;

[0013] Based on the encryption matrix and the mask, the linear regression coefficients are obtained through the cooperation of the first computing node and the second computing node.

[0014] Furthermore, at least two user terminals encrypt their own private data samples using their own masks to obtain an encryption matrix and an encryption product. The private data samples include an independent variable column vector and a dependent variable, and the mask includes a masked independent variable column vector and a masked dependent variable, including:

[0015] Based on the column vector of independent variables of the private data sample and the column vector of independent variables of the mask, each of at least two user terminals obtains a first matrix;

[0016] Each of at least two user terminals multiplies the first matrix and a preset matrix to obtain the encryption matrix;

[0017] Each of at least two user terminals calculates the product of the independent variable column vector and the dependent variable to obtain a first product;

[0018] Each of at least two user terminals calculates the product of the mask independent variable column vector and the mask dependent variable to obtain a second product;

[0019] Each of at least two user terminals obtains the encrypted product based on the first product and the second product.

[0020] Furthermore, the preset matrix includes:

[0021] Orthogonal matrix, or,

[0022] A matrix obtained by multiplying an orthogonal matrix by a coefficient.

[0023] Furthermore, the at least two user terminals sending, or equivalently sending, the encryption matrix and the encryption product to the first computing node includes:

[0024] At least two user terminals will send the first preset index information to the first computing node;

[0025] The first computing node obtains the encryption matrix and the encryption product based on the preset index information;

[0026] or,

[0027] At least two user terminals will send the preset column and the encrypted product to the first computing node;

[0028] The first computing node obtains the encryption matrix based on the preset columns.

[0029] Furthermore, the at least two user terminals sending the mask, or an equivalent transmission, to the second computing node includes:

[0030] The second computing node sends the mask to at least two user terminals;

[0031] or,

[0032] At least two user terminals send the second preset index information to the second computing node;

[0033] The second computing node obtains the mask based on the second preset index information;

[0034] or,

[0035] The second computing node obtains the mask, which is preset by the second computing node and the at least two user terminals.

[0036] Furthermore, the step of obtaining the linear regression coefficients based on the encryption matrix and the mask through the cooperation of the first computing node and the second computing node includes:

[0037] The first computing node obtains the target encrypted regularized symmetric matrix and the target product based on the encryption matrix and the encryption product;

[0038] The second computing node obtains a first target mask matrix and a second target mask vector based on the mask.

[0039] The first computing node and the second computing node obtain the linear regression coefficients based on the target encryption regularized symmetric matrix and the target product, as well as the first target mask matrix and the second target mask vector, using secure multi-party computation techniques.

[0040] Furthermore, the first computing node obtains the target encryption regularized symmetric matrix and the target product based on the encryption matrix and the encryption product, including:

[0041] The first computing node performs symmetry processing on the encryption matrix to generate a symmetric matrix;

[0042] The first computing node sums the symmetric matrix to obtain the symmetric matrix sum;

[0043] The first computing node obtains a preset regularization coefficient and a preset identity matrix;

[0044] The first computing node generates the target encrypted regularized symmetric matrix based on the symmetric matrix, the preset regularized symmetric coefficients, and the preset identity matrix;

[0045] The first computing node sums the encrypted product to obtain the target product.

[0046] Furthermore, based on the mask, the second computing node obtains a first target mask matrix and a second target mask vector. The mask includes a column vector of mask independent variables and a mask dependent variable, including:

[0047] The second computing node performs symmetric processing on the column vector of the mask independent variables to obtain a symmetric mask matrix;

[0048] The second computing node accumulates the mask symmetric matrix to obtain the first target mask matrix;

[0049] The second computing node calculates the product of the mask independent variable column vector and the mask dependent variable to obtain an intermediate product;

[0050] The second computing node accumulates the intermediate products to obtain the second target mask vector.

[0051] Furthermore, in the case where at least two target user terminals contain multiple private data samples, each of the multiple private data samples includes multiple independent variables, and the mask includes one or more mask independent variable vectors, the method further includes:

[0052] Each of at least two user terminals calculates the sum of the number of the plurality of private data samples and the number of the one or more mask independent variable vectors to obtain the sample mask number;

[0053] Each of at least two user terminals compares the number of sample masks and the number of the plurality of independent variables, and sets the transmission format identifier to a transmission vector multiplied by an orthogonal transformation or a transmission symmetric matrix. Further, each of the at least two user terminals comparing the number of samples and masks and the number of the plurality of independent variables, and setting the transmission format identifier to a transmission vector multiplied by an orthogonal transformation or a transmission symmetric matrix, includes:

[0054] Each of at least two user terminals calculates the difference between the number of sample masks and the number of the plurality of independent variables to obtain a standard value;

[0055] When the standard value is less than a preset specific value, the transmission format identifier is set to the transmission vector multiplied by an orthogonal transformation;

[0056] When the standard value is greater than or equal to the preset specific value, the transmission format identifier is set as a transmission symmetric matrix.

[0057] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application.

[0058] Implementing this application will have the following beneficial effects:

[0059] This application embodiment involves at least two user terminals encrypting their own private data samples using their own masks to obtain an encryption matrix and an encryption product. The at least two user terminals then send the encryption matrix and the encryption product, or equivalently, to a first computing node. The at least two user terminals also send the mask, or equivalently, to a second computing node. Based on the encryption matrix and the mask, the first and second computing nodes collaboratively calculate the linear regression coefficients. This reduces the complexity of encryption and the amount of data transmitted.

[0060] Other features and aspects of this application will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0061] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this application together with the specification and serve to explain the principles of this application.

[0062] Figure 1 This is a schematic diagram of the implementation environment of a privacy-preserving linear regression method provided in an embodiment of the present invention.

[0063] Figure 2 This is a flowchart illustrating a privacy-preserving linear regression method provided in an embodiment of the present invention.

[0064] Figure 3 This is a flowchart illustrating a privacy-preserving linear regression method provided in an embodiment of the present invention. Figure 1 . Detailed Implementation

[0065] Various exemplary embodiments, features, and aspects of this application will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0066] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.

[0067] Furthermore, to better illustrate this application, numerous specific details are provided in the following detailed embodiments. Those skilled in the art should understand that this application can be implemented without certain specific details. In some instances, methods, means, components, and circuits well-known to those skilled in the art have not been described in detail in order to highlight the main points of this application.

[0068] Figure 1 This is a schematic diagram illustrating the implementation environment of a data regression processing method provided in an embodiment of this application. For example... Figure 1 As shown, this implementation environment may include at least two user terminals 01-1 and 01-2 and at least two computing nodes 02-1 and 02-2. The at least two user terminals 01-1 and 01-2 and the at least two computing nodes 02-1 and 02-2 can be directly or indirectly connected via wired or wireless communication, which is not limited herein. For example, the at least two user terminals 01-1 and 01-2 can send encrypted data to the at least two computing nodes 02-1 and 02-2 via wired or wireless communication.

[0069] For example, at least two user terminals 01-1 and 01-2 can be smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, etc., but are not limited to these. At least two computing nodes 02-1 and 02-2 can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. In particular, at least two computing nodes 02-1 and 02-2 can also be two computing nodes within a communication base station in a wireless communication system.

[0070] It should be noted that, Figure 1 This is just one example.

[0071] Example 1

[0072] Figure 2This is a flowchart illustrating a linear regression data processing method provided in an embodiment of this application. This method can be used for... Figure 1 In the implementation environment described herein, the method operation steps are as shown in the embodiments or flowcharts, but based on conventional or non-inventive labor, more or fewer operation steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many steps and does not represent the only execution order. In actual system or server products, the method can be executed in the order shown in the embodiments or drawings or in parallel (e.g., in a parallel processor or multi-threaded processing environment).

[0073] In Example 1, each user in at least two user terminals contains only one private data sample.

[0074] For example, such as Figure 2 As shown, the method may include:

[0075] S101. Each of at least two user terminals encrypts a private data sample using its own mask to obtain an encryption matrix and an encryption product.

[0076] The mask held by the user terminal is usually a random mask generated by the user terminal. A random mask is a mask generated by the user terminal that cannot be predicted or known by other user terminals or participating parties.

[0077] For example, the encryption matrix can be C i The encrypted product can be Suppose there are a total of n user terminals, where n is a positive integer greater than or equal to 2. The i-th user terminal is called user terminal i, where i is a positive integer greater than or equal to 1 and less than or equal to n.

[0078] Optionally, the at least two user terminals can be multiple participants in a horizontal federated learning process.

[0079] User terminals may choose not to add other encryption schemes, or they may add other encryption schemes themselves, including but not limited to homomorphic encryption.

[0080] For example, S101 may include:

[0081] Based on private data samples and masks, each of at least two user terminals obtains a first matrix;

[0082] A private data sample held by user terminal i includes d independent variables, that is, a column vector of independent variables composed of d features. as well as The corresponding dependent variable y i yi Also known as output variables. The mask generated by user terminal i also contains d independent variables, that is, a column vector composed of d features. And one dependent variable v i . It can be called the masking independent variable vector, v i It can be called the mask dependent variable.

[0083] Each user terminal i stores the column vector of independent variables from the private data sample. and the column vector of independent variables in the mask The first matrix is ​​obtained by splicing.

[0084] Each of at least two user terminals multiplies the first matrix and the preset matrix to obtain the encryption matrix.

[0085] The preset matrix includes

[0086] Orthogonal matrix, or,

[0087] A matrix obtained by multiplying an orthogonal matrix by a coefficient.

[0088] For example, this orthogonal matrix can be written as Θ i The first matrix multiplied by the orthogonal matrix Θ i Obtain the encryption matrix This encryption matrix can be used in C. i This means that the encryption matrix is ​​obtained by multiplying the first matrix and the preset orthogonal matrix. The specific operation process of the encryption matrix can be as follows:

[0089] By multiplying an orthogonal matrix by the matrix formed by concatenating the independent variable vectors in the data sample and the independent variable vectors in the mask, the independent variable vectors of the input data in the data to be processed are encrypted, instead of the homomorphic encryption scheme commonly used in current technology. This reduces the complexity of encryption and decreases the computational difficulty for at least two user terminals.

[0090] Each of at least two user terminals calculates the product of the column vector of independent variables and the dependent variable to obtain the first product;

[0091] Each of at least two user terminals calculates the product of the masked independent variable column vector and the masked dependent variable to obtain the second product;

[0092] Each of at least two user terminals obtains an encrypted product based on the first product and the second product.

[0093] This encrypted product can be but The specific calculation process can be

[0094] Therefore, the encrypted product is the sum of the first product and the second product.

[0095] S103. Each of at least two user terminals sends or equivalently sends the encrypted data to the first computing node. The encrypted data includes an encryption matrix C. i and encrypted product

[0096] The equivalent transmission methods described above include, but are not limited to, any one of the following methods:

[0097] The first computing node and user i pre-agreed upon a fixed codebook, which includes multiple C... i and The value of is determined by user i sending a specific C from the codebook to the first computing node. i The corresponding sequence number (or index), and a certain The first computing node finds the corresponding C in the codebook based on the received sequence number (or index). i and

[0098] User i uses a set mask variable vector. Ensure that the calculated encryption matrix C i Column 1 (e.g., column 1) is a specific column whose value is pre-agreed upon by compute node 1 and user i, so user i only needs to input C. i Columns not defined in the code (e.g., column 2) and Send to compute node 1; at the same time, user i sets its own mask. and v i Send to compute node 2. For example, assume that compute node 1 and user i have agreed in advance that a specific column is column 1, C. i (:,1), while orthogonal transformation Θ i The first column is Then the mask independent variable vector set by user i The calculation is done through This is by This was derived. Correspondingly, user i only needs to change column C in the second column. i (:,2) and Send the mask to compute node one. and v i Send to compute node two.

[0099] S105. At least two user terminals send the mask, or an equivalent one, to the second computing node. The mask consists of a column vector of d independent variables (i.e., d features) and one dependent variable.

[0100] The equivalent transmission methods described above include, but are not limited to, any one of the following methods:

[0101] By the second computing node and v i Send to user i;

[0102] The second computing node and user i pre-agreed upon a fixed codebook, which includes multiple... and v i The value of . User i sends a certain value from the codebook to the second computing node. The corresponding sequence number (or index), and a certain v i The second computing node finds the corresponding sequence number (or index) in the codebook based on the received sequence number (or index). and v i Alternatively, the second computing node can send the aforementioned sequence number (or index) to user i, and user i can find the corresponding sequence number (or index) in the codebook based on the received sequence number (or index). and v i .

[0103] The second computing node and user i have agreed on a fixed value beforehand. and v i The value is such that nothing actually needs to be transmitted.

[0104] In the equivalent transmission scenario described above, compute node 2 may not necessarily need to utilize [the data / method / feature]. and v i To calculate and and It is known in advance by compute node 2, or and There are only a few possible values, and the second computation node can select one of them.

[0105] S107. Based on the encrypted data and the mask, the linear regression coefficients are obtained through the cooperation of the first computing node and the second computing node.

[0106] For example, S107 may include:

[0107] The first computing node performs symmetric processing on the encryption matrix of user terminal i to generate a symmetric matrix.

[0108] The encryption matrix can be C. i Then the symmetric matrix can be C. i and C i The product of the transposes of the matrices, i.e.

[0109] The first computing node accumulates a symmetric matrix of n users. Get the sum of n user symmetric matrices

[0110] The first computing node obtains the preset regularization coefficients and the preset identity matrix.

[0111] The regularization coefficient can be represented as λ, and the preset identity matrix can be represented as I.

[0112] The first computing node is based on the sum of n user symmetric matrices. Generate the target encryption regularized symmetric matrix by setting the preset regularization coefficients and the preset identity matrix.

[0113] The first computing node performs masked encryption on the product of the independent variable vector and the dependent variable from the n received user terminals. Summing is performed to obtain the target product.

[0114] The first computing node and the second computing node must not collude, meaning they cannot conspire to spy on the data of at least two user terminals.

[0115] The mask independent variable column vector in the mask of the second computing node for user i Symmetric processing is performed to obtain the mask symmetric matrix. The symmetric processing results of n user terminals are then summed to obtain the first target mask matrix.

[0116] The second computing node calculates the column vector of mask independent variables in the user i mask. The mask dependent variable v in the mask i product This is called the intermediate product, which is the product of the column vectors of independent variables and the dependent variable in the mask. The intermediate products of n users are then summed. get This is called the second target mask vector.

[0117] The first computing node holds the target encrypted regularized symmetric matrix. And the product of the target encrypted independent variable vector and dependent variable The second computing node holds the first target mask matrix μ. A Second target mask vector μ b The first and second computing nodes use any Secure Multi-Party Computation technique to compute the result from the original data of computing node one and computing node two. During the process, the original data of computing node one... and It will not leak to compute node two, and vice versa; that is, the original data μ of compute node two will not be leaked. A and μ b It will not be leaked to computing node one. There are various secure multi-party computation techniques; the process of secure multi-party computation can be as follows:

[0118] Obtained through secure multi-party computation

[0119] Obtained through secure multi-party computation

[0120] Using secure multi-party computation based on the equation The linear regression coefficient vector β is obtained and output to the party that needs β. This can be calculated from node 1 or 2, or other nodes.

[0121] The principle behind the above calculations can be derived as follows:

[0122] Easy to obtain Substitute it into Finally obtained Substitute again get You can Substitution get This verifies the results obtained using secure multi-party computation. It is indeed equal to A.

[0123] On the other hand, we can deduce that Substitute it into get Substitution get This verifies the results obtained using secure multi-party computation. Indeed equals

[0124] There are many secure multi-party computation techniques that can be used for the first and second computing nodes, one of which is secure multi-party computation based on obfuscated circuits. One such secure multi-party computation technique based on obfuscated circuits is as follows:

[0125] The first computing node holds the target encrypted regularized symmetric matrix. And the product of the target encrypted independent variable vector and dependent variable The second computing node holds the rights.

[0126] 1. The first computing node generates a pre-defined obfuscation circuit. This obfuscation circuit can be a Yao-style obfuscation circuit, the core technology of which is to compile the secure computation function involving both parties into a Boolean circuit form and encrypt and scramble the truth table, thereby achieving normal circuit output without leaking the private information of the two parties involved in the computation. Since any secure computation function can be converted into a corresponding Boolean circuit form, it has higher versatility compared to other secure computation methods.

[0127] The first computing node is based on the target encrypted regularized symmetric matrix. And the product of the target encrypted independent variable vector and dependent variable Generate the first obfuscation circuit input result corresponding to the target encryption regularization symmetric matrix and the product of the target encryption independent variable vector and dependent variable. This first obfuscation circuit input result can be expressed as...

[0128] The first computing node inputs the preset obfuscation circuit and the first obfuscation circuit results. Send to the second computing node.

[0129] 2. The second computing node is based on the first target mask matrix μ A Second target mask vector μ b Through unintentional transmission with the first computing node, the first target mask matrix μ is obtained. A Second target mask vector μ b The corresponding input result of the second obfuscation circuit. This second obfuscation circuit input result can be represented as GI(μA; μb).

[0130] 3. The second computing node inputs the results of the first obfuscation circuit described above. The input results GI(μA; μb) from the second obfuscation circuit are then input into the aforementioned preset obfuscation circuit. The aforementioned preset obfuscation circuit... and Subtract (μA; μb) from the middle to recover Then according to the equation The input result GI(β) of the obfuscated circuit corresponding to the linear regression coefficient vector β is obtained. Finally, the obfuscated circuit is used to decode GI(β) to obtain β.

[0131] Example 2

[0132] Figure 3 This is a flowchart illustrating a linear regression data processing method provided in an embodiment of this application. This method can be used for... Figure 1In the implementation environment described herein, the method operation steps are as shown in the embodiments or flowcharts, but based on conventional or non-inventive labor, more or fewer operation steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many steps and does not represent the only execution order. In actual system or server products, the method can be executed in the order shown in the embodiments or drawings or in parallel (e.g., in a parallel processor or multi-threaded processing environment).

[0133] In Example 2, each of the at least two user terminals can contain multiple private data samples.

[0134] For example, such as Figure 3 As shown, the method may include:

[0135] S301. Each of at least two user terminals encrypts a private data sample using its own mask to obtain an encryption matrix and an encryption product.

[0136] The mask held by the user terminal is usually a random mask generated by the user terminal. A random mask is a mask generated by the user terminal that cannot be predicted or known by other user terminals or participating parties.

[0137] For example, the encryption matrix can be C i The encrypted product can be Suppose there are a total of n user terminals, where n is a positive integer greater than or equal to 2. The i-th user terminal is called user terminal i, where i is a positive integer greater than or equal to 1 and less than or equal to n.

[0138] Optionally, the at least two user terminals can be multiple participants in a horizontal federated learning process.

[0139] User terminals may choose not to add other encryption schemes, or they may add other encryption schemes themselves, including but not limited to homomorphic encryption.

[0140] For example, S301 may include:

[0141] Based on private data samples and masks, each of at least two user terminals obtains a first matrix;

[0142] User terminal i holds ki (ki is greater than or equal to 1) private data samples, each of which includes d independent variables, i.e., a column vector composed of d features. as well as The corresponding dependent variable y i,j y i,jAlso known as output variables. The mask generated by user terminal i includes mi (mi ≥ 1) mask independent variable vectors and corresponding mask dependent variables. Each mask independent variable vector and its corresponding mask dependent variable also contains d independent variables, which are column vectors composed of d features. And one dependent variable v i,l (l=1,2,…,mi).

[0143] Each user terminal i concatenates the ki column vectors of independent variables from the ki private data samples with the mi column vectors of independent variables from the mask to obtain the first matrix.

[0144] In an optional embodiment, S301 may include the following steps:

[0145] S30101 User terminal i compares the size of ki+mi with d. If (ki+mi) minus d is less than a specific value (e.g., 0), proceed to step S30101-1; if (ki+mi) minus d is greater than or equal to a specific value (e.g., 0), proceed to step S30101-2.

[0146] User terminal i in S30101-1 sets the transmission format identifier to the transmission vector multiplied by an orthogonal transform. The transmission format identifier can be set to 1.

[0147] Each of at least two user terminals will use the first matrix Π i Multiplying it by a preset matrix yields the encryption matrix.

[0148] The preset matrix includes

[0149] An orthogonal matrix, or a matrix obtained by multiplying an orthogonal matrix by a coefficient.

[0150] For example, this orthogonal matrix can be written as Θ i It includes (ki+mi) rows and (ki+mi) columns. The first matrix Π i Multiply by an orthogonal matrix Θ i Obtain the encryption matrix C i =Π i Θ i This encryption matrix can be used in C. i This means that the encryption matrix is ​​obtained by multiplying the first matrix and the preset orthogonal matrix. The specific operation process of the encryption matrix can be C. i =Π i Θ i Proceed to step S30102.

[0151] User terminal i of S30101-2 sets the transmission format identifier to a transmission symmetric matrix. The transmission format identifier can be set to 0.

[0152] Each of at least two user terminals will use the first matrix Π i and its transpose (The data is usually real numbers, so this is a transpose matrix) If the data is complex, then it is the conjugate transpose matrix. Multiply by , and you get the encryption matrix. This encryption matrix can be used in C. i This indicates that the process will proceed to step S30102.

[0153] S30102 For each user terminal i in at least two user terminals, calculate the product of the column vectors of ki independent variables and the corresponding dependent variables, and sum them to obtain the sum of the first product. Right now

[0154] For each of at least two user terminals, calculate the product of the mi column vectors of the mask with their corresponding dependent variables and sum them to obtain the sum of the second product.

[0155] Right now

[0156] Each of at least two user terminals obtains an encrypted product based on the sum of the first product and the sum of the second product. Right now

[0157] Therefore, the encrypted product is the sum of the first product and the sum of the second product.

[0158] S303. Each of at least two user terminals sends, or equivalently sends, the encrypted data to the first computing node. The encrypted data includes an encryption matrix C. i and encrypted product Furthermore, each of at least two user terminals sends, or equivalently sends, a transmission format identifier to the first computing node. This equivalent transmission can be implemented, but is not limited to, in the following manner: each user terminal sends its own transmitted encryption matrix C. i The dimensions are sent to the first compute node; if C i If it is a d-row, d-column matrix, then the transmitted C i It is a symmetric matrix; otherwise, the transmitted C... i It is a vector multiplication or orthogonal transformation.

[0159] S305. At least two user terminals send the mask, or an equivalent one, to the second computing node. The mask comprises mi column vectors of independent variables and corresponding dependent variables, as described above. and v i,l (where l = 1, 2, ..., mi).

[0160] S307. Based on the encrypted data and the mask, the linear regression coefficients are obtained through the cooperation of the first computing node and the second computing node.

[0161] For example, S307 may include:

[0162] For each user terminal, i.e., user terminal i, the first computing node receives the transmission format identifier sent by user terminal i.

[0163] If the transmission format identifier indicates that a vector multiplication or orthogonal transformation is being transmitted, then the first computing node performs symmetric processing on the encryption matrix of user terminal i to generate a symmetric matrix, which can be C. i Then the symmetric matrix can be C. i and C i The product of the transposes of the matrices, i.e. On the other hand, if the transmission format identifier indicates that a symmetric matrix is ​​being transmitted, then the first computation node directly uses the received C... i As a symmetric matrix, i.e. Ξ i =C i .

[0164] The first computing node accumulates a symmetric matrix Ξ of n users. i (i = 1, 2, ..., n), obtain the sum of n user symmetric matrices.

[0165] The first computing node obtains the preset regularization coefficients and the preset identity matrix.

[0166] The regularization coefficient can be represented as λ, and the preset identity matrix can be represented as I.

[0167] The first computing node is based on the sum of n user symmetric matrices. Generate the target encryption regularized symmetric matrix by setting the preset regularization coefficients and the preset identity matrix.

[0168] The first computing node performs masked encryption on the product of the independent variable vector and the dependent variable from the n received user terminals. Summing is performed to obtain the target product.

[0169] The first computing node and the second computing node must not collude, meaning they cannot conspire to spy on the data of at least two user terminals.

[0170] The second computing node obtains the target mask; the target mask is generated based on the masks corresponding to at least two user terminals.

[0171] The second computing node performs symmetric processing on the mi column vectors of independent variables in the mask of user i to obtain a symmetric matrix. Right now The symmetric processing results of n user terminals are then summed to obtain the first target mask matrix.

[0172] The second computing node calculates the encrypted product of the sum of the products of the mi column vectors of independent variables and the corresponding mi dependent variables in the user i mask. Right now The sum η of the products of the column vectors of independent variables and the dependent variable in the mask is called η. i The sum η of the products of the independent variable column vectors and the dependent variable in the above masks for n users. i get This is called the second target mask vector.

[0173] The first computing node holds the target encrypted regularized symmetric matrix. And the product of the target encrypted independent variable vector and dependent variable The second computing node holds the first target mask matrix μ. A Second target mask vector μ b The first and second computing nodes use any Secure Multi-Party Computation technique to compute the result from the original data of computing node one and computing node two. During the process, the original data of computing node one... and It will not leak to compute node two, and vice versa; that is, the original data μ of compute node two will not be leaked. A and μ b It will not be leaked to computing node one. There are various secure multi-party computation techniques; the process of secure multi-party computation can be as follows:

[0174] Obtained through secure multi-party computation

[0175] Obtained through secure multi-party computation

[0176] Using secure multi-party computation based on the equation The linear regression coefficient vector β is obtained and output to the party that needs β. This can be calculated from node 1 or 2, or other nodes.

[0177] This application may be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this application.

[0178] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0179] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0180] The computer program instructions used to perform the operations of this application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), are personalized by utilizing the status information of the computer-readable program instructions. These electronic circuits can execute the computer-readable program instructions to implement various aspects of this application.

[0181] Various aspects of this application are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0182] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0183] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0184] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternatives...

[0185] In implementation, the functions marked in the boxes may occur in a different order than those shown in the accompanying drawings. For example, two consecutive boxes may actually be executed in substantially parallel order, or sometimes in reverse order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as combinations of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or actions, or using a combination of dedicated hardware and computer instructions.

[0186] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technological improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A privacy-preserving linear regression method, characterized in that, The method includes: At least two user terminals encrypt their own private data samples using their own masks to obtain an encryption matrix and an encryption product; the private data samples include independent variable column vectors and dependent variables, and the mask includes mask independent variable column vectors and mask dependent variables; The at least two user terminals encrypt their own private data samples using their own masks to obtain an encryption matrix and encryption product, including: Based on the column vector of independent variables in the private data sample and the column vector of independent variables in the mask, each of the at least two user terminals obtains a first matrix; the first matrix is ​​obtained by each of the at least two user terminals concatenating the column vector of independent variables in the private data sample and the column vector of independent variables in the mask. Each of at least two user terminals multiplies the first matrix and a preset matrix to obtain the encryption matrix; the preset matrix includes: an orthogonal matrix, or a matrix obtained by multiplying an orthogonal matrix with a coefficient; Each of at least two user terminals calculates the product of the independent variable column vector and the dependent variable to obtain a first product; Each of at least two user terminals calculates the product of the mask independent variable column vector and the mask dependent variable to obtain a second product; Each of at least two user terminals obtains the encrypted product based on the first product and the second product; the at least two user terminals send or equivalently send the encryption matrix and the encrypted product to the first computing node; the encrypted product is the sum of the first product and the second product; At least two user terminals send the mask or equivalently send it to the second computing node; Based on the encryption matrix and the mask, the linear regression coefficients are obtained through the cooperation of the first computing node and the second computing node.

2. The method according to claim 1, characterized in that, The at least two user terminals sending, or equivalently sending, the encryption matrix and the encryption product to the first computing node includes: At least two user terminals will send the first preset index information to the first computing node; The first computing node obtains the encryption matrix and the encryption product based on the preset index information.

3. The method according to claim 1, characterized in that, The at least two user terminals sending or equivalently sending the mask to the second computing node includes: The second computing node sends the mask to at least two user terminals; or, At least two user terminals send the second preset index information to the second computing node; The second computing node obtains the mask based on the second preset index information; or, The second computing node obtains the mask, which is preset by the second computing node and the at least two user terminals.

4. The method according to claim 1, characterized in that, The step of obtaining the linear regression coefficients based on the encryption matrix and the mask through the cooperation of the first computing node and the second computing node includes: The first computing node obtains the target encrypted regularized symmetric matrix and the target product based on the encryption matrix and the encryption product; The second computing node obtains a first target mask matrix and a second target mask vector based on the mask. The first computing node and the second computing node obtain the linear regression coefficients based on the target encryption regularized symmetric matrix and the target product, as well as the first target mask matrix and the second target mask vector, using secure multi-party computation techniques. The first computing node obtains the target encrypted regularized symmetric matrix and the target product based on the encryption matrix and the encryption product, including: The first computing node performs symmetry processing on the encryption matrix to generate a symmetric matrix; The first computing node sums the symmetric matrix to obtain the symmetric matrix sum; The first computing node obtains a preset regularization coefficient and a preset identity matrix; The first computing node generates the target encrypted regularized symmetric matrix based on the symmetric matrix, the preset regularized symmetric coefficients, and the preset identity matrix; The first computing node sums the encrypted products to obtain the target product; The second computing node obtains a first target mask matrix and a second target mask vector based on the mask. The mask includes a column vector of mask independent variables and a mask dependent variable, including: The second computing node performs symmetric processing on the column vector of the mask independent variables to obtain a symmetric mask matrix; The second computing node accumulates the mask symmetric matrix to obtain the first target mask matrix; The second computing node calculates the product of the mask independent variable column vector and the mask dependent variable to obtain an intermediate product; The second computing node accumulates the intermediate products to obtain the second target mask vector.

5. The method according to claim 1, wherein in the case where at least two target user terminals contain multiple private data samples, each of the multiple private data samples includes multiple independent variables, the mask includes one or more mask independent variable vectors, and the method further includes: Each of at least two user terminals calculates the sum of the number of the plurality of private data samples and the number of the one or more mask independent variable vectors to obtain the sample mask number; Each of at least two user terminals compares the number of sample masks and the number of the plurality of independent variables, and sets the transmission format identifier to the transmission vector multiplied by an orthogonal transformation or a transmission symmetric matrix.

6. The method according to claim 5, wherein each of the at least two user terminals compares the number of sample masks and the number of the plurality of independent variables, and sets the transmission format identifier to a transmission vector multiplied by an orthogonal transformation or a transmission symmetric matrix, comprising: Each of at least two user terminals calculates the difference between the number of sample masks and the number of the plurality of independent variables to obtain a standard value; When the standard value is less than a preset specific value, the transmission format identifier is set to the transmission vector multiplied by an orthogonal transformation; When the standard value is greater than or equal to the preset specific value, the transmission format identifier is set as a transmission symmetric matrix.