Matrix transmission method and computer equipment

By generating a binary second plaintext matrix and plaintext vector instead of dense bucket matrix transmission, the problem of large traffic is solved, efficient and secure data transmission is achieved, and data transmission is suitable for data transmission in machine learning model training.

CN120433922APending Publication Date: 2025-08-05ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510530413.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

When multi-party joint training of machine learning model, the transmission of the dense bucket matrix in the prior art occupies more communication resources, especially when the purchaser's data privacy requirements are higher, resulting in large and insecure communication volume.

Method used

The first party generates a binary second plaintext matrix with the same size as the first plaintext matrix, and generates a corresponding first plaintext vector. By sending the second ciphertext matrix and the plaintext vector, the second party recovers the first ciphertext matrix based on the received data.

Benefits of technology

It reduces the traffic volume, improves the security and efficiency of data transmission, and reduces the communication pressure, especially in scenarios with high privacy requirements, which significantly reduces the data volume, improves the communication pressure and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120433922A_ABST
    Figure CN120433922A_ABST
Patent Text Reader

Abstract

The invention provides a matrix transmission method and computer equipment. The first party firstly generates a second plaintext matrix with the same size according to the size of the first plaintext matrix. The first party has f first plaintext matrixes, and the f first plaintext matrixes respectively represent binning results of f types of feature data owned by the first party. Correspondingly, the first plaintext matrix and the second plaintext matrix are binary matrixes, each column comprises a first value, and elements of the first values represent binning results of samples corresponding to the column. The first party then generates a first plaintext vector corresponding to each first plaintext matrix. The first plaintext vector is used for representing the position difference of the first value in the column corresponding to the element of the first plaintext matrix and the second plaintext matrix. And then the first party encrypts the second plaintext matrix to obtain a second ciphertext matrix, and sends the second ciphertext matrix and the f first plaintext vectors to the second party. And the second party recovers the first ciphertext matrixes respectively corresponding to the f first plaintext matrixes according to the received data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification belong to the field of computer application technology, and more particularly, to a matrix transmission method and computer equipment. Background Art

[0002] With the advancement of information technology, it's becoming increasingly common for multiple parties to jointly train machine learning models. For example, in the financial sector, a second party might want to train a binary classification model to determine whether a user is at risk of default. However, this second party only possesses the labels for multiple training samples, which indicate whether a user is at risk of default, but lacks the feature data necessary for model training. To complete model training, the second party must either procure the feature data or collaborate with the first party, which possesses the feature data.

[0003] In this case, to reduce unnecessary computation, it's necessary to first determine the predictive power of each feature the first party intends to use for training for the label value. In other words, it's necessary to determine whether a particular feature is useful for the training objectives of the machine learning model. Both parties then need to determine the ratio of positive and negative samples within each bucket based on the feature's bucketing results, and based on this ratio, determine whether the feature's predictive power for the label value meets the requirements. For example, the Information Value (IV) can be calculated based on the ratio of positive and negative samples within each bucket, and the IV value can be used to determine whether the feature should be used for model training.

[0004] In some cases, however, the privacy level of the tag values held by the second party is higher, and even encrypted tag values cannot be sent to the first party. For example, in a scenario where the second party needs to purchase feature data from the first party, the privacy level of the second party's tag values is higher than the privacy level of the feature data held by the first party. Therefore, to determine the ratio of positive and negative samples in each bucket, the first party needs to send a secret bucket matrix to the second party.

[0005] The schematic diagram of the dense bucket matrix is as follows Figure 1 As shown, the size of the dense bucket matrix is b*n, and each element in the dense bucket matrix is homomorphically encrypted. Figure 1 The green elements in the middle represent the homomorphic encryption. Figure 1 The size of each element in the dense bucket matrix shown in is the size of the corresponding plaintext. Each column of the dense bucket matrix represents a training example, and each row represents a bucket. In each column, only one element has a value of 1, indicating that the training example falls into the bucket corresponding to the element with a value of 1. The remaining elements have a value of 0.

[0006] Furthermore, if Figure 1As shown, the second party can perform the ciphertext multiplication of the ciphertext and plaintext between the dense bucket matrix and the label vector corresponding to each feature to obtain the ciphertext vector representing the number of positive samples in each bucket (i.e. Figure 1 ). That is, [2,0,2,0] means that the first bucket has two positive samples, the third bucket has two positive samples, and the remaining buckets have no positive samples.

[0007] In the above process, if the first party has multiple types of feature data, it needs to send multiple dense bucket matrices to the second party. However, the transmission of dense bucket matrices occupies a lot of communication resources. Summary of the Invention

[0008] The purpose of this specification is to provide a matrix transmission method and computer equipment.

[0009] A first aspect of the present disclosure provides a matrix transmission method, applied to a first party having f first plaintext matrices, wherein each column of the first plaintext matrices includes a first value; the method comprising:

[0010] generating a second binarized plaintext matrix of the same size as the first plaintext matrix, wherein each column of the second plaintext matrix includes a first value;

[0011] For each first plaintext matrix, generate a corresponding first plaintext vector; the first plaintext vector includes the same number of elements as the number of columns of the first plaintext matrix, and the value of each element is used to represent the position difference between the first value in the column corresponding to the element in the first plaintext matrix and the second plaintext matrix;

[0012] The second ciphertext matrix corresponding to the second plaintext matrix and the f first plaintext vectors are sent to the second party, so that the second party obtains the first ciphertext matrices corresponding to the f first plaintext matrices respectively according to the second ciphertext matrix and the f first plaintext vectors.

[0013] A second aspect of this specification provides a matrix transmission method, applied to a second party, comprising:

[0014] Receive a second ciphertext matrix and f first plaintext vectors from a first party; wherein the first party has f first plaintext matrices, a second plaintext matrix corresponding to the second ciphertext matrix is a binary matrix of the same size as the first plaintext matrix, and each column in the first plaintext matrix and the second plaintext matrix includes a first value; the first plaintext vectors and the first plaintext matrices have a one-to-one correspondence, the number of elements included in the first plaintext vector is the same as the number of columns in the first plaintext matrix, and the value of each element of any first plaintext vector is used to represent the positional difference of the first values in the columns corresponding to the element in the first plaintext matrix and the second plaintext matrix;

[0015] Obtain f first ciphertext matrices according to the second ciphertext matrix and f first plaintext vectors; the f first ciphertext matrices correspond to the f first plaintext matrices.

[0016] A third aspect of the present specification provides a matrix transmission device, applied to a first party having f first plaintext matrices, wherein each column of the first plaintext matrices includes a first value; the device includes:

[0017] a matrix generation module, configured to generate a second binary plaintext matrix of the same size as the first plaintext matrix; each column of the second plaintext matrix includes a first value;

[0018] a vector generation module, configured to generate, for each first plaintext matrix, a corresponding first plaintext vector; the first plaintext vector includes the same number of elements as the number of columns of the first plaintext matrix, and the value of each element represents a positional difference between a first value in a column corresponding to the element in the first plaintext matrix and a first value in a column corresponding to the element in the second plaintext matrix;

[0019] A sending module is used to send a second ciphertext matrix corresponding to the second plaintext matrix and f first plaintext vectors to the second party, so that the second party obtains first ciphertext matrices corresponding to the f first plaintext matrices respectively according to the second ciphertext matrix and the f first plaintext vectors.

[0020] A fourth aspect of this specification provides a matrix transmission device, applied to a second party, comprising:

[0021] A receiving module is configured to receive a second ciphertext matrix and f first plaintext vectors from a first party; wherein the first party has f first plaintext matrices, and the second plaintext matrix corresponding to the second ciphertext matrix is a binary matrix of the same size as the first plaintext matrix, and each column in the first plaintext matrix and the second plaintext matrix includes a first value; the first plaintext vectors and the first plaintext matrices have a one-to-one correspondence, the number of elements included in the first plaintext vectors is the same as the number of columns in the first plaintext matrix, and the value of each element of any first plaintext vector is used to represent the positional difference of the first values in the columns corresponding to the element in the first plaintext matrix and the second plaintext matrix;

[0022] A recovery module is configured to obtain f first ciphertext matrices based on the second ciphertext matrix and f first plaintext vectors; the f first ciphertext matrices correspond to the f first plaintext matrices.

[0023] A fifth aspect of this specification provides a computer program product, comprising a computer program / instruction, which implements the steps of the above-mentioned matrix transmission method when executed by a processor.

[0024] A sixth aspect of the present specification provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed in a computer, the computer is caused to execute the above-mentioned matrix transmission method.

[0025] A seventh aspect of this specification provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the above-mentioned matrix transmission method is implemented.

[0026] In order to solve the problem of large communication volume in the related art, this specification provides a matrix transmission method. The first party first generates a second plaintext matrix of the same size based on the size of the first plaintext matrix. The first party has f first plaintext matrices, and the f first plaintext matrices respectively represent the binning results of f types of feature data owned by the first party. Correspondingly, the first plaintext matrix and the second plaintext matrix are binary matrices, and each column includes a first value, and the elements of the first value represent the binning results of the samples corresponding to the column. Then the first party generates a first plaintext vector corresponding to each first plaintext matrix. The first plaintext vector is used to characterize the position difference of the first value in the column corresponding to the element of the first plaintext matrix and the second plaintext matrix. Then the first party encrypts the second plaintext matrix to obtain a second ciphertext matrix, and sends the second ciphertext matrix and the f first plaintext vectors to the second party.

[0027] The second party recovers the first ciphertext matrices corresponding to the f first plaintext matrices based on the received data.

[0028] Compared to the prior art method of sending f encrypted bucket matrices when there are f features, in this specification, only a second ciphertext matrix of the same size as the encrypted bucket matrix and f plaintext vectors are sent in this case. Because the size of f plaintext vectors is smaller than the size of f-1 encrypted bucket matrices, the communication volume between the two parties is reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0030] Figure 1 It is a schematic diagram of the matrix-vector multiplication performed by two parties in the process of joint modeling;

[0031] Figure 2 It is a flow chart of a matrix transmission method;

[0032] Figure 3It is a schematic diagram of a matrix transmission method;

[0033] Figure 4 It is a schematic diagram of another matrix transmission method;

[0034] Figure 5 It is a block diagram of a matrix transmission device;

[0035] Figure 6 It is a block diagram of another matrix transmission device. DETAILED DESCRIPTION

[0036] To help those skilled in the art better understand the technical solutions in this specification, the following will provide a clear and complete description of the technical solutions in the embodiments of this specification, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. All other embodiments derived by those skilled in the art based on the embodiments in this specification without creative effort shall fall within the scope of protection of this specification.

[0037] In order to solve the problem of large communication volume in the related art, this specification provides a matrix transmission method. The first party first generates a second plaintext matrix of the same size based on the size of the first plaintext matrix. The first party has f first plaintext matrices, and the f first plaintext matrices respectively represent the binning results of f types of feature data owned by the first party. Correspondingly, the first plaintext matrix and the second plaintext matrix are binary matrices, and each column includes a first value, and the elements of the first value represent the binning results of the samples corresponding to the column. Then the first party generates a first plaintext vector corresponding to each first plaintext matrix. The first plaintext vector is used to characterize the position difference of the first value in the column corresponding to the element of the first plaintext matrix and the second plaintext matrix. Then the first party encrypts the second plaintext matrix to obtain a second ciphertext matrix, and sends the second ciphertext matrix and the f first plaintext vectors to the second party.

[0038] The second party recovers the first ciphertext matrices corresponding to the f first plaintext matrices based on the received data.

[0039] Compared to the prior art method of sending f encrypted bucket matrices when there are f features, in this specification, only a second ciphertext matrix of the same size as the encrypted bucket matrix and f plaintext vectors are sent in this case. Because the size of f plaintext vectors is smaller than the size of f-1 encrypted bucket matrices, the communication volume between the two parties is reduced.

[0040] Next, we will combine Figure 2 A matrix transmission method shown in this specification is described below. Figure 2There are two parties involved, the first party and the second party. The first party has f first plaintext matrices, where each column in the first plaintext matrices contains a first value. The second party needs to obtain f first ciphertext matrices corresponding to each of the f first plaintext matrices.

[0041] Next, the first party and the second party will be described in conjunction with a specific scenario. This scenario description does not limit this specification.

[0042] In a scenario where two parties jointly train a machine learning model, the first party is the party that owns the feature data, and the second party is the party that owns the label value. The two parties need to use the methods described in this specification to determine whether the various types of data owned by the first party are suitable for training the machine learning model. For example, in the financial field, the second party needs to train a binary classification model to determine whether the user has a risk of financial default. The second party only has the label value, and in order to improve the quality of the model, the second party needs to purchase feature data. In this case, the second party is also called the purchaser, and the first party is also called the data party.

[0043] In this scenario, it is necessary to calculate the multiplication of the bucket matrix and the label vector for each feature. To ensure data security, the first party needs to send multiple encrypted bucket matrices to the second party, or the second party needs to send the encrypted label vector to the first party. However, since the purchaser has higher data privacy requirements in the above-mentioned procurement scenario, the label vector of the purchaser (second party) cannot be sent to the data party (first party) even after encryption. Therefore, in order for the purchaser to obtain data quality indicators (that is, data indicating whether various types of feature data are suitable for model training in the current scenario), the data party needs to send the ciphertext of the bucket data to the purchaser, so that the purchaser can quickly perform homomorphic multiplication between the encrypted bucket matrix (first ciphertext matrix) and the label vector.

[0044] In this case, the first party has f types of features. Accordingly, the first party can bucket each of these f types of features. Bucketing involves dividing the samples into a number of buckets based on the values of the feature data. Specific bucketing methods include equal-width bucketing, equal-frequency bucketing, and decision tree bucketing, and are not limited in this specification.

[0045] Furthermore, the first party can obtain f plaintext bucket matrices, which are also the first plaintext matrices mentioned above. Each column in the first plaintext matrix represents a sample, and each row represents a bucket. The first plaintext matrix is a binary matrix, and the value of each element is a first value and a second value. In an optional embodiment, the first value is 1 and the second value is 0. Similar to the dense bucket matrix mentioned above, each column of the first plaintext matrix includes a first value and several second values. The meaning of the first value is: the sample falls into the bucket corresponding to the first value. For example, Figure 1 In the dense bucket matrix shown, the element at the first position in the first column has a value of 1 (first value), which means that the first sample falls into the first bucket.

[0046] In other words, the f first plaintext matrices represent the bucketing results of the f types of feature data held by the first party. Each column in the first plaintext matrix corresponds to a sample, and the position of the element with the first value in each column represents the bucket into which the element in that column falls. In other words, the first plaintext matrix is a 0 / 1 matrix with b rows and n columns, with each column containing a single 1 and all other elements being 0. If the element in the i-th row and j-th column is 1, it means that the j-th sample is assigned to the i-th bucket.

[0047] Furthermore, the first and second parties may perform Figure 2 The steps shown are performed so that the second party can obtain f first ciphertext matrices. Figure 2 In the steps shown, 21x represents the steps performed by the first party, and 22x represents the steps performed by the second party. Figure 2 As shown, the method includes:

[0048] Step 211: The first party generates a binarized second plaintext matrix of the same size as the first plaintext matrix.

[0049] Each column of the second plaintext matrix includes a first value.

[0050] Specifically, the second plaintext matrix is a matrix with the same characteristics and size as the first plaintext matrix. The position of the first value in each column of the second plaintext matrix can be randomly generated.

[0051] In step 213 , the first party generates a corresponding first plaintext vector for each first plaintext matrix.

[0052] The number of elements included in the first plaintext vector is the same as the number of columns in the first plaintext matrix, and the value of each element is used to represent the position difference of the first value in the column corresponding to the element in the first plaintext matrix and the second plaintext matrix.

[0053] Specifically, the first party generates f first plaintext vectors based on the positional relationship between the first values in the first plaintext matrix and the second plaintext matrix. The elements in the first plaintext vectors correspond one-to-one to the columns in the second ciphertext matrix, indicating what operation needs to be performed on each column of the second ciphertext matrix to obtain the first ciphertext matrix. The second party can then recover the first ciphertext matrix corresponding to the first plaintext vector based on the first plaintext matrix and the first plaintext vector.

[0054] The specific form of the first plaintext vector will be described here through a specific embodiment. It should be understood that this embodiment does not limit this specification.

[0055] In an optional implementation, the value i of the j-th element in any first plaintext vector indicates that each element in the j-th column of the second plaintext matrix is shifted by i bits in a preset direction to obtain the j-th column element of the first plaintext matrix corresponding to the first plaintext vector.

[0056] The preset direction may be a downward direction or an upward direction. In addition, if an element reaches the last row after being shifted downward by a number of bits, the remaining bits may be used to shift the element to the position of the first element in the column.

[0057] Here we will use a specific example to illustrate the first plaintext vector. Figure 3 As shown, the first party originally has f secret bucket matrices, and the plaintext bucket matrices corresponding to the multiple secret bucket matrices (also called the first plaintext matrix). In order to distinguish between plaintext and ciphertext, and Figure 1 similar, Figure 3 The green portion represents the ciphertext, and the white portion represents the plaintext. The encrypted bucket matrix is the matrix obtained by encrypting the first plaintext matrix, and is hereafter referred to as the first ciphertext matrix. Following the above steps, the encrypted bucket matrix is converted into a encrypted random matrix and f column rotation factors. The encrypted random matrix is the matrix obtained by encrypting the second plaintext matrix, and is hereafter referred to as the second ciphertext matrix. The column rotation factors are the first plaintext vectors.

[0058] The f column rotation factors correspond one-to-one to the f features. For each column rotation factor, the jth element has a value of i, which means that each element in the jth column of the second ciphertext matrix needs to be rotated by i offset values. After rotating each column, the first ciphertext matrix corresponding to that column rotation factor is obtained.

[0059] for example, Figure 3 In the example, the first element of the first column rotation factor is 2, which means that the first column of the dense random matrix is rotated down by two positions. The vector corresponding to the first column of the dense random matrix is originally [0,0,1,0] T , in this example, rotating downward means rotating to the right, and after rotation we can get [1,0,0,0] T .

[0060] In step 215 , the first party sends the second ciphertext matrix corresponding to the second plaintext matrix and the f first plaintext vectors to the second party.

[0061] In this way, the second party can obtain the first ciphertext matrices corresponding to the f first plaintext matrices according to the second ciphertext matrix and the f first plaintext vectors.

[0062] Correspondingly, the second party executes step 225 (not shown in the figure) to receive the second ciphertext matrix and f first plaintext vectors from the first party.

[0063] Specifically, the second plaintext matrix can be encrypted to obtain a second ciphertext matrix, and the second ciphertext matrix and f first plaintext vectors can be sent to the second party. Compared to the prior art method of sending f encrypted bucket matrices of the same size as the second ciphertext matrix, since the size occupied by a single plaintext is smaller than the size of the ciphertext, and instead of sending f matrices, one matrix and f vectors are sent, this greatly reduces the amount of data that needs to be sent, reducing communication pressure.

[0064] The encryption method for the second plaintext matrix can be any encryption method. Each element in the second plaintext matrix is encrypted separately. Furthermore, in the aforementioned two-party joint modeling scenario, where the second party needs to calculate the multiplication of the recovered first ciphertext matrix and the plaintext ciphertext of the label vector, homomorphic encryption can be used to encrypt each element of the first plaintext matrix separately to obtain the second ciphertext matrix.

[0065] In addition, the above method encrypts each element separately. Since even the same value will show different results in different encryptions, the encrypted second ciphertext matrix will not leak any information.

[0066] In step 227 , the second party obtains f first ciphertext matrices according to the second ciphertext matrix and the f first plaintext vectors.

[0067] The f first ciphertext matrices correspond to the f first plaintext matrices.

[0068] Specifically, the second party ultimately needs to obtain the f first ciphertext matrices corresponding to the f first plaintext matrices. Therefore, after obtaining the f first plaintext vectors and the second ciphertext matrix, the f first ciphertext matrices must be recovered from this data. This recovery method is the inverse of the process for generating the first plaintext vectors described above. That is, for each first plaintext vector, the order of the elements in the corresponding column of the second ciphertext matrix is adjusted based on the values of the elements in the first plaintext vector to recover the first ciphertext matrix.

[0069] The implementation of step 227 will be described here with reference to the example given in step 213. Specifically, as described above, the value i of the j-th element in any first plaintext vector represents shifting each element in the j-th column of the second plaintext matrix by i bits in a preset direction to obtain the j-th column element of the first plaintext matrix corresponding to the first plaintext vector.

[0070] Correspondingly, step 227 specifically includes: for each j-th element of each first plaintext vector, shifting each element in the j-th column of the second ciphertext matrix by i bits in a preset direction to obtain the j-th column of the first plaintext matrix corresponding to the first plaintext vector, where i is the value of the j-th element in the first plaintext vector.

[0071] Here we still use the above example, Figure 3 As shown, the first column rotation factor is [2,0,1,1]. The plaintext corresponding to the secret random matrix is: After rotating according to the column rotation factor, the plaintext corresponding to the first ciphertext matrix is:

[0072] In addition to the first plaintext vector mentioned above, in another optional embodiment, the first party may also generate a corresponding second plaintext vector for each first plaintext matrix; the second plaintext vector has the same number of elements as the first plaintext vector, and the value of each element in the second plaintext vector is used to represent the position of the column of the second plaintext matrix corresponding to the element of the first plaintext vector corresponding to the element in the second plaintext matrix; and the f second plaintext vectors are sent to the second party.

[0073] Correspondingly, the second party can also receive f second plaintext vectors from the first party; the second plaintext vector has the same number of elements as the first plaintext vector, and the value of each element in the second plaintext vector is used to represent the position of the column of the second plaintext matrix corresponding to the element of the first plaintext vector corresponding to the element in the second plaintext matrix.

[0074] Correspondingly, step 227 may specifically include obtaining the value i1 of the j-th element of each first plaintext vector, and the value i2 of the j-th element of the second plaintext vector corresponding to the first plaintext vector; and shifting each element of the i2-th column in the second ciphertext matrix by i1 bits in a preset direction as the vector of the j-th column in the first ciphertext matrix corresponding to the first plaintext vector.

[0075] Specifically, if only the first plaintext vector is generated, the second party may be able to determine the relationship between the bucketing results for each sample for different features based on the element sizes for the same sample in multiple first plaintext vectors. For example, if the first element of the first plaintext vector is 0, and the first element of the second plaintext vector is 2, then the difference between the bucketing results for the first sample under the first feature and the bucketing results under the second feature can be determined to be 2.

[0076] Based on this, this specification also generates a second plaintext vector, and the value of each element in the second plaintext vector is used to represent the position of the column of the second plaintext matrix corresponding to the element of the first plaintext vector in the second plaintext matrix.

[0077] Here we will target Figure 4 The method provided in this specification is specifically described with reference to examples. Figure 4 Omitted Figure 3 The left side of the figure represents the solution of the prior art, which only shows the solution provided in this specification. Figure 3 Compared with the scheme, Figure 4 The scheme generates an additional column index for each dense bucket matrix, which is the second plaintext vector. In order to distinguish between plaintext and ciphertext, and Figure 1 similar, Figure 4 The green part represents the ciphertext, and the white part represents the plaintext

[0078] For the meaning of the second plaintext vector, for example, the first element 3 of the first second plaintext vector refers to the result of rotating the elements in the fourth column of the second ciphertext matrix (i.e., the secret random matrix) according to the value of the first element of the first plaintext vector, as the corresponding element in the first column of the first ciphertext matrix. For example, the original plaintext corresponding to the element in the first column of the second ciphertext matrix is [0,0,1,0] T , and the first element of the first plaintext vector is 2, and the first second plaintext vector is 3. Then the result of rotating the elements in the fourth column of the second ciphertext matrix by 2 offsets is used as the elements in the first column of the first ciphertext matrix. Figure 4 The plaintext in the fourth column of the second ciphertext matrix is [0,0,0,1] T , after rotating the vector by two offsets, we get [0,1,0,0] T The corresponding ciphertext can then be [0,1,0,0] T The corresponding ciphertext is used as the element of the first column of the first ciphertext matrix.

[0079] According to the above method, the plaintext corresponding to the second ciphertext matrix (secret random matrix) is The first plaintext vector (column rotation factor) is [2,0,1,1], and the first second plaintext vector (column index) is [3,1,2,0]. Then the second party recovers the plaintext of the first ciphertext matrix based on these matrices and vectors.

[0080] In this way, the second party can obtain less data and security is further guaranteed.

[0081] like Figure 1 As shown in the figure, in the scenario of joint modeling between two parties, the second party also needs to calculate the homomorphic multiplication of the dense bucket matrix and the label vector for each feature to obtain the proportion of various samples in each bucket. Figure 1In this article, we use a binary classification model as an example. This means we need to find the number of positive samples in each bucket to determine whether the feature is suitable for model training.

[0082] In other words, the second party holds a plaintext label vector, each element of which represents the label value of each sample in the first plaintext matrix. The second ciphertext matrix is encrypted using a homomorphic encryption algorithm. After executing step 227, the second party also determines a ciphertext result vector for each feature based on the product of the first ciphertext matrix corresponding to the feature and the plaintext label vector, and determines the contribution of the feature to the label value based on the ciphertext result vector.

[0083] The product of the first ciphertext matrix and the plaintext label vector refers to the product of the first ciphertext matrix and the plaintext label vector obtained through homomorphic multiplication. Since homomorphic encryption supports: E(a)*b=E(a*b), E(a)+E(b)=E(a+b), the first ciphertext matrix obtained through homomorphic encryption can support matrix-vector multiplication. Here, E(a) represents homomorphic encryption of a, E(a)*b refers to the homomorphic multiplication between the two, and E(a)+E(b) represents the homomorphic addition between the two.

[0084] Next, the communication traffic reduction effect of the above method will be further explained.

[0085] Specifically, in a two-party joint modeling scenario, assume the first plaintext matrix has b rows and n columns, meaning the number of buckets is b and the number of samples is n. The second party holds a total of f types of features. In the prior art, a bucket matrix contains nb elements, and f bucket matrices contain nbf elements. The data size of each element after homomorphic encryption is logq. Therefore, the total amount of data to be transmitted is nbflogq.

[0086] Correspondingly, Figure 3 In the scheme shown, it is necessary to pass a second ciphertext matrix of size nblogq and f first plaintext vectors, each of which contains n elements and each element has a size of logb (since the maximum value of each element is b, the size it occupies is logb). Then Figure 3 The method shown needs to transfer data of nblogq+nflogb.

[0087] Figure 4 In the scheme shown, it is necessary to pass a second ciphertext matrix, f first plaintext vectors and f second plaintext vectors. For the second plaintext vector, it includes n elements, and the size of each element is logn (since the maximum value of each element is n, the size it occupies is logb). Then Figure 4The method shown needs to transfer data of nblogq+nflogb+nflogn.

[0088] For security reasons, q is generally large, so the above two methods in this specification can reduce the overall data storage volume and data transmission volume.

[0089] Generally, assume that the sample size n = 6*10 5 , feature type f = 6*10 4 , the number of buckets b = 10, the maximum encrypted value q = 2 64 The amount of data transmitted by the existing technology solution is about 3TB. Figure 3 The amount of data transmitted by the scheme shown is about 20GB, which is 160 times higher than the scheme of the prior art. Figure 4 The amount of data transmitted by the scheme shown is about 117GB, which is 26 times higher than the scheme of the prior art.

[0090] like Figure 5 As shown, this specification also shows a matrix transmission device, applied to a first party having f first plaintext matrices, each column of the first plaintext matrices including a first value; the device includes:

[0091] A matrix generation module 510 is configured to generate a second binary plaintext matrix of the same size as the first plaintext matrix; each column of the second plaintext matrix includes a first value;

[0092] A vector generation module 520 is configured to generate, for each first plaintext matrix, a corresponding first plaintext vector; the first plaintext vector includes the same number of elements as the number of columns in the first plaintext matrix, and the value of each element represents the position difference between the first value in the column corresponding to the element in the first plaintext matrix and the second plaintext matrix;

[0093] The sending module 530 is configured to send the second ciphertext matrix corresponding to the second plaintext matrix and the f first plaintext vectors to the second party, so that the second party obtains the first ciphertext matrices corresponding to the f first plaintext matrices respectively based on the second ciphertext matrix and the f first plaintext vectors.

[0094] In an optional implementation, the value i of the j-th element in any first plaintext vector indicates that each element in the j-th column of the second plaintext matrix is shifted by i bits in a preset direction to obtain the j-th column element of the first plaintext matrix corresponding to the first plaintext vector.

[0095] In an optional embodiment, the sending module 530 is further used to generate a corresponding second plaintext vector for each first plaintext matrix; the second plaintext vector has the same number of elements as the first plaintext vector, and the value of each element in the second plaintext vector is used to represent the position of the column of the second plaintext matrix corresponding to the element of the first plaintext vector corresponding to the element in the second plaintext matrix; and the f second plaintext vectors are sent to the second party.

[0096] like Figure 6 As shown, this specification also shows another matrix transmission device, which is applied to a second party and includes:

[0097] Receiving module 610 is configured to receive a second ciphertext matrix and f first plaintext vectors from a first party; wherein the first party has f first plaintext matrices, and the second plaintext matrix corresponding to the second ciphertext matrix is a binary matrix of the same size as the first plaintext matrix, and each column in the first plaintext matrix and the second plaintext matrix includes a first value; the first plaintext vectors and the first plaintext matrices have a one-to-one correspondence, the number of elements included in the first plaintext vectors is the same as the number of columns in the first plaintext matrix, and the value of each element of any first plaintext vector represents the positional difference between the first values in the columns corresponding to the element in the first plaintext matrix and the second plaintext matrix;

[0098] The recovery module 620 is configured to obtain f first ciphertext matrices according to the second ciphertext matrix and f first plaintext vectors; the f first ciphertext matrices correspond to the f first plaintext matrices.

[0099] In an optional embodiment, the value i of the j-th element in any first plaintext vector represents shifting each element in the j-th column of the second plaintext matrix by i bits in a preset direction to obtain the j-th column element of the first plaintext matrix corresponding to the first plaintext vector. The recovery module 620 is specifically configured to, for each j-th element of the first plaintext vector, shift each element in the j-th column of the second ciphertext matrix by i bits in a preset direction to obtain the j-th column of the first plaintext matrix corresponding to the first plaintext vector, where i is the value of the j-th element in the first plaintext vector.

[0100] In an optional embodiment, the receiving module 610 is further used to receive f second plaintext vectors from the first party; the second plaintext vector has the same number of elements as the first plaintext vector, and the value of each element in the second plaintext vector is used to represent the position of the column of the second plaintext matrix corresponding to the element of the first plaintext vector corresponding to the element in the second plaintext matrix.

[0101] Correspondingly, in an optional embodiment, the recovery module 620 is specifically used to obtain, for each j-th element of the first plaintext vector, its value i1, and the value i2 of the j-th element of the second plaintext vector corresponding to the first plaintext vector; and offset each element of the i2-th column in the second ciphertext matrix by i1 bits in a preset direction as the vector of the j-th column in the first ciphertext matrix corresponding to the first plaintext vector.

[0102] In an optional embodiment, the f first plaintext matrices are used to represent the bucketing results of the f types of feature data held by the first party; each column in the first plaintext matrix corresponds to a sample, and the position of the element taking the first value in each column is used to represent the bucket into which the element of the column falls; the second party holds a plaintext label vector, and each element in the vector is used to represent the label value of each sample in the first plaintext matrix.

[0103] In an optional embodiment, the second ciphertext matrix is encrypted by a homomorphic encryption algorithm; the device also includes a multiplication module 630 (not shown in the figure), which is used to determine the ciphertext result vector for each feature based on the product of the first ciphertext matrix corresponding to the feature and the plaintext label vector, and determine the contribution of the feature to the label value based on the ciphertext result vector.

[0104] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD through their own programming, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.

[0105] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.

[0106] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a server system. Of course, this specification does not exclude that with the future development of computer technology, the computer that implements the functions of the above embodiments may be, for example, a personal computer, a laptop computer, an in-vehicle human-computer interaction device, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0107] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flow charts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way of executing the order of many steps and does not represent the only execution order. When the device or terminal product in practice is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, a parallel processor or a multi-threaded processing environment, or even a distributed data processing environment). The term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or equipment including a series of elements includes not only those elements, but also includes other elements that are not clearly listed, or also includes elements inherent to such process, method, product or equipment. In the absence of more restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or equipment including the elements. For example, if the words first, second, etc. are used to represent the name, they do not represent any particular order.

[0108] For the convenience of description, the above devices are described in terms of functions divided into various modules. Of course, when implementing one or more of the present specifications, the functions of each module can be implemented in the same or multiple software and / or hardware, or the module that implements the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0109] This specification is described with reference to the flowcharts and / or block diagrams of the methods, apparatus (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0110] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0111] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0112] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0113] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0114] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage, graphene storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0115] Those skilled in the art will appreciate that one or more embodiments of this specification may be provided as a method, system, or computer program product. Thus, one or more embodiments of this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0116] One or more embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0117] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between the various embodiments can be referenced across them. Each embodiment focuses on the differences from the other embodiments. In particular, since the system embodiments are generally similar to the method embodiments, their description is relatively simple. For relevant parts, reference can be made to the description of the method embodiments. Throughout this specification, reference to the terms "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. Furthermore, those skilled in the art may combine and integrate the different embodiments or examples, and features of different embodiments or examples, described in this specification, without conflict.

[0118] The foregoing description is merely an example of one or more embodiments of this specification and is not intended to limit the one or more embodiments of this specification. Those skilled in the art will appreciate that various modifications and variations of one or more embodiments of this specification are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this specification are intended to be included within the scope of the claims.

Claims

1. A matrix transmission method, applied to a first party having f first plaintext matrices, wherein each column of the first plaintext matrices includes a first value; the method comprising: generating a second binarized plaintext matrix of the same size as the first plaintext matrix; Each column of the second plaintext matrix includes a first value; For each first plaintext matrix, generate a corresponding first plaintext vector; The number of elements included in the first plaintext vector is the same as the number of columns in the first plaintext matrix, and the value of each element is used to represent the position difference of the first value in the column corresponding to the element in the first plaintext matrix and the second plaintext matrix; The second ciphertext matrix corresponding to the second plaintext matrix and the f first plaintext vectors are sent to the second party, so that the second party obtains the first ciphertext matrices corresponding to the f first plaintext matrices respectively according to the second ciphertext matrix and the f first plaintext vectors.

2. According to the method of claim 1, the value i of the j-th element in any first plaintext vector represents shifting each element in the j-th column of the second plaintext matrix by i bits in a preset direction to obtain the j-th column element of the first plaintext matrix corresponding to the first plaintext vector.

3. The method according to claim 1, further comprising: For each first plaintext matrix, generate a corresponding second plaintext vector; The second plaintext vector has the same number of elements as the first plaintext vector, and the value of each element in the second plaintext vector is used to represent the position of the column of the second plaintext matrix corresponding to the element of the first plaintext vector in the second plaintext matrix; Send the f second plaintext vectors to the second party.

4. A matrix transmission method, applied to a second party, comprising: Receive a second ciphertext matrix and f first plaintext vectors from a first party; wherein the first party has f first plaintext matrices, a second plaintext matrix corresponding to the second ciphertext matrix is a binary matrix of the same size as the first plaintext matrix, and each column in the first plaintext matrix and the second plaintext matrix includes a first value; the first plaintext vectors and the first plaintext matrices have a one-to-one correspondence, the number of elements included in the first plaintext vector is the same as the number of columns in the first plaintext matrix, and the value of each element of any first plaintext vector is used to represent the positional difference of the first values in the columns corresponding to the element in the first plaintext matrix and the second plaintext matrix; Obtain f first ciphertext matrices according to the second ciphertext matrix and f first plaintext vectors; the f first ciphertext matrices correspond to the f first plaintext matrices.

5. The method according to claim 4, wherein the value i of the j-th element in any first plaintext vector represents shifting each element in the j-th column of the second plaintext matrix by i bits in a preset direction to obtain the j-th column element of the first plaintext matrix corresponding to the first plaintext vector; The obtaining f first ciphertext matrices according to the second ciphertext matrix and f first plaintext vectors includes: For each j-th element of each first plaintext vector, each element in the j-th column of the second ciphertext matrix is shifted by i bits in a preset direction to obtain the j-th column of the first plaintext matrix corresponding to the first plaintext vector, where i is the value of the j-th element in the first plaintext vector.

6. The method according to claim 5, further comprising: Receive f second plaintext vectors from the first party; The second plaintext vector has the same number of elements as the first plaintext vector, and the value of each element in the second plaintext vector is used to represent the position of the column of the second plaintext matrix corresponding to the element of the first plaintext vector in the second plaintext matrix.

7. The method according to claim 6, wherein obtaining f first ciphertext matrices according to the second ciphertext matrix and f first plaintext vectors comprises: For each j-th element of the first plaintext vector, obtain its value i1 and the value i2 of the j-th element of the second plaintext vector corresponding to the first plaintext vector; Each element of the i2th column in the second ciphertext matrix is shifted by i1 bits in a preset direction to serve as the vector of the jth column in the first ciphertext matrix corresponding to the first plaintext vector.

8. According to the method of claim 4, the f first plaintext matrices are used to represent the bucketing results of the f types of feature data held by the first party; each column in the first plaintext matrix corresponds to a sample, and the position of the element taking the first value in each column is used to represent the bucket into which the element in the column falls; the second party holds a plaintext label vector, and each element in the vector is used to represent the label value of each sample in the first plaintext matrix.

9. The method according to claim 8, wherein the second ciphertext matrix is encrypted using a homomorphic encryption algorithm; The method further comprises: For each feature, a ciphertext result vector is determined based on the product of the first ciphertext matrix corresponding to the feature and the plaintext label vector, and the contribution of the feature to the label value is determined based on the ciphertext result vector.

10. A computing device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 9 is implemented.