Data processing method, medium and program product for secure computation
By using homomorphic encryption and trusted execution environment devices in multi-party secure computation, efficient collaborative operation of feature and tag fragments is achieved, solving the problem of excessive communication volume in feature fragmentation, improving data processing efficiency and ensuring data security.
Patent Information
- Application Number
- CN202411397407.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-08
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-10-08
AI Technical Summary
In multi-party secure computation, the communication volume during feature segmentation is too large, especially the communication overhead between the side with large data volume and the side with small data volume, which affects the computation efficiency.
Homomorphic encryption technology is used to encrypt first-party data using trusted execution environment devices within the local area network, and to perform feature and tag fragmentation operations in collaboration with the second party, thereby reducing communication volume.
By reducing communication volume, the efficiency of data processing tasks is improved, especially in cases involving large and small data volumes, significantly reducing the burden on wide area network communication and protecting data security.
Smart Images

Figure CN119382873B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of secure multi-party computation, and more specifically, to a data processing method, medium, and program product for secure computation. Background Technology
[0002] Secure multi-party computation (MPC), also known as secure multi-party computation, allows multiple parties to collaboratively compute the result of a function without disclosing the input data of each party. The result is then made public to one or more of the parties. Typical applications of secure MPC include joint statistical analysis of privacy-preserving multi-party data and machine learning. Here, the function is a statistical operation, a machine learning algorithm, etc. During secure MPC, to prevent the disclosure of data and intermediate results, the data or intermediate results can be shared among the parties. Each party holds a data fragment, and the fragments held by all parties are merged to reconstruct the corresponding data. Typically, the computation is performed in a shared state. Therefore, the number of data communications and the amount of data exchanged in secure MPC are significant factors affecting the efficiency of secure computation. Summary of the Invention
[0003] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] In a first aspect, this disclosure provides a data processing method for secure computing, wherein the participants in the secure computing include a first party, a second party, and a third party. The first party holds a first identifier set, and the first identifier in the first identifier set corresponds to an h-dimensional feature. The second party holds a second identifier set, and the second identifier in the second identifier set corresponds to a tag. The third party is a trusted execution environment device, and the third party and the first party are located in the same local area network, where h≥1. The method is applied to the first party and includes: sending the first identifier set and the target feature corresponding to the first identifier set to the third party, so that the third party uses a target key to homomorphically encrypt the first identifier set and the target feature corresponding to the first identifier set to obtain a first ciphertext, and sending the first ciphertext to the first party, wherein the target feature includes at least one of the h-dimensional features. Partial dimensional features; in response to receiving the first ciphertext and the second ciphertext sent by the second party, based on the first ciphertext and the second ciphertext, generate ciphertext of the intersection of the first identifier set and the second identifier set, ciphertext of the first feature, and ciphertext of the target tag, wherein the second ciphertext is obtained by the second party using the target key to homomorphically encrypt the second identifier set and the tags corresponding to the second identifier set, the first feature is the target feature corresponding to the intersection, and the target tag is the tag corresponding to the intersection; based on the ciphertext of the first feature, perform a feature fragmentation operation with the second party to obtain a first fragment of the first feature, and based on the ciphertext of the target tag, perform a tag fragmentation operation with the second party to obtain a second fragment of the target tag; based on the first fragment and the second fragment, perform a target data processing task.
[0005] Secondly, this disclosure provides a data processing method for secure computing, wherein the participants in the secure computing include a first party, a second party, and a third party. The first party holds a first identifier set, and the first identifiers in the first identifier set correspond to h-dimensional features. The second party holds a second identifier set, and the second identifiers in the second identifier set correspond to tags. The third party is a trusted execution environment device, and the third party is located in the same local area network as the first party, where h≥1. The method is applied to the second party and includes: homomorphically encrypting the second identifier set and the tags corresponding to the second identifier set using a target key to obtain a second ciphertext; sending the second ciphertext to the first party so that the second party can use the first ciphertext and the second ciphertext as a basis for processing. The ciphertext is generated by the third party using the target key to homomorphically encrypt the first identifier set and the target feature corresponding to the first identifier set. The target feature includes at least some of the h-dimensional features. The first feature is the target feature corresponding to the intersection, and the target label is the label corresponding to the intersection. A feature sharding operation is performed with the first party to obtain a third shard of the first feature, and a label sharding operation is performed with the first party to obtain a fourth shard of the target label. Based on the third shard and the fourth shard, a target data processing task is performed.
[0006] Thirdly, this disclosure provides a data processing method for secure computing, wherein the participants in the secure computing include a first party, a second party, and a third party. The first party holds a first identifier set, and the first identifier in the first identifier set corresponds to an h-dimensional feature. The second party holds a second identifier set, and the second identifier in the second identifier set corresponds to a tag. The third party is a trusted execution environment device, and the third party and the first party are located in the same local area network, where h≥1. The method is applied to the third party and includes: in response to receiving the first identifier set and the target feature corresponding to the first identifier set sent by the first party, homomorphically encrypting the first identifier set and the target feature corresponding to the first identifier set using a target key to obtain a first ciphertext; wherein the target feature includes at least some of the h-dimensional features; sending the first ciphertext to the first party so that the first party generates a first fragment of the first feature and a second fragment of the target tag based on the first ciphertext, wherein the first feature is the target feature corresponding to the intersection of the first identifier set and the second identifier set, and the target tag is the tag corresponding to the intersection.
[0007] Fourthly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the data processing method for secure computing provided in the first aspect of this disclosure, or the steps of the data processing method for secure computing provided in the second aspect of this disclosure, or the steps of the data processing method for secure computing provided in the third aspect of this disclosure.
[0008] Fifthly, this disclosure provides an electronic device, comprising: a storage device having a computer program stored thereon; and a processing device for executing the computer program in the storage device to implement the steps of the data processing method for secure computing provided in the first aspect of this disclosure, or the steps of the data processing method for secure computing provided in the second aspect of this disclosure, or the steps of the data processing method for secure computing provided in the third aspect of this disclosure.
[0009] In a sixth aspect, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the data processing method for secure computing provided in the first aspect of this disclosure, or the steps of the data processing method for secure computing provided in the second aspect of this disclosure, or the steps of the data processing method for secure computing provided in the third aspect of this disclosure.
[0010] In the above technical solution, firstly, assuming the third party (TEE device) is honest with the first party, the first party sends its own first identifier set and at least some of the corresponding dimensional features (i.e., target features) to the third party located on the same local area network. Then, the third party uses the target key to homomorphically encrypt the plaintext data received from the first party, obtaining the first ciphertext, and sends the first ciphertext to the first party. Simultaneously, the second party uses the target key to homomorphically encrypt its own second identifier set and the tags corresponding to the second identifier set, obtaining the second ciphertext, and sends the second ciphertext to the first party. Next, the first party... Based on the first ciphertext received from the third party and the second ciphertext received from the second party, ciphertext of the intersection of the first and second identifier sets, ciphertext of the target feature (i.e., the first feature) corresponding to the intersection, and ciphertext of the label (i.e., the target label) corresponding to the intersection are generated. Then, based on the ciphertext of the first feature, the first party and the second party jointly perform a feature sharding operation to obtain a shard of the first feature respectively. Based on the ciphertext of the target label, the first party and the second party jointly perform a label sharding operation to obtain a shard of the target label respectively. Finally, the first party and the second party respectively perform the target data processing task based on the shards they hold. In this way, based on the assumption that the third party is honest with the first party, the first party can send its plaintext data to the third party located on the same local area network. The third party can then encrypt the plaintext data held by the first party using a target key pre-agreed with the second party. This avoids encrypting the data held by the first party and sending it to the second party via a wide area network. Furthermore, when the first party and the second party perform feature tag fragmentation and tag fragmentation based on the ciphertext of the first feature and the ciphertext of the target tag, only the ciphertext of the features corresponding to the intersection of the first and second identifier sets is involved, and the ciphertext of the data held by the first party is not involved. By encrypting the data, the first party can avoid sending the encrypted version of all its data to the second party. Since this avoids sending the encrypted version of all the data held by the first party, the communication volume between the two parties (which falls under wide area network communication) is reduced. This significantly reduces communication volume during feature fragmentation and label fragmentation, especially when the first party's data volume is much larger than the second party's, thereby improving the processing efficiency of the target data processing task. Therefore, this solution can protect data security and reduce communication volume in scenarios such as model training and SQL queries.
[0011] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:
[0013] Figure 1 This is a schematic diagram illustrating a data processing system for secure computing according to an exemplary embodiment.
[0014] Figure 2 This is a flowchart illustrating a data processing method for secure computation applied to a first party, according to an exemplary embodiment.
[0015] Figure 3 This is a flowchart illustrating a data processing method for secure computing applied to a second party, according to an exemplary embodiment.
[0016] Figure 4 This is a flowchart illustrating a data processing method for secure computing applied to a third party, according to an exemplary embodiment.
[0017] Figure 5 This is a timing diagram illustrating a data processing method for secure computing according to an exemplary embodiment.
[0018] Figure 6 This is a timing diagram illustrating a data processing method for secure computing according to another exemplary embodiment.
[0019] Figure 7 This is a block diagram illustrating a data processing apparatus for secure computing applied to a first party, according to an exemplary embodiment.
[0020] Figure 8 This is a block diagram illustrating a data processing apparatus for secure computing applied to a second party, according to an exemplary embodiment.
[0021] Figure 9 This is a block diagram illustrating a data processing apparatus for secure computing applied to a third party, according to an exemplary embodiment.
[0022] Figure 10 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment. Detailed Implementation
[0023] Before introducing specific embodiments of this disclosure, the terminology involved in this disclosure and the specific application scenarios of multi-party secure computation will be explained first.
[0024] A ring is a set that defines two operations: addition and multiplication. It forms an abelian group for addition and a semigroup for multiplication for all elements except zero. Multiplication satisfies the distributive property of addition.
[0025] Secret sharing, also known as secret partitioning or secret sharing, works by dividing a secret (such as a key or private data) into multiple shares, each held by a different data holder. The secret can only be recovered when more than a certain number of parties merge their shares; shares obtained from fewer than the threshold cannot recover any information from the secret. In multi-party secure computation, the threshold number is usually the same as the number of participating parties, and the shares into which the secret is divided can be called fragments. The private data refers to the data that the parties do not want to know in a multi-party secure computation.
[0026] Homomorphic encryption is a technique that allows computation on encrypted data (i.e., ciphertext) followed by decryption to obtain the result. The result of homomorphic encryption is the same as the result of direct computation on the original data (i.e., plaintext), but the entire computation process is performed on the encrypted data. There are many encryption algorithms that can be used to implement homomorphic encryption, and BFV is one of them, referred to here as BFV homomorphic encryption.
[0027] Private Set Intersection (PSI) is a proprietary protocol in the field of secure multi-party computation that allows participating parties to input private sets and jointly compute the intersection of sets, while guaranteeing that no additional element information is leaked except for the result of the intersection.
[0028] Communication volume: Since the data of the participants in secure computation are on different machines, they need to communicate over a network to complete the interaction. During the computation process, encrypted data will be transmitted over the network, and the amount of data transmitted is the communication volume.
[0029] A Trusted Execution Environment (TEE) is a secure hardware or software environment used to provide protection for sensitive data and code against unauthorized access and modification. The design goal of a TEE is to ensure that data and code are protected during execution, remaining secure even if the operating system or application is attacked.
[0030] TEE typically includes the following main features and functions:
[0031] • Security Isolation: TEE provides an isolated execution environment, ensuring that sensitive data and code are protected during execution and are not interfered with by the operating system or other processes.
[0032] • Secure Boot: TEE verifies its integrity and authenticity during startup to ensure that the environment has not been tampered with or compromised by malware.
[0033] • Secure Storage: TEE provides a secure storage mechanism that can be used to store encryption keys, certificates, sensitive data, etc., to prevent data leakage and tampering.
[0034] • Secure Communication: TEE supports secure communication mechanisms to ensure data is protected during transmission and prevent man-in-the-middle attacks and data theft.
[0035] • Secure operating environment: TEE provides a trusted execution environment in which sensitive code, such as encryption algorithms and digital signatures, can be run to ensure the security and integrity of the code.
[0036] TEE is widely used in security chips, smartphones, cloud computing environments, and other fields to protect user data, financial transaction information, and digital copyright content. The development and application of TEE technology provides strong support for data security and privacy protection.
[0037] In the field of secure computing, the terms "semi-honest" and "honest" describe different behavioral characteristics of participants when executing a protocol. Semi-honest participants adhere to the protocol rules while exhibiting curiosity; that is, they act according to the protocol's rules and do not intentionally violate its steps or operations. Despite adhering to the rules, semi-honest participants may attempt to extract additional information from the protocol's execution process, such as analyzing received information or trying to infer other participants' inputs from the protocol's output. Honest participants not only adhere to the protocol's rules but also do not exhibit curiosity about other participants' inputs or the protocol's internal state; that is, honest participants act entirely according to the protocol's requirements and do not perform any operations beyond what is stipulated in the protocol, such as attempting to extract any additional information from the protocol's execution process.
[0038] ElGamal encryption is a public-key encryption algorithm based on the discrete logarithm problem. ElGamal encryption based on elliptic curves (hereinafter referred to as elliptic curve encryption) is a variant of the ElGamal encryption algorithm that applies elliptic curves to achieve a more efficient and secure encryption mechanism.
[0039] The implementation method of the ElGamal encryption mechanism based on elliptic curves is as follows:
[0040] S1, Parameter Settings: Select appropriate elliptic curve parameters, including curve equation, base point, order, etc.
[0041] S2, Key Generation: Randomly select a private key d, and calculate the public key Q = d * G, where G is the base point of the elliptic curve.
[0042] S3, Encryption: Select a random number k, calculate points C1 = k*G and C2 = B + k*Q, where B is the message to be encrypted, and C1 and C2 are the encrypted ciphertext.
[0043] S4, Decryption: Based on the private key d and the ciphertexts C1 and C2, the original message B = C2 - d * C1 can be calculated.
[0044] Elliptic curve cryptography (ELGamal) offers higher security and shorter key lengths compared to traditional ELGamal cryptography, while also providing greater encryption efficiency. It is widely used in secure communication, digital signatures, and authentication, providing a reliable encryption solution for protecting the confidentiality and integrity of data.
[0045] Furthermore, the ElGamal encryption mechanism based on elliptic curves possesses the ±1 homomorphic property, meaning that for any point H on the elliptic curve, Dec((-1)) = 1. w Enc(H))=(-1) w H, For a binary field consisting of 0 and 1, Dec represents elliptic curve decryption, and Enc(H) represents elliptic curve encryption of H.
[0046] In practical applications, for privacy protection purposes, secure multi-party computation (SMC) algorithms are typically black-box algorithms, meaning the data transmission behavior between the various computing nodes running the algorithm is opaque. As discussed in the background section, typical applications of secure multi-party computation include machine learning. In the inference and training phases of machine learning, secure multi-party computation can be used to protect privacy data, primarily involving the protection of model parameters and the data of each participating party during training.
[0047] Currently, common privacy-preserving machine learning solutions based on secure multi-party computation include: privacy-preserving machine learning protocols based on technologies such as obfuscated circuits and unintentional transmission, which execute secure multi-party computation protocols to perform nonlinear operations such as activation function calculations. Secret-sharing technology allows multiple parties to participate in the training or prediction of machine learning network models without revealing data or model information.
[0048] In addition to the above application areas, multi-party secure computation can also be applied to network security testing to protect privacy, joint statistical analysis of privacy-protected multi-party data, spam filtering of encrypted emails, and ad conversion.
[0049] Among them, secure computation by both parties is usually used for joint statistical analysis of data from two parties with privacy protection, that is, querying across the databases of two parties while protecting the private data of both parties.
[0050] For example, consider the following Structured Query Language (SQL) statement:
[0051] select avg(a.key),sum(bs=0)from a join b on a.id=b.id;
[0052] This SQL statement aligns tables a and b by the id column, calculates the average value of a.key based on the aligned tables, and counts the number of times bs = 0, thus obtaining the query results. Table a is stored on the first party P0, and includes an id column and at least one feature column. Table b is stored on the second party P1, and includes an id column and one feature column (i.e., bs, which can be a label column with a value of 0 or 1). This feature column has a different dimension from at least one feature column in table a. P0 and P1 perform the SQL query using secure computation techniques, without exposing their own private data. After the query, only the query results are exposed to the querying party. The query results are stored in fragmented form on P0 and P1.
[0053] Specifically, P0 and P1 align tables a and b according to the id column in the following way (which can be called PSI toshare): First, both parties obtain the intersection of their id columns through PSI, where the intersection is held exclusively by P1 and not exposed to P0; then, feature sharding is performed, that is, the features held by P0 that correspond to the intersection are stored in the form of shards on both P0 and P1, and the features held by P1 that correspond to the intersection are stored in the form of shards on both P0 and P1, so that both parties can perform SQL queries based on their respective feature shards.
[0054] In related technologies, feature sharding is usually implemented based on the DDH assumption and homomorphic encryption algorithm. The amount of data communication involved in this process is proportional to the data size of the big data side. When the data size of the big data side is large, it will lead to huge communication overhead for feature sharding.
[0055] For example, P0 has N = 1 billion data entries, each containing h = 1000 features and an identifier (e.g., id); P1 has d = 1 million data entries, each containing an identifier (e.g., id). P0 represents the large data volume. The communication volume involved in feature sharding based on the DDH assumption and homomorphic encryption algorithm is on the order of O(hN). It is evident that the communication overhead of feature sharding is enormous.
[0056] In view of this, this disclosure provides a data processing method, medium, and program product for secure computing to reduce the amount of communication between the two parties during feature fragmentation.
[0057] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0058] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0059] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0060] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0061] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0062] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0063] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0064] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require access to and use of their information. This allows the user to autonomously choose, based on the prompt message, whether to provide information to the software or hardware, such as the electronic device, application program, server, or storage medium performing the operations of this disclosed technical solution.
[0065] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide information to the electronic device.
[0066] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0067] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0068] Figure 1 This is a schematic diagram illustrating a data processing system for secure computing according to an exemplary embodiment. Figure 1 As shown, the data processing system may include a first party, a second party, and a third party. The third party is a TEE device. The third party and the first party are in the same local area network (LAN). That is, the first party and the third party communicate through the LAN, the second party and the third party communicate through the wide area network (WAN), and the first party and the second party also communicate through the WAN.
[0069] The participants in secure computation include a first party, a second party, and a third party. The first party holds a first set of identifiers, where each identifier in the first set corresponds to an h-dimensional feature. The second party holds a second set of identifiers, where each identifier in the second set corresponds to a label. h ≥ 1, meaning each identifier in the first set corresponds to at least one-dimensional feature. Both the first and second sets contain at least one identifier, and the first set differs from the second set. In one implementation, the number of first identifiers in the first set is significantly greater than the number of second identifiers in the second set; in this case, the first party represents the party with the largest data volume. The label feature can be a binary classification label, with values of 0 or 1. For example, the first party is P0 as described above, and the second party is P1 as described above.
[0070] The data processing methods used for secure computation are based on the semi-honest plus assumption (a security model), in which the third party satisfies the following semi-honest plus assumption:
[0071] A. The third party is honest with the first party: The third party only communicates with the outside world through the first party, and the first party is willing to send private data to the third party;
[0072] B. The third party is semi-honest with the second party: The second party believes that the third party will faithfully perform the agreement, but is unwilling to send its private data to the third party;
[0073] C. The first party also does not want an adversary to be able to recover the first party's private data based on the data within the third party and the data sent by the first party to the third party after the adversary has compromised the third party.
[0074] Specifically, the second party and the third party communicate only once via the WAN to agree on a key for encryption, namely the target key. The target key can be generated by the third party and sent to the second party via the WAN, or the target key can be generated by the second party and sent to the third party via the WAN.
[0075] The core idea of the data processing method for secure computing is as follows: A first party sends its own first set of identifiers and at least some of the corresponding dimensional features to a third party. The third party then uses a pre-agreed target key to homomorphically encrypt the data sent by the first party and sends the resulting first ciphertext back to the first party. Simultaneously, the second party uses the pre-agreed target key to homomorphically encrypt its own second set of identifiers and the tags corresponding to the second set, and sends the resulting second ciphertext back to the first party. The first party receives the first ciphertext sent by the third party and the second ciphertext sent by the third party. After the two ciphertexts are encrypted, based on the first and second ciphertexts, ciphertexts of the intersection of the first and second identifier sets, ciphertexts of at least some dimensional features corresponding to the intersection, and ciphertexts of labels corresponding to the intersection are generated. Then, based on the ciphertexts of at least some dimensional features corresponding to the intersection, the first and second parties jointly perform feature sharding operations to obtain a shard of at least some dimensional features corresponding to the intersection, and based on the ciphertexts of labels corresponding to the intersection, the first and second parties jointly perform label sharding operations to obtain a shard of labels corresponding to the intersection. Finally, the first and second parties each perform the target data processing task based on their respective shards.
[0076] Figure 2 This is a flowchart illustrating a data processing method for secure computation applied to a first party, according to an exemplary embodiment. (e.g.) Figure 2As shown, the data processing method for secure computing applied to the first party may include S101 to S104.
[0077] In S101, the first identifier set and the target feature corresponding to the first identifier set are sent to a third party, so that the third party can use the target key to perform homomorphic encryption on the first identifier set and the target feature corresponding to the first identifier set to obtain the first ciphertext, and send the first ciphertext to the first party.
[0078] In this disclosure, the target features include at least some of the h-dimensional features.
[0079] In one implementation, the target feature includes at least some of the h-dimensional features. In this case, the first party can send the first identifier set and the corresponding partial dimensional features to the third party, so that the third party can use the target key to perform homomorphic encryption on the first identifier set and the corresponding partial dimensional features to obtain the first ciphertext, and send the first ciphertext to the first party.
[0080] In another implementation, the target features include h-dimensional features. In this case, the first party can send the first identifier set and the h-dimensional features corresponding to the first identifier set to the third party, so that the third party can use the target key to homomorphically encrypt the first identifier set and the h-dimensional features corresponding to the first identifier set to obtain the first ciphertext, and send the first ciphertext to the first party.
[0081] When the first party executes the above data processing method, it first sends its own first identifier set and the target features corresponding to the first identifier set to the third party, based on the assumption that the third party is honest with the first party. After receiving the first identifier set and the target features corresponding to the first identifier set sent by the first party, the third party uses the target key agreed upon with the second party in advance to homomorphically encrypt the first identifier set and the target features corresponding to the first identifier set to obtain the first ciphertext, and then sends the first ciphertext to the first party. The second party receives the third ciphertext sent by the third party.
[0082] In S102, in response to receiving the first ciphertext and the second ciphertext sent by the second party, the ciphertext of the intersection of the first identifier set and the second identifier set, the ciphertext of the first feature, and the ciphertext of the target tag are generated based on the first ciphertext and the second ciphertext.
[0083] In this disclosure, the second ciphertext is obtained by the second party using the target key to homomorphically encrypt the second identifier set and the tags corresponding to the second identifier set. The first feature is the target feature corresponding to the intersection of the first identifier set and the second identifier set, and the target tag is the tag corresponding to the intersection.
[0084] When the second party executes the above data processing method, it first assumes that the third party is semi-honest to the second party. In order to avoid the leakage of the second party's data, the second party performs homomorphic encryption on its own data locally. That is, it uses the target key agreed upon with the third party in advance to homomorphically encrypt the second set of identifiers and the tags corresponding to the second set of identifiers to obtain the second ciphertext, and sends the second ciphertext to the first party; the first party receives the second ciphertext.
[0085] After receiving the first ciphertext sent by the third party and the second ciphertext sent by the second party, the first party generates, based on the first ciphertext and the second ciphertext, the ciphertext of the intersection of the first identifier set and the second identifier set, the ciphertext of the first feature, and the ciphertext of the target tag.
[0086] In S103, based on the ciphertext of the first feature, a feature fragmentation operation is performed with the second party to obtain the first fragment of the first feature, and based on the ciphertext of the target label, a label fragmentation operation is performed with the second party to obtain the second fragment of the target label.
[0087] In this disclosure, after obtaining the ciphertext of the first feature and the ciphertext of the target label, the first party can, based on the ciphertext of the first feature, jointly perform a feature segmentation operation with the second party to obtain a segment of the first feature (i.e., the first segment and the third segment), and simultaneously, based on the ciphertext of the target label, jointly perform a label segmentation operation with the second party to obtain a segment of the target label (i.e., the second segment and the fourth segment).
[0088] In S104, the target data processing task is executed based on the first and second fragments.
[0089] The aforementioned target data processing task can be an SQL query task or a machine learning model training task.
[0090] In one implementation, when the target data processing task is an SQL query task, the first party and the second party can jointly perform SQL queries based on their respective feature fragments and tag fragments to obtain query result fragments. Then, the two parties will send their respective query result fragments back to the querying party. After receiving the query result fragments sent by the two parties, the querying party will merge them to obtain the final query result.
[0091] Specifically, the first party can perform an SQL query with the second party based on the first and second shards to obtain the first query result shard, and then return the first query result shard to the querying party; correspondingly, the second party can perform an SQL query with the first party based on the third and fourth shards to obtain the second query result shard, and then return the second query result shard to the querying party; finally, the querying party merges the first and second query result shards to obtain the final query result.
[0092] In another implementation, when the target data processing task is machine learning model training, the second party can utilize the features of the first party for model training. Specifically, both the first and second parties can perform model training using MPC (Multi-Processing) based on their respective feature slices and label slices to obtain model parameter slices. The first party then sends its own model parameter slices to the second party, which merges these slices with its own to obtain the model parameters, thus completing the second party's model training. This approach allows the second party to utilize the features of the first party for model training when it lacks or has insufficient training data, improving training accuracy while protecting the first party's data privacy.
[0093] Specifically, the first party can train the model together with the second party based on the first and second slices to obtain the first model parameter slice, and then feed the first model parameter slice back to the second party; correspondingly, the second party can train the model together with the first party based on the third and fourth slices to obtain the second model parameter slice, and then merge the second model parameter slice with the first model parameter slice received from the first party to obtain the complete model parameters.
[0094] In the above technical solution, firstly, assuming the third party (TEE device) is honest with the first party, the first party sends its own first identifier set and at least some of the corresponding dimensional features (i.e., target features) to the third party located on the same local area network. Then, the third party uses the target key to homomorphically encrypt the plaintext data received from the first party, obtaining the first ciphertext, and sends the first ciphertext to the first party. Simultaneously, the second party uses the target key to homomorphically encrypt its own second identifier set and the tags corresponding to the second identifier set, obtaining the second ciphertext, and sends the second ciphertext to the first party. Next, the first party... Based on the first ciphertext received from the third party and the second ciphertext received from the second party, ciphertext of the intersection of the first and second identifier sets, ciphertext of the target feature (i.e., the first feature) corresponding to the intersection, and ciphertext of the label (i.e., the target label) corresponding to the intersection are generated. Then, based on the ciphertext of the first feature, the first party and the second party jointly perform a feature sharding operation to obtain a shard of the first feature respectively. Based on the ciphertext of the target label, the first party and the second party jointly perform a label sharding operation to obtain a shard of the target label respectively. Finally, the first party and the second party respectively perform the target data processing task based on the shards they hold. In this way, based on the assumption that the third party is honest with the first party, the first party can send its plaintext data to the third party located on the same local area network. The third party can then encrypt the plaintext data held by the first party using a target key pre-agreed with the second party. This avoids encrypting the data held by the first party and sending it to the second party via a wide area network. Furthermore, when the first party and the second party perform feature fragmentation and tag fragmentation based on the ciphertext of the first feature and the ciphertext of the target tag, only the ciphertext of the features corresponding to the intersection of the first and second identifier sets is involved, and not all the data held by the first party is involved. By encrypting the data, the first party can avoid sending the encrypted version of all its data to the second party. Since it's unnecessary to send the encrypted version of all the data held by the first party to the second party, the communication volume between the two parties (which falls under wide area network communication) is reduced. This significantly reduces communication volume during feature fragmentation and label fragmentation, especially when the first party's data volume is much larger than the second party's, thereby improving the processing efficiency of the target data processing task. Therefore, this solution can protect data security and reduce communication volume in scenarios such as model training and SQL queries.
[0095] In one possible implementation, the first ciphertext may include a first sub-ciphertext and a second sub-ciphertext of the target feature corresponding to each first identifier. The second sub-ciphertext includes a first random number and a feature ciphertext of the target feature corresponding to the first identifier. The first sub-ciphertext is obtained by a third party using a target key to perform pseudo-random encryption on the first identifier set. The feature ciphertext is obtained by a third party using a target key to perform symmetric encryption on the first random number to obtain a third ciphertext, and the third ciphertext is added to the target feature corresponding to the first identifier.
[0096] Specifically, after receiving the first identifier set and the target feature corresponding to the first identifier set from the first party, the third party can use the target key to perform homomorphic encryption on the first identifier set and the target feature corresponding to the first identifier set through the following steps (a1) to (a3):
[0097] Step (a1): Use the target key to perform pseudo-random encryption on the first identifier set to obtain the first sub-ciphertext.
[0098] Step (a2): Generate a first random number, and use the target key to symmetrically encrypt the first random number to obtain the third ciphertext.
[0099] In this disclosure, the first random number can be encrypted using a target key and symmetric encryption algorithms such as Advanced Encryption Standard (AES) and Data Encryption Standard (DES).
[0100] Step (a3): For each target feature corresponding to the first identifier in the first identifier set, add the third ciphertext to the target feature to obtain the feature ciphertext of the target feature.
[0101] For example, the first ciphertext T0 = {prf k1 (i),E k1 (x): i∈I0, x=F0(i)}. Where, prf k1 (i) The ciphertext obtained by pseudo-randomly encrypting the first identifier i in the first identifier set I0 using the target key k1, and the ciphertexts corresponding to each first identifier in the first identifier set constitute the first sub-ciphertext; F0:I0→R m Where m is the dimension of the target feature, and R is a finite commutative ring. I represents the identifier space containing the first and second identifier sets, and E represents the identifier space containing the first and second identifier sets. k1 (x) is the second sub-ciphertext of the target feature F0(i) (i.e. x) corresponding to the first identifier i in the first identifier set, and E k1 (x)=(u,x+g k1 (u)), where g k1(u) represents the third ciphertext obtained by symmetric encryption of the first random number u using the target key k1; x+g k1 (u) is the ciphertext of the target feature x corresponding to the first identifier i in the first identifier set.
[0102] Using the encryption methods shown in steps (a1) to (a3) above, E can be made k1 (x) possesses the properties of homomorphic encryption, but the above encryption method is more efficient than ordinary homomorphic encryption, which can improve the encryption efficiency of first-party private data, thereby improving the processing efficiency of target data processing tasks.
[0103] In one possible implementation, the second ciphertext may include a third sub-ciphertext and a fourth sub-ciphertext of the tag corresponding to each second identifier. The third sub-ciphertext is obtained by the second party using a target key to perform pseudo-random encryption on the second identifier set. The fourth sub-ciphertext is obtained by the second party using a target key to perform elliptic curve encryption on the tag corresponding to the second identifier and the target point. The target point is any point on a preset elliptic curve.
[0104] Specifically, the second party can use the target key to homomorphically encrypt the second identifier set and the tags corresponding to the second identifier set through the following steps (b1) and (b2) to obtain the second ciphertext:
[0105] Step (b1): Use the target key to perform pseudo-random encryption on the second identifier set to obtain the third sub-ciphertext.
[0106] Step (b2): For each tag corresponding to the second identifier in the second identifier set, use the target key to perform elliptic curve encryption on the tag and the target point to obtain the fourth sub-ciphertext of the tag.
[0107] In this disclosure, the target point is any point on a preset elliptic curve, and the second party can randomly select a point on the preset elliptic curve as the target point.
[0108] For example, the second ciphertext T1 = {(prf k1 (j),Enc k1 ((-1) y P):j∈I1,y=F1(j)}. Where, prf k1 (j) is the ciphertext obtained by pseudo-randomly encrypting the second identifier j in the second identifier set I1 using the target key k1. These ciphertexts corresponding to each second identifier in the second identifier set constitute the third sub-ciphertext; F1: Enc k1 ((-1) yP)) is the fourth sub-ciphertext obtained by using the target key k1 to perform elliptic curve encryption on the tag F1(j) (i.e. y) corresponding to the second identifier j in the second identifier set and the target point P.
[0109] In this case, a method similar to S1 to S3 in the ElGamal encryption mechanism based on elliptic curves can be used. The target key k1 is used to perform elliptic curve encryption on the tag F1(j) (i.e., y) corresponding to the second identifier j in the second identifier set and the target point P. In this case, (-1) y P is the message to be encrypted, and k1 is the private key, i.e., d = k1.
[0110] The tags held by the second party are binary tags. Using elliptic curve encryption can be more efficient and produce smaller ciphertext, thereby reducing the amount of communication between the first and second parties.
[0111] The following is a detailed description of the specific implementation method for generating the ciphertext of the intersection of the first identifier set and the second identifier set, the ciphertext of the first feature, and the ciphertext of the target tag based on the first ciphertext and the second ciphertext in S102 above. Specifically, this can be achieved through steps (c1) and (c2):
[0112] Step (c1): Calculate the intersection of the first sub-ciphertext and the third sub-ciphertext to obtain the ciphertext of the intersection of the first identifier set and the second identifier set.
[0113] Step (c2): Determine the ciphertext corresponding to the intersection of all second sub-ciphertexts to obtain the ciphertext of the first feature, and determine the ciphertext corresponding to the intersection of all fourth sub-ciphertexts to obtain the ciphertext of the target label.
[0114] The following is a detailed description of the specific implementation method for obtaining the first fragment of the first feature by performing a feature fragmentation operation with the second party based on the ciphertext of the first feature in S103 above. Specifically, it can be achieved through steps (d1) and (d2):
[0115] Step (d1): Obtain the first random number in the ciphertext of the first feature, and send the first random number to the second party so that the second party can use the target key to perform symmetric encryption on the first random number to obtain the third fragment of the first feature.
[0116] After obtaining the ciphertext of the first feature, the first party can first extract the first random number from the ciphertext of the first feature, and then send the first random number to the second party; after receiving the first random number, the second party uses the target key to symmetrically encrypt the first random number to obtain the fourth ciphertext, and then performs an inversion operation on the fourth ciphertext to obtain the third fragment of the first feature.
[0117] The second party can use the target key to encrypt the first random number using symmetric encryption algorithms such as AES and DES. The encryption method used by the second party to encrypt the first random number using the target key is the same as the encryption method used by the first party to encrypt the first random number using the target key.
[0118] Step (d2): Determine the ciphertext of the first feature as the first fragment of the first feature.
[0119] For example, the ciphertext of the target feature x corresponding to the first identifier i in the first identifier set is x+g. k1 (u), the fourth ciphertext is g k1 (u), after inverting the fourth ciphertext, we get -g k1 (u), then the first segment of the first feature is x+g k1 (u), the second segment of the first feature is -g k1 (u), where x=F0(q)and q∈I0∩I1.
[0120] In this implementation, the first party can, based on the properties of homomorphic encryption, only send a first random number to the second party. In this case, during the step of performing feature fragmentation with the second party based on the ciphertext of the first feature, the communication volume between the first and second parties is λn, where n is the number of second identifiers in the second identifier set, and the first random number is λbit. It is evident that this communication volume is independent of the dimension m of the target feature, thereby reducing the communication volume between the first and second parties during feature fragmentation and tag fragmentation.
[0121] The following is a detailed description of the specific implementation method for obtaining the second fragment of the target tag by performing a tag fragmentation operation with the second party based on the ciphertext of the target tag in S103 above. Specifically, the above tag is a binary classification tag, which can be achieved through the following steps (e1) to (e3):
[0122] Step (e1): Generate a second random number and use the second random number to mask the ciphertext of the target tag to obtain masked data.
[0123] The second random number is either 0 or 1.
[0124] For example, masking data yg=Enc k1 ((-1) y-r P)), where Enc k1 ((-1) y P):y=F1(q) is the ciphertext of the target label, and r is the second random number.
[0125] Step (e2): The masking data is sent to the second party, which uses the target key to perform homomorphic decryption on the masking data and generates the fourth fragment of the target tag based on the homomorphic decryption result.
[0126] Step (e3): Use the second random number as the second slice of the target label.
[0127] After obtaining the masking data, the first party sends the masking data to the second party. After receiving the masking data, the second party uses the target key to perform homomorphic decryption on the masking data. If the homomorphic decryption result is the target point, then 0 is determined as the fourth fragment of the target label. If the homomorphic decryption result is not the target point, then 1 is determined as the fourth fragment of the target label. At the same time, the first party uses the second random number as the second fragment of the target label.
[0128] For example, masking data Enc k1 ((-1) y-r P)=(-1) r Enc k1 ((-1) y Using the ±1 homomorphic property described above, we can perform homomorphic decryption on P), and the resulting homomorphic decryption result is (-1). y-r P, will (-1) y-r Compare P with P, if (-1) y-r If P equals P, then 0 is determined as the fourth segment of the target label. If (-1) y-r If P is not equal to P, then 1 is determined as the fourth segment of the target label.
[0129] In steps (e1) to (e3) above, the first party and the second party only need to transmit masking data. The amount of communication involved is 512n, where the size of the masking data is 512 bits.
[0130] In one possible implementation, S101 may include:
[0131] The first identifier set and the target feature corresponding to the first identifier set are sent to a third party, which uses the target key to homomorphically encrypt the first identifier set and the target feature corresponding to the first identifier set to obtain a first ciphertext. The order of the first ciphertext is then shuffled to obtain a fifth ciphertext, which is then sent to the first party. At this time, the above S102 may include: in response to receiving the fifth ciphertext and the second ciphertext sent by the second party, generating the ciphertext of the intersection of the first identifier set and the second identifier set, the ciphertext of the first feature, and the ciphertext of the target tag based on the fifth ciphertext and the second ciphertext.
[0132] Specifically, when the first party executes the above data processing method, it first sends its own first identifier set and the target features corresponding to the first identifier set to the third party, based on the assumption that the third party is honest with the first party. After receiving the first identifier set and the target features corresponding to the first identifier set sent by the first party, the third party uses the target key agreed upon with the second party in advance to homomorphically encrypt the first identifier set and the target features corresponding to the first identifier set to obtain the first ciphertext. Then, it shuffles the order of the first ciphertext to obtain the fifth ciphertext and sends the fifth ciphertext to the first party. The second party receives the fifth ciphertext sent by the third party and, based on the fifth ciphertext and the second ciphertext sent by the second party, generates the ciphertext of the intersection of the first identifier set and the second identifier set, the ciphertext of the first feature, and the ciphertext of the target tag.
[0133] In the above implementation, after homomorphically encrypting the first identifier set and the target features corresponding to the first identifier set using the target key, the homomorphic encryption result is not directly sent to the first party. Instead, the order of the homomorphic encryption result (i.e., the fifth ciphertext) is shuffled first, and the ciphertext obtained after shuffling (i.e., the first ciphertext) is sent to the first party. In this way, the first party can avoid deducing the target key by associating the first ciphertext with the data it holds, thereby avoiding the problem of data leakage of the second party due to knowing the target key and cracking the second ciphertext.
[0134] In one possible implementation, the second ciphertext is obtained by the second party using the target key to homomorphically encrypt the second identifier set and the tags corresponding to the second identifier set, and then shuffling the order of the homomorphic encryption results.
[0135] Specifically, when the second party executes the above data processing method, it first assumes that the third party is semi-honest towards the second party. In order to avoid data leakage, the second party performs homomorphic encryption on its own data locally. That is, it uses the target key agreed upon with the third party in advance to homomorphically encrypt the second set of identifiers it holds and the tags corresponding to the second set of identifiers to obtain the second ciphertext. Then, it shuffles the order of the second ciphertext to obtain the sixth ciphertext, and determines the sixth ciphertext as the second ciphertext. That is, the homomorphic encryption result obtained after shuffling the order (i.e., the sixth ciphertext) is determined as the new second ciphertext, and the new second ciphertext is sent to the first party; the first party receives the second ciphertext.
[0136] In the above implementation, after homomorphically encrypting the second identifier set and the units corresponding to the second identifier set using the target key, the homomorphic encryption result is not directly sent to the first party. Instead, the order of the homomorphic encryption result is shuffled, and the ciphertext obtained after shuffling (i.e., the sixth ciphertext) is sent to the first party. In this way, the first party can avoid obtaining the order of the plaintext corresponding to the second ciphertext, thereby avoiding data leakage of the second party.
[0137] Figure 3This is a flowchart illustrating a data processing method for secure computation applied to a second party, according to an exemplary embodiment. Figure 3 As shown, the data processing method for secure computing applied to the second party may include S201 to S204.
[0138] In S201, the target key is used to homomorphically encrypt the second identifier set and the tags corresponding to the second identifier set to obtain the second ciphertext.
[0139] In S202, the second ciphertext is sent to the first party so that the second party can generate, based on the first ciphertext and the second ciphertext, the ciphertext of the intersection of the first identifier set and the second identifier set, the ciphertext of the first feature, and the ciphertext of the target tag.
[0140] The first ciphertext is obtained by a third party using the target key to homomorphically encrypt the first identifier set and the target features corresponding to the first identifier set. The target features include at least some of the h-dimensional features. The first feature is the target feature corresponding to the intersection, and the target label is the label corresponding to the intersection.
[0141] In S203, a feature segmentation operation is performed with the first party to obtain a third segment of the first feature, and a label segmentation operation is performed with the first party to obtain a fourth segment of the target label.
[0142] In S204, the target data processing task is executed based on the third and fourth shards.
[0143] In the above technical solution, firstly, assuming the third party (TEE device) is honest with the first party, the first party sends its own first identifier set and at least some of the corresponding dimensional features (i.e., target features) to the third party located on the same local area network. Then, the third party uses the target key to homomorphically encrypt the plaintext data received from the first party, obtaining the first ciphertext, and sends the first ciphertext to the first party. Simultaneously, the second party uses the target key to homomorphically encrypt its own second identifier set and the tags corresponding to the second identifier set, obtaining the second ciphertext, and sends the second ciphertext to the first party. Next, the first party... Based on the first ciphertext received from the third party and the second ciphertext received from the second party, ciphertext of the intersection of the first and second identifier sets, ciphertext of the target feature (i.e., the first feature) corresponding to the intersection, and ciphertext of the label (i.e., the target label) corresponding to the intersection are generated. Then, based on the ciphertext of the first feature, the first party and the second party jointly perform a feature sharding operation to obtain a shard of the first feature respectively. Based on the ciphertext of the target label, the first party and the second party jointly perform a label sharding operation to obtain a shard of the target label respectively. Finally, the first party and the second party respectively perform the target data processing task based on the shards they hold. In this way, based on the assumption that the third party is honest with the first party, the first party can send its plaintext data to the third party located on the same local area network. The third party can then encrypt the plaintext data held by the first party using a target key pre-agreed with the second party. This avoids encrypting the data held by the first party and sending it to the second party via a wide area network. Furthermore, when the first party and the second party perform feature fragmentation and tag fragmentation based on the ciphertext of the first feature and the ciphertext of the target tag, only the ciphertext of the features corresponding to the intersection of the first and second identifier sets is involved, and not all the data held by the first party is involved. By encrypting the data, the first party can avoid sending the encrypted version of all its data to the second party. Since it's unnecessary to send the encrypted version of all the data held by the first party to the second party, the communication volume between the two parties (which falls under wide area network communication) is reduced. This significantly reduces communication volume during feature fragmentation and label fragmentation, especially when the first party's data volume is much larger than the second party's, thereby improving the processing efficiency of the target data processing task. Therefore, this solution can protect data security and reduce communication volume in scenarios such as model training and SQL queries.
[0144] Optionally, the first ciphertext includes a first sub-ciphertext and a second sub-ciphertext of the target feature corresponding to each first identifier. The second sub-ciphertext includes a first random number and a feature ciphertext of the target feature corresponding to the first identifier. The first sub-ciphertext is obtained by a third party using a target key to perform pseudo-random encryption on the first identifier set. The feature ciphertext is obtained by a third party using a target key to perform symmetric encryption on the first random number to obtain a third ciphertext. The third ciphertext is then added to the target feature corresponding to the first identifier.
[0145] Optionally, a feature segmentation operation is performed with the first party to obtain a third segment of the first feature, including:
[0146] In response to receiving the first random number sent by the first party, the first random number is symmetrically encrypted using the target key to obtain the fourth ciphertext;
[0147] Invert the fourth ciphertext to obtain the third fragment of the first feature.
[0148] Optionally, the second identifier set and the tags corresponding to the second identifier set are homomorphically encrypted using the target key to obtain the second ciphertext, including:
[0149] The second identifier set is pseudo-randomly encrypted using the target key to obtain the third sub-ciphertext;
[0150] For each tag corresponding to the second identifier, the tag and the target point are encrypted using elliptic curve cryptography with the target key to obtain the fourth sub-ciphertext of the tag, where the target point is any point on the preset elliptic curve.
[0151] The second ciphertext includes the third sub-ciphertext and the fourth sub-ciphertext of the tag corresponding to each second identifier.
[0152] Optionally, the labels are binary labels;
[0153] Perform tag fragmentation with the first party to obtain the fourth fragment of the target tag, including:
[0154] In response to receiving the masking data sent by the first party, the masking data is homomorphically decrypted using the target key, wherein the masking data is obtained by the first party masking the ciphertext of the target tag using a second random number, which is 0 or 1;
[0155] If the homomorphic decryption result is the target point, then 0 is determined as the fourth segment;
[0156] If the homomorphic decryption result is not the target point, then 1 is determined as the fourth segment.
[0157] Optionally, prior to the step of sending the second ciphertext to the first party, the data processing method for secure computation applied to the second party further includes:
[0158] By shuffling the order of the second ciphertext, the sixth ciphertext is obtained;
[0159] The sixth ciphertext is identified as the second ciphertext.
[0160] Optionally, the target data processing task is a structured query language query task or a machine learning model training task.
[0161] The specific implementation of each step in the data processing method for secure computing applied to a second party according to the embodiments of this disclosure has been described in detail in the data processing method for secure computing applied to a first party according to the embodiments of this disclosure, and will not be repeated here.
[0162] Figure 4 This is a flowchart illustrating a data processing method for secure computing applied to a third party, according to an exemplary embodiment. For example... Figure 4 As shown, the data processing method for secure computing applied to a third party may include S301 and S302.
[0163] In S301, in response to receiving the first identifier set and the target feature corresponding to the first identifier set sent by the first party, the first identifier set and the target feature corresponding to the first identifier set are homomorphically encrypted using the target key to obtain the first ciphertext.
[0164] The target features include at least some of the h-dimensional features.
[0165] In S302, the first ciphertext is sent to the first party so that the first party can generate a first fragment of the first feature and a second fragment of the target label based on the first ciphertext.
[0166] Wherein, the first feature is the target feature corresponding to the intersection of the first identifier set and the second identifier set, and the target label is the label corresponding to the intersection.
[0167] In the above technical solution, firstly, assuming the third party (TEE device) is honest with the first party, the first party sends its own first identifier set and at least some of the corresponding dimensional features (i.e., target features) to the third party located on the same local area network. Then, the third party uses the target key to homomorphically encrypt the plaintext data received from the first party, obtaining the first ciphertext, and sends the first ciphertext to the first party. Simultaneously, the second party uses the target key to homomorphically encrypt its own second identifier set and the tags corresponding to the second identifier set, obtaining the second ciphertext, and sends the second ciphertext to the first party. Next, the first party... Based on the first ciphertext received from the third party and the second ciphertext received from the second party, ciphertext of the intersection of the first and second identifier sets, ciphertext of the target feature (i.e., the first feature) corresponding to the intersection, and ciphertext of the label (i.e., the target label) corresponding to the intersection are generated. Then, based on the ciphertext of the first feature, the first party and the second party jointly perform a feature sharding operation to obtain a shard of the first feature respectively. Based on the ciphertext of the target label, the first party and the second party jointly perform a label sharding operation to obtain a shard of the target label respectively. Finally, the first party and the second party respectively perform the target data processing task based on the shards they hold. In this way, based on the assumption that the third party is honest with the first party, the first party can send its plaintext data to the third party located on the same local area network. The third party can then encrypt the plaintext data held by the first party using a target key pre-agreed with the second party. This avoids encrypting the data held by the first party and sending it to the second party via a wide area network. Furthermore, when the first party and the second party perform feature fragmentation and tag fragmentation based on the ciphertext of the first feature and the ciphertext of the target tag, only the ciphertext of the features corresponding to the intersection of the first and second identifier sets is involved, and not all the data held by the first party is involved. By encrypting the data, the first party can avoid sending the encrypted version of all its data to the second party. Since it's unnecessary to send the encrypted version of all the data held by the first party to the second party, the communication volume between the two parties (which falls under wide area network communication) is reduced. This significantly reduces communication volume during feature fragmentation and label fragmentation, especially when the first party's data volume is much larger than the second party's, thereby improving the processing efficiency of the target data processing task. Therefore, this solution can protect data security and reduce communication volume in scenarios such as model training and SQL queries.
[0168] Optionally, the first identifier set and the target features corresponding to the first identifier set are homomorphically encrypted using the target key to obtain the first ciphertext, including:
[0169] The first identifier set is pseudo-randomly encrypted using the target key to obtain the first sub-ciphertext;
[0170] Generate a first random number, and then use the target key to symmetrically encrypt the first random number to obtain the third ciphertext;
[0171] For each target feature corresponding to the first identifier, the third ciphertext is added to the target feature to obtain the feature ciphertext of the target feature;
[0172] The first ciphertext includes a first sub-ciphertext and a second sub-ciphertext of the target feature corresponding to each first identifier. The second sub-ciphertext includes a first random number and the feature ciphertext of the target feature corresponding to the first identifier.
[0173] Optionally, prior to the step of sending the first ciphertext to the first party, the data processing method for secure computation applied to the third party further includes:
[0174] By shuffling the order of the first ciphertext, the fifth ciphertext is obtained;
[0175] The first ciphertext is sent to the first party, which then generates a first fragment of the first feature and a second fragment of the target label based on the first ciphertext, including:
[0176] The fifth ciphertext is sent to the first party, which then generates a first fragment of the first feature and a second fragment of the target label based on the fifth ciphertext.
[0177] The specific implementation of each step in the data processing method for secure computing applied to a third party according to the embodiments of this disclosure has been described in detail in the data processing method for secure computing applied to a first party according to the embodiments of this disclosure, and will not be repeated here.
[0178] Figure 5 This is a timing diagram illustrating a data processing method for secure computing according to an exemplary embodiment. For example... Figure 5 As shown, the data processing method for secure computing may include the following S401 to S410.
[0179] In S401, the first identifier set and the target features corresponding to the first identifier set are sent to a third party.
[0180] In S402, the third party uses the target key to homomorphically encrypt the first identifier set and the target features corresponding to the first identifier set to obtain the first ciphertext.
[0181] In S403, the third party sends the first ciphertext to the first party.
[0182] In S404, the second party uses the target key to homomorphically encrypt the second identifier set and the tags corresponding to the second identifier set to obtain the second ciphertext.
[0183] In S405, the second party sends the second ciphertext to the first party.
[0184] In S406, the first party generates, based on the first ciphertext and the second ciphertext, the ciphertext of the intersection of the first identifier set and the second identifier set, the ciphertext of the first feature, and the ciphertext of the target tag.
[0185] In S407, the first party performs a feature fragmentation operation with the second party based on the ciphertext of the first feature to obtain a first fragment of the first feature, and performs a tag fragmentation operation with the second party based on the ciphertext of the target tag to obtain a second fragment of the target tag.
[0186] In S408, the second party and the first party perform a feature segmentation operation to obtain a third segment of the first feature, and perform a label segmentation operation to obtain a fourth segment of the target label.
[0187] In S409, the first party performs the target data processing task based on the first and second fragments.
[0188] In S410, the second party performs the target data processing task based on the third and fourth fragments.
[0189] The specific implementations of S401 to S410 described above have been described in detail in the data processing method for secure computing applied to a first party according to the embodiments of this disclosure, and will not be repeated here.
[0190] Figure 6 This is a timing diagram illustrating a data processing method for secure computing according to another exemplary embodiment. (e.g.) Figure 6 As shown, the data processing method for secure computing may include the following S501 to S522.
[0191] In S501, the first party sends the first identifier set and the target features corresponding to the first identifier set to the third party.
[0192] In S502, a third party uses the target key to perform pseudo-random encryption on the first identifier set to obtain the first sub-ciphertext.
[0193] In S503, a third party generates a first random number, and then uses the target key to symmetrically encrypt the first random number to obtain the third ciphertext.
[0194] In S504, for each target feature corresponding to the first identifier, the third party adds the third ciphertext to the target feature to obtain the feature ciphertext of the target feature.
[0195] In S505, the third party sends the first ciphertext to the first party, wherein the first ciphertext includes a first sub-ciphertext and a second sub-ciphertext of the target feature corresponding to each first identifier, and the second sub-ciphertext includes a first random number and the feature ciphertext of the target feature corresponding to the first identifier.
[0196] In S506, the second party uses the target key to perform pseudo-random encryption on the second identifier set to obtain the third sub-ciphertext.
[0197] In S507, the second party uses the target key to perform elliptic curve encryption on the tag and the target point for each tag corresponding to the second identifier, and obtains the fourth sub-ciphertext of the tag.
[0198] In S508, the second party sends the second ciphertext to the first party, wherein the second ciphertext includes the third sub-ciphertext and the fourth sub-ciphertext of the tag corresponding to each second identifier.
[0199] In S509, the first party calculates the intersection of the first sub-ciphertext and the third sub-ciphertext to obtain the ciphertext of the intersection of the first identifier set and the second identifier set.
[0200] In S510, the first party determines the ciphertext corresponding to the intersection of the ciphertexts from all the second sub-ciphertexts to obtain the ciphertext of the first feature, and determines the ciphertext corresponding to the intersection of the ciphertexts from all the fourth sub-ciphertexts to obtain the ciphertext of the target label.
[0201] In S511, the first party obtains the first random number in the ciphertext of the first feature.
[0202] In S512, the first party sends a first random number to the second party.
[0203] In S513, the first party determines the ciphertext of the first feature as the first fragment of the first feature.
[0204] In S514, the second party uses the target key to symmetrically encrypt the first random number to obtain the fourth ciphertext.
[0205] In S515, the second party inverts the fourth ciphertext to obtain the third fragment of the first feature.
[0206] In S516, the first party generates a second random number and uses the second random number to mask the ciphertext of the target tag to obtain masked data.
[0207] The second random number is either 0 or 1.
[0208] In S517, the first party sends masked data to the second party.
[0209] In S518, the first party uses the second random number as the second slice of the target label.
[0210] In S519, the second party uses the target key to perform homomorphic decryption of the masked data.
[0211] In S520, if the second party determines that the homomorphic decryption result is the target point, then 0 is determined as the fourth segment; if it determines that the homomorphic decryption result is not the target point, then 1 is determined as the fourth segment.
[0212] In S521, the first party performs the target data processing task based on the first and second fragments.
[0213] In S522, the second party performs the target data processing task based on the third and fourth fragments.
[0214] In S501, the LAN communication between the first party and the third party is (logI|+64m)N, with a communication round of 1, where N is the number of first identifiers in the first identifier set, and the target feature is 64 bits; in S505, the LAN communication between the first party and the third party is (λ+64m)N, with a communication round of 1; in S508, the WAN communication between the first party and the second party is (λ+512)n, with a communication round of 1, where the size of the third sub-ciphertext is 512n bits, and the size of the fourth sub-ciphertext is λ bits; in S512, the WAN communication between the first party and the second party is λn, with a communication round of 1; in S517, the WAN communication between the first party and the second party is 512n bits, with a communication round of 1.
[0215] Therefore, the total LAN communication volume between the first party and the third party is (logI|+64m)N+(λ+64m)N, with a total of 2 communication rounds; the total WAN communication volume between the first party and the second party is (2λ+1024)n, with a total of 3 communication rounds. It is evident that the total WAN communication volume between the first party and the second party is independent of both m and N.
[0216] The specific implementations of S501 to S522 described above have been described in detail in the data processing method for secure computing applied to a first party according to the embodiments of this disclosure, and will not be repeated here.
[0217] Furthermore, in this disclosure, the second party and the third party only communicate once when synchronizing the target key, and do not communicate at other times. Also, the third party only participates in the computation when encrypting the first party's private data, and does not participate in the computation at other stages.
[0218] In addition, in order to improve the processing efficiency of the target data processing task, the above-mentioned S101, S301~S302, S401~S402 and S501~S505 can all be performed offline.
[0219] Figure 7This is a block diagram illustrating a data processing apparatus for secure computation applied to a first party, according to an exemplary embodiment. The parties involved in the secure computation include a first party, a second party, and a third party. The first party holds a first set of identifiers, where each identifier in the first set corresponds to an h-dimensional feature. The second party holds a second set of identifiers, where each identifier in the second set corresponds to a label. The third party is a trusted execution environment device, and the third party is located within the same local area network as the first party, where h ≥ 1. Figure 7 As shown, a data processing apparatus 600 for secure computation applied to a first party includes: a first sending module 601, configured to send a first identifier set and a target feature corresponding to the first identifier set to the third party, so that the third party uses a target key to homomorphically encrypt the first identifier set and the target feature corresponding to the first identifier set to obtain a first ciphertext, and sends the first ciphertext to the first party, wherein the target feature includes at least some of the h-dimensional features; and a generation module 602, configured to, in response to receiving the first ciphertext and a second ciphertext sent by the second party, generate, based on the first ciphertext and the second ciphertext, a ciphertext of the intersection of the first identifier set and the second identifier set, and a ciphertext of the second identifier set. The system comprises: a ciphertext of a feature and a ciphertext of a target tag, wherein the second ciphertext is obtained by the second party using the target key to homomorphically encrypt the second identifier set and the tag corresponding to the second identifier set; the first feature is the target feature corresponding to the intersection; and the target tag is the tag corresponding to the intersection. A first sharding module 603 is configured to perform a feature sharding operation with the second party based on the ciphertext of the first feature to obtain a first shard of the first feature, and to perform a tag sharding operation with the second party based on the ciphertext of the target tag to obtain a second shard of the target tag. A first execution module 604 is configured to execute a target data processing task based on the first shard and the second shard.
[0220] In the above technical solution, firstly, assuming the third party (TEE device) is honest with the first party, the first party sends its own first identifier set and at least some of the corresponding dimensional features (i.e., target features) to the third party located on the same local area network. Then, the third party uses the target key to homomorphically encrypt the plaintext data received from the first party, obtaining the first ciphertext, and sends the first ciphertext to the first party. Simultaneously, the second party uses the target key to homomorphically encrypt its own second identifier set and the tags corresponding to the second identifier set, obtaining the second ciphertext, and sends the second ciphertext to the first party. Next, the first party... Based on the first ciphertext received from the third party and the second ciphertext received from the second party, ciphertext of the intersection of the first and second identifier sets, ciphertext of the target feature (i.e., the first feature) corresponding to the intersection, and ciphertext of the label (i.e., the target label) corresponding to the intersection are generated. Then, based on the ciphertext of the first feature, the first party and the second party jointly perform a feature sharding operation to obtain a shard of the first feature respectively. Based on the ciphertext of the target label, the first party and the second party jointly perform a label sharding operation to obtain a shard of the target label respectively. Finally, the first party and the second party respectively perform the target data processing task based on the shards they hold. In this way, based on the assumption that the third party is honest with the first party, the first party can send its plaintext data to the third party located on the same local area network. The third party can then encrypt the plaintext data held by the first party using a target key pre-agreed with the second party. This avoids encrypting the data held by the first party and sending it to the second party via a wide area network. Furthermore, when the first party and the second party perform feature fragmentation and tag fragmentation based on the ciphertext of the first feature and the ciphertext of the target tag, only the ciphertext of the features corresponding to the intersection of the first and second identifier sets is involved, and not all the data held by the first party is involved. By encrypting the data, the first party can avoid sending the encrypted version of all its data to the second party. Since it's unnecessary to send the encrypted version of all the data held by the first party to the second party, the communication volume between the two parties (which falls under wide area network communication) is reduced. This significantly reduces communication volume during feature fragmentation and label fragmentation, especially when the first party's data volume is much larger than the second party's, thereby improving the processing efficiency of the target data processing task. Therefore, this solution can protect data security and reduce communication volume in scenarios such as model training and SQL queries.
[0221] Optionally, the first ciphertext includes a first sub-ciphertext and a second sub-ciphertext corresponding to the target feature for each of the first identifiers. The second sub-ciphertext includes a first random number and a feature ciphertext of the target feature corresponding to the first identifier. The first sub-ciphertext is obtained by the third party using the target key to perform pseudo-random encryption on the first identifier set. The feature ciphertext is obtained by the third party using the target key to perform symmetric encryption on the first random number to obtain a third ciphertext, and the third ciphertext is added to the target feature corresponding to the first identifier.
[0222] Optionally, the first fragmentation module 603 includes: an acquisition submodule, configured to acquire the first random number in the ciphertext of the first feature and send the first random number to the second party, so that the second party can use the target key to perform symmetric encryption on the first random number to obtain a third fragment of the first feature; and a first determination submodule, configured to determine the feature ciphertext of the first feature as the first fragment.
[0223] Optionally, the tag is a binary classification tag; the first fragmentation module 603 includes: a first generation submodule, used to generate a second random number, and use the second random number to mask the ciphertext of the target tag to obtain masked data, wherein the second random number is 0 or 1; a sending submodule, used to send the masked data to the second party, so that the second party can use the target key to perform homomorphic decryption on the masked data, and generate a fourth fragment of the target tag according to the homomorphic decryption result; and a second determination submodule, used to use the second random number as the second fragment.
[0224] Optionally, the second ciphertext includes a third sub-ciphertext and a fourth sub-ciphertext of the tag corresponding to each of the second identifiers, wherein the third sub-ciphertext is obtained by the second party using the target key to perform pseudo-random encryption on the second identifier set, and the fourth sub-ciphertext is obtained by the second party using the target key to perform elliptic curve encryption on the tag corresponding to the second identifier and the target point, wherein the target point is any point on a preset elliptic curve.
[0225] Optionally, the generation module 602 includes: a calculation submodule, configured to calculate the intersection of the first sub-ciphertext and the third sub-ciphertext to obtain the ciphertext of the intersection of the first identifier set and the second identifier set; and a third determination submodule, configured to determine the ciphertext corresponding to the ciphertext of the intersection from all the second sub-ciphertexts to obtain the ciphertext of the first feature, and to determine the ciphertext corresponding to the ciphertext of the intersection from all the fourth sub-ciphertexts to obtain the ciphertext of the target tag.
[0226] Optionally, the first sending module 601 is configured to send the first identifier set and the target feature corresponding to the first identifier set to the third party, so that the third party can use the target key to homomorphically encrypt the first identifier set and the target feature corresponding to the first identifier set to obtain a first ciphertext, and shuffle the order of the first ciphertext to obtain a fifth ciphertext, and send the fifth ciphertext to the first party; the generation module 602 is configured to, in response to receiving the fifth ciphertext and the second ciphertext sent by the second party, generate the ciphertext of the intersection of the first identifier set and the second identifier set, the ciphertext of the first feature, and the ciphertext of the target tag according to the fifth ciphertext and the second ciphertext.
[0227] Optionally, the second ciphertext is obtained by the second party using the target key to homomorphically encrypt the second identifier set and the tags corresponding to the second identifier set, and then shuffling the order of the homomorphic encryption results.
[0228] Optionally, the target data processing task is a structured query language query task or a machine learning model training task.
[0229] Figure 8 This is a block diagram illustrating a data processing apparatus for secure computation applied to a second party, according to an exemplary embodiment. The parties involved in the secure computation include a first party, a second party, and a third party. The first party holds a first set of identifiers, where each identifier in the first set corresponds to an h-dimensional feature. The second party holds a second set of identifiers, where each identifier in the second set corresponds to a label. The third party is a trusted execution environment device, and the third party is located within the same local area network as the first party, where h ≥ 1. Figure 8As shown, the data processing device 700 for secure computation applied to a second party includes: a first encryption module 701, used to perform homomorphic encryption on the second identifier set and the tag corresponding to the second identifier set using a target key to obtain a second ciphertext; a second sending module 702, used to send the second ciphertext to the first party so that the second party can generate ciphertext of the intersection of the first identifier set and the second identifier set, ciphertext of a first feature, and ciphertext of a target tag based on the first ciphertext and the second ciphertext, wherein the first ciphertext is obtained by the third party using the target key to perform homomorphic encryption on the first identifier set and the target feature corresponding to the first identifier set, the target feature includes at least some of the h-dimensional features, the first feature is the target feature corresponding to the intersection, and the target tag is the tag corresponding to the intersection; a second sharding module 703, used to perform feature sharding operation with the first party to obtain a third shard of the first feature, and perform tag sharding operation with the first party to obtain a fourth shard of the target tag; and a second execution module 704, used to execute a target data processing task based on the third shard and the fourth shard.
[0230] In the above technical solution, firstly, assuming the third party (TEE device) is honest with the first party, the first party sends its own first identifier set and at least some of the corresponding dimensional features (i.e., target features) to the third party located on the same local area network. Then, the third party uses the target key to homomorphically encrypt the plaintext data received from the first party, obtaining the first ciphertext, and sends the first ciphertext to the first party. Simultaneously, the second party uses the target key to homomorphically encrypt its own second identifier set and the tags corresponding to the second identifier set, obtaining the second ciphertext, and sends the second ciphertext to the first party. Next, the first party... Based on the first ciphertext received from the third party and the second ciphertext received from the second party, ciphertext of the intersection of the first and second identifier sets, ciphertext of the target feature (i.e., the first feature) corresponding to the intersection, and ciphertext of the label (i.e., the target label) corresponding to the intersection are generated. Then, based on the ciphertext of the first feature, the first party and the second party jointly perform a feature sharding operation to obtain a shard of the first feature respectively. Based on the ciphertext of the target label, the first party and the second party jointly perform a label sharding operation to obtain a shard of the target label respectively. Finally, the first party and the second party respectively perform the target data processing task based on the shards they hold. In this way, based on the assumption that the third party is honest with the first party, the first party can send its plaintext data to the third party located on the same local area network. The third party can then encrypt the plaintext data held by the first party using a target key pre-agreed with the second party. This avoids encrypting the data held by the first party and sending it to the second party via a wide area network. Furthermore, when the first party and the second party perform feature fragmentation and tag fragmentation based on the ciphertext of the first feature and the ciphertext of the target tag, only the ciphertext of the features corresponding to the intersection of the first and second identifier sets is involved, and not all the data held by the first party is involved. By encrypting the data, the first party can avoid sending the encrypted version of all its data to the second party. Since it's unnecessary to send the encrypted version of all the data held by the first party to the second party, the communication volume between the two parties (which falls under wide area network communication) is reduced. This significantly reduces communication volume during feature fragmentation and label fragmentation, especially when the first party's data volume is much larger than the second party's, thereby improving the processing efficiency of the target data processing task. Therefore, this solution can protect data security and reduce communication volume in scenarios such as model training and SQL queries.
[0231] Optionally, the first ciphertext includes a first sub-ciphertext and a second sub-ciphertext corresponding to the target feature for each of the first identifiers. The second sub-ciphertext includes a first random number and a feature ciphertext of the target feature corresponding to the first identifier. The first sub-ciphertext is obtained by the third party using the target key to perform pseudo-random encryption on the first identifier set. The feature ciphertext is obtained by the third party using the target key to perform symmetric encryption on the first random number to obtain a third ciphertext, and the third ciphertext is added to the target feature corresponding to the first identifier.
[0232] Optionally, the second sharding module 703 includes: a first encryption submodule, configured to, in response to receiving the first random number sent by the first party, perform symmetric encryption on the first random number using the target key to obtain a fourth ciphertext; and an inversion submodule, configured to invert the fourth ciphertext to obtain a third shard of the first feature.
[0233] Optionally, the first encryption module 701 includes: a second encryption submodule, configured to perform pseudo-random encryption on the second identifier set using a target key to obtain a third sub-ciphertext; and a third encryption submodule, configured to perform elliptic curve encryption on the tag corresponding to each second identifier and a target point using the target key to obtain a fourth sub-ciphertext of the tag, wherein the target point is any point on a preset elliptic curve; wherein the second ciphertext includes the third sub-ciphertext and the fourth sub-ciphertext of the tag corresponding to each second identifier.
[0234] Optionally, the tag is a binary classification tag; the second fragmentation module 703 includes: a decryption submodule, configured to, in response to receiving the masking data sent by the first party, perform homomorphic decryption on the masking data using the target key, wherein the masking data is obtained by the first party masking the ciphertext of the target tag using a second random number, the second random number being 0 or 1; a fourth determination submodule, configured to, if the homomorphic decryption result is the target point, determine 0 as the fourth fragment; and a fifth determination submodule, configured to, if the homomorphic decryption result is not the target point, determine 1 as the fourth fragment.
[0235] Optionally, the data processing device 700 for secure computing applied to the second party further includes: a first scrambling module, configured to scramble the order of the second ciphertext to obtain a sixth ciphertext before the second sending module 702 sends the second ciphertext to the first party; and a determining module, configured to determine the sixth ciphertext as the second ciphertext.
[0236] Optionally, the target data processing task is a structured query language query task or a machine learning model training task.
[0237] Figure 9 This is a block diagram illustrating a data processing apparatus for secure computing applied to a third party, according to an exemplary embodiment. The participants in the secure computing include a first party, a second party, and a third party. The first party holds a first set of identifiers, where each identifier in the first set corresponds to an h-dimensional feature. The second party holds a second set of identifiers, where each identifier in the second set corresponds to a label. The third party is a trusted execution environment device, and the third party is located within the same local area network as the first party, where h ≥ 1. Figure 9 As shown, a data processing device 800 for secure computing applied to a third party includes: a second encryption module 801, configured to, in response to receiving a first identifier set and a target feature corresponding to the first identifier set sent by the first party, perform homomorphic encryption on the first identifier set and the target feature corresponding to the first identifier set using a target key to obtain a first ciphertext; wherein the target feature includes at least some of the h-dimensional features; and a third sending module 802, configured to send the first ciphertext to the first party, so that the first party generates a first fragment of the first feature and a second fragment of the target label based on the first ciphertext, wherein the first feature is the target feature corresponding to the intersection of the first identifier set and the second identifier set, and the target label is the label corresponding to the intersection.
[0238] In the above technical solution, firstly, assuming the third party (TEE device) is honest with the first party, the first party sends its own first identifier set and at least some of the corresponding dimensional features (i.e., target features) to the third party located on the same local area network. Then, the third party uses the target key to homomorphically encrypt the plaintext data received from the first party, obtaining the first ciphertext, and sends the first ciphertext to the first party. Simultaneously, the second party uses the target key to homomorphically encrypt its own second identifier set and the tags corresponding to the second identifier set, obtaining the second ciphertext, and sends the second ciphertext to the first party. Next, the first party... Based on the first ciphertext received from the third party and the second ciphertext received from the second party, ciphertext of the intersection of the first and second identifier sets, ciphertext of the target feature (i.e., the first feature) corresponding to the intersection, and ciphertext of the label (i.e., the target label) corresponding to the intersection are generated. Then, based on the ciphertext of the first feature, the first party and the second party jointly perform a feature sharding operation to obtain a shard of the first feature respectively. Based on the ciphertext of the target label, the first party and the second party jointly perform a label sharding operation to obtain a shard of the target label respectively. Finally, the first party and the second party respectively perform the target data processing task based on the shards they hold. In this way, based on the assumption that the third party is honest with the first party, the first party can send its plaintext data to the third party located on the same local area network. The third party can then encrypt the plaintext data held by the first party using a target key pre-agreed with the second party. This avoids encrypting the data held by the first party and sending it to the second party via a wide area network. Furthermore, when the first party and the second party perform feature fragmentation and tag fragmentation based on the ciphertext of the first feature and the ciphertext of the target tag, only the ciphertext of the features corresponding to the intersection of the first and second identifier sets is involved, and not all the data held by the first party is involved. By encrypting the data, the first party can avoid sending the encrypted version of all its data to the second party. Since it's unnecessary to send the encrypted version of all the data held by the first party to the second party, the communication volume between the two parties (which falls under wide area network communication) is reduced. This significantly reduces communication volume during feature fragmentation and label fragmentation, especially when the first party's data volume is much larger than the second party's, thereby improving the processing efficiency of the target data processing task. Therefore, this solution can protect data security and reduce communication volume in scenarios such as model training and SQL queries.
[0239] Optionally, the second encryption module 802 includes: a fourth encryption submodule, configured to perform pseudo-random encryption on the first identifier set using a target key to obtain a first sub-ciphertext; a second generation submodule, configured to generate a first random number and perform symmetric encryption on the first random number using the target key to obtain a third ciphertext; and an addition submodule, configured to add the third ciphertext to the target feature corresponding to each first identifier to obtain the feature ciphertext of the target feature; wherein the first ciphertext includes the first sub-ciphertext and a second sub-ciphertext of the target feature corresponding to each first identifier, and the second ciphertext includes the first random number and the feature ciphertext of the target feature corresponding to the first identifier.
[0240] Optionally, the data processing device 800 for secure computing applied to a third party further includes: a second scrambling module, configured to scramble the order of the first ciphertext to obtain a fifth ciphertext before the third sending module 802 sends the first ciphertext to the first party; the third sending module 802 is configured to send the fifth ciphertext to the first party so that the first party can generate a first fragment of the first feature and a second fragment of the target tag based on the fifth ciphertext.
[0241] This disclosure also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the data processing method for secure computing provided herein, applied to a first party, applied to a second party, or applied to a third party.
[0242] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the data processing method for secure computing provided above, applied to a first party, applied to a second party, or applied to a third party.
[0243] The following is for reference. Figure 10 The diagram illustrates a structural schematic of an electronic device (e.g., a terminal device or a server) 900 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 10 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0244] like Figure 10 As shown, electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 902 or a program loaded from storage device 908 into random access memory (RAM) 903. RAM 903 also stores various programs and data required for the operation of electronic device 900. Processing device 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.
[0245] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 10 An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0246] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of embodiments of this disclosure.
[0247] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0248] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0249] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0250] The aforementioned computer-readable medium carries one or more programs. When the electronic device executes the aforementioned one or more programs, the electronic device causes the electronic device to: send a first identifier set and the target features corresponding to the first identifier set to a third party, so that the third party uses a target key to homomorphically encrypt the first identifier set and the target features corresponding to the first identifier set to obtain a first ciphertext, and send the first ciphertext to the first party. The participants in the secure computation include the first party, the second party, and the third party. The first party holds a first identifier set, and the first identifier in the first identifier set corresponds to h-dimensional features. The second party holds a second identifier set, and the second identifier in the second identifier set corresponds to a tag. The third party is a trusted execution environment device, and the third party and the first party are located in the same local area network, where h ≥ 1. The target features include at least one of the h-dimensional features. A few dimensional features; in response to receiving the first ciphertext and the second ciphertext sent by the second party, based on the first ciphertext and the second ciphertext, generate the ciphertext of the intersection of the first identifier set and the second identifier set, the ciphertext of the first feature, and the ciphertext of the target tag, wherein the second ciphertext is obtained by the second party using the target key to homomorphically encrypt the second identifier set and the tags corresponding to the second identifier set, the first feature is the target feature corresponding to the intersection, and the target tag is the tag corresponding to the intersection; based on the ciphertext of the first feature, perform a feature fragmentation operation with the second party to obtain a first fragment of the first feature, and based on the ciphertext of the target tag, perform a tag fragmentation operation with the second party to obtain a second fragment of the target tag; based on the first fragment and the second fragment, perform a target data processing task.
[0251] Alternatively, the aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: homomorphically encrypt a second identifier set and the tags corresponding to the second identifier set using a target key to obtain a second ciphertext; wherein the participants in the secure computation include a first party, a second party, and a third party; the first party holds a first identifier set, where the first identifier in the first identifier set corresponds to h-dimensional features; the second party holds a second identifier set, where the second identifier in the second identifier set corresponds to a tag; and the third party is a trusted execution environment device located on the same local area network as the first party, where h ≥ 1; and sends the second ciphertext to the first party so that the second party can, based on the first ciphertext and the... The second ciphertext generates the ciphertext of the intersection of the first identifier set and the second identifier set, the ciphertext of the first feature, and the ciphertext of the target label. The first ciphertext is obtained by the third party using the target key to homomorphically encrypt the first identifier set and the target feature corresponding to the first identifier set. The target feature includes at least some of the h-dimensional features. The first feature is the target feature corresponding to the intersection, and the target label is the label corresponding to the intersection. A feature sharding operation is performed with the first party to obtain a third shard of the first feature, and a label sharding operation is performed with the first party to obtain a fourth shard of the target label. Based on the third and fourth shards, a target data processing task is performed.
[0252] Alternatively, the aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: in response to receiving a first identifier set and a target feature corresponding to the first identifier set sent by a first party, perform homomorphic encryption on the first identifier set and the target feature corresponding to the first identifier set using a target key to obtain a first ciphertext. The participants in the secure computation include a first party, a second party, and a third party. The first party holds a first identifier set, where the first identifier in the first identifier set corresponds to h-dimensional features. The second party holds a second identifier set, where the second identifier in the second identifier set corresponds to a tag. The third party is a trusted execution environment device, and the third party is located within the same local area network as the first party, where h ≥ 1. The target feature includes at least some of the h-dimensional features. The first ciphertext is sent to the first party, so that the first party generates a first fragment of the first feature and a second fragment of the target tag based on the first ciphertext. The first feature is the target feature corresponding to the intersection of the first identifier set and the second identifier set, and the target tag is the tag corresponding to the intersection.
[0253] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0254] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0255] The modules described in the embodiments of this disclosure can be implemented in software or in hardware. The names of the modules do not necessarily limit the module itself; for example, the first execution module can also be described as "a module that executes a target data processing task based on the first and second slices".
[0256] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0257] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0258] According to one or more embodiments of this disclosure, Example 1 provides a data processing method for secure computing. The participants in the secure computing include a first party, a second party, and a third party. The first party holds a first identifier set, where each first identifier in the first identifier set corresponds to an h-dimensional feature. The second party holds a second identifier set, where each second identifier in the second identifier set corresponds to a tag. The third party is a trusted execution environment device, and the third party and the first party are located on the same local area network, where h ≥ 1. The method is applied to the first party and includes: sending the first identifier set and the target feature corresponding to the first identifier set to the third party, so that the third party uses a target key to homomorphically encrypt the first identifier set and the target feature corresponding to the first identifier set to obtain a first ciphertext, and sending the first ciphertext to the first party, wherein the target feature includes the h-dimensional feature. The features include at least some of the following dimensions: In response to receiving the first ciphertext and the second ciphertext sent by the second party, based on the first ciphertext and the second ciphertext, ciphertext of the intersection of the first identifier set and the second identifier set, ciphertext of the first feature, and ciphertext of the target tag are generated, wherein the second ciphertext is obtained by the second party using the target key to homomorphically encrypt the second identifier set and the tags corresponding to the second identifier set; the first feature is the target feature corresponding to the intersection, and the target tag is the tag corresponding to the intersection; based on the ciphertext of the first feature, a feature fragmentation operation is performed with the second party to obtain a first fragment of the first feature, and based on the ciphertext of the target tag, a tag fragmentation operation is performed with the second party to obtain a second fragment of the target tag; based on the first fragment and the second fragment, a target data processing task is performed.
[0259] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, wherein the first ciphertext includes a first sub-ciphertext and a second sub-ciphertext corresponding to the target feature of each first identifier, wherein the second sub-ciphertext includes a first random number and a feature ciphertext of the target feature corresponding to the first identifier, the first sub-ciphertext is obtained by the third party using the target key to perform pseudo-random encryption on the first identifier set, the feature ciphertext is obtained by the third party using the target key to perform symmetric encryption on the first random number to obtain a third ciphertext, and the third ciphertext is added to the target feature corresponding to the first identifier.
[0260] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 2, wherein performing a feature fragmentation operation with a second party based on the ciphertext of the first feature to obtain a first fragment of the first feature includes: obtaining a first random number in the ciphertext of the first feature and sending the first random number to the second party, so that the second party uses the target key to symmetrically encrypt the first random number to obtain a third fragment of the first feature; and determining the feature ciphertext of the first feature as the first fragment.
[0261] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 1, wherein the tag is a binary classification tag; the step of performing tag fragmentation operation with the second party based on the ciphertext of the target tag to obtain a second fragment of the target tag includes: generating a second random number; using the second random number to mask the ciphertext of the target tag to obtain masked data, wherein the second random number is 0 or 1; sending the masked data to the second party, so that the second party can perform homomorphic decryption on the masked data using the target key, and generate a fourth fragment of the target tag based on the homomorphic decryption result; and using the second random number as the second fragment.
[0262] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 2, wherein the second ciphertext includes a third sub-ciphertext and a fourth sub-ciphertext of the tag corresponding to each second identifier, wherein the third sub-ciphertext is obtained by the second party using the target key to perform pseudo-random encryption on the second identifier set, and the fourth sub-ciphertext is obtained by the second party using the target key to perform elliptic curve encryption on the tag corresponding to the second identifier and the target point, wherein the target point is any point on a preset elliptic curve.
[0263] According to one or more embodiments of this disclosure, Example 6 provides the method of Example 5, wherein generating ciphertext of the intersection of the first identifier set and the second identifier set, ciphertext of the first feature, and ciphertext of the target tag based on the first ciphertext and the second ciphertext includes: calculating the intersection of the first sub-ciphertext and the third sub-ciphertext to obtain the ciphertext of the intersection of the first identifier set and the second identifier set; determining the ciphertext corresponding to the ciphertext of the intersection from all the second sub-ciphertexts to obtain the ciphertext of the first feature; and determining the ciphertext corresponding to the ciphertext of the intersection from all the fourth sub-ciphertexts to obtain the ciphertext of the target tag.
[0264] According to one or more embodiments of this disclosure, Example 7 provides a method of Example 1, wherein sending the first identifier set and the target feature corresponding to the first identifier set to the third party, so that the third party uses a target key to homomorphically encrypt the first identifier set and the target feature corresponding to the first identifier set to obtain a first ciphertext, and sending the first ciphertext to the first party, includes: sending the first identifier set and the target feature corresponding to the first identifier set to the third party, so that the third party uses a target key to homomorphically encrypt the first identifier set and the target feature corresponding to the first identifier set to obtain a first ciphertext, and sending the first ciphertext to the first party. Obtaining a first ciphertext and shuffling its order to obtain a fifth ciphertext, then sending the fifth ciphertext to the first party; the step of generating, in response to receiving the first ciphertext and the second ciphertext sent by the second party, the ciphertext of the intersection of the first identifier set and the second identifier set, the ciphertext of the first feature, and the ciphertext of the target tag based on the first ciphertext and the second ciphertext, includes: in response to receiving the fifth ciphertext and the second ciphertext sent by the second party, generating, in response to receiving the fifth ciphertext and the second ciphertext, the ciphertext of the intersection of the first identifier set and the second identifier set, the ciphertext of the first feature, and the ciphertext of the target tag based on the fifth ciphertext and the second ciphertext.
[0265] According to one or more embodiments of this disclosure, Example 8 provides the method of Example 1, wherein the second ciphertext is obtained by the second party using the target key to homomorphically encrypt the second identifier set and the tag corresponding to the second identifier set, and shuffling the order of the homomorphic encryption results.
[0266] According to one or more embodiments of this disclosure, Example 9 provides a method of any one of Examples 1-8, wherein the target data processing task is a structured query language query task or a machine learning model training task.
[0267] According to one or more embodiments of this disclosure, Example 10 provides a data processing method for secure computing. The participants in the secure computing include a first party, a second party, and a third party. The first party holds a first identifier set, where each identifier in the first identifier set corresponds to an h-dimensional feature. The second party holds a second identifier set, where each identifier in the second identifier set corresponds to a tag. The third party is a trusted execution environment device located on the same local area network as the first party, where h ≥ 1. The method is applied to the second party and includes: homomorphically encrypting the second identifier set and the tags corresponding to the second identifier set using a target key to obtain a second ciphertext; sending the second ciphertext to the first party so that the second party can use the first key to perform data processing based on the first ciphertext. The first ciphertext and the second ciphertext are used to generate ciphertext of the intersection of the first identifier set and the second identifier set, ciphertext of the first feature, and ciphertext of the target label. The first ciphertext is obtained by the third party using the target key to homomorphically encrypt the first identifier set and the target feature corresponding to the first identifier set. The target feature includes at least some of the h-dimensional features. The first feature is the target feature corresponding to the intersection, and the target label is the label corresponding to the intersection. A feature sharding operation is performed with the first party to obtain a third shard of the first feature, and a label sharding operation is performed with the first party to obtain a fourth shard of the target label. Based on the third and fourth shards, a target data processing task is performed.
[0268] According to one or more embodiments of this disclosure, Example 11 provides the method of Example 10, wherein the first ciphertext includes a first sub-ciphertext and a second sub-ciphertext corresponding to the target feature for each first identifier, wherein the second sub-ciphertext includes a first random number and a feature ciphertext of the target feature corresponding to the first identifier, the first sub-ciphertext is obtained by the third party using the target key to perform pseudo-random encryption on the first identifier set, the feature ciphertext is obtained by the third party using the target key to perform symmetric encryption on the first random number to obtain a third ciphertext, and the third ciphertext is added to the target feature corresponding to the first identifier.
[0269] According to one or more embodiments of this disclosure, Example 12 provides a method of Example 11, wherein performing a feature fragmentation operation with the first party to obtain a third fragment of the first feature includes: in response to receiving a first random number sent by the first party, symmetrically encrypting the first random number using the target key to obtain a fourth ciphertext; and inverting the fourth ciphertext to obtain a third fragment of the first feature.
[0270] According to one or more embodiments of this disclosure, Example 13 provides the method of Example 10, wherein the second identifier set and the tags corresponding to the second identifier set are homomorphically encrypted using a target key to obtain a second ciphertext, comprising: performing pseudo-random encryption on the second identifier set using the target key to obtain a third sub-ciphertext; for each tag corresponding to the second identifier, performing elliptic curve encryption on the tag and a target point using the target key to obtain a fourth sub-ciphertext of the tag, wherein the target point is any point on a preset elliptic curve; wherein the second ciphertext includes the third sub-ciphertext and the fourth sub-ciphertext of the tag corresponding to each second identifier.
[0271] According to one or more embodiments of this disclosure, Example 14 provides the method of Example 13, wherein the tag is a binary classification tag; the step of performing tag fragmentation operation with the first party to obtain a fourth fragment of the target tag includes: in response to receiving masking data sent by the first party, performing homomorphic decryption on the masking data using the target key, wherein the masking data is obtained by the first party masking the ciphertext of the target tag using a second random number, the second random number being 0 or 1; if the homomorphic decryption result is the target point, then 0 is determined as the fourth fragment; if the homomorphic decryption result is not the target point, then 1 is determined as the fourth fragment.
[0272] According to one or more embodiments of this disclosure, Example 15 provides the method of Example 10, which, prior to the step of sending the second ciphertext to the first party, further includes: shuffling the order of the second ciphertext to obtain a sixth ciphertext; and determining the sixth ciphertext as the second ciphertext.
[0273] According to one or more embodiments of this disclosure, Example 16 provides a data processing method for secure computing, wherein the participants in the secure computing include a first party, a second party, and a third party. The first party holds a first identifier set, where the first identifier in the first identifier set corresponds to an h-dimensional feature. The second party holds a second identifier set, where the second identifier in the second identifier set corresponds to a tag. The third party is a trusted execution environment device, and the third party and the first party are located in the same local area network, where h≥1. The method is applied to the third party and includes: in response to receiving the first identifier set and the target feature corresponding to the first identifier set sent by the first party, homomorphically encrypting the first identifier set and the target feature corresponding to the first identifier set using a target key to obtain a first ciphertext; wherein the target feature includes at least some of the h-dimensional features; sending the first ciphertext to the first party so that the first party generates a first fragment of the first feature and a second fragment of the target tag based on the first ciphertext, wherein the first feature is the target feature corresponding to the intersection of the first identifier set and the second identifier set, and the target tag is the tag corresponding to the intersection.
[0274] According to one or more embodiments of this disclosure, Example 17 provides the method of Example 16, wherein the method of homomorphically encrypting the first identifier set and the target feature corresponding to the first identifier set using a target key to obtain a first ciphertext includes: performing pseudo-random encryption on the first identifier set using the target key to obtain a first sub-ciphertext; generating a first random number and symmetrically encrypting the first random number using the target key to obtain a third ciphertext; and adding the third ciphertext to the target feature corresponding to each first identifier to obtain a feature ciphertext of the target feature; wherein the first ciphertext includes the first sub-ciphertext and a second sub-ciphertext of the target feature corresponding to each first identifier, and the second sub-ciphertext includes the first random number and the feature ciphertext of the target feature corresponding to the first identifier.
[0275] According to one or more embodiments of this disclosure, Example 18 provides the method of Example 16, which, prior to the step of sending the first ciphertext to the first party, further includes: shuffling the order of the first ciphertext to obtain a fifth ciphertext; the step of sending the first ciphertext to the first party so that the first party generates a first fragment of a first feature and a second fragment of a target tag based on the first ciphertext includes: sending the fifth ciphertext to the first party so that the first party generates a first fragment of a first feature and a second fragment of a target tag based on the fifth ciphertext.
[0276] According to one or more embodiments of the present disclosure, Example 19 provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the method described in any one of Examples 1-18.
[0277] According to one or more embodiments of the present disclosure, Example 20 provides a computer program product including a computer program that, when executed by a processor, implements the steps of the method described in any one of Examples 1-18.
[0278] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0279] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0280] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. A data processing method for secure computing, characterized in that, The participants in the secure computation include a first party, a second party, and a third party. The first party holds a first set of identifiers, where each identifier in the first set corresponds to an h-dimensional feature. The second party holds a second set of identifiers, where each identifier in the second set corresponds to a label. The third party is a trusted execution environment device located on the same local area network as the first party, where h ≥ 1. The method is applied to the first party and includes: The first identifier set and the target feature corresponding to the first identifier set are sent to the third party, so that the third party can use the target key to homomorphically encrypt the first identifier set and the target feature corresponding to the first identifier set to obtain the first ciphertext, and send the first ciphertext to the first party, wherein the target feature includes at least some of the h-dimensional features, and wherein the target key is agreed upon in advance by the second party and the third party; In response to receiving the first ciphertext and the second ciphertext sent by the second party, based on the first ciphertext and the second ciphertext, ciphertext of the intersection of the first identifier set and the second identifier set, ciphertext of the first feature, and ciphertext of the target tag are generated, wherein the second ciphertext is obtained by the second party using the target key to homomorphically encrypt the second identifier set and the tag corresponding to the second identifier set, the first feature is the target feature corresponding to the intersection, and the target tag is the tag corresponding to the intersection; Based on the ciphertext of the first feature, a feature fragmentation operation is performed with the second party to obtain a first fragment of the first feature, and based on the ciphertext of the target tag, a tag fragmentation operation is performed with the second party to obtain a second fragment of the target tag; Based on the first and second fragments, execute the target data processing task; The first ciphertext includes a first sub-ciphertext and a second sub-ciphertext corresponding to the target feature of each first identifier. The second sub-ciphertext includes a first random number and the feature ciphertext of the target feature corresponding to the first identifier. The second ciphertext includes a third sub-ciphertext and a fourth sub-ciphertext corresponding to the tag of each second identifier. The step of generating ciphertext of the intersection of the first identifier set and the second identifier set, ciphertext of the first feature, and ciphertext of the target tag based on the first ciphertext and the second ciphertext includes: calculating the intersection of the first sub-ciphertext and the third sub-ciphertext to obtain the ciphertext of the intersection of the first identifier set and the second identifier set; determining the ciphertext corresponding to the ciphertext of the intersection from all the second sub-ciphertexts to obtain the ciphertext of the first feature; and determining the ciphertext corresponding to the ciphertext of the intersection from all the fourth sub-ciphertexts to obtain the ciphertext of the target tag. The step of performing feature fragmentation operation with the second party based on the ciphertext of the first feature to obtain a first fragment of the first feature includes: obtaining the first random number in the ciphertext of the first feature and sending the first random number to the second party, so that the second party can use the target key to perform symmetric encryption on the first random number to obtain a third fragment of the first feature; and determining the feature ciphertext of the first feature as the first fragment. The tag is a binary classification tag; the step of performing tag fragmentation operation with the second party based on the ciphertext of the target tag to obtain a second fragment of the target tag includes: generating a second random number, using the second random number to mask the ciphertext of the target tag to obtain masked data, wherein the second random number is 0 or 1; sending the masked data to the second party, so that the second party can use the target key to perform homomorphic decryption on the masked data, and generate a fourth fragment of the target tag based on the homomorphic decryption result; and using the second random number as the second fragment.
2. The method according to claim 1, characterized in that, The first sub-ciphertext is obtained by the third party using the target key to perform pseudo-random encryption on the first identifier set. The feature ciphertext is obtained by the third party using the target key to perform symmetric encryption on the first random number to obtain the third ciphertext. The third ciphertext is then added to the target feature corresponding to the first identifier.
3. The method according to claim 1, characterized in that, The third sub-ciphertext is obtained by the second party using the target key to perform pseudo-random encryption on the second identifier set, and the fourth sub-ciphertext is obtained by the second party using the target key to perform elliptic curve encryption on the tag and target point corresponding to the second identifier, wherein the target point is any point on a preset elliptic curve.
4. The method according to claim 1, characterized in that, The step of sending the first identifier set and the target feature corresponding to the first identifier set to the third party, so that the third party can use the target key to homomorphically encrypt the first identifier set and the target feature corresponding to the first identifier set to obtain a first ciphertext, and then sending the first ciphertext to the first party, includes: The first identifier set and the target feature corresponding to the first identifier set are sent to the third party, so that the third party can use the target key to homomorphically encrypt the first identifier set and the target feature corresponding to the first identifier set to obtain a first ciphertext, and shuffle the order of the first ciphertext to obtain a fifth ciphertext, and send the fifth ciphertext to the first party. In response to receiving the first ciphertext and the second ciphertext sent by the second party, the method generates, based on the first ciphertext and the second ciphertext, the ciphertext of the intersection of the first identifier set and the second identifier set, the ciphertext of the first feature, and the ciphertext of the target tag, including: In response to receiving the fifth ciphertext and the second ciphertext sent by the second party, the ciphertext of the intersection of the first identifier set and the second identifier set, the ciphertext of the first feature, and the ciphertext of the target tag are generated based on the fifth ciphertext and the second ciphertext.
5. The method according to claim 1, characterized in that, The second ciphertext is obtained by the second party using the target key to homomorphically encrypt the second identifier set and the tags corresponding to the second identifier set, and then shuffling the order of the homomorphic encryption results.
6. The method according to any one of claims 1-5, characterized in that, The target data processing task is either a structured query language query task or a machine learning model training task.
7. A data processing method for secure computing, characterized in that, The participants in the secure computation include a first party, a second party, and a third party. The first party holds a first set of identifiers, where each identifier in the first set corresponds to an h-dimensional feature. The second party holds a second set of identifiers, where each identifier in the second set corresponds to a label. The third party is a trusted execution environment device located on the same local area network as the first party, where h ≥ 1. The method is applied to the second party and includes: The second identifier set and the corresponding tag of the second identifier set are homomorphically encrypted using the target key to obtain the second ciphertext, wherein the target key is agreed upon in advance by the second party and the third party; The second ciphertext is sent to the first party, so that the first party can generate, based on the first ciphertext and the second ciphertext, the ciphertext of the intersection of the first identifier set and the second identifier set, the ciphertext of the first feature, and the ciphertext of the target label. The first ciphertext is obtained by the third party using the target key to homomorphically encrypt the first identifier set and the target feature corresponding to the first identifier set. The target feature includes at least some of the h-dimensional features. The first feature is the target feature corresponding to the intersection, and the target label is the label corresponding to the intersection. Perform a feature segmentation operation with the first party to obtain a third segment of the first feature, and perform a tag segmentation operation with the first party to obtain a fourth segment of the target tag; Based on the third and fourth slices, execute the target data processing task; The first ciphertext includes a first sub-ciphertext and a second sub-ciphertext corresponding to each of the first identifiers and the target features. The second ciphertext includes a first random number and the feature ciphertext of the target features corresponding to the first identifier. Based on the first ciphertext and the second ciphertext, the first party generates the ciphertext of the intersection of the first identifier set and the second identifier set, the ciphertext of the first feature, and the ciphertext of the target tag in the following manner: calculating the intersection of the first sub-ciphertext and the third sub-ciphertext to obtain the ciphertext of the intersection of the first identifier set and the second identifier set, wherein the second ciphertext includes the third sub-ciphertext and the fourth sub-ciphertext of the tag corresponding to each second identifier; determining the ciphertext corresponding to the ciphertext of the intersection from all the second sub-ciphertexts to obtain the ciphertext of the first feature, and determining the ciphertext corresponding to the ciphertext of the intersection from all the fourth sub-ciphertexts to obtain the ciphertext of the target tag; The step of performing feature fragmentation with the first party to obtain a third fragment of the first feature includes: in response to receiving the first random number sent by the first party, symmetrically encrypting the first random number using the target key to obtain a fourth ciphertext; and inverting the fourth ciphertext to obtain a third fragment of the first feature. The tag is a binary classification tag; the tag fragmentation operation performed with the first party to obtain the fourth fragment of the target tag includes: in response to receiving the masking data sent by the first party, homomorphically decrypting the masking data using the target key, wherein the masking data is obtained by the first party masking the ciphertext of the target tag using a second random number, the second random number being 0 or 1; if the homomorphic decryption result is a target point, then 0 is determined as the fourth fragment, wherein the target point is any point on a preset elliptic curve.
8. The method according to claim 7, characterized in that, The first sub-ciphertext is obtained by the third party using the target key to perform pseudo-random encryption on the first identifier set. The feature ciphertext is obtained by the third party using the target key to perform symmetric encryption on the first random number to obtain the third ciphertext. The third ciphertext is then added to the target feature corresponding to the first identifier.
9. The method according to claim 7, characterized in that, The second ciphertext is obtained by homomorphically encrypting the second identifier set and the corresponding tags using the target key, including: The second identifier set is pseudo-randomly encrypted using the target key to obtain the third sub-ciphertext; For each tag corresponding to the second identifier, the tag and the target point are encrypted using elliptic curve cryptography with the target key to obtain the fourth sub-ciphertext of the tag.
10. The method according to claim 7, characterized in that, The label is a binary category label; The step of performing tag fragmentation with the first party to obtain the fourth fragment of the target tag further includes: If the homomorphic decryption result is not the target point, then 1 is determined to be the fourth segment.
11. The method according to claim 7, characterized in that, Prior to the step of sending the second ciphertext to the first party, the method further includes: By shuffling the order of the second ciphertext, the sixth ciphertext is obtained; The sixth ciphertext is identified as the second ciphertext.
12. A data processing method for secure computing, characterized in that, The participants in the secure computation include a first party, a second party, and a third party. The first party holds a first set of identifiers, where each identifier in the first set corresponds to an h-dimensional feature. The second party holds a second set of identifiers, where each identifier in the second set corresponds to a label. The third party is a trusted execution environment device located on the same local area network as the first party, where h ≥ 1. The method is applied to the third party, including: In response to receiving the first identifier set and the target feature corresponding to the first identifier set sent by the first party, the first identifier set and the target feature corresponding to the first identifier set are homomorphically encrypted using the target key to obtain the first ciphertext; wherein, the target feature includes at least some of the h-dimensional features, and the target key is pre-agreed upon by the second party and the third party; The first ciphertext is sent to the first party, so that the first party generates a first fragment with a first feature and a second fragment with a target tag based on the first ciphertext, wherein the first feature is the target feature corresponding to the intersection of the first identifier set and the second identifier set, and the target tag is the tag corresponding to the intersection; The first ciphertext includes a first sub-ciphertext and a second sub-ciphertext corresponding to each of the first identifiers and the target features. The second ciphertext includes a first random number and the feature ciphertext of the target features corresponding to the first identifier. The first party generates the first fragment and the second fragment based on the first ciphertext in the following manner: obtaining the first random number in the ciphertext of the first feature, and sending the first random number to the second party, so that the second party can use the target key to perform symmetric encryption on the first random number to obtain the third fragment of the first feature; determining the feature ciphertext of the first feature as the first fragment; generating a second random number, and using the second random number to mask the ciphertext of the target tag to obtain masked data, wherein the second random number is 0 or 1; sending the masked data to the second party, so that the second party can use the target key to perform homomorphic decryption on the masked data, and generating the fourth fragment of the target tag according to the homomorphic decryption result; and using the second random number as the second fragment.
13. The method according to claim 12, characterized in that, The first ciphertext is obtained by homomorphically encrypting the first identifier set and the target feature corresponding to the first identifier set using the target key, including: The first identifier set is pseudo-randomly encrypted using the target key to obtain the first sub-ciphertext; Generate the first random number, and symmetrically encrypt the first random number using the target key to obtain the third ciphertext; For each target feature corresponding to the first identifier, the third ciphertext is added to the target feature to obtain the feature ciphertext of the target feature.
14. The method according to claim 12, characterized in that, Prior to the step of sending the first ciphertext to the first party, the method further includes: By shuffling the order of the first ciphertext, the fifth ciphertext is obtained; The step of sending the first ciphertext to the first party, so that the first party can generate a first fragment of the first feature and a second fragment of the target tag based on the first ciphertext, includes: The fifth ciphertext is sent to the first party, so that the first party can generate a first fragment of the first feature and a second fragment of the target label based on the fifth ciphertext.
15. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by a processing device, the computer program performs the steps of the method described in any one of claims 1-14.
16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-14.
Citation Information
Patent Citations
Longitudinal federated learning linear regression and logistic regression model training method and device
CN113505894A
Data processing method and device for federal feature engineering, equipment and medium
CN113722744A