Privacy-preserving computation methods, devices, equipment, storage media, and program products

By receiving computation requests and performing lexical expansion and Merkel path localization, target encrypted data is obtained for privacy computation, solving the data leakage problem in multi-party secure computation in existing technologies and improving the security of privacy data.

CN115994371BActive Publication Date: 2025-10-31CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310181921.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-21
Publication Date
2025-10-31
Estimated Expiration
2043-02-21

AI Technical Summary

Technical Problem

In existing multi-party secure computation technologies, a large amount of user privacy data is acquired during multi-party data interaction, but not all of this data is used in privacy computation, leading to privacy data leakage.

Method used

By receiving computation requests, performing lexical expansion, using Merkel paths and identification information to obtain target encrypted data, and performing privacy computations, it ensures that only the necessary privacy data is obtained.

Benefits of technology

It improves the security of privacy data, ensures that only the data needed for privacy computing is obtained, and reduces unnecessary data leaks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115994371B_ABST
    Figure CN115994371B_ABST
Patent Text Reader

Abstract

This application discloses a privacy computation method, apparatus, device, storage medium, and program product. The method includes: receiving a first computation request, the first computation request including a target privacy keyword and privacy computation rules; performing lexical expansion on the target privacy keyword to obtain a target encrypted lexical; determining a target encrypted lexical inverted item associated with the target encrypted lexical from multiple preset encrypted lexical inverted items, the target encrypted lexical inverted item including the Merkle path of the target encrypted lexical and identification information of the target encrypted data; obtaining target encrypted data based on the Merkle path of the target encrypted lexical and the identification information of the target encrypted data; and performing privacy computation according to the privacy computation rules and the target encrypted data to obtain a target computation result. Therefore, based on the target keyword in the user's computation request, the privacy data required for privacy computation can be accurately obtained, improving the data security of the privacy data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data security technology, and in particular relates to a privacy computing method, apparatus, device, storage medium and program product. Background Technology

[0002] Secure multi-party computation is a core area of ​​cryptographic research. It solves the problem of collaborative computation with privacy protection among a group of distrustful participants. It can provide data requesters with multi-party collaborative computation capabilities without disclosing the original data, and provide them with an overall data profile after computation by all parties. Therefore, it can complete the analysis, processing and result publication of data without the data leaving the data holding node, and provide data access control and consistency guarantee for data exchange.

[0003] Existing secure multi-party computation technologies require the acquisition of a large amount of user privacy data during multi-party data interaction. However, not all of this acquired user privacy data is used in privacy computation, which may result in the acquisition of privacy data that is not needed for multi-party computation, leading to the leakage of privacy data. Summary of the Invention

[0004] This application provides a privacy computing method, apparatus, device, storage medium, and program product that can improve the security of privacy data.

[0005] In a first aspect, embodiments of this application provide a privacy computing method, the method comprising:

[0006] Receive a first computation request, the first computation request including target privacy keywords and privacy computation rules;

[0007] The target privacy keywords are subjected to lexical expansion to obtain target encrypted lexical units;

[0008] Based on the target encrypted token, a target encrypted token inverted item associated with the target encrypted token is determined from multiple preset encrypted token inverted items. The target encrypted token inverted item includes the Merkel path of the target encrypted token and the identification information of the target encrypted data.

[0009] The target encrypted data is obtained based on the Merkel path of the target encrypted token and the identification information of the target encrypted data;

[0010] Privacy calculations are performed based on the privacy calculation rules and the target encrypted data to obtain the target calculation result.

[0011] Secondly, embodiments of this application provide a privacy computing device, the device comprising:

[0012] The receiving module is configured to receive a first computation request, wherein the first computation request includes a target privacy keyword and a privacy computation rule;

[0013] The lexical expansion module is used to perform lexical expansion processing on the target privacy keyword to obtain the target encrypted lexical;

[0014] The determining module is used to determine, based on the target encrypted token, a target encrypted token inverted item associated with the target encrypted token from a plurality of preset encrypted token inverted items, wherein the target encrypted token inverted item includes the Merkel path of the target encrypted token and the identification information of the target encrypted data;

[0015] The acquisition module is used to acquire the target encrypted data based on the Merkel path of the target encrypted token and the identification information of the target encrypted data;

[0016] The calculation module is used to perform privacy calculations based on the privacy calculation rules and the target raw data to obtain the target calculation result. The target raw data is obtained by decrypting the target encrypted data.

[0017] Thirdly, embodiments of this application provide a privacy computing device, the device including: a processor and a memory storing computer program instructions;

[0018] The processor implements the above privacy computing method when executing computer program instructions.

[0019] Fourthly, embodiments of this application provide a computer storage medium storing computer program instructions, which, when executed by a processor, implement the privacy computing method described above.

[0020] Fifthly, embodiments of this application provide a computer program product, the computer program product including computer program instructions, which, when executed by a processor, implement the privacy computing method described above.

[0021] The privacy computation method provided in this application receives a first computation request, which includes a target privacy keyword and privacy computation rules; it performs lexical expansion on the target privacy keyword to obtain a target encrypted lexical; based on the target encrypted lexical, it determines a target encrypted lexical inverted item associated with the target encrypted lexical from multiple preset encrypted lexical inverted items, the target encrypted lexical inverted item including the Merkle path of the target encrypted lexical and the identification information of the target encrypted data; based on the Merkle path of the target encrypted lexical and the identification information of the target encrypted data, it obtains the target encrypted data, decrypts the target data to obtain the target original data; and performs privacy computation according to the privacy computation rules and the target original data to obtain the target computation result. Therefore, by obtaining the target encrypted lexical based on the target keyword in the user's computation request and obtaining the target original data based on the encrypted lexical inverted item of the target encrypted lexical, the privacy data required for privacy computation can be accurately obtained, improving the security of privacy data. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of a privacy computing process provided in an embodiment of this application;

[0024] Figure 2 This is a schematic diagram of a computational binary tree provided in another embodiment of this application;

[0025] Figure 3 This is a flowchart illustrating word expansion provided in another embodiment of this application;

[0026] Figure 4 This is a schematic diagram of a privacy computing process provided in another embodiment of this application;

[0027] Figure 5 This is a schematic diagram of the structure of a privacy computing device provided in an embodiment of this application;

[0028] Figure 6 This is a schematic diagram of the structure of a privacy computing device provided in an embodiment of this application. Detailed Implementation

[0029] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples of this application.

[0030] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0031] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.

[0032] The internet has transitioned from the IT era to the DT (Data Technology) era, and data has become a core competitive advantage for businesses in this era. Data, as a new energy source, only generates value when it flows. However, most companies are very cautious about data sharing due to concerns about data security and personal privacy.

[0033] The basic concept of privacy computing is a technology and system in which two or more parties collaborate to perform joint machine learning and joint analysis on their data without disclosing their own data, and the computation results can be verified.

[0034] A Merkle tree, also known as a hash binary tree, is a data structure used to efficiently summarize and verify the integrity of large datasets. It can be used to prove the existence of a specific transaction among thousands of transactions in a very large block. To prove that a block contains a specific transaction, a node only needs to calculate a hash value of log2(N)32 bytes, forming an authentication path or Merkle path from the specific transaction to the root of the tree.

[0035] Privacy-preserving computation, based on cryptographic technologies, includes fully homomorphic encryption (FHE), multi-party secure computation (MPC), zero-knowledge proofs, and federated learning (FL). Hardware-based solutions primarily consist of trusted execution environments (TEEs). (Blockchain will become a key technology for privacy-preserving computation, enabling secure, compliant, and reasonable use of data while ensuring data trustworthiness.)

[0036] Privacy-preserving computation is currently mainly divided into three categories: 1) Secure multi-party computation: privacy-preserving computation technology based on cryptography. 2) Federated learning: a technology derived from the integration of artificial intelligence and privacy protection technologies. 3) Trusted execution technology (TEE): representing privacy-preserving computation technology based on trusted hardware.

[0037] Secure Multi-Party Computation (MPC, also known as SMC or SMPC) addresses the privacy-preserving collaborative computation problem among a group of mutually distrustful participants. SMC must ensure input independence, computational correctness, and decentralization, while preventing the leakage of input values ​​to other participants. For example, when an MPC computation task is initiated, the hub node transmits network and signaling control. Each data holder can initiate a collaborative computation task. The hub node performs routing and addresses, selecting other data holders with similar data types for secure collaborative computation. The MPC nodes of the participating data holders, according to the computation logic, query the required data from their local databases and collaboratively perform computation on the MPC task across data streams. While ensuring input privacy, each party receives correct data feedback, and local data is not leaked to any other participant throughout the process.

[0038] However, existing secure multi-party computation (SMC) emphasizes data protection and computation synchronization, but lacks detailed regulations on the consistency of data among multiple parties. Furthermore, it lacks the backing of blockchain technology for data traceability and computation verification, resulting in insufficient transparency. Moreover, existing MSC technologies require the acquisition of large amounts of user privacy data during multi-party data interaction; however, not all of this acquired user privacy data is utilized in privacy-preserving computation, potentially leading to the acquisition of unnecessary privacy data and resulting in privacy leaks.

[0039] To address the problems of existing technologies, embodiments of this application provide a privacy computing method, apparatus, device, storage medium, and program product. The privacy computing method provided in this application embodiment will be described first below.

[0040] Figure 1 A flowchart illustrating a privacy computing method according to an embodiment of this application is shown. The method includes the following steps S101 to S105:

[0041] S101, Receive the first calculation request.

[0042] The first computation request includes the target privacy keywords and privacy computation rules. The first service request can be a computation request input by the user.

[0043] Privacy keywords are the privacy keywords corresponding to the target data required for the calculation, that is, keywords related to user privacy in the target data, such as keywords like infection, contagion, massive bleeding, amniotic fluid embolism, and cerebral palsy in medical data.

[0044] Privacy-preserving computation rules are the specific computational rules that need to be applied to privacy-preserving data. For example, calculating the average of data A and data B.

[0045] Specifically, after receiving a computation request from a user, the privacy keywords and privacy computation rules in the computation request can be determined by parsing the computation request.

[0046] For example, a privacy computation requirement is to filter cerebral palsy data from users in the Beijing area and calculate the average age of the patients. In this case, "Beijing" and "cerebral palsy" are privacy keywords, and calculating the "average age" of "cerebral palsy patients in the Beijing area" is the privacy computation rule.

[0047] S102, perform lexical expansion on the target privacy keywords to obtain the target encrypted lexical.

[0048] Specifically, since target privacy keywords are often very short strings, in order to facilitate subsequent searches for target encrypted data based on target keywords and improve the accuracy of data acquisition,

[0049] After determining the target privacy keywords, the target privacy keywords need to be subjected to word expansion processing to obtain target encrypted words.

[0050] The term expansion process here can be used to augment the data of the target privacy keyword, increasing its size. For example, additional characters can be added after the characters of the target privacy keyword; these characters could be invalid characters or other identifiers.

[0051] S103, Based on the target encrypted word, determine the target encrypted word inverted item associated with the target encrypted word from multiple preset encrypted word inverted items.

[0052] The inverted items of the target encrypted token include the Merkel path of the target encrypted token and the identification information of the target encrypted data.

[0053] Here, the Merkle path of the target encrypted token is used to locate the position of the target encrypted token in the target encrypted data, and the identification information of the target encrypted data is used to find the target encrypted data.

[0054] Specifically, based on the relationship between the pre-set encrypted tokens and the pre-set inverted encrypted tokens, the pre-set inverted encrypted tokens associated with the target encrypted token can be found from multiple pre-set inverted encrypted tokens, i.e., the target encrypted token inverted token.

[0055] S104. Obtain the target encrypted data based on the Merkel path of the target encrypted token and the identification information of the target encrypted data.

[0056] Specifically, the Merkle path of the target encrypted token can determine its position within the target encrypted data. The identifier information of the target encrypted data allows it to be retrieved from the database. After retrieving the target encrypted data from the database, its existence within the target encrypted data can be determined using the Merkle path, thereby verifying the accuracy of the target encrypted data.

[0057] S105, perform privacy calculations based on privacy calculation rules and the target raw data to obtain the target calculation result.

[0058] The original target data is obtained by decrypting the encrypted target data.

[0059] Specifically, the target encrypted data is decrypted to obtain the target original data, and privacy calculations are performed on the target original data according to privacy calculation rules.

[0060] The privacy computation method provided in this application receives a first computation request, which includes a target privacy keyword and privacy computation rules; it performs lexical expansion on the target privacy keyword to obtain a target encrypted lexical; based on the target encrypted lexical, it determines a target encrypted lexical inverted item associated with the target encrypted lexical from multiple preset encrypted lexical inverted items, the target encrypted lexical inverted item including the Merkle path of the target encrypted lexical and the identification information of the target encrypted data; based on the Merkle path of the target encrypted lexical and the identification information of the target encrypted data, it obtains the target encrypted data, decrypts the target data to obtain the target original data; and performs privacy computation according to the privacy computation rules and the target original data to obtain the target computation result. Therefore, by obtaining the target encrypted lexical based on the target keyword in the user's computation request and obtaining the target original data based on the encrypted lexical inverted item of the target encrypted lexical, the privacy data required for privacy computation can be accurately obtained, improving the security of privacy data.

[0061] In some embodiments, after S101, the following steps may also be included:

[0062] Construct a computational binary tree based on the first computation request;

[0063] Based on the computational binary tree, the target privacy keywords and privacy computation rules are determined.

[0064] Specifically, privacy computation requests are converted into privacy computation expressions by computational binary trees. Computational binary trees are then constructed based on these expressions. Specifically, according to the computational expressions, from left to right, non-terminal nodes are used as operands (addition, subtraction, multiplication, division, etc.), and the terminal nodes of the binary trees are the original data participating in the privacy computation (i.e., the name of the target original data).

[0065] In one example, such as Figure 2 The example shown is a computational binary tree of meta-matrixes. A privacy computation request aims to filter cerebral palsy data from users in the Beijing area and calculate the average age of patients. "Beijing" and "cerebral palsy" are privacy keywords, and calculating the "average age" of "cerebral palsy patients in Beijing" is the privacy computation rule. The resulting privacy computation expression, derived from a computational binary tree of meta-matrixes, is shown below:

[0066] if ((data.find(“cerebral palsy”))and(data.user.find(Beijing)))

[0067] then result=sum(data.useruser.age()) / user.Num

[0068] Among them, result is the average value of patients with cerebral palsy.

[0069] In some embodiments, the step of performing lexical expansion on the target privacy keyword to obtain the target encrypted lexical in S102 above may include the following steps:

[0070] Data is populated based on the target privacy keywords to obtain target encrypted keywords. The data size of the target encrypted keywords is a preset value.

[0071] Specifically, because the target keyword data is inconsistent in size and relatively small, making it difficult to determine the target encrypted data, the target privacy keywords undergo word expansion processing to obtain target encrypted words with a preset data size. For example, the data size of the target keyword data can be increased by adding characters to the end of the target keyword data.

[0072] In some embodiments, lexical expansion processing may include the following steps:

[0073] The file names and dates of the original files corresponding to the target privacy keywords are obtained and arranged in order to obtain the augmented data. Multiple augmented data are added to the end of the target privacy keyword data. Data exceeding the preset value is truncated to obtain the target encrypted tokens of the preset value size.

[0074] In this embodiment of the application, the target keyword can be expanded according to a preset word expansion rule to obtain a target encrypted word with a preset data size.

[0075] In some embodiments, before determining the target encrypted token inverted item associated with the target encrypted token from a plurality of preset encrypted token inverted items based on the target encrypted token, the following steps may be included:

[0076] The preset privacy keywords are subjected to word expansion processing to obtain preset encrypted words. The data size of the preset encrypted words is a preset value. The preset privacy keywords are determined according to the data type of the original data.

[0077] The original data is divided into multiple sub-data, and the size of each sub-data is a preset value;

[0078] A Merkle tree is constructed based on preset encrypted terms and multiple encrypted sub-data. The multiple encrypted sub-data are obtained by encrypting each sub-data.

[0079] Preset encrypted token inverted items are generated based on the Merkel path of the preset encrypted token and the identification information of the preset encrypted data. The preset encrypted token inverted items are associated with the preset encrypted token. The Merkel path is determined according to the Merkel tree. The preset encrypted data includes the preset encrypted token and multiple encrypted sub-data.

[0080] Among them, determining the preset privacy keywords based on the data type of the original data can be achieved by constructing a keyword library for different data types, and then determining the preset privacy keywords from the keyword library for different data types.

[0081] The preset encrypted data is the encrypted data obtained by encrypting the original data corresponding to preset privacy keywords. This can be an encrypted file generated based on preset encrypted tokens and multiple encrypted sub-data.

[0082] Encrypted term inverted items are used to find and locate encrypted data associated with privacy keywords.

[0083] Specifically, after determining the preset privacy keywords of the original data, the preset privacy keywords are then subjected to lexical expansion processing to obtain preset privacy encrypted lexical units (the lexical expansion processing of the preset privacy keywords here requires the same processing method and lexical expansion rules as the lexical expansion processing of the target privacy keywords mentioned above, and the size of the encrypted lexical unit data is consistent); the original data is segmented, that is, the original data is divided into multiple sub-data of preset size, and the sub-data is then encrypted to obtain multiple encrypted sub-data; based on the preset encrypted lexical units of the same size and the multiple encrypted sub-data, a Merkle tree is constructed, from which the Merkle path of the preset encrypted lexical unit (i.e., its position in the preset encrypted data) can be obtained from the Merkle tree; finally, the preset encrypted lexical unit inverted index is generated according to the Merkle path of the preset encrypted lexical unit and the identification information of the preset encrypted data.

[0084] In this embodiment, the association between encrypted words and encrypted data can be pre-constructed based on encrypted words and original data, the encrypted data can be stored, and the encrypted data can be found through the inverted items of the encrypted words associated with the encrypted words. The encrypted data can then be obtained for subsequent privacy calculations, thereby improving data security.

[0085] In one example, building a library of primitive data type keywords involves the following steps:

[0086] (1) Divide the user's raw data into three types: traditional enterprise data, machine and sensor data, and social data. Traditional enterprise data includes consumer data, traditional ERP data, inventory data, accounting data, medical data, education data, and other data types. Machine and sensor data includes call logs, smart meters, IoT data, industrial equipment sensors, equipment logs, transaction data, and other data types. Social data includes user image data, diary data, chat data, and other data types. Construct a keyword library for raw data types based on the different data types. (2) Using the keyword extraction module, extract privacy keywords related to the raw data types from the raw data based on the keyword library for raw data types and write them into privacy keywords. (For example: privacy keywords for medical raw data include infection, contagion, massive hemorrhage, amniotic fluid embolism, cerebral palsy, etc. Privacy keywords for IoT raw data include sensor, RFID, radio frequency, Zigbee, Bluetooth, etc. Image raw data includes photo time, photo location, owner name, etc.)

[0087] In one example, constructing an inverted index of cryptographic terms may include the following steps:

[0088] During privacy-preserving computations, it's necessary to read users' private data. Since this data is sealed within encrypted sectors, it's essential to locate and retrieve it. The encrypted lexicon inverted index module performs this function. This module constructs an encrypted lexicon inverted index database, containing fields such as: encrypted lexicon, encrypted lexicon ID, encrypted lexicon privacy permissions, encrypted lexicon inverted index document number, and encrypted lexicon occurrence position.

[0089] Encryption tokens: 2k encryption tokens.

[0090] Encryption token ID: The serial number of the encryption token.

[0091] Cryptograph privacy permissions: Which computations the data owner allows cryptographs to participate in (cryptograph attribute, which can be set by the user).

[0092] Encryption token inverted item document number: Content identifier (CID) of the encrypted file (32G in size) containing the encryption token (i.e., the identification information of the preset encrypted data).

[0093] Location of encrypted token: The Merkle number position of the encrypted token in the sealed sector (i.e., the Merkle path of the preset encrypted token).

[0094] During privacy-preserving computations, the system locates and positions encrypted tokens using an inverted index database. Based on privacy permissions, it decrypts the 32GB encrypted data file using the user's private key and then performs the computation. The 32GB encrypted data file is generated by dividing the original data into 2KB fragments, encrypting each 2KB fragment using the user's private key, generating a series of encrypted data fragments, and then combining these fragments with the encrypted tokens to form a 32GB encrypted data file.

[0095] Encrypted encapsulated sectors with lexical inflation padding, and an encrypted lexical inverted index database for privacy data search.

[0096] In this embodiment, attributes such as privacy permissions and calculation conditions can be added to the encrypted tokens, thereby narrowing the scope of privacy data acquisition in subsequent privacy calculations and further improving the security of privacy data.

[0097] In some embodiments, the construction of the Merkle tree based on preset encrypted terms and multiple encrypted sub-data may further include:

[0098] Calculate the first hash value of the preset encrypted token and the second hash value of each encrypted sub-data;

[0099] A Merkle tree is constructed based on the first hash value and each of the second hash values, wherein the Merkle path of the preset encrypted token is the path from the first hash value in the Merkle tree to the Merkle root of the Merkle tree.

[0100] In one example, such as Figure 3 As shown: After performing lexical expansion on the original data, the output is encrypted data suitable for privacy computation containing encrypted lexical terms (i.e., the aforementioned target encrypted data or preset encrypted data). The specific steps are as follows:

[0101] Step 301: Obtain the privacy keyword A (several tens of bytes). After processing with token inflation padding technology, generate a 2KB encrypted token A. Specifically, perform a SHA256 hash calculation on the privacy keyword A to obtain 256 bits of data a. Arrange the original data's filename and date in order and perform a SHA256 hash calculation to obtain 256 bits of data b. Append several bits of data b to the end of data a. Data exceeding 2KB is truncated to form 2KB data, which is called the encrypted token A.

[0102] Step 302: Divide the original data into 2k data fragments, encrypt each 2k data fragment using the user's private key to generate a series of encrypted data fragments, and write all the encrypted data fragments and encrypted tokens into an encrypted data file B. The size of an encrypted data file B is 32G.

[0103] Step 303: Construct a Merkle tree from all the encrypted data fragments and encrypted tokens.

[0104] Step 304: Then, construct the encrypted term inverted index, which is used to locate the position of the privacy keyword within the encrypted term. The encrypted term inverted index contains the Merkle path of encrypted term A and the content identifier (CID) of encrypted data file B. The content identifier (CID) of encrypted data file B represents the storage location of encrypted data file B.

[0105] Step 305: Write the encrypted term inverted index into the blockchain to facilitate subsequent traceability and verification of each step of privacy computing, thereby improving data security.

[0106] In some embodiments, after performing privacy calculations based on privacy calculation rules and target encrypted data to obtain the target calculation result, S104 may further include:

[0107] Record the actions of acquiring target encrypted data and the target calculation results on the blockchain.

[0108] Specifically, the data acquisition and query behaviors and calculation results in each step of the above privacy computation are recorded on the blockchain. This allows for the traceability and verification of data and calculation results in each step during privacy computation and secure computation by the other party, thereby improving the security of privacy data.

[0109] To facilitate understanding, the above privacy computation methods will be explained below with specific scenarios. For example... Figure 4 As shown: Based on the above privacy computing method, this application embodiment provides a privacy computing solution system based on lexical expansion and padding encapsulation technology. The system includes a privacy data preset module and a privacy data computing module.

[0110] The privacy computing preset module completes the construction of the encrypted lexical inverted index. The specific steps are as follows:

[0111] S401. Determine the keywords of the original data, privacy calculation conditions, and data owner signature authorization.

[0112] (1) Using the keyword extraction module, based on the keyword library for the original data type, extract privacy keywords related to the original data type and write them into privacy keywords. (For example: privacy keywords for medical data include infection, contagion, massive hemorrhage, amniotic fluid embolism, cerebral palsy, etc. Privacy keywords for IoT data include sensor, RFID, radio frequency, Zigbee, Bluetooth, etc. Privacy keywords for image data include photo time, photo location, owner name, etc.)

[0113] Construction of the Raw Data Type Keyword Library: User raw data is categorized into three types: traditional enterprise data, machine and sensor data, and social data. Traditional enterprise data includes consumer data, traditional ERP data, inventory data, accounting data, medical data, and educational data. Machine and sensor data includes call logs, smart meters, IoT data, industrial equipment sensors, equipment logs, and transaction data. Social data includes user image data, diary data, and chat data. A raw data type keyword library is then constructed based on these different data types.

[0114] (2) Write the privacy computation conditions into the privacy terminology.

[0115] Privacy-preserving computation conditions are categorized into three types: fully permitted, conditionally permitted, and not permitted. (For example: User A's image data is searchable by users in Beijing. User B's medical data is not searchable by anyone. User C's blood pressure record data is searchable and computed by anyone.)

[0116] (3) The owner of the original data signs and authorizes the privacy term metadata.

[0117] In S401, a keyword library for raw data types is constructed based on different data types. Privacy terms are then built, and encrypted terms are constructed based on these privacy terms and encrypted data. After pre-setting, encrypted data is given encrypted term attributes (i.e., the storage location of the encrypted terms and the privacy keywords of the encrypted terms) and privacy authorization attributes (allow, disallow, conditionally allowed), which can be used by the privacy calculation module to participate in privacy calculations. For example, (data of user A containing the privacy keyword "cerebral palsy," stored in the erdddddddddd2333 directory, allows users with a geographical location in Beijing to perform search-based privacy calculations.)

[0118] S402. Generate privacy lexical units based on the original data keywords, privacy calculation conditions, and data owner signature authorization.

[0119] S403. Perform lexical expansion on privacy lexical units and construct an inverted index of encrypted lexical units.

[0120] (1) Convert privacy terms into encrypted terms through the term expansion and filling module. That is, expand the original data keywords (privacy keywords) to obtain encrypted terms (with the same attributes as privacy terms, including privacy calculation conditions and data owner signature authorization).

[0121] (2) Construct an inverted index of encrypted tokens based on the storage locations of the encrypted tokens and the original data. The inverted index of encrypted tokens contains the storage location of the encrypted document and the Merkle path of the encrypted tokens. The location of the encrypted tokens can be located through the inverted index of encrypted tokens.

[0122] S404. Store the encrypted data of the original data in the blockchain.

[0123] After performing token inflation on the original data, the output is encrypted data suitable for privacy computation containing encrypted tokens. The privacy keyword A (tens of bytes) is processed using token inflation and padding techniques to generate an encrypted token A of size 2KB. The original data is divided into 2KB data fragments, and each 2KB fragment is encrypted using the user's private key, generating a series of encrypted data fragments. All encrypted data fragments and encrypted tokens are written into an encrypted data file B, which is 32GB in size. A Merkle tree is constructed from all the encrypted data fragments and encrypted tokens, and then an inverted index of encrypted tokens is built. The inverted index is used to locate the privacy keyword within the encrypted tokens. The inverted index contains the Merkle path of encrypted token A and the content identifier (CID) of encrypted data file B. The content identifier (CID) of encrypted data file B represents the storage location of encrypted data file B.

[0124] In summary, the privacy data pre-setting module involves the data owner providing the original data, which is then used by the system to generate privacy tokens (generated by the keyword extraction module, containing privacy keywords, privacy computation conditions, the data owner's signature, and authorization). After the data owner authorizes the privacy token signature (private key), the privacy data pre-setting module collects information submitted by the original data owner, constructs encrypted tokens and inverted encrypted token terms from the privacy tokens, and establishes an encrypted data index. The inverted encrypted token terms are used to locate the positions of the encrypted data tokens for privacy computation. The privacy data pre-setting module is responsible for processing the privacy data, performing token expansion on the original data, and outputting encrypted data with encrypted tokens suitable for privacy computation.

[0125] The privacy data calculation module, after the user inputs the calculation requirements, searches for suitable encrypted keywords and performs privacy data calculations. The specific steps are as follows:

[0126] S405, User Revenue Calculation Requirements.

[0127] S406. The computational meta-language parser parses the user's computation request into a computational meta-language binary tree expression. The computational meta-language parser is responsible for constructing the privacy-preserving computational meta-language binary tree. When the user (the privacy-preserving computation delegate) inputs the privacy-preserving computation request, the system constructs a computational meta-language binary tree and parses it within the privacy-preserving computation sandbox. According to the computational expression, from left to right, non-terminal nodes are used as operands (addition, subtraction, multiplication, division, etc.), and the terminal nodes of the binary tree are the original data participating in the privacy-preserving computation.

[0128] S407. In the privacy computing sandbox, based on the privacy terms and privacy computing conditions preset by the data owner, encrypted data is extracted from the relevant encrypted data sectors according to the Merkel path of the privacy terms, and restored to the original data.

[0129] S408. Based on the raw data and following the computational rules in the privacy-preserving computation binary tree, the computation conditions, requirements, and rules are sent to the collaborating computer. The collaborating computer distributes the computation requirements to the computing machines, which perform the computation according to the computation interface, return the results to the collaborating computer, and store them in the blockchain. Computing machines requiring collaborative computation read the results from the blockchain and perform collaborative computation. After the privacy-preserving computation collaboration is completed, the hash values ​​of the computation process and results are uploaded to the blockchain, signifying the completion of the computation.

[0130] Specifically, in steps S407 to S408 above, the original stored data undergoes zero-knowledge proof computation, is encrypted, and the zero-knowledge proof is written to the blockchain. This hides the original stored data itself while establishing the relationships between the original stored data. After the encrypted stored data undergoes slicing, encryption, and encapsulation, a sealed encrypted data and a spacetime proof file are calculated. The spacetime proof file, after compression and extraction, is written into a block on the blockchain. The encrypted stored data is sliced ​​into 2K-sized slices, packed into data boxes, and these data boxes, along with other information, are recorded. A dynamic hash list (DHT) is generated using a Merkle tree method. The data boxes and the hash list DHT are then labeled, and the labeled data boxes and the hash list DHT are used in a proof circuit for encryption computation to generate a sealed data file and a spacetime proof file. After the spacetime proof file is compressed and extracted, a 5K-sized data (proof) is generated and submitted to the blockchain for accounting. The blockchain-based blockchain archive storage encryption device randomly selects a block on the blockchain every second to verify whether the proof-of-space time and the sealed data are consistent. If they are consistent, it means that the sealed data has been effectively stored.

[0131] The privacy computation sandbox extracts encrypted data from relevant encrypted data sectors according to the Merkel path of privacy terms, based on the privacy terms and privacy computation conditions preset by the data owner, and restores the original data. When searching for encrypted data by privacy terms, it first searches for encrypted data through the encrypted data pool based on the encrypted keywords. After finding the encrypted data, it finds the encrypted data owner, notifies the owner, and after the owner agrees to decrypt the data, it decrypts the data according to the encrypted storage blockchain and performs privacy computation.

[0132] S409. Write the reading and privacy calculation process of the encrypted sealed sector into the blockchain.

[0133] S410. Write the hash value of the privacy computation result into the blockchain.

[0134] S411. Return the privacy calculation results to the user.

[0135] In summary, after the user (privacy computation entruster) inputs privacy computation requirements in the privacy data computation module, the system constructs a computational meta-tree for the privacy computation requirements. Within the privacy computation sandbox, the system parses the computational meta-tree and, based on the privacy terms and computation conditions preset by the data owner, extracts encrypted data from the relevant encrypted data sectors according to the Merkel path of the privacy terms, restoring the original data. Collaborative computation is then performed on the system's computing power. The reading of the encrypted sealed sectors, the privacy computation process, and the hash value of the privacy computation result are written to the blockchain, and the privacy computation result is returned to the user.

[0136] The introduction of a privacy-preserving computation solution system based on lexical inflation and padding encapsulation technology in this embodiment can verify mapping relationships, calculate these relationships, and verify them. Protection based on encrypted relationships enhances data privacy. By coupling encrypted relationships with time-dependent operations, algorithms with constraints can be used to prove that a storage system indeed stores a copy of a certain data, and that each copy uses different physical storage, thus protecting against three common attacks in decentralized systems. The privacy-preserving computation solution system based on lexical inflation and padding encapsulation technology enables computation and solution of privacy-preserving data while protecting privacy.

[0137] Based on the privacy computing method provided in the above embodiments, this application also provides specific implementations of a privacy computing device. Please refer to the following embodiments.

[0138] First see Figure 5 The privacy computing device 500 provided in this application embodiment includes the following modules:

[0139] The receiving module 501 is used to receive a first calculation request, which includes a target privacy keyword and a privacy calculation rule;

[0140] The lexical expansion module 502 is used to perform lexical expansion processing on the target privacy keyword to obtain the target encrypted lexical;

[0141] The determination module 503 is used to determine the target encrypted word inverted item associated with the target encrypted word from multiple preset encrypted word inverted items based on the target encrypted word. The target encrypted word inverted item includes the Merkel path of the target encrypted word and the identification information of the target encrypted data.

[0142] The acquisition module 504 is used to acquire the target encrypted data based on the Merkel path of the target encrypted token and the identification information of the target encrypted data;

[0143] The calculation module 505 is used to perform privacy calculations based on privacy calculation rules and the target raw data to obtain the target calculation result. The target raw data is obtained by decrypting the target encrypted data.

[0144] In some embodiments, the word expansion module 502 described above can be used to: fill data for target privacy keywords to obtain target encrypted words, wherein the data size of the target encrypted words is a preset value.

[0145] In some embodiments, the determining module 503 may include:

[0146] The word expansion processing unit is used to perform word expansion processing on preset privacy keywords to obtain preset encrypted words. The data size of the preset encrypted words is a preset value. The preset privacy keywords are determined according to the data type of the original data.

[0147] The segmentation unit is used to divide the original data into multiple sub-data, and the data size of each sub-data is a preset value;

[0148] The processing unit is used to construct a Merkle tree based on preset encrypted tokens and multiple encrypted sub-data. The multiple encrypted sub-data are obtained by performing data encryption processing on each sub-data.

[0149] The generation unit is used to generate an inverted list of preset encrypted words based on the Merkel path of the preset encrypted words and the identification information of the preset encrypted data. The inverted list of preset encrypted words is associated with the preset encrypted words. The Merkel path is determined according to the Merkel tree. The preset encrypted data includes preset encrypted words and multiple encrypted sub-data.

[0150] In some embodiments, the above processing unit is specifically used to: calculate the first hash value of the preset encrypted token and the second hash value of each encrypted sub-data respectively;

[0151] A Merkle tree is constructed based on the first hash value and each of the second hash values, wherein the Merkle path of the preset encrypted token is the path from the first hash value in the Merkle tree to the Merkle root of the Merkle tree.

[0152] In some embodiments, the device 500 may further include:

[0153] Create a module to construct a computational primitive binary tree based on the first computation request;

[0154] The analysis module is used to determine the target privacy keywords and privacy computation rules based on the computational meta-word binary tree.

[0155] In some embodiments, the device 500 may further include:

[0156] The recording module is used to record the actions of acquiring target encrypted data and the target computation results on the blockchain.

[0157] Figure 6 A schematic diagram of the hardware structure of the privacy computing device provided in an embodiment of this application is shown.

[0158] A privacy computing device may include a processor 601 and a memory 602 storing computer program instructions.

[0159] Specifically, the processor 601 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0160] Memory 602 may include mass storage for data or instructions. For example, and not limitingly, memory 602 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 602 may include removable or non-removable (or fixed) media. Where appropriate, memory 5602 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 502 is non-volatile solid-state memory.

[0161] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the methods according to one aspect of this disclosure.

[0162] The processor 601 implements any of the privacy computing methods described in the above embodiments by reading and executing computer program instructions stored in the memory 602.

[0163] In one example, the privacy computing device may also include a communication interface 603 and a bus 610. Wherein, as Figure 6 As shown, the processor 601, memory 602, and communication interface 603 are connected through bus 610 and complete communication with each other.

[0164] The communication interface 603 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0165] Bus 610 includes hardware, software, or both, that couples components of a privacy computing device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 610 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.

[0166] This privacy computing device can be based on the above embodiments to achieve a combination of... Figures 1 to 5 The privacy-preserving computation method and apparatus described.

[0167] Furthermore, in conjunction with the privacy computing methods in the above embodiments, this application embodiment can provide a computer storage medium for implementation. This computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the privacy computing methods in the above embodiments and achieve the same technical effect. To avoid repetition, further details are omitted here. The aforementioned computer-readable storage medium may include non-transitory computer-readable storage media, such as read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, etc., and is not limited thereto.

[0168] In addition, this application also provides a computer program product, including computer program instructions, which, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.

[0169] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0170] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0171] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0172] The foregoing flowcharts and / or block diagrams of methods, apparatus, and computer program products according to embodiments of the present disclosure have described various aspects of the present disclosure. It should be understood that each block in the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable privacy computing device to create a machine such that these instructions, executable via the processor of the computer or other programmable privacy computing device, enable the implementation of the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0173] The above are merely specific embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. A privacy-preserving computation method, characterized in that, The method includes: Receive a first computation request, the first computation request including target privacy keywords and privacy computation rules; The target privacy keywords are subjected to lexical expansion to obtain target encrypted lexical units; Based on the target encrypted token, a target encrypted token inverted item associated with the target encrypted token is determined from multiple preset encrypted token inverted items. The target encrypted token inverted item includes the Merkel path of the target encrypted token and the identification information of the target encrypted data. The target encrypted data is obtained based on the Merkel path of the target encrypted token and the identification information of the target encrypted data; Privacy calculations are performed based on the privacy calculation rules and the target raw data to obtain the target calculation result. The target raw data is obtained by decrypting the target encrypted data. The step of performing lexical expansion on the target privacy keyword to obtain the target encrypted lexical includes: The data for the target privacy keyword is filled to obtain the target encrypted token, and the data size of the target encrypted token is a preset value.

2. The method according to claim 1, characterized in that, Before determining the target encrypted word inverted item associated with the target encrypted word from a plurality of preset encrypted word inverted items based on the target encrypted word, the method further includes: Lexical expansion is performed on preset privacy keywords to obtain preset encrypted lexical units. The data size of the preset encrypted lexical units is a preset value. The preset privacy keywords are determined according to the data type of the original data. The original data is divided into multiple sub-data, and the data size of each sub-data is a preset value; A Merkle tree is constructed based on the preset encrypted tokens and multiple encrypted sub-data, wherein the multiple encrypted sub-data are obtained by performing data encryption processing on each of the sub-data. A preset encrypted term inverted list is generated based on the Merkel path of the preset encrypted term and the identification information of the preset encrypted data. The preset encrypted term inverted list is associated with the preset encrypted term. The Merkel path is determined according to the Merkel tree. The preset encrypted data includes the preset encrypted term and the plurality of encrypted sub-data.

3. The method according to claim 1, characterized in that, The step of constructing a Merkle tree based on the preset encrypted tokens and multiple encrypted sub-data includes: Calculate the first hash value of the preset encrypted token and the second hash value of each encrypted sub-data; A Merkle tree is constructed based on the first hash value and each of the second hash values, wherein the Merkle path of the preset encrypted token is the path from the first hash value in the Merkle tree to the Merkle root of the Merkle tree.

4. The method according to claim 1, characterized in that, After receiving the first calculation request, the method further includes: Construct a computational binary tree based on the first computation request; Based on the computational binary tree, the target privacy keyword and privacy computation rules are determined.

5. The method according to claim 1, characterized in that, After performing privacy calculations based on the privacy calculation rules and the target raw data to obtain the target calculation result, the process further includes: The act of acquiring the target encrypted data and the target calculation result are recorded on the blockchain.

6. A privacy computing device, characterized in that, The device includes: The receiving module is configured to receive a first computation request, wherein the first computation request includes a target privacy keyword and a privacy computation rule; The lexical expansion module is used to perform lexical expansion processing on the target privacy keyword to obtain the target encrypted lexical; The determining module is used to determine, based on the target encrypted token, a target encrypted token inverted item associated with the target encrypted token from a plurality of preset encrypted token inverted items, wherein the target encrypted token inverted item includes the Merkel path of the target encrypted token and the identification information of the target encrypted data; The acquisition module is used to acquire the target encrypted data based on the Merkel path of the target encrypted token and the identification information of the target encrypted data; The calculation module is used to perform privacy calculations based on the privacy calculation rules and the target raw data to obtain the target calculation result, wherein the target raw data is obtained by decrypting the target encrypted data; The step of performing lexical expansion on the target privacy keyword to obtain the target encrypted lexical includes: The data for the target privacy keyword is filled to obtain the target encrypted token, and the data size of the target encrypted token is a preset value.

7. An electronic device, characterized in that, The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the steps of the privacy computing method as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the steps of the privacy computation method as described in any one of claims 1-5.

9. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device performs the steps of the privacy computing method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Citizen privacy protection method and system based on zero knowledge proof and storage medium

    CN110336672A

  • Distributed transaction propagation and verification system

    US20190147438A1