Privacy information retrieval method and system

By introducing trusted third-party and multi-level encryption processing in data privacy protection, the trade-off between efficiency and security in the prior art is solved, and efficient and secure data privacy retrieval and flexible key management are achieved.

CN120145433APending Publication Date: 2025-06-13BUBI (BEIJING) NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510149540.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art has the trade-off between efficiency and security in data privacy protection, especially in large-scale data query scenarios. Comprehensive encryption leads to low query efficiency, pseudonymization processing has the risk of being cracked, and there is a lack of efficient and secure public-private key management methods.

Method used

By introducing trusted third parties, multi-level data encryption and obfuscation processing is adopted, including the generation and distribution of public and private keys, the disruption of keyword lists, the generation of random values ​​and obfuscation encryption, ensuring secure key exchange and data transmission between the data requester and the data provider.

Benefits of technology

It improves the security and efficiency of data privacy retrieval, provides a flexible and scalable key management mechanism, solves the shortcomings of secure interaction and key management in the prior art, and ensures the security of data during transmission and use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145433A_ABST
    Figure CN120145433A_ABST
Patent Text Reader

Abstract

The invention discloses a privacy information retrieval method, which comprises the following steps that: a data requester initiates a query request, and a trusted third party distributes public and private keys to the data requester and a data provider; the data provider obtains corresponding data according to the query request and returns a keyword list to the data requester; the data requester disorganizes the keyword list and records the ith item to be queried; the data provider selects corresponding data according to the keyword list, generates n random values x and sends the n random values x to the data requester; the data requester generates a random value k, the k is mixed with x of the ith item, a ciphertext v is generated through encryption, and the ciphertext v is sent to the data provider; the data provider sequentially decrypts the ciphertext v, encrypts the data by using the generated k value, and sends the encrypted data to the data requester; and the data requester decrypts the ith item of data to obtain a query result. According to the method, the security and the efficiency in the privacy information retrieval process are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information retrieval and encryption, and particularly to a method and system for retrieving private information. Background Art

[0002] With the rapid development of the Internet, the explosive growth of data and the high degree of information sharing have brought great convenience to society. However, the accompanying problem of privacy leakage has become increasingly serious. Data privacy protection has become a research hotspot in the current information technology field. Especially in the context of the wide application of technologies such as big data, cloud computing, and blockchain, how to ensure the security of users' private information while sharing and utilizing data has become an urgent problem to be solved. Existing data privacy protection technologies mainly include means such as data encryption, anonymization processing, and access control. However, these technologies often face the trade-off problem between efficiency and security in practical applications.

[0003] In the prior art, data privacy protection usually relies on the comprehensive encryption of data or the pseudonymization processing of data. However, these methods have certain defects in practical applications. First, although comprehensive encryption can effectively protect the security of data, its computational and storage overheads are relatively large. Especially in the scenario of large-scale data queries, the query efficiency is often not high, which affects the overall performance of the system. Second, although pseudonymization processing protects privacy to a certain extent, there is still a risk of being cracked because there may be a potential connection between the pseudonym and the original data. In addition, most of the existing technologies adopt a single encryption or processing method. For data exchange and queries involving multiple parties, there is a lack of a method that can efficiently, securely process and manage public and private keys, resulting in loopholes in the secure interaction process between data requesters and data providers.

[0004] In view of the above problems, the present invention proposes a method for retrieving private information. By introducing a trusted third party and adopting multi-level data encryption and obfuscation processing, efficient and secure retrieval of data is achieved. The technical solution of the present invention can effectively solve several defects in the prior art.

[0005] First, the present invention generates and distributes public and private key pairs through a trusted third party, realizing the secure key exchange between data requesters and data providers, and avoiding the time and complexity problems of key generation and distribution in traditional methods. Second, during the data query process, the present invention shuffles the keyword list and reduces the total amount of data processing based on a custom indistinguishability parameter, improving the query efficiency. At the same time, through the design of the uniqueness and unpredictability of the generated random value x, the security of the data query process is enhanced. In addition, the present invention ensures the security of data during transmission and use through the multiple encryption processing of data by data providers and the modular operation decryption of query results by data requesters, avoiding potential data leakage risks.

[0006] In summary, the present invention not only improves the security and efficiency of the privacy information retrieval process, but also provides a flexible and scalable key management mechanism, solving the deficiencies in security interaction and key management in the prior art. Therefore, the present invention has broad application prospects in the field of data privacy protection. Summary of the invention

[0007] The purpose of the present invention is to provide a privacy information retrieval method to solve the technical problems in the prior art that data privacy protection usually relies on comprehensive encryption of data or pseudonymization of data, resulting in low query efficiency and affecting the overall performance of the system; the processing method is single; and there are loopholes in the security interaction process between the data requester and the data provider.

[0008] To achieve the above-mentioned purpose, the present invention provides a privacy information retrieval method on the one hand, comprising the following steps: step S10, a data requester initiates a query request, and a trusted third party distributes public and private keys to the data requester and the data provider; step S20, the data provider obtains corresponding data according to the query request and returns a keyword list to the data requester; step S30, the data requester scrambles the keyword list and records the i-th item to be queried; step S40, the data provider selects corresponding data according to the keyword list and generates n random values ​​x, and sends them to the data requester; step S50, the data requester generates a random value k, uses k to confuse the i-th item x, encrypts and generates a ciphertext v, and sends it to the data provider; step S60, the data provider decrypts the ciphertext v in turn, encrypts the data using the generated k value, and sends the encrypted data to the data requester; step S70, the data requester decrypts the i-th item data to obtain the query result.

[0009] The beneficial effects of this technical solution are: first, through the intervention and management of a trusted third party, the security and privacy protection level of the data requester and the data provider in the entire query process is greatly improved. Specifically, the trusted third party is responsible for generating and distributing public and private keys to ensure that the data requester and the data provider complete data interaction without knowing the content of each other's data. In this way, even if there is a potential network attack during the data transmission process, or there is malicious behavior between the data provider and the requester, the security of the data can still be effectively guaranteed. In addition, the introduction of a trusted third party can also prevent privacy leakage caused by information asymmetry between the data requester and the data provider, further improving the overall security and reliability of the system.

[0010] Secondly, in step S30, the data requester shuffles the keyword list and records the i-th item to be queried, effectively avoiding the possibility that the data provider infers the data requester's query intention based on the keyword list. This design greatly enhances the indistinguishability of the query and protects the user's query privacy. In traditional private information retrieval methods, the data provider can often indirectly infer the user's query intention by analyzing the query request and the returned results, thus threatening the user's privacy. However, this technical solution ensures that the query behavior of the data requester remains highly confidential throughout the process by introducing the keyword list shuffling and recording mechanism, greatly enhancing the protection of user privacy.

[0011] Thirdly, during the query process, this technical solution effectively reduces the computational overhead and significantly optimizes the overall computational efficiency by performing encryption and decryption operations on the random values k and x in steps S50 and S60. Compared with traditional homomorphic encryption technologies, this solution reduces the consumption of computing resources by adopting a more lightweight encryption and decryption method. Especially when dealing with large-scale data, it shows higher efficiency and shorter response time. This improvement in computational efficiency not only speeds up the data requester's acquisition of the query results but also reduces the resource consumption of the system under high load, enhancing the scalability and stability of the system.

[0012] In addition, the design of step S70 ensures that the data requester can accurately obtain the required query results even when the data of the data provider changes, greatly guaranteeing the stability and accuracy of data queries. In traditional private information retrieval schemes, when the data structure or content of the data provider changes, it often affects the accuracy of the query results and even leads to query failures. However, in this solution, by pre-defining the order of the data and indexing it with the query keywords, the data requester can still obtain accurate query results through decryption operations regardless of how the data of the data provider changes, thus ensuring the stability and consistency of the query results.

[0013] Finally, this technical solution provides users with a high degree of flexibility, enabling them to independently choose the balance between privacy protection and query efficiency according to specific application scenarios. Through the indistinguishability parameter in step S30, users can flexibly adjust the level of privacy protection to meet the requirements of different scenarios. For example, in scenarios with high privacy protection requirements, users can choose a higher indistinguishability to maximize the protection of personal privacy; while in scenarios with high query efficiency requirements, users can reduce the indistinguishability to improve the query speed and efficiency. This flexibility not only enables this technical solution to adapt to different usage scenarios but also provides users with a more personalized service experience.

[0014] In some alternative embodiments, in step S30, the shuffling process of the keyword list by the data requester is based on a custom indistinguishability parameter.

[0015] The beneficial effect of this technical solution is that in step S30, the shuffling process of the keyword list by the data requester is based on a custom indistinguishability parameter. This design aims to enhance the flexibility of privacy protection and the security of the query process. By introducing the indistinguishability parameter, users can adjust the unpredictability of the keyword list according to their own needs, thereby preventing the data provider from inferring the user's query intention by analyzing the keyword order. This mechanism effectively enhances the indistinguishability of the query, making it difficult for the data provider to identify the specific query target even if they have access to some data content. In addition, users can independently select different indistinguishability levels according to the actual application scenario to find the best balance between privacy protection and query efficiency. This flexibility not only improves the adaptability of the system but also gives users more control, enabling this solution to meet diverse privacy protection requirements.

[0016] In some alternative embodiments, the trusted third party generates and manages multiple public-private key pairs in step S10 and distributes appropriate public-private key pairs to the data requester and the data provider at the start of the query.

[0017] The beneficial effect of this technical solution is that the trusted third party generates and manages multiple public-private key pairs in step S10 and distributes appropriate public-private key pairs to the data requester and the data provider at the start of the query. The beneficial effects of this technical solution are mainly reflected in the following aspects:

[0018] First, the intervention of the trusted third party ensures the secure generation and distribution process of the public-private key pairs, effectively reducing potential security risks during the key generation process. Since the generation of the public and private keys is performed by an independent and trusted third party, it can prevent malicious operations or information theft by the data requester or the data provider during the key generation process, thereby further enhancing the security of the entire system.

[0019] Second, the trusted third party can pre-generate and manage multiple public-private key pairs and select an appropriate key pair for distribution according to the specific situation at the start of the query, which can significantly reduce the query startup time. In traditional methods, the data requester and the data provider usually need to generate key pairs by themselves during the query, which not only increases the computational overhead but also prolongs the query response time. By pre-managing the key pairs, the trusted third party can quickly respond to query requests, improving the system efficiency and user experience.

[0020] In addition, the key management of the trusted third party can also ensure that different queries use different public-private key pairs, further preventing potential security risks caused by key reuse. Key reuse may, under certain conditions, enable attackers to infer sensitive information by analyzing the encrypted data in multiple queries. By allocating independent public-private key pairs for each query, this solution effectively reduces the likelihood of such attacks and enhances the security and privacy protection capabilities of the entire system.

[0021] In some alternative embodiments, as in step S40, after the data provider selects data according to the received keyword list, it generates n random values x and records the number of data items n. 。

[0022] The beneficial effects of this technical solution are as follows: in step S40, after the data provider selects the corresponding data according to the received keyword list, it generates n random values x and records the number of data items n. The beneficial effects of this technical solution are that by generating the random values x, the confusion between the data and the query request is increased, preventing the data requester from inferring the content of other data items from the returned data. This design ensures that even if the data requester obtains some encrypted data, it is still unable to decrypt and obtain the complete data information, thus further protecting the privacy of the data provider. In addition, recording the number of data items n enables the data provider to effectively manage and track each query operation, ensuring that each piece of data in the query process is processed, and preventing data omission or duplicate processing. Overall, this design not only improves the security of the system, but also enhances the integrity and accuracy of the data query process, ensuring the reliability of the query results.

[0023] In some alternative embodiments, in step S50, after the data requester confuses the random value k and the x of the i-th item, it generates the ciphertext v through public-key encryption, and the public key is provided by the trusted third party.

[0024] The beneficial effects of this technical solution are as follows: In step S50, after the data requester confuses the random value k and the x of the i-th item, a ciphertext v is generated through public key encryption, and the public key is provided by a trusted third party. The beneficial effect of this technical solution is that by introducing the public key encryption mechanism provided by a trusted third party, the security and confidentiality of the ciphertext v generated by the data requester during the query process are effectively ensured. The confusion operation combined with public key encryption enables the ciphertext v generated by the data requester to have a high anti-attack ability. Even if the attacker obtains the ciphertext v, it is difficult to decrypt the random value k or the x of the i-th item through reverse engineering or other means. This design further enhances the security of data transmission during the query process and prevents the leakage of sensitive information. At the same time, by providing the public key by a trusted third party, the risk of directly sharing the encryption key between the data requester and the data provider is avoided, and the security vulnerabilities in key management are reduced. In addition, the confusion operation also increases the unpredictability of the query results, making it impossible for the data provider to easily infer the query intention of the data requester during the decryption process, thereby further protecting user privacy.

[0025] In a specific embodiment, in step S60, the data provider sequentially performs AES symmetric encryption on the data by using multiple k values.

[0026] The beneficial effects of this technical solution are as follows: In step S60, the data provider sequentially performs AES symmetric encryption on the data by using multiple k values. The beneficial effect of this technical solution is that by adopting the AES symmetric encryption technology and combining multiple k values for encryption, the security and privacy protection capabilities of the data are significantly enhanced. The AES encryption algorithm is known for its high efficiency and strong encryption ability, and can effectively prevent data from being stolen or tampered with during transmission and storage. The encryption operation of each data block corresponding to a k value ensures that even if the attacker obtains part of the encrypted data, the entire data set cannot be decrypted, thereby preventing the leakage of sensitive information.

[0027] In addition, using multiple k values for encryption further improves the complexity and anti-attack ability of encryption, making it impossible for the data provider to easily infer the query content or intention of the data requester. This design also ensures that even if a certain k value is cracked, other data blocks remain secure, enhancing the overall security and robustness of the system. Through the combination of AES symmetric encryption and multiple k values, the technical solution not only improves the security level of encryption, but also ensures the integrity and confidentiality of the query results, providing a highly secure data retrieval service for the data requester.

[0028] In a specific embodiment, in step S70, the data requester decrypts the i-th item of data by using modular arithmetic and ensures the security of other data through the cooperation of the public key exponent and the private key exponent.

[0029] The beneficial effects of this technical solution are as follows: In step S70, the data requester decrypts the i-th item of data using modular arithmetic and ensures the security of other data through the cooperation of the public key exponent and the private key exponent. The beneficial effects of this technical solution are that by using modular arithmetic, the data requester can accurately decrypt the required i-th item of data while avoiding attempts to decrypt other data, which effectively guarantees the accuracy and targetability of the data request. In addition, the combined use of the public key exponent and the private key exponent makes it impossible for the data requester to obtain unauthorized data content through inference or calculation even if the data requester has the public key, thus further protecting the privacy and security of the data provider.

[0030] This design ensures that only specific data items are decrypted during the query process, and the remaining data remains encrypted, preventing the data requester from decrypting all data through repeated attempts. This precise decryption mechanism not only improves the security and reliability of the query but also reduces unnecessary computational burdens and improves the overall efficiency of the system. At the same time, this solution manages the separation of keys, avoiding the risk of global data leakage caused by the leakage of a single key, thereby enhancing the system's anti-attack ability and data protection level. Ultimately, this technical solution effectively realizes the provision of accurate and reliable data query services while ensuring data security.

[0031] In some alternative embodiments, the data requester records the association relationship between the i-th keyword and the index in step S30.

[0032] The beneficial effects of this technical solution are as follows: The data requester records the association relationship between the i-th keyword and the index in step S30. The beneficial effects of this technical solution are that by recording the association relationship between the keyword and the index, the data requester can accurately locate and retrieve the required specific data during subsequent queries and decryption processes. This design ensures that when the data requester receives the encrypted data returned by the data provider, it can quickly and accurately identify the target data, avoiding query errors caused by index confusion or data order changes.

[0033] In addition, the recording of this association relationship helps to enhance the reliability of the query. Even when the database structure of the data provider changes or the data entries are updated, the data requester can still accurately obtain the required information through the pre-recorded index. This not only improves the stability of the data query but also reduces the risk of query interruption or error caused by data updates or order changes. Through this precise association recording, the technical solution significantly improves the efficiency and accuracy of data retrieval, ensuring that users can obtain fast and reliable query results under the premise of privacy protection.

[0034] In some alternative embodiments, the random value x generated by the data provider in step S40 has uniqueness and unpredictability.

[0035] The beneficial effect of this technical solution is that: the random value x generated by the data provider in step S40 has uniqueness and unpredictability. The beneficial effect of this technical solution is that by generating a random value x with uniqueness and unpredictability, the security and anti-attack ability during data transmission are further enhanced. The unique random value x ensures that each query operation is independent, preventing the risk of data leakage that may be caused by reusing the same random value. This design means that even if an attacker successfully intercepts part of the communication content, it is difficult to speculate or decrypt other data through the known random value, protecting user privacy.

[0036] The unpredictable random value x increases the complexity of data interaction between the data provider and the requester, making it difficult for any external attacker to infer the generation rule of the random value or the next query result by analyzing the query pattern or communication content. This mechanism effectively prevents various common network security threats such as replay attacks and obfuscation attacks, ensuring the confidentiality and integrity of data during transmission.

[0037] By combining the uniqueness and unpredictability of the random value x, the technical solution not only ensures high data security but also improves the robustness of the system, enabling it to exhibit stronger defensive capabilities when dealing with potential security threats. Ultimately, this design ensures that the data requester can securely and reliably perform data retrieval in an environment where privacy and security are fully guaranteed.

[0038] In some alternative embodiments, the trusted third party simultaneously sends public-private key information to the data requester and the data provider in step S10.

[0039] The beneficial effect of this technical solution is that: the trusted third party simultaneously sends public-private key information to the data requester and the data provider in step S10. The beneficial effect of this technical solution is that by having the trusted third party send the public-private key information simultaneously, the security and synchronization of the key distribution process are ensured, effectively preventing potential security hazards caused by key distribution delays or inconsistencies. Since the generation and distribution of the public-private key pair are managed by an independent trusted third party, the data requester and the data provider do not need to worry about the key being tampered with or intercepted by malicious attackers when receiving the key, thus greatly enhancing the overall security of the system.

[0040] In addition, sending the public and private key information simultaneously helps to ensure that the data requester and the data provider have consistent encryption and decryption capabilities before the query operation starts. This synchronization avoids operation failures or data errors caused by key desynchronization, ensuring the smoothness and reliability of the entire query process. Through this mechanism, the data requester and the data provider can quickly enter the query and data interaction phases, reducing waiting time and improving the system's response speed and efficiency.

[0041] Finally, the intervention of a trusted third party can also provide transparency and fairness in key management, avoiding potential trust issues that may occur during the key generation or use process by the data requester and the data provider. By centrally managing and distributing the public and private key pairs by the trusted third party, the system is significantly enhanced in terms of privacy protection and data security, making the entire data retrieval process more secure, trustworthy, and effective.

[0042] In some alternative embodiments, in step S70, after decryption, the data requester converts the query result into a format and outputs it to an Excel spreadsheet.

[0043] On the other hand, the present invention also provides a privacy information retrieval system, including means for performing the privacy information retrieval method described in any one of the foregoing, the system including a data requester, a data provider, and a trusted third party, the trusted third party being used to manage and distribute the public and private key pairs, the data requester being used to initiate a query request and decrypt the final result, and the data provider being used to receive the query request and return encrypted data.

[0044] The beneficial effects of this technical solution are as follows: By introducing a complete privacy information retrieval system, integrating the data requester, the data provider, and the trusted third party, the security, efficiency, and reliability of the entire privacy information retrieval process are significantly improved. First, the addition of the trusted third party ensures the security and fairness of the public and private key management and distribution process, avoiding direct key exchange between the data requester and the data provider, thereby reducing the risk of key leakage. At the same time, the trusted third party can also dynamically generate and distribute the public and private key pairs as needed, ensuring the security and independence of each query.

[0045] Secondly, the data requester undertakes the role of initiating the query request and decrypting the final result in the system. The system design enables the data requester to fully control the privacy protection level of the query throughout the process. By encrypting using the public key provided by the trusted third party, the security of the query request and the decryption result is ensured. This functional design of the data requester enables the system to efficiently process and decrypt large-scale data, provide a fast-response query service, and at the same time ensure that user privacy is fully protected.

[0046] The data provider is responsible for receiving query requests and returning encrypted data. In the system, the data provider can return corresponding encrypted data according to the query content of the requester without worrying about the data being stolen or tampered with during transmission. The data provider only needs to handle the transmission and generation of encrypted data and perform necessary encryption operations using the private key provided by a trusted third party, which greatly simplifies its burden in terms of privacy protection while ensuring the integrity and security of the data.

[0047] By integrating the data requester, data provider, and trusted third party into a unified system, the privacy information retrieval system provided by the present invention not only improves the collaboration efficiency among all parties but also significantly enhances the privacy protection level during the data query process, ensuring that the user's private data will not be leaked or misused while enabling efficient retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0050] Figure 1 is a flowchart of a privacy information retrieval method according to an embodiment of the present invention.

[0051] Figure 2 is a detailed flowchart of a privacy information retrieval method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.

[0053] Existing technologies currently available include searchable encryption technology: Searchable encryption technology allows users to search encrypted data without disclosing the plaintext data. It generates a pair of public and private keys (public key (PK), private key (SK)). The public key is used to encrypt data and generate search tokens, and the private key is used for decryption. The user encrypts each item in the dataset using the public key (PK). Each data item (d_i) is encrypted as (C_i = Enc(PK, d_i)). At the same time, an encrypted index is generated for each data item. The user encrypts the keyword set using the public key (PK) to generate an index. For example, the keyword (w_j) is encrypted as (I_j = Enc(PK, w_j)). These indexes are stored on the server together with the encrypted data. When the user needs to search for a keyword (w), the user uses the public key (PK) to generate a search token (T_w = Enc(PK, w)). After receiving the encrypted search token (T_w), the server uses this token to match in the encrypted index and finds all matching encrypted data items (C_i). The server returns the matching encrypted data items (C_i) to the user. The user decrypts these data items using the private key (SK) to obtain the plaintext. This enables the data to be stored in an encrypted state, protecting the privacy rights of the data owner, and allowing users to effectively search the encrypted data, ensuring the confidentiality and integrity of the data.

[0054] Appendix Figure 1 Shows a privacy information retrieval method according to an embodiment of the present invention, including the following steps:

[0055] Step S10, the data requester initiates a query request, and a trusted third party distributes the public and private keys to the data requester and the data provider;

[0056] Step S20, the data provider obtains the corresponding data according to the query request and returns a keyword list to the data requester;

[0057] Step S30, the data requester shuffles the keyword list and records the i-th item to be queried;

[0058] Step S40, the data provider selects the corresponding data according to the keyword list and generates n random values x, and sends them to the data requester;

[0059] Step S50, the data requester generates a random value k, uses k to confuse with the x of the i-th item, encrypts to generate a ciphertext v, and sends it to the data provider;

[0060] Step S60, the data provider decrypts the ciphertext v in sequence, encrypts the data using the generated k value, and sends the encrypted data to the data requester;

[0061] Step S70, the data requester decrypts the i-th data to obtain the query result.

[0062] As shown Figure 2 in the figure, in step S10, the data requester starts the query operation and first requests a public-private key pair from a trusted third party. The role of the trusted third party is to generate and manage these key pairs to ensure that both the data requester and the data provider can use keys corresponding to the security protocol for encryption and decryption operations. The public key is used to encrypt data, while the private key is used to decrypt data, thus ensuring the security during data transmission. The intervention of the trusted third party ensures that the generation and distribution of keys are secure and fair, preventing the keys from being intercepted or tampered with by malicious attackers, thereby improving the security and credibility of the entire system.

[0063] In step S20, after receiving the query request from the data requester, the data provider extracts the relevant keyword list from its data storage. These keywords are related to the requested content, and the data provider returns these keyword lists to the data requester for use in subsequent steps. By providing the keyword list, the data provider helps the data requester narrow down the query scope, improve the query efficiency, and prepare for subsequent encryption and data transmission steps.

[0064] In step S30, in this step, to protect privacy, the data requester shuffles the keyword list received from the data provider. This shuffling prevents the data provider from inferring the data requester's query intention from the order of the keywords. At the same time, the data requester records the i-th keyword it is actually interested in for use in subsequent steps. Also, by shuffling the keyword list, the unpredictability of the query operation is increased, preventing the data provider from speculating on the data requester's query target, thereby improving the privacy protection level of the query.

[0065] In step S40, the data provider extracts the corresponding data items from its data storage according to the shuffled keyword list. To further protect the security of the data, the data provider generates a random value x for each data item and sends these random values together with the data to the data requester. The generation and use of the random value x increase the confusion of the data, making it difficult for attackers to infer the actual data content even if the data is intercepted, thereby enhancing the security of data transmission.

[0066] In step S50, the data requester generates a random value k for the i-th item of data that it is interested in, and confuses this random value with the i-th item x received previously. Then, the obfuscated value is encrypted using the public key to generate a ciphertext v, which is sent back to the data provider. The obfuscation and encryption operations of the random value k and x ensure that even if the data provider obtains the encrypted data, it is difficult to infer the actual data content through reverse engineering or other means, thereby further protecting the query privacy of the data requester.

[0067] In step S60, after receiving the ciphertext v returned by the data requester, the data provider will use its own private key to decrypt these ciphertexts in turn, thereby generating the corresponding k value. Then, the data provider will use these k values ​​to perform AES symmetric encryption on the corresponding data items, and send the encrypted data to the data requester. By encrypting the data using the k value, it is ensured that even if the data is intercepted during the data transmission process, it is difficult for the attacker to decrypt and obtain the actual data content, further improving the security and privacy of the data.

[0068] In step S70, finally, after receiving the encrypted data, the data requester will use modular operation and the combination of public key exponent and private key exponent to decrypt the i-th data item and obtain the final result of its query. Through precise decryption operation, the data requester can accurately obtain the target data of its query while ensuring that other data items remain encrypted, thereby effectively protecting the privacy of the data provider and ensuring the accuracy and security of the query results.

[0069] Overall, these steps together constitute a highly secure and privacy-protecting privacy information retrieval process. Through the intervention of a trusted third party, the management of public and private keys, the generation and use of random values, and the careful design of encryption and decryption operations, the security of data and the protection of user privacy are ensured throughout the query process.

[0070] In one embodiment, in step S10, the data requester initiates a query request and transmits some basic task information, including the following steps:

[0071] Step S101, the data requester obtains query parameters and obtains relevant information such as the data to be queried;

[0072] Step S102, obtaining the service address of the data provider and constructing the data provider parameter URL;

[0073] Step S103, creating a new empty list for storing subsequent result data;

[0074] Step S103: assemble these parameters and initiate a request to the data provider and the third-party trusted platform.

[0075] In step S104, the third-party trusted platform distributes the corresponding public and private key exponents and modulus to the data provider and the data requester.

[0076] The beneficial effects of this technical solution are as follows. By structuring and step-by-step executing the initiation and management operations of the query request, the rigor and effectiveness of the query process are ensured. First, the process of obtaining the query parameters (step S101) and the service address of the data provider (step S102) enables the query request to clearly point to the source of the required data, avoiding query failures or incorrect data detections caused by unclear targets. This step also ensures that the system fully prepares all necessary information before the query, reducing possible unexpected errors during the query process.

[0077] Creating a new empty list (step S103) to store the query result data further ensures that the data requester has sufficient resources for management and processing when the query results are returned. This pre-prepared method improves the efficiency of the system when dealing with a large amount of data, preventing data loss or delays caused by insufficient memory or improper processing.

[0078] Finally, assembling all parameters and initiating the request (step S104) ensures the consistency and coordination of the query process. At the same time, through the intervention of a trusted third party and the management of public and private keys, the confidentiality and integrity of the data during transmission are ensured. This solution not only improves the security of data transmission, avoiding the risk of data being intercepted or tampered with during transmission, but also ensures that the query operations between the data requester and the data provider are completed efficiently and accurately. Generally speaking, this technical solution effectively improves the overall reliability and security of the private information retrieval system through a carefully designed set of steps and operation sequences.

[0079] In one embodiment, in step S20, the data provider obtains the data to be queried according to the query request and returns a keyword list to the data requester, including the following steps:

[0080] In step S201, upon receiving the request parameters from the data requester, the data source information to be queried is obtained, then the corresponding data source is queried, and all keyword lists in the data source are extracted according to the keyword information to be queried.

[0081] In step S202, the keyword list is checked to see if the indistinguishability can be satisfied. If not, an error message is returned.

[0082] In step S203, the entire keyword list is returned to the data requester.

[0083] The beneficial effects of this technical solution are as follows. By performing the operations of generating, checking, and returning the keyword list step by step, the accuracy and privacy protection of the query process are ensured. First, in step S201, the data provider accurately obtains and extracts the keywords related to the query request, ensuring that the returned data is highly relevant, thereby improving the effectiveness of the query results. At the same time, through the efficient query and processing of the data source, the data provider can quickly respond to the query needs of the data requester, improving the overall response speed of the system.

[0084] The indistinguishability check in step S202 further enhances the privacy protection ability of the system. The design of indistinguishability aims to prevent the data requester from inferring the sensitive information of the data provider through the returned keyword list. This mechanism ensures that even if the data requester conducts a detailed analysis of the returned keyword list, it is difficult to obtain additional information beyond its query scope, thereby improving the privacy security of the data provider.

[0085] Finally, the keyword list return operation in step S203 ensures that the data requester can obtain complete and privacy-protected keyword information, ensuring the smooth progress of subsequent queries and data processing. The design of this process not only ensures the accuracy and effectiveness of data queries but also prevents potential information leakage risks through strict privacy protection measures. Overall, this technical solution effectively improves the privacy protection level and query efficiency of the system in data queries through an orderly and rigorous step arrangement, enabling the data requester to obtain the required information in a secure environment.

[0086] In one embodiment, in step S30, the data requester shuffles the keyword list to be queried according to the indistinguishability, then sends it to the data provider, and records the i-th item it needs to query. It includes the following steps:

[0087] In step S301, the data requester obtains the entire keyword list returned by the data provider. Here, the data requester defines a parameter of indistinguishability (this indistinguishability is less than the total number of data of the data provider), deletes it to a keyword list that meets the indistinguishability (this can reduce the total amount of data to be processed subsequently), and then shuffles it to prevent the data provider from obtaining additional information;

[0088] In step S302, the shuffled keyword list is passed to the data provider again, and the i-th item to be queried by itself is recorded.

[0089] The beneficial effects of this technical solution are as follows. By performing indistinguishability processing and shuffling operations on the keyword list, the privacy protection during the query process is significantly enhanced, while ensuring the accuracy and efficiency of the query. First, the setting of the indistinguishability parameter in step S301 makes the keyword list have a certain degree of ambiguity during transmission and processing, preventing the data provider from inferring the specific query intention of the data requester by analyzing the content and order of the keyword list. This fuzzy processing not only effectively protects the privacy of the data requester but also reduces the additional information that the data provider may obtain, thus ensuring the security of the entire query process.

[0090] In addition, the operation of shuffling the keyword list further increases the difficulty for the data provider to obtain information during decryption and processing. By randomizing the order of the keywords, the data requester can effectively prevent the data provider from identifying or analyzing the query content based on specific patterns, and this design greatly improves the concealment and privacy of the query operation.

[0091] In step S302, the data requester not only sends the processed keyword list back to the data provider but also records the i-th keyword that it is actually interested in. This recording operation ensures that after the data requester receives the encrypted data returned by the data provider, it can accurately decrypt and obtain the target data, avoiding query failures or data confusion caused by disordered order or processing errors.

[0092] In one embodiment, in step S40, the data provider selects the corresponding data according to the keyword list, records the number of data items n, then generates n random values x, and sends them to the data requester. It includes the following steps:

[0093] In step S401, the data provider sequentially selects the corresponding data according to the keyword list returned by the data requester, denoted as (m1, m2, m3....mn), which ensures the association relationship between the keywords and the indexes of the data. No matter how the original data of the data provider changes, it still does not affect the result of this query.

[0094] In step S402, then the data provider generates n random values (x1, x2...xn) according to the number of data items n, then records and sends these n random values to the data requester.

[0095] The beneficial effects of this technical solution are as follows. By precisely selecting and associating the keyword list with data items and generating a unique and unpredictable random value x, the security and privacy protection level of data during transmission and processing are effectively improved. First, the operation of the data provider accurately selecting the corresponding data according to the keyword list in step S401 ensures the accuracy and consistency of data queries. Even if the data items in the data provider's database change, since the index relationship between the keywords and the data items remains unchanged, the data requester can still accurately obtain the target data it queries. This mechanism greatly improves the reliability and stability of the query operation, avoiding query errors or data misalignment caused by data changes.

[0096] In step S402, the data provider further enhances privacy protection during data transmission by generating n random values x and associating them with the corresponding data items. The uniqueness and unpredictability of each random value x ensure that even if the data is intercepted during transmission, attackers cannot infer its original content by directly analyzing the data. The introduction of random values increases the confusion of the data, making subsequent encryption and decryption operations more complex and secure, thus effectively preventing potential information leakage or tampering risks.

[0097] In one embodiment, in step S50, the data requester generates a random value k, uses k to confuse the x of the i-th item, and then encrypts it according to the public key to generate a ciphertext v and sends v to the data provider. It includes the following steps:

[0098] Step S501, the data requester receives the n random values x sent by the data provider, and finds the i-th random value (x_i) according to the i-th item of data it wants to query.

[0099] Step S502, then the data requester generates another random number k and uses k to confuse x_i.

[0100] Step S503, then encrypt: v = (x_i + k^e) mod N, add the selected random value (x_i) and the encrypted (k^e), e is the public key exponent provided by a third-party trusted platform, and then take the modulus of the modulus (N). The obtained ciphertext v is sent to the data provider.

[0101] The beneficial effects of this technical solution are as follows. By the confusion processing of the random values k and x_i and the public key encryption operation, the security and privacy protection level during data transmission are significantly improved. First, the accurate identification of the target random value x_i by the data requester in step S501 ensures the effectiveness and accuracy of subsequent encryption operations, enabling the data requester to concentrate resources for encryption without having to perform unnecessary operations on all random values, thus improving the efficiency of the system.

[0102] The generation of the random number k and the confusion processing of x_i in step S502 add an unpredictable layer to the encryption process. The confusion operation makes it difficult for attackers to infer the confused value through simple analysis even if they obtain x_i or k, which greatly enhances the data privacy protection ability. The confusion processing not only increases the difficulty of data cracking but also ensures the uniqueness and complexity of the encrypted data, making the ciphertext generated by each query unique and impossible to be speculated or attacked through historical data.

[0103] In step S503, the data requester encrypts the confused value using the public key exponent e provided by a trusted third party and generates the ciphertext v through modular arithmetic. This encryption process ensures that even if the ciphertext v is intercepted, the attacker cannot decrypt x_i or k because the decryption operation requires the cooperation of the private key, which is in the hands of the data provider. The use of modular arithmetic further enhances the security of encryption, making decryption more difficult and complex, thus preventing unauthorized access.

[0104] In one embodiment, in step S60, after receiving the ciphertext v, the data provider decrypts it using the private key in sequence and generates n values of k. Then, the data provider encrypts the data using these n values of k in sequence and sends all the encrypted ciphertexts to the data requester, including the following steps:

[0105] Step S601, after receiving the ciphertext v, the data provider decrypts v using its n random values (x1, x2... xn). The calculation method is: [k = (v - x)^d mod N], where d is the private key exponent provided by a third-party trusted platform. Since the data provider does not know which item of data the data requester wants to query, it cannot determine the value of x_i and can only decrypt it in sequence. In this way, multiple k values will be obtained, k_1 = (v - x_1)^d mod N, k_2 = (v - x_2)^d mod N, k_3 = (v - x_3)^d mod N... Among the multiple k values obtained, one is the corresponding k generated by the data requester and should be the i-th item, denoted as k_i;

[0106] Step S602, the data provider performs AES symmetric encryption on the n data m to be queried using k in sequence. First, the plaintext data is divided into blocks of a fixed size (usually 128 bits), and (m) is divided into (n) blocks (M_1, M_2... M_n). Then, the encryption process is initialized using the key and the initial vector. Assuming the initial state matrix is (S): [S_0 = M_i ⊕ K_0]

[0107] For each round (r): [S_r = SubBytes(ShiftRows(MixColumns(S_r-1))) ⊕ K_r]

[0108] Final round:

[0109] [C_i = SubBytes(ShiftRows(S_Nr-1)) ⊕ K_Nr]

[0110] Multiple rounds of substitution, permutation, and confusion operations are performed on each data block. Each round of operation involves a different key round. The last data block may need to be padded to meet the block size requirements, and then all the encrypted data is sent to the data requester in sequence.

[0111] The beneficial effect of this technical solution is that it ensures the security and privacy protection of data during transmission through rigorous decryption and encryption operations. First, in step S601, the data provider must decrypt all random values to obtain multiple k values. This operation increases the complexity and security of the data. Since the data provider does not know the specific data items queried by the data requester, it must try all possible random values in sequence. This process not only prevents the data provider from guessing the query intention of the data requester but also ensures that the privacy of the data requester is fully protected.

[0112] The decryption operation using the private key exponent d ensures that even if the ciphertext v is intercepted, an unauthorized third party cannot easily decrypt and obtain the data content. This key-based decryption mechanism ensures the security of the data and avoids the risk of information leakage during transmission.

[0113] In step S602, the data provider uses the k value to perform AES symmetric encryption on the data, further enhancing the security of the data. The AES encryption algorithm is widely used due to its high efficiency and strong encryption ability. Its multiple rounds of substitution, permutation, and confusion operations ensure that the encrypted data is difficult to crack. Using the k value for AES encryption makes each piece of data undergo complex encryption processing. Even if a certain k value is cracked, the attacker cannot easily infer the content of other data items. This multi-level encryption mechanism greatly improves the security during data transmission and prevents the possibility of information leakage or tampering.

[0114] In one embodiment, in step S70, after the data requester obtains all the encrypted data, it decrypts the i-th item it wants to query to obtain the result, and finally generates the query result, including the following steps:

[0115] Step S701, after receiving all the encrypted data, select the i-th ciphertext that you want to query: (m_i' = (m_i + k_i) mod N). Use the property of modular arithmetic: [(a + b) mod N = c], then: [(c - b) mod N = a]; thus, (m_i = (m_i' - k_i) mod N) is obtained. Then, the data m_i to be queried can be decrypted using modular arithmetic. Moreover, since the data requester only has the public key exponent e and does not have the private key exponent, it cannot calculate other k values, and thus cannot obtain other data through (m_n = (m_n' - k_n) mod N), thereby protecting the security of other data of the data provider.

[0116] Step S702, after the data requester obtains the data to be queried, perform appropriate data format conversion, and finally write the data into an Excel table to obtain the final query result.

[0117] The beneficial effect of this technical solution is that through rigorous decryption operations and data processing, it ensures that the data requester can obtain the required query results safely and accurately, while maximizing the protection of the security and privacy of other data of the data provider. First, the modular arithmetic decryption process in Step S701 enables the data requester to recover the original data from the ciphertext through simple and effective mathematical operations. This decryption method utilizes the basic properties of modular arithmetic, which not only ensures the effectiveness of the decryption operation but also prevents the data requester from accessing non-target data, further protecting the privacy of the data provider.

[0118] Since the data requester only has the public key exponent e and does not have the private key exponent d, it is impossible to deduce other k values, and thus impossible to decrypt and obtain information about other data items. This design ensures that the data requester can only decrypt and access the data it has legally requested, preventing possible privacy leakage or data abuse. Through this strict key management and access control, the system further enhances the security and privacy protection level of the data.

[0119] The data format conversion and result generation operations in Step S702 ensure that the data requester can save and analyze the data in a suitable form after decryption. By writing the data into an Excel table, the data requester can further analyze and process the query results. This intuitive format is also convenient for the display and use of the results. Generally speaking, this step not only ensures the secure transmission and decryption of the data but also improves the usability and practicality of the query results, making the entire private information retrieval process more secure, reliable, and user-friendly.

[0120] Generally speaking, in the present invention, after the data requester selects from the keyword list sent by the data provider and forms a new keyword list, it is then sent to the data provider. The data provider then obtains the corresponding data in sequence according to this new keyword list, thus avoiding the situation where the above-mentioned data serial number changes and leads to incorrect query results. In this way, after positioning using keywords and sequence, no matter what changes occur to the data of the data provider subsequently, the accuracy of this query can still be guaranteed; the data requester puts forward the concept of indistinguishability, further narrowing the query scope and reducing the calculation overhead. While ensuring the established privacy, the query efficiency is improved, and the user can be provided with the right of customization. The user can choose to use complete privacy or faster query efficiency according to his usage scenario, whether he needs to ensure that privacy is not leaked to the greatest extent or hopes to query the result faster; a third-party trusted platform is used. The third-party trusted platform can generate many pairs of public and private keys in advance. Then when the query starts, the third-party trusted platform can immediately send the public and private keys to the corresponding data requester and data provider, which can save the time for generating public and private keys and speed up the query efficiency; secondly, using a third-party trusted platform can also ensure the privacy of both parties and avoid the situation where one party can obtain the data of the other party; for the privacy information retrieval using the OT protocol, 3n encryption and decryption operations are performed during the implementation process. While for the privacy information retrieval using homomorphic encryption, 3n homomorphic encryption operations are required in the same situation, and the calculation overhead of homomorphic encryption is much greater than that of ordinary encryption and decryption operations. Therefore, the query efficiency of constructing privacy information retrieval using OT is higher than that of ordinary homomorphic encryption privacy information retrieval.

[0121] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A privacy information retrieval method, characterized in that: The following steps are involved: Step S10: The data requester initiates a query request, and a trusted third party distributes the public and private keys to the data requester and the data provider; Step S20, the data provider obtains corresponding data according to the query request and returns a keyword list to the data requester; Step S30, the data requesting party shuffles the keyword list and records the i-th item to be queried; Step S40, the data provider selects corresponding data according to the keyword list and generates n random values ​​x, and sends them to the data requester; Step S50, the data requester generates a random value k, uses k to confuse with the i-th item x, encrypts to generate a ciphertext v, and sends it to the data provider; Step S60, the data provider decrypts the ciphertext v in turn, encrypts the data using the generated k value, and sends the encrypted data to the data requester; Step S70: the data requesting party decrypts the i-th item of data to obtain a query result.

2. The method according to claim 1, characterized in that In step S30, the data requesting party shuffles the keyword list based on a user-defined indistinguishability parameter.

3. The method according to claim 1, characterized in that The trusted third party generates and manages a plurality of public-private key pairs in step S10, and distributes appropriate public-private key pairs to the data requester and the data provider when a query starts.

4. The method according to claim 1, characterized in that: In the step S40, the data provider selects data according to the received keyword list, generates n random values ​​x, and records the number of data items n.

5. The method according to claim 1, characterized in that In the step S50, the data requesting party confuses the random value k and the i-th item x, and generates a ciphertext v through public key encryption, where the public key is provided by a trusted third party.

6. The method according to claim 1, characterized in that In the step S60, the data provider performs AES symmetric encryption on the data in sequence by using multiple k values.

7. The method according to claim 1, characterized in that In the step S70, the data requesting party uses modular operation to decrypt the i-th item of data, and ensures the security of other data through the cooperation of the public key index and the private key index.

8. The method according to claim 1, characterized in that The data requesting party records the association relationship between the i-th keyword and the index in the step S30.

9. The method according to claim 1, characterized in that: The random value x generated by the data provider in step S40 is unique and unpredictable.

10. The method according to claim 1, characterized in that In step S10, the trusted third party sends public and private key information to the data requester and the data provider at the same time.

11. The method according to claim 1, characterized in that: In the step S70, the data requesting party converts the query result into a format after decryption and outputs it into an Excel spreadsheet.

12. A privacy information retrieval system, used to execute the privacy information retrieval method according to any one of claims 1 to 11, characterized in that: The system includes a data requester, a data provider and a trusted third party, wherein the trusted third party is used to manage and distribute public-private key pairs, the data requester is used to initiate a query request and decrypt the final result, and the data provider is used to receive the query request and return the encrypted data.