Privacy information retrieval method and system based on privacy information rewriting

Through entity recognition and noise addition methods, query requests are rewritten and vector mapped, which solves the problems of insufficient privacy-preserving retrieval efficiency and accuracy in existing technologies, and realizes privacy information protection and query intent retention in cross-domain data interaction.

CN120670449APending Publication Date: 2025-09-19GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510670106.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing privacy-preserving retrieval technologies find it difficult to simultaneously guarantee query privacy, query efficiency, and result accuracy when processing natural language queries. They also suffer from high computational complexity and high communication overhead. Especially in cross-domain data interaction, query intent is ambiguous and private information is easily snooped.

Method used

Through entity recognition, dependency parsing, and noise addition, the query request is rewritten to preserve the query intent and mapped into a query embedding vector to prevent private information leakage.

Benefits of technology

It achieves effective rewriting of privacy information in a cross-domain environment, ensures the retention of query intent and privacy protection, and improves query efficiency and result accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670449A_ABST
    Figure CN120670449A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, in particular to a privacy information retrieval method and system based on privacy information rewriting. The method comprises the steps of obtaining a first query request; performing entity identification on the first query request to obtain an entity identification result; rewriting the first query request according to the entity recognition result to obtain a second query request; performing vector mapping and noise addition on the second query request to obtain a query embedded vector; and sending the query embedded vector to a responder for information retrieval. According to the method, entity identification is carried out after the first query request, then the entity containing privacy information is rewritten according to the entity identification result obtained through identification, and the noise vector is added after rewriting is completed. According to the method, the first query request is rewritten through query rewriting and noise adding, privacy information elimination and core query intention reservation are achieved, and therefore a reliable solution is provided for safety information acquisition in a cross-domain environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing technology, and in particular to a method and system for retrieving private information based on private information rewriting. Background Art

[0002] Currently, the demand for cross-domain data retrieval and sharing is growing. In application scenarios such as enterprise knowledge management, scientific research collaborative innovation, and government-enterprise information exchange, cross-domain data interaction has important value. At the same time, it faces severe security and privacy challenges: the inquirer is worried that the query content may be inferred and analyzed, resulting in sensitive or private information being spied on.

[0003] Current privacy-preserving retrieval technologies mainly include cryptography-based secure multi-party computation, fully homomorphic encryption, and differential privacy. Although they can protect data privacy to a certain extent, they usually have disadvantages such as high computational complexity, high communication overhead, and difficulty in application in actual production environments. Especially when processing natural language queries, existing technologies cannot simultaneously guarantee query privacy, query efficiency, and result accuracy.

[0004] Existing cross-domain retrieval privacy protection technologies have the following shortcomings:

[0005] 1. The original query submitted by the user usually contains private information. Simply blocking or replacing sensitive information may lead to ambiguous query intent and affect the relevance and accuracy of the retrieval results.

[0006] 2. The encrypted or replaced information in the query request may be reverse-derived or decrypted by the cracking algorithm.

[0007] 3. It has low adaptability to requests in multiple language modes and is prone to rewriting failures.

[0008] In order to rewrite private information and retain the query intent, the present application provides a private information retrieval method and system based on private information rewriting. Summary of the Invention

[0009] To overcome the problems existing in the related art, the first aspect of the present application provides a private information retrieval method based on private information rewriting, comprising:

[0010] Get the first query request;

[0011] Performing entity recognition on the first query request to obtain an entity recognition result;

[0012] Rewriting the first query request according to the entity recognition result to obtain a second query request;

[0013] Performing vector mapping and noise addition on the second query request to obtain a query embedding vector;

[0014] The query embedding vector is sent to the responder for information retrieval.

[0015] In one embodiment, performing entity recognition on the first query request to obtain an entity recognition result specifically includes:

[0016] Performing entity recognition on the first query request to obtain N entity recognition results, where N is an integer greater than or equal to 1;

[0017] The entity recognition results are classified to determine entity sensitivity levels of N entity recognition results; the entity sensitivity levels are used to represent the privacy leakage risk of the first query request.

[0018] In one embodiment, rewriting the first query request according to the entity recognition result to obtain a second query request includes:

[0019] Performing entity replacement on the first query request according to the entity sensitivity level;

[0020] The first query request after entity replacement is structurally adjusted based on dependency syntactic analysis.

[0021] In one embodiment, the entity sensitivity level includes a low risk level, a medium risk level, and a high risk level;

[0022] Performing entity replacement on the first query request according to the entity sensitivity level specifically includes:

[0023] Replace high-risk entities with generic representations;

[0024] Replace entities with medium risk levels with superordinate expressions;

[0025] Replace low-risk entities with vague representations.

[0026] In one embodiment, the structural adjustment of the first query request after entity replacement based on dependency syntactic analysis specifically includes:

[0027] Performing dependency syntax analysis on the first query request to obtain a syntax tree structure of the first query request;

[0028] Transforming the syntax tree structure according to a preset structure conversion rule library;

[0029] The second query request is generated according to the transformed syntax tree structure.

[0030] In one embodiment, after generating the second query request according to the transformed syntax tree structure, the method further includes:

[0031] evaluating semantic similarity between the first query request and the second query request;

[0032] Determine whether the semantic similarity is greater than or equal to a preset similarity; if so, output the second query request; if not, modify the second query request.

[0033] In one embodiment, modifying the second query request specifically includes:

[0034] Performing intent analysis on the first query request to obtain an intent analysis result;

[0035] Modify the entity of the second query request according to the intent analysis result.

[0036] In one embodiment, modifying the entity of the second query request according to the intent analysis result specifically includes:

[0037] Determine the intent key entity in the second query request according to the intent recognition result;

[0038] Perform semantic expansion on the key entities of the intent.

[0039] In one embodiment, performing vector mapping and noise addition on the second query request to obtain a query embedding vector specifically includes:

[0040] Mapping the second query request to a vector space to obtain a vector representation of the second query request;

[0041] generating a noise vector according to the dimension of the vector representation of the second query request;

[0042] The noise vector and the vector representation of the second query request are fused to obtain the query embedding vector.

[0043] The second aspect of the present application provides a private information retrieval system based on private information rewriting, which is implemented based on the steps in the private information retrieval method described in the first aspect of the present application.

[0044] The technical solution provided by this application may have the following beneficial effects:

[0045] In this application, after obtaining the first query request sent by the search party, multiple entities in the first query request are identified, and the entity recognition result indicates whether the entity contains private information, and the entity containing private information is rewritten. A noise vector is added to the rewritten second query request to prevent a third party from obtaining the original query content by restoring the query embedding vector. This application provides a reliable solution for secure information acquisition in a cross-domain environment. It rewrites the first query request through query rewriting and noise addition to eliminate private information and retain the core query intent.

[0046] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The above and other objects, features and advantages of the present application will become more apparent through a more detailed description of exemplary embodiments of the present application in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components in the exemplary embodiments of the present application.

[0048] Figure 1 A flowchart of a privacy information retrieval method according to an embodiment of the present application;

[0049] Figure 2 This is a system architecture diagram of the privacy information retrieval system shown in an embodiment of the present application. DETAILED DESCRIPTION

[0050] The preferred embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.

[0051] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0052] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0053] Example 1

[0054] Current privacy-preserving retrieval technologies mainly include cryptography-based secure multi-party computation, fully homomorphic encryption, and differential privacy. Although they can protect data privacy to a certain extent, they usually have disadvantages such as high computational complexity, high communication overhead, and difficulty in application in actual production environments. Especially when processing natural language queries, existing technologies cannot simultaneously guarantee query privacy, query efficiency, and result accuracy.

[0055] In response to the above technical problems, an embodiment of the present application provides a private information retrieval method based on private information rewriting, which can rewrite private information and retain the query intent.

[0056] Figure 1 This is a flowchart of the privacy information retrieval method shown in an embodiment of the present application.

[0057] The following combination Figure 1 The privacy information retrieval method is described in detail.

[0058] like Figure 1 As shown, the privacy information retrieval method includes the following steps:

[0059] S1. Obtain a first query request;

[0060] S2. Perform entity recognition on the first query request to obtain an entity recognition result;

[0061] S3. Rewrite the first query request according to the entity recognition result to obtain a second query request;

[0062] S4. Perform vector mapping and noise addition on the second query request to obtain a query embedding vector;

[0063] S5. Send the query embedding vector to the responder for information retrieval.

[0064] In step S1 , the querying party first receives a first query request submitted by a user, and performs format standardization processing on the text content of the first query request.

[0065] Specifically, the following steps are included:

[0066] S101: Identify the natural language type of the first query request using a language detection tool;

[0067] S102: Remove non-standard characters in the first query request;

[0068] S103: uniformly formatting punctuation marks in the first query request;

[0069] S104: Convert Latin characters in the first query request to lowercase;

[0070] S105. Encode the first query request;

[0071] S106: Segment the first query request using a word segmenter to remove words with insufficient semantics.

[0072] When the querying party receives the first query request, it parses the text content, identifies the natural language type of the text content through language detection tools such as langdetect, fastText or CLD3, and then gradually executes steps S102 to S106 for formatting.

[0073] In steps S102 and S103, for example, extra spaces, control characters, line breaks, and non-standard characters in different languages ​​are removed from the text content. Punctuation is uniformly formatted, with full-width punctuation replaced with half-width for Chinese, and standard ASCII punctuation used for English and other Western languages.

[0074] In step S105, according to the language type contained in the text content, a suitable word segmentation tool is selected to perform lemmatization on the text. Chinese text uses dictionary- and statistics-based word segmenters such as Jieba, THULAC or LTP, which can effectively identify independent words in sentences. For Western languages ​​such as English and Spanish, tools such as spaCy and BERT Tokenizer are used for processing. For languages ​​such as Japanese, Korean, and Arabic, localized word segmenters such as MeCab, Sudachi, and Camel Tools are used respectively. For queries in uncertain or mixed languages ​​in a multilingual environment, word segmenters such as WordPiece or SentencePiece that come with multilingual pre-training models such as mBERT or LaBSE can also be used for compatible word segmentation.

[0075] In step S106, after word segmentation, the corresponding stop word list is selected according to the language, and words with low semantic load in the text content are removed. For example, "the", "is", and "of" in English text; "的", "是", and "了" in Chinese text. The word segmentation operation in step S2 lays a clean and language-unified corpus foundation for entity recognition.

[0076] In step S2, it specifically includes:

[0077] S201. Perform entity recognition on the first query request to obtain N entity recognition results; N is an integer greater than or equal to 1;

[0078] S202. Classify the entity recognition results to determine the entity sensitivity level of the N entity recognition results; the entity sensitivity level is used to represent the privacy leakage risk of the first query request.

[0079] In step S201, a pre-trained named entity recognition model is used to perform entity recognition on the first query request. The named entity recognition model is an existing pre-trained neural network model, such as spaCy, Stanza, mBERT, and XLM-RoBERTa. The named entity recognition model can recognize typical entities such as person names, geographical locations, organization names, time, contact information, etc. in the text content.

[0080] Furthermore, to compensate for the possible problem of insufficient recall rate of the NER model, in entity recognition, the embodiment of the present application also integrates a pattern matching rule based on regular expressions to identify sensitive content with discriminable format features such as email, phone number, ID number, IP address, precise time, etc.

[0081] After step S201 is executed, all recognized entities will be uniformly labeled as structured records, including entity text, category, location in the original text, confidence, etc. information, providing a complete and accurate entity list for subsequent determination of entity sensitivity level and query rewriting.

[0082] In step S202, first, a sensitivity weight is assigned to each entity type, then the frequency of entity occurrence and category are accumulated to form an entity sensitivity score, and finally, the entities are divided into different levels of entity sensitivity levels through the entity sensitivity score.

[0083] Exemplarily, the sensitivity weight of personal identity types such as person names and ID numbers is 1.0. The sensitivity weight of contact information such as email and phone is 0.9. The sensitivity weight of time information is 0.5, and the sensitivity weight of general geographical locations is 0.6.

[0084] In an embodiment of the present application, whether the query needs to be rewritten is determined based on the sensitivity score. If the entity sensitivity level exceeds a set privacy risk threshold, a query rewrite will be triggered for the entity.

[0085] Furthermore, step S3 specifically includes:

[0086] S301: Perform entity replacement on the first query request according to the entity sensitivity level;

[0087] S302: Perform structural adjustment on the first query request after entity replacement based on dependency syntax analysis.

[0088] In step S301, the entity sensitivity level includes low risk level, medium risk level and high risk level. S301 specifically includes:

[0089] S3011. Replace entities with high-risk levels with general descriptions;

[0090] S3012. Replace entities with medium risk levels with superordinate expressions;

[0091] S3013. Replace entities with low risk levels with vague expressions.

[0092] For high-risk entities, such as a person's full name, precise address, and contact information, replace them with semantically irrelevant placeholders or generic expressions. For example, replace "Zhang San" with "Someone" and "13812345678" with "Contact Information."

[0093] For medium-risk entities, such as region names and organization names, generalize the entities to superordinate concepts. For example, generalize "Peking University" to "university" and generalize "Zhangjiang Hi-Tech Park, Pudong New Area, Shanghai" to "a city's science and technology park."

[0094] For entities with low risk levels, such as dates, partial information retention is performed. For example, "October 15, 2024" is blurred into "October 2024" or "recently."

[0095] In step S302, it specifically includes:

[0096] S3021: Perform dependency syntax analysis on the first query request to obtain a syntax tree of the first query request;

[0097] S3022, transforming the syntax tree according to a preset structure conversion rule library;

[0098] S3023: Generate the second query request according to the transformed syntax tree.

[0099] A syntax tree is a tree structure that represents a sentence structure, and is used to represent the grammatical components of a sentence and their hierarchical relationships. In an embodiment of the present application, step S302 uses a structured reorganization technology based on syntax tree transformation. First, dependency parsing is used to parse the text content into a structured syntax tree, and core grammatical components such as subject, predicate, and object and their relationships are identified. The syntax tree is transformed according to a preset structure transformation rule library, and finally a natural language expression is regenerated based on the transformed syntax tree. The structure transformation rule library contains a variety of sentence transformation modes. Exemplary:

[0100] 1. Active to passive transformation mode: Convert user-centered active statements into objective passive statements, such as converting "I need to solve the problem of black screen on my computer" into "the solution to the problem of black screen on my computer."

[0101] 2. Question to statement transformation mode: Convert direct questions into indirect demand statements, such as converting "Why is my account frozen?" into "Reasons and solutions for account freezing."

[0102] 3. Conditional reconstruction mode: Convert complex queries containing personal conditions into general conditional expressions, such as converting "I am 30 years old and have high blood pressure. Can I eat chili peppers?" into "Dietary taboos for patients with high blood pressure."

[0103] 4. Core demand extraction model: Extract core information needs from complex narratives, such as converting "Yesterday I bought a piece of clothing at Beijing Road Mall and found that it had quality problems and wanted a refund" into "Refund process for product quality issues."

[0104] 5. Tense standardization mode: Convert the narrative of a specific time point into a universal tense expression, such as converting "My online payment failed yesterday" into "The reasons and solutions for online payment failure."

[0105] The embodiment of the present application adopts a contextual structure conversion engine to select the most suitable combination of structure conversion rules from a structure conversion rule library.

[0106] ACSTE uses a structural importance scoring algorithm based on the attention mechanism to automatically identify the core structural components and secondary components in the query, ensuring that the core information requirements are not deleted or weakened during the reconstruction process. For example, the structural importance scoring formula is:

[0107] StructuralImportance(component)

[0108] =α×SyntacticCentrality(component)+β

[0109] ×SemanticRelevance(component)+γ

[0110] ×PrivacyRisk(component)

[0111] Among them, SyntacticCentrality(·) represents the measurement of the centrality of the component in the syntax tree, SemanticRelevance(·) is used to evaluate the semantic relevance between the component and the query topic, PrivacyRisk(·) evaluates the privacy risk of the component, and α, β, and γ are adjustable weight parameters.

[0112] In order to prevent the second query request from distorting the user's query intent, the embodiment of the present application evaluates the semantic preservation degree of the rewriting result after step S3023, including the following steps:

[0113] S3024. Evaluate the semantic similarity between the first query request and the second query request;

[0114] S3025: Determine whether the semantic similarity is greater than or equal to a preset similarity; if so, output the second query request; if not, modify the second query request.

[0115] In this embodiment of the present application, semantic similarity is determined by calculating the cosine similarity between the first query request and the second query request. Specifically, the calculation formula for semantic similarity is:

[0116]

[0117] Where v1 is the mapping vector for the first query request, and v2 is the mapping vector for the second query request. Semantic similarity ranges from [0 to 1], with higher semantic similarity indicating closer semantics. If the similarity is higher than a preset threshold, the semantics are considered well-preserved. If it is lower than the preset threshold, semantic drift or misunderstanding may have occurred during the rewriting process, and the second query request should be modified accordingly.

[0118] When modifying the second query request, the following steps are specifically included:

[0119] S3026. Perform intent analysis on the first query request to obtain an intent analysis result.

[0120] S3027. Determine the intent key entity in the second query request according to the intent recognition result;

[0121] S3028. Perform semantic expansion on the key intent entity.

[0122] Intent analysis uses a deep learning-based intent analysis model to decompose query intent into primary and secondary intents. Fine-tuned based on pre-trained models such as BERT, the intent analysis model can identify multiple intent types and their hierarchical relationships, including information acquisition, problem solving, opinion consultation, and service requests.

[0123] In step S3027, the retention priority is set for different intent components based on the importance of the intent and the privacy relevance. The priority calculation formula is:

[0124] IntentPreservationPriority(i)=Importance(i)×(1-PrivacyRisk(i))

[0125] In which, IntentPreservationPriority(i) is the priority, Importance(i) represents the importance of intent, PrivacyRisk(i) represents the privacy risk, and i represents the entity of the second query request.

[0126] Using the above priority calculation formula, the entities with the highest priority in the intent analysis results are regarded as the key intent entities. Then, the following operations are performed on the key intent entities:

[0127] 1. Keyword Enhancement: Identify and retain or strengthen keywords related to core intent;

[0128] 2. Semantic extension: Appropriately introduce supplementary concepts related to the core intent to fill the semantic gap caused by the removal of private information;

[0129] 3. Clarify expressions: Convert vague or ambiguous expressions into clear and specific ones to reduce ambiguity.

[0130] After the modification of the second query request is completed, step S3024 is executed again.

[0131] To further improve privacy protection, step S4 specifically includes:

[0132] S401: Map the second query request to a vector space to obtain a vector representation of the second query request;

[0133] S402: Generate a noise vector according to the dimension of the vector representation of the second query request;

[0134] S403: Fusing the noise vector and the vector representation of the second query request to obtain the query embedding vector.

[0135] In an embodiment of the present application, the vector representation v of the second query request has the same dimension as the generated noise vector n. For example, the representation of the generated noise vector n is:

[0136] n~N(0,σ 2 I d )

[0137] Where σ is the standard deviation, which controls the noise amplitude; I d is a d×d identity matrix, where d is the dimension of the noise vector; 0 is the mean vector; and N(·) represents a multivariate normal distribution, ensuring that each dimension is sampled independently.

[0138] In this embodiment of the present application, the noise vector and the second query request are fused. For example, the query embedding vector after the final perturbation is:

[0139] In step S5, the query is embedded into the vector The second query request is encoded in Base64URL format and constructed to conform to the API specification. The second query request includes necessary authentication information, query parameters, and other metadata. The second query request is transmitted to the responder via HTTPS. After receiving the request, the responder performs a similarity search in the vector database to find the document embedding that is most similar to the second query request. The responder returns the raw data of the search results to the queryer.

[0140] In an embodiment of the present application, after obtaining a first query request sent by a search party, multiple entities in the first query request are identified, and an entity recognition result is obtained to indicate whether the entity contains private information. The entity containing private information is rewritten. A noise vector is added to the rewritten second query request to prevent a third party from obtaining the original query content by restoring the query embedding vector. This embodiment of the present application provides a reliable solution for secure information acquisition in a cross-domain environment. It rewrites the first query request through query rewriting and noise addition to eliminate private information and retain the core query intent.

[0141] Example 2

[0142] The embodiment of the present application further provides a privacy information retrieval system, corresponding to the privacy information retrieval method of the first embodiment. Figure 2 This is a system architecture diagram of the privacy information retrieval system shown in the embodiment of this application, such as Figure 2 As shown, the private information retrieval system executes the steps of the private information retrieval method described in the first embodiment.

[0143] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated again here.

[0144] The solution of the present application has been described in detail above with reference to the accompanying drawings. In the above embodiments, the description of each embodiment has its own focus. For parts not described in detail in a particular embodiment, please refer to the relevant description of other embodiments. Those skilled in the art should also be aware that the actions and modules mentioned in the description are not necessarily required for this application.

[0145] In addition, it can be understood that the steps in the method of the embodiment of the present application can be adjusted in order, merged and deleted according to actual needs, and the modules in the device of the embodiment of the present application can be merged, divided and deleted according to actual needs.

[0146] In addition, the method according to the present application may also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing some or all of the steps in the above method of the present application.

[0147] Alternatively, the present application can also be implemented as a non-transitory machine-readable storage medium (or computer-readable storage medium, or machine-readable storage medium) on which executable code (or computer program, or computer instruction code) is stored. When the executable code (or computer program, or computer instruction code) is executed by a processor of an electronic device (or electronic device, server, etc.), the processor executes part or all of the steps of the above-mentioned method according to the present application.

[0148] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the application herein may be implemented as electronic hardware, computer software, or combinations of both.

[0149] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems and methods according to multiple embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or code, and the part of the module, program segment or code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0150] The embodiments of the present application have been described above. The above description is illustrative and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to the technology in the market, or to enable other persons skilled in the art to understand the embodiments disclosed herein.

Claims

1. A private information retrieval method based on private information rewriting, characterized in that: include: Get the first query request; Performing entity recognition on the first query request to obtain an entity recognition result; Rewriting the first query request according to the entity recognition result to obtain a second query request; Performing vector mapping and noise addition on the second query request to obtain a query embedding vector; The query embedding vector is sent to the responder for information retrieval.

2. A method for retrieving private information based on private information rewriting according to claim 1, characterized in that: Performing entity recognition on the first query request to obtain an entity recognition result specifically includes: Performing entity recognition on the first query request to obtain N entity recognition results, where N is an integer greater than or equal to 1; The entity recognition results are classified to determine entity sensitivity levels of N entity recognition results; the entity sensitivity levels are used to represent the privacy leakage risk of the first query request.

3. The method for retrieving private information based on private information rewriting according to claim 2, characterized in that: Rewriting the first query request according to the entity recognition result to obtain a second query request, including: Performing entity replacement on the first query request according to the entity sensitivity level; The first query request after entity replacement is structurally adjusted based on dependency syntactic analysis.

4. The method for retrieving private information based on private information rewriting according to claim 3, characterized in that: The entity sensitivity level includes low risk level, medium risk level and high risk level; Performing entity replacement on the first query request according to the entity sensitivity level specifically includes: Replace high-risk entities with generic representations; Replace entities with medium risk levels with superordinate expressions; Replace low-risk entities with vague representations.

5. The method for retrieving private information based on private information rewriting according to claim 3, characterized in that: The structural adjustment of the first query request after entity replacement based on dependency syntax analysis specifically includes: Performing dependency syntax analysis on the first query request to obtain a syntax tree structure of the first query request; Transforming the syntax tree structure according to a preset structure conversion rule library; The second query request is generated according to the transformed syntax tree structure.

6. The method for retrieving private information based on private information rewriting according to claim 5, characterized in that: After generating the second query request according to the transformed syntax tree structure, the method further includes: evaluating semantic similarity between the first query request and the second query request; Determine whether the semantic similarity is greater than or equal to a preset similarity; if so, output the second query request; if not, modify the second query request.

7. The method for retrieving private information based on private information rewriting according to claim 6, characterized in that: Modifying the second query request specifically includes: Performing intent analysis on the first query request to obtain an intent analysis result; Modify the entity of the second query request according to the intent analysis result.

8. The method for retrieving private information based on private information rewriting according to claim 7, characterized in that: Modifying the entity of the second query request according to the intent analysis result specifically includes: Determine the intent key entity in the second query request according to the intent recognition result; Perform semantic expansion on the key entities of the intent.

9. The method for retrieving private information based on private information rewriting according to claim 1, characterized in that: Performing vector mapping and noise addition on the second query request to obtain a query embedding vector specifically includes: Mapping the second query request to a vector space to obtain a vector representation of the second query request; generating a noise vector according to the dimension of the vector representation of the second query request; The noise vector and the vector representation of the second query request are fused to obtain the query embedding vector.

10. A private information retrieval system based on private information rewriting, characterized in that: The method is implemented based on the steps in the privacy information retrieval method according to any one of claims 1 to 9.