Retrieval enhancement method and device for question and answer scene, equipment and storage medium
By employing an adaptive selection desensitization algorithm to desensitize sensitive data in retrieval enhancement scenarios of large language models, the contradiction between privacy protection and retrieval generation is resolved. This enables the semantic restoration of desensitized documents stored in the knowledge base, thereby improving the accuracy and completeness of question-and-answer results.
Patent Information
- Application Number
- CN202510942230.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-11-28
AI Technical Summary
Existing technologies cannot effectively balance privacy protection and the accuracy of retrieval generation in large language model retrieval enhancement scenarios. Static desensitization leads to semantic loss, rule replacement lacks consistency, and full encryption cannot participate in semantic retrieval.
By adaptively selecting a desensitization algorithm, a reversible desensitization algorithm is used to desensitize sensitive data based on its secrecy level and data type, and the desensitized documents are stored in a knowledge base to ensure that the original semantics can be restored during retrieval.
It improves the accuracy of search results, ensures the completeness and readability of question-and-answer results, and maintains the effectiveness of search while protecting data privacy.
Smart Images

Figure CN121029922A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data security technology, and in particular to retrieval enhancement methods, apparatus, devices, and storage media for question-and-answer scenarios. Background Technology
[0002] Large language models possess powerful text understanding and generation capabilities. Their knowledge originates from snapshots of training data, but they lack access to internal data, private data, confidential information, or personal documents beyond the training data, leading to a gap in proprietary knowledge at times. In such cases, retrieval-enhanced generation methods can be used. Before or simultaneously with generating text content (such as answering questions, writing articles, or engaging in dialogues), large language models can retrieve relevant information from external knowledge sources and provide this information to the language model, enabling it to acquire domain-specific knowledge. To meet regulatory and privacy protection requirements, many internal documents contain highly sensitive information such as names, ID numbers, contact information, and addresses, requiring anonymization before being uploaded to knowledge sources.
[0003] Related technologies employ static anonymization, direct rule replacement, or full encryption to anonymize document data in search enhancement scenarios. Static anonymization results in a loss of contextual semantics in the knowledge source, affecting the accuracy of retrieval and result generation. While rule replacement preserves semantic structure to some extent, the replaced data lacks consistency and reproducibility, making it difficult to provide effective information in the response phase. Full encryption prevents privacy leaks, but the encrypted data cannot be directly used in vectorization or semantic retrieval processes, rendering it largely ineffective in search enhancement scenarios. Summary of the Invention
[0004] The main objective of this application is to propose a retrieval enhancement method, apparatus, device, and storage medium for question-and-answer scenarios, thereby improving the accuracy of retrieval enhancement results and question-and-answer results.
[0005] To achieve the above objectives, a first aspect of this application proposes a retrieval enhancement method for question-answering scenarios, comprising:
[0006] The system acquires user question data and sends it to a knowledge base for retrieval. The knowledge base stores de-identified documents in the target domain. These de-identified documents are obtained by de-identifying the original documents. During the de-identification process, after acquiring the sensitive data of the original documents, a de-identification algorithm is adaptively determined based on the confidentiality level and data type of the sensitive data. The sensitive data is then de-identified according to the algorithm to obtain the corresponding target de-identified data. The de-identified documents are obtained based on the target de-identified data. At least one type of target de-identified data is obtained using a reversible de-identification algorithm.
[0007] Receive the de-identified search results returned by the knowledge base, and locate the corresponding target de-identified data from the de-identified search results;
[0008] Based on the corresponding desensitization algorithm, the target desensitized data is restored to obtain at least the sensitive data corresponding to the reversible desensitization algorithm and the restored data corresponding to the other target desensitized data.
[0009] The question and answer results corresponding to the question data are obtained based on the de-identified search results, the sensitive data, and the restored data.
[0010] In some embodiments, obtaining sensitive data from the original document includes the following steps:
[0011] Based on the document type of the original document, the original document is converted into text data;
[0012] The text data is input into a pre-trained document parsing model for data parsing to obtain at least one sensitive position. Based on the sensitive position, the sensitive data is determined from the original document. The sensitive data includes a corresponding secret level, and the secret level includes at least a high encryption level, a medium encryption level, and a low encryption level.
[0013] In some embodiments, the step of adaptively determining the desensitization algorithm based on the secrecy level and data type of the sensitive data includes:
[0014] For the sensitive data with a high level of encryption, an irreversible desensitization algorithm is selected as the desensitization algorithm based on the data type.
[0015] For the sensitive data with a security level of medium encryption, a reversible desensitization algorithm is selected as the desensitization algorithm based on the data type and the target domain.
[0016] For the sensitive data with a low encryption level of secrecy, a generalized desensitization algorithm is selected as the desensitization algorithm based on the data type.
[0017] In some embodiments, a reversible desensitization algorithm is selected as the desensitization algorithm based on the data type and the target domain. The sensitive data is then desensitized using the desensitization algorithm to obtain the corresponding target desensitized data, including:
[0018] When the data type is a number, obtain the encryption key, and use the encryption key and encryption parameters to perform format-preserving encryption on the sensitive data to obtain the target desensitized data;
[0019] When the data type is string, obtain the encryption key, obtain the field characteristics of the sensitive data, determine the encrypted field and encryption format of the sensitive data based on the field characteristics, obtain the index data of the encrypted field in the preset dictionary, encrypt the encrypted field with pseudo-random numbers according to the user identifier, the field characteristics, and the encryption key to obtain the corresponding pseudo-random offset, and obtain the target desensitized data based on the pseudo-random offset and the index data;
[0020] When the data type is an image, the encrypted region and semantic region are determined from the sensitive data according to the target domain. The encrypted region is erased to obtain the encryption key. Based on the encryption key, the semantic region is encrypted while preserving its format to obtain the target desensitized data.
[0021] In some embodiments, obtaining the encryption key includes:
[0022] The process that calls the key generation interface is determined to be a target process with only de-identification and restoration permissions. Based on the target process, the key generation interface is called to generate an initial key.
[0023] Based on the data type, the initial key is specified for either format-preserving encryption or pseudo-random number encryption to obtain encryption purpose parameters;
[0024] Determine the rotation strategy for the initial key, and generate the encryption key based on the rotation strategy, the encryption purpose parameter, and the initial key.
[0025] In some embodiments, the step of restoring the target de-identified data based on the corresponding de-identification algorithm to obtain at least the sensitive data corresponding to the reversible de-identification algorithm and the restored data corresponding to the other target de-identified data includes:
[0026] When the target de-identified data is obtained based on the sensitive data with the high encryption level, the corresponding restored data is obtained according to the corresponding de-identification algorithm;
[0027] When the target de-identified data is obtained based on the sensitive data of the encryption level, the corresponding sensitive data is obtained according to the corresponding de-identification algorithm;
[0028] When the target de-identified data is obtained based on the sensitive data with the low encryption level, the corresponding restored data is obtained according to the corresponding de-identification algorithm.
[0029] In some embodiments, obtaining the question-and-answer result corresponding to the question data based on the de-identified retrieval result, the sensitive data, and the restored data includes:
[0030] Obtain the data other than the target sensitive data from the de-identified search results as the basic search data;
[0031] The sensitive data and the restored data are filled into the retrieval base data according to the positions of the corresponding target sensitive data to obtain the retrieval results;
[0032] The search results and the question data are input together into a large language model to generate the question and answer, thus obtaining the question and answer results.
[0033] To achieve the above objectives, a second aspect of this application provides a retrieval enhancement device for question-and-answer scenarios, comprising:
[0034] Knowledge base construction module: used to acquire user question data and send the question data to the knowledge base for retrieval and query. The knowledge base stores de-identified documents in the target domain. The de-identified documents are obtained by de-identifying the original documents. During the de-identification process, after acquiring the sensitive data of the original documents, the de-identification algorithm is adaptively determined according to the confidentiality level and data type of the sensitive data. The sensitive data is de-identified according to the de-identification algorithm to obtain the corresponding target de-identified data. The de-identified documents are obtained from the target de-identified data. At least one type of target de-identified data is obtained by a reversible de-identification algorithm.
[0035] Restoration and positioning module: used to receive the de-identified retrieval results returned by the knowledge base, and locate the corresponding target de-identified data from the de-identified retrieval results;
[0036] Restoration module: used to restore the target de-identified data based on the corresponding de-identification algorithm, to obtain at least the sensitive data corresponding to the reversible de-identification algorithm and the restored data corresponding to the other target de-identified data;
[0037] Question and answer result generation module: used to obtain the question and answer results corresponding to the question data based on the de-identified search results, the sensitive data, and the restored data.
[0038] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.
[0039] To achieve the above objectives, a fourth aspect of the present application provides a storage medium that stores a computer program, which, when executed by a processor, implements the method described in the first aspect.
[0040] The retrieval enhancement method, apparatus, device, and storage medium for question-and-answer scenarios proposed in this application acquire user question data and send the question data to a knowledge base for retrieval. The knowledge base stores de-identified documents in the target domain. These de-identified documents are obtained by de-identifying the original documents. During the de-identification process, after acquiring the sensitive data of the original documents, a de-identification algorithm is adaptively determined based on the confidentiality level and data type of the sensitive data. The sensitive data is then de-identified using the algorithm to obtain corresponding target de-identified data. De-identified documents are obtained from the target de-identified data, with at least one type of target de-identified data obtained using a reversible de-identification algorithm. The method receives de-identification retrieval results returned from the knowledge base and locates the corresponding target de-identified data from these results. The target de-identified data is then restored using the corresponding de-identification algorithm, yielding at least the sensitive data corresponding to the reversible de-identification algorithm and restored data corresponding to other target de-identified data. Finally, question-and-answer results corresponding to the question data are obtained based on the de-identification retrieval results, the sensitive data, and the restored data. This application utilizes a knowledge base to store anonymized documents for a specific domain, protecting data privacy while incorporating domain-specific knowledge. During the anonymization process, non-sensitive context is preserved, maintaining the vectorization quality of the original documents and improving retrieval enhancement. Furthermore, the anonymization algorithm is dynamically selected based on the secrecy level and data type, preventing retrieval failures caused by full encryption and avoiding semantic loss associated with static anonymization. Simultaneously, a reversible anonymization algorithm is employed for sensitive data required for retrieval, ensuring the original semantics can be restored subsequently. This allows the knowledge base to participate in vectorized retrieval and support semantic understanding in generating large language models, avoiding semantic defects caused by irreversible anonymization in retrieval enhancement scenarios. This improves the accuracy of retrieval enhancement results, thereby enhancing the accuracy of question-answering results. Attached Figure Description
[0041] Figure 1 This is a flowchart of a retrieval enhancement method for question-and-answer scenarios provided in an embodiment of this application.
[0042] Figure 2 This is a flowchart of obtaining sensitive data from the original document, provided in an embodiment of this application.
[0043] Figure 3 This is a flowchart illustrating the adaptive determination of a desensitization algorithm based on the confidentiality level and data type of sensitive data, as provided in an embodiment of this application.
[0044] Figure 4 This is a flowchart provided in an embodiment of the present application, which describes the selection of a reversible desensitization algorithm based on the data type and target domain, and the desensitization of sensitive data using the desensitization algorithm to obtain the corresponding target desensitized data.
[0045] Figure 5 This is a flowchart of obtaining the encryption key provided in an embodiment of this application.
[0046] Figure 6 This is a flowchart provided in the embodiments of this application, which describes the restoration of target desensitized data based on the corresponding desensitization algorithm, to obtain at least the sensitive data corresponding to the reversible desensitization algorithm and the restored data corresponding to other target desensitized data.
[0047] Figure 7 This is an overall flowchart of the retrieval enhancement method for question-and-answer scenarios provided in the embodiments of this application.
[0048] Figure 8 This is a structural block diagram of a retrieval enhancement device for a question-and-answer scenario provided in another embodiment of this application.
[0049] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0051] It should be noted that although functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart.
[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0053] First, let's analyze some of the terms used in this application:
[0054] Artificial Intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.
[0055] Large language models possess powerful text understanding and generation capabilities. However, their knowledge originates from snapshots of training data, limiting access to internal data, private data, confidential information, or personal documents beyond the training data. This can lead to gaps in proprietary knowledge. In such cases, retrieval-enhanced generation methods can be used. Before or simultaneously with generating text content (such as answering questions, writing articles, or engaging in dialogues), the large language model retrieves relevant information from external knowledge sources and provides this information to the language model, enabling it to acquire domain-specific knowledge. Retrieval-enhanced generation is applied in enterprise-level intelligent question-answering systems, knowledge base system construction, and customer service automation. A typical process involves users uploading structured or unstructured knowledge documents to the cloud, the system vectorizing these documents and storing them in a knowledge base, and then using a combined retrieval and generation process to provide intelligent responses in subsequent question-answering sessions. This technology improves the accuracy and contextual relevance of question-answering, making it suitable for handling complex or domain-specific problems. To meet regulatory and privacy protection requirements, many internal documents contain highly sensitive information such as names, ID numbers, contact information, and addresses, requiring desensitization before being uploaded to knowledge sources.
[0056] Related technologies employ static anonymization, direct rule replacement, or full encryption to anonymize document data in retrieval enhancement scenarios. Static anonymization leads to a loss of contextual semantics in the knowledge source, disrupting the natural language flow within the document and resulting in missing semantic representation. This causes errors or invalid results in the retrieval stage, severely impacting the quality of subsequent generation. While rule replacement preserves semantic structure to some extent, both static and rule replacement are essentially one-way processes, unable to restore the anonymized information in the generated results to the original information. The final answer returned to the user lacks contextual completeness, affecting user understanding and decision-making, and failing to provide effective information in the answering process. Full encryption, while preventing privacy leaks, renders encrypted content ineffective in the retrieval enhancement generation framework because it cannot be used for semantic modeling or retrieval. Due to the inability to restore sensitive content, the generated results often lack original entities (such as names, company names, product models, etc.), reducing the completeness and usability of the question-and-answer format. Therefore, the retrieval enhancement generation framework of these technologies presents a significant contradiction between privacy protection and the quality of retrieval enhancement services, failing to achieve high-quality, reversible anonymization processing.
[0057] Maintaining the original semantics and format of the text is crucial for retrieval enhancement generation technology. This is because, during the retrieval enhancement generation process, the system first "vectorizes" the user-uploaded document, converting natural language content into a semantic vector representation that the machine can understand. This representation is highly dependent on the language structure, contextual semantics, and entity information in the original text. Subsequently, when a user asks a question, the system retrieves the most relevant document fragments based on semantic similarity and provides answers based on these fragments during the generation phase. Therefore, if a document contains a large number of information fragments that have been "statically anonymized" or whose semantic structure has been disrupted, the system will have difficulty understanding the true content expressed in the document, resulting in an inability to accurately match the user's question. For example, anonymizing "Zhang San joined a company in 2023" to "XXX joined a company in XXXX" not only disrupts the semantic flow but also loses key retrieval points such as time and names. Such anonymized text cannot be effectively recalled and will result in vague, incomplete, or even incorrect answers.
[0058] In short, the "search" step in search enhancement generation is not a simple keyword search, but relies on deep semantic understanding and matching. The more complete the original semantics and format of the text, the more accurate the system's understanding and response will be. Therefore, the search enhancement generation process needs to protect sensitive information while maintaining the consistency of the text's semantic structure and contextual information as much as possible, ensuring search accuracy and answer readability, and making the data processing transparent and imperceptible to the user, so that the intelligent question-answering results are both safe and useful.
[0059] Based on this, embodiments of this application provide a retrieval enhancement method, apparatus, device, and storage medium for question-and-answer scenarios. By utilizing a knowledge base to store de-identified documents from a specific domain, data privacy is protected while introducing domain-specific knowledge. Non-sensitive context is preserved during the de-identification process, maintaining the vectorization quality of the original documents and improving retrieval enhancement effects. Furthermore, the de-identification algorithm is dynamically selected based on the confidentiality level and data type, preventing retrieval failures caused by full encryption and avoiding semantic loss due to static de-identification. Simultaneously, a reversible de-identification algorithm is used for sensitive data that needs to be retrieved, ensuring that the original semantics can be restored subsequently. This allows the knowledge base to participate in vectorized retrieval and support semantic understanding for generating large language models, avoiding semantic defects caused by irreversible de-identification in retrieval enhancement scenarios. This improves the accuracy of retrieval enhancement results, thereby improving the accuracy of question-and-answer results.
[0060] This application provides a retrieval enhancement method, apparatus, device, and storage medium for question-and-answer scenarios, which are specifically described through the following embodiments. First, the retrieval enhancement method for question-and-answer scenarios in this application embodiment is described.
[0061] This application's embodiments can acquire and process relevant data based on artificial intelligence (AI) technology. AI is the theory, methods, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can react in a way similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.
[0062] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0063] The retrieval enhancement method for question-and-answer scenarios provided in this application relates to the field of data security technology. This method can be applied to a terminal, a server, or a computer program running on either the terminal or the server. For example, the computer program can be a native program or software module in an operating system; it can be a native application (APP), i.e., a program that needs to be installed in the operating system to run, such as a client supporting retrieval enhancement for question-and-answer scenarios, i.e., a program that only needs to be downloaded to a browser environment to run; or it can be a mini-program that can be embedded in any APP. In short, the above-mentioned computer program can be any form of application, module, or plugin. The terminal communicates with the server via a network. This retrieval enhancement method for question-and-answer scenarios can be executed by the terminal or the server, or by the terminal and the server working together.
[0064] In some embodiments, the terminal may be a smartphone, tablet, laptop, desktop computer, or smartwatch, etc. Additionally, the terminal may also be a smart in-vehicle device. This smart in-vehicle device applies the question-and-answer scenario retrieval enhancement method of this embodiment to provide relevant services and improve the driving experience. The server may be an independent server, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms; it may also be a service node in a blockchain system, where the service nodes in the blockchain system form a peer-to-peer network, and the peer-to-peer network protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). The terminal and the server can be connected via Bluetooth, Universal Serial Bus (USB), or network communication methods, which are not limited in this embodiment.
[0065] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0066] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0067] The following describes a retrieval enhancement method for question-and-answer scenarios in embodiments of this application.
[0068] Figure 1 This is an optional flowchart of the retrieval enhancement method for question-answering scenarios provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps 110 to 130. It is also understood that this embodiment... Figure 1 The order of steps 110 to 130 is not specifically limited. The order of steps can be adjusted or some steps can be reduced or added according to actual needs.
[0069] Step 110: Obtain the user's question data and send the question data to the knowledge base for retrieval and query.
[0070] In one embodiment, in the interaction flow of a question-and-answer system (e.g., a large language model), the user's question data is first acquired. This user question data can be acquired through multimodal methods such as text input, speech recognition, and image upload. For example, the user could directly input "How to prevent pneumonia?", or use speech-to-text technology to convert "What are the early symptoms of pneumonia?" into text, or the user could upload a lung CT image with the text "Please analyze whether this image is abnormal."
[0071] Next, we need to enhance the search within the relevant professional domain. This requires sending the question data to a knowledge base for retrieval. The knowledge base is pre-built before the question-and-answer process and stores anonymized documents related to the target domain. The target domain can be set according to actual needs, such as finance, healthcare, government affairs, or transportation. The anonymized documents are obtained by de-identifying the original documents. The specific anonymization process is described below.
[0072] In one embodiment, the local user terminal performs the data masking operation using an on-device data masking system. During the masking process, it is first necessary to obtain the sensitive data to be masked from the original document. (Refer to...) Figure 2 , Figure 2 This is a flowchart of obtaining sensitive data from the original document provided in an embodiment of this application, specifically including the following steps:
[0073] Step 210: Based on the original document's document type, convert the original document into text data.
[0074] In one embodiment, the original document can be of either text or image type. If it is a text document, the text data is extracted directly. If it is an image document, an OCR recognition method can be used to identify the image and obtain the text content contained therein. Simultaneously, the image content is input into a relevant neural network model for image processing to abstract and summarize the image content, such as image size, image content, and the position of elements in the image, to obtain a text description. The text content and text description are then used together as text data.
[0075] Step 220: Input the text data into the pre-trained document parsing model to parse the data and obtain at least one sensitive position. Based on the sensitive position, determine the sensitive data from the original document.
[0076] In one embodiment, when text data is input into a large-scale document parsing model, the model first performs semantic encoding on the text data through a multi-layer Transformer architecture to capture the contextual dependencies between text. For example, in the medical field, when the text data is a medical report, the document parsing model analyzes the positional relationships of words such as "patient name," "medical record number," and "diagnosis result," and combines this with a medical knowledge graph learned during pre-training to locate paragraphs and sentences containing personal privacy or medically sensitive information. It is understandable that text data from different target domains can be input into corresponding domain-specific large-scale document parsing models. These models are trained using training data from the corresponding domain, thereby improving the accuracy of data parsing.
[0077] Next, based on at least one sensitive location in the output of the document parsing model, the corresponding sensitive data is extracted from the original document. This sensitive data may range from a field (such as "ID number: 1234567"), a sentence (such as "the patient was diagnosed with late-stage lung cancer"), or a piece of text.
[0078] In one embodiment, the document parsing model outputs not only sensitive locations but also the security level of the corresponding sensitive data. Each piece of sensitive data is assigned a specific security level. This classification is based on the risk and severity of data leakage, and the security levels include at least three encryption levels: high, medium, and low. It is understood that different target domains may have different security level settings.
[0079] Specifically, highly encrypted sensitive data can be the most sensitive information involving confidentiality, core business data, or personal privacy, such as coordinate data in military deployment documents, user account passwords in financial institutions, and gene sequencing results in medical records. Medium encryption levels include general business data and personally identifiable information, such as the amount of a company's procurement contract, employee ID numbers, and summaries of a user's medical records. Low encryption levels are for data that is less sensitive but still requires protection, such as agenda documents for public meetings, contact information for ordinary employees, and non-confidential product manuals.
[0080] Taking a corporate financial report as the original document as an example, the document parsing model will first locate sensitive locations such as "net profit," "bank account," and "shareholder shareholding ratio," and then extract data based on these sensitive locations. For example, "bank account: 123456789" involves fund security and is determined to be at a high encryption level; "2024 net profit: 120 million yuan" is corporate operating data and is determined to be at a medium encryption level; and "finance department contact person: Zhang San, telephone: 13800000000" is a common contact method and is determined to be at a low encryption level.
[0081] In this embodiment of the application, a corresponding confidentiality level is set for sensitive data in each target area, which serves as the basis for subsequent desensitization.
[0082] In one embodiment, once the sensitive data and its corresponding secrecy level are known, a desensitization algorithm can be determined. Then, the sensitive data is desensitized according to the algorithm to obtain the corresponding target desensitized data. (Refer to...) Figure 3 , Figure 3 This is a flowchart illustrating the adaptive determination of a de-identification algorithm based on the confidentiality level and data type of sensitive data, provided in an embodiment of this application. The flowchart specifically includes the following steps:
[0083] Step 310: For sensitive data with a high level of encryption, select an irreversible desensitization algorithm as the desensitization algorithm according to the data type.
[0084] In one embodiment, for sensitive data with a high level of encryption, since leakage would cause irreparable damage, an irreversible desensitization algorithm is required to completely destroy the original readability of the data. The core of the irreversible desensitization algorithm is to make the desensitized data unable to be restored to the original sensitive data through one-way hashing or encryption transformation, while ensuring that there is no inverse operation path in the transformation process.
[0085] Different irreversible de-identification algorithms can be selected for sensitive data of different data types. For example, for sensitive text data, cryptographic hash algorithms such as SHA-256 and MD5 can be used as irreversible de-identification algorithms, with a random salt value added to obtain the corresponding target de-identified data, completely losing the original information. For sensitive numeric data, Message Authentication Code (MAC) or Key Hash (HMAC) can be used as irreversible de-identification algorithms to obtain the target de-identified data. The numerical value can then be truncated or moduloed according to business rules. For example, "1,000,000 yuan" can be de-identified as "1***000 yuan", retaining only the magnitude information.
[0086] Understandably, the irreversible desensitization algorithm here can be set according to the actual situation.
[0087] Step 320: For sensitive data with a secret level of medium encryption, select a reversible desensitization algorithm as the desensitization algorithm according to the data type and target domain.
[0088] In one embodiment, sensitive data with a medium encryption level (secret level) needs to retain its data recovery capability after de-identification to meet the semantic requirements of enhanced retrieval. Specifically, a reversible de-identification algorithm controls the de-identification and recovery process through key control, ensuring that the data can be restored to its original value under authorized conditions. (Refer to...) Figure 4 , Figure 4 This application provides a flowchart illustrating how a reversible desensitization algorithm is selected based on data type and target domain, and how sensitive data is desensitized using this algorithm to obtain the corresponding target desensitized data. The flowchart specifically includes the following steps:
[0089] Step 410: When the data type is numeric, obtain the encryption key, and use the encryption key and encryption parameters to perform format-preserving encryption on the sensitive data to obtain the target desensitized data.
[0090] In one embodiment, if the data type is numeric, such as sensitive data that needs to maintain a numerical format, such as financial transaction amounts, ID card numbers, or phone numbers, a corresponding encryption key needs to be obtained for desensitization. (Refer to...) Figure 5 , Figure 5 This is a flowchart of obtaining an encryption key provided in an embodiment of this application, which specifically includes the following steps:
[0091] Step 510: Determine that the process calling the key generation interface is a target process with only de-identification and restoration permissions. Based on the target process calling the key generation interface, generate the initial key.
[0092] In one embodiment, when a user calls a key generation interface on the user's end, either locally or in the cloud, the user key management module of the end-side de-identification system first generates a process to call the key generation interface. At this point, it is necessary to determine whether this process is a target process "only having de-identification and restoration permissions" for secure access control. Then, the target process is used to call the key generation interface to generate an initial key.
[0093] The user key management module is responsible for the full lifecycle management of encryption keys, including generation, storage, rotation, access control and auditing, to ensure that the keys are always under the control of the client side and that the service provider that generates the enhanced retrieval cannot obtain them in plaintext.
[0094] Step 520: Specify the initial key for either format encryption or pseudo-random number encryption based on the data type to obtain the encryption purpose parameter.
[0095] In one embodiment, when generating the initial key, it is necessary to specify whether the initial key is used for format-preserving encryption or pseudo-random number encryption according to the data type, thereby obtaining the encryption purpose parameter. In other words, the encryption purpose parameter is used to determine whether the encryption method is format-preserving encryption or pseudo-random number encryption.
[0096] Step 530: Determine the rotation strategy for the initial key, and generate an encryption key based on the rotation strategy, encryption purpose parameters, and the initial key.
[0097] In one embodiment, a rotation strategy for the initial key also needs to be generated. This rotation strategy determines how to periodically change the encryption key to reduce the risk of key leakage, such as short-cycle (e.g., daily / weekly) rotation, session-level dynamic rotation, etc. During rotation, the old encryption key is retained for historical data restoration, while the new encryption key undergoes a de-identification process. With the rotation strategy in place, the encryption key is obtained by using encryption purpose parameters, version number, and other information as relevant attribute parameters for the initial key.
[0098] In one embodiment, after the encryption key is generated, it is sent to the edge-side de-identification system via a local access / secure channel. The encryption key is not stored on disk but exists only briefly in memory, and the memory copy is cleared immediately after the call is completed. At the same time, all key operations, such as generation, rotation, access, and deletion, are written to the audit log, and an alarm is automatically triggered when abnormal access occurs.
[0099] In one embodiment, after obtaining the encryption key, sensitive data of the digital type is desensitized. Specifically, the sensitive data is encrypted with the encryption key and encryption parameters while preserving its format to obtain the target desensitized data. The desensitized data has the same format rules as the sensitive data, such as the number of bits and the checksum.
[0100] In one embodiment, the FPE_FF1 algorithm from the NIST SP 800-38G standard can be selected as the format-preserving encryption algorithm, and the desensitization process is represented as follows:
[0101] token=FPE_FF1 K (plain,domain)
[0102] Wherein, token represents the target de-identified data, FPE_FF1 K This indicates a format-preserving encryption algorithm based on encryption key K, where K represents the encryption key, plain represents sensitive data, and domain represents encryption parameters. These are numeric field parameters used to define the character set and length (e.g., "0–9", length = 18) to ensure the consistency of the encrypted content with the original content format.
[0103] The following examples use ID card number, mobile phone number, and bank card number as examples respectively.
[0104] For example, sensitive data is: 11xxxxxxxxxxxxxx4, encryption parameter is domain = 18-bit number, encryption key is: K = 0x3f2a…, target de-identified data is: token = 89xxxxxxxxxxxxxx6. For example, sensitive data is: 13000000000, encryption parameter is domain = 11-bit number, encryption key is: K = 0x3f2a…, target de-identified data is: token = 57980000000. For example, sensitive data is: 62000000000000000, encryption parameter is domain = 16-bit number, encryption key is: K = 0x3f2a…, target de-identified data is: token = 7845120000000000.
[0105] Step 420: When the data type is string, obtain the encryption key, obtain the field characteristics of the sensitive data, determine the encrypted field and / or encryption format of the sensitive data based on the field characteristics, obtain the index data of the encrypted field in the preset dictionary, encrypt the encrypted field with pseudo-random numbers according to the user identifier, field characteristics, and encryption key, obtain the corresponding pseudo-random offset, and obtain the target desensitized data based on the pseudo-random offset and index data.
[0106] In one embodiment, for sensitive data of string type, it is necessary to obtain the field characteristics of the sensitive data. For example, the field characteristics could be name / address, email, or free text. As can be seen, field characteristics are used to indicate the nature of the sensitive data content. With the field characteristics, the sensitive data can be processed accordingly to obtain the corresponding encrypted fields and encryption formats that need to be de-identified. For example, for sensitive data of name / address type, the encrypted field is the entire string, and the encryption format is: maintaining the basic structure, where the basic structure could be: surname + given name, province + city + district, etc. For sensitive data of email type, the encrypted field is the part before "@", excluding the suffix. For free text, the encrypted field is the entity obtained using word segmentation.
[0107] After obtaining the field features, we can also obtain the user identifier. This user identifier is a unique identifier set by the user to distinguish different application environments or different search enhancement generation services. It is set according to the actual situation to improve the randomness and security isolation of the de-identification results. Then, we set the pseudo-random number encryption to a pseudo-random function PRF, and then perform pseudo-random number encryption on the encrypted field according to the user identifier, field features, and encryption key to obtain the corresponding pseudo-random offset, represented as:
[0108] r = PRF K (type_id, nonce)
[0109] Where r represents the pseudo-random offset, type_id represents the field feature, and PRF K This indicates a pseudo-random number encryption algorithm based on encryption key K, and nonce represents the user identifier.
[0110] Next, the index data of the encrypted fields in the preset dictionary is obtained. This preset dictionary contains custom data such as commonly used names, addresses, and English words pre-defined according to actual needs. Each encrypted field can find its corresponding index data in the preset dictionary. Then, based on the pseudo-random offset and the index data, the target de-identified data is obtained, represented as:
[0111] token = dict[(idx+r)mod N]
[0112] Where N represents the size of the preset dictionary, idx represents the index data, mod represents the modulo operation, dict represents the preset dictionary, and token represents the target de-identified data.
[0113] In one embodiment, taking a name as an example, the sensitive data is: Zhang San, the index data is: idx("Zhang San") = 125, the field type is: type_id = 1, the encryption key is: K = 0x3f2a…, and the user identifier is: nonce = "abc 123". In this case, the calculated pseudo-random offset is represented as: r = 47, and the further desensitized target data is represented as: token = Li Si. Alternatively, taking an email address as an example, if the field feature indicates that the sensitive data is an email address, the target desensitized data needs to be further concatenated to token@xxx.com. Assuming the sensitive data is: alice@example.com, the index data is: idx("alice") = 200, the field type is: type_id = 3, the encryption key is: K = 0x3f2a…, and the user identifier is: nonce = "xyz-789". In this case, the calculated pseudo-random offset is represented as: r = 345, and the further desensitized target data is represented as: token = maria@example.com. Taking free text as an example, the sensitive data is: xx city xx district sports road, the index data is: idx("xx city xx district sports road") = 210, the field type is: type_id = 2, the encryption key is: K = 0x3f2a…, the user identifier is: nonce = "addr-xyz". At this time, the calculated pseudo-random offset is represented as: r = 87, and the further obtained target desensitized data is represented as: token = Dream City Future District Xingchen Road.
[0114] Step 430: When the data type is an image, determine the encrypted region and semantic region from the sensitive data according to the target domain, erase the encrypted region, obtain the encryption key, and encrypt the semantic region with preserved format based on the encryption key to obtain the target desensitized data.
[0115] In one embodiment, for sensitive data of the image type, it is necessary to determine the encrypted region and semantic region based on the business rules of the target domain. For example, taking a lung CT scan in the medical field as an example, the encrypted region can be: text-annotated areas such as patient name, medical record number, and age, as well as biometric feature areas such as face and body contours, while the semantic region can be: lung tissue, lesions (such as nodules and shadows), key anatomical structures, and other areas related to diagnosis. Taking a bank card photo in the financial field as an example, the encrypted region can be: payment-sensitive information areas such as card number, CVV code, and cardholder name, while the semantic region can be non-sensitive areas used for card type identification, such as the bank logo and card design.
[0116] Next, after establishing the encrypted area, irreversible physical masking techniques are applied to ensure that the privacy information cannot be recovered. For example, Gaussian blur, mosaic, or other filters can be applied to the encrypted area to destroy pixel details, or a solid color block (such as a black rectangle) can be used to cover the encrypted area, or a pattern similar to the background can be generated for filling. Alternatively, semantic erasure can be performed. For text areas, image restoration algorithms are used to reconstruct the content based on surrounding pixels, preserving background texture while eliminating text information. The erasure method here can be set according to the actual situation. Semantic areas require format-preserving desensitization, treating them as character-type sensitive data and performing desensitization according to the reversible desensitization process for string-type sensitive data described above, which will not be elaborated further here.
[0117] Step 330: For sensitive data with a low encryption level, select a generalized desensitization algorithm as the desensitization algorithm based on the data type.
[0118] In one embodiment, the risk of leakage of sensitive data with low encryption levels is relatively low, but it is still necessary to avoid the exposure of individual information or detailed characteristics. Therefore, a generalized de-identification algorithm is used to eliminate individual identifiability in sensitive data while retaining group characteristics, so that the resulting de-identified data can be used for non-sensitive business scenarios such as statistical analysis and trend judgment. At the same time, the generalized de-identification algorithm can reduce computational overhead.
[0119] In one embodiment, different generalized desensitization algorithms can be used for sensitive data of different data types. For text-type sensitive data, entity substitution generalization can be used, replacing specific individual information with abstract representations of similar groups. Sensitive entities are located through Named Entity Recognition (NER) and replaced with a preset generalization template. Alternatively, keyword fuzzy generalization can be used, preserving the text structure while partially masking sensitive words. For example, "Li Si in the R&D department is responsible for the AI project" becomes "An employee in the R&D department is responsible for the AI project" after desensitization. For numeric-type sensitive data, range aggregation generalization can be used, mapping precise values to a wider range while preserving order-of-magnitude features. For example, "The price of a certain product is 2999 yuan" becomes "The price of this product is 2000-3000 yuan." In this process, the range span can be set according to the business scenario. Rounding generalization can also be used, discarding the lower significant digits of the value and retaining the higher significant digits. For example, "User rating 4.87 points" becomes "User rating is 4.9 points" after desensitization. It is understandable that the above generalized desensitization algorithms are just examples, and the appropriate algorithm can be selected based on the actual data type and reverse engineering.
[0120] As can be seen, through the above process, the corresponding desensitization algorithm can be adaptively determined according to the secret level and data type of the sensitive data in the embodiments of this application. Then, the sensitive data is desensitized according to the desensitization algorithm to obtain the corresponding target desensitized data. At least one type of target desensitized data is obtained by a reversible desensitization algorithm.
[0121] In one embodiment, after obtaining the target de-identified data, the sensitive data in the original document is replaced with the target de-identified data to obtain the de-identified document. The de-identified document is then uploaded to the cloud, where a cloud-based search enhancement generation service vectorizes the uploaded de-identified document and constructs an indexable knowledge base structure based on these vectors.
[0122] Step 120: Receive the de-identified search results returned by the knowledge base, and locate the corresponding target de-identified data from the de-identified search results.
[0123] In one embodiment, after the retrieval enhancement generation service is deployed, users can query the knowledge base content by asking questions in natural language. At this time, the user sends question data to the knowledge base for retrieval, performs retrieval enhancement generation, and obtains anonymized retrieval results.
[0124] Understandably, since the search enhancement generated by the system is based on anonymized knowledge data, the anonymized search results returned by the system are also anonymized text. To enable end users to see the restored, readable answers, the restoration component needs to be called locally on the user's device to restore the anonymized search results returned from the cloud to their original form.
[0125] In one embodiment, in order to restore the data, it is first necessary to locate the corresponding target de-identified data from the de-identification retrieval results. The target de-identified data can be directly identified in the de-identification retrieval results, or it can be obtained by extracting the content using a corresponding large model. This embodiment does not limit this approach.
[0126] Step 130: Restore the target desensitized data based on the corresponding desensitization algorithm, and at least obtain the sensitive data corresponding to the reversible desensitization algorithm and the restored data corresponding to the other target desensitized data.
[0127] In one embodiment, reference is made to Figure 6 , Figure 6 This application provides a flowchart for restoring target de-identified data based on a corresponding de-identification algorithm, to obtain at least the sensitive data corresponding to the reversible de-identification algorithm and the restored data corresponding to other target de-identified data. The flowchart specifically includes the following steps:
[0128] Step 610: When the target de-identified data is obtained from highly encrypted sensitive data, the corresponding restored data is obtained according to the corresponding de-identification algorithm.
[0129] Step 620: When the target de-identified data is obtained from the sensitive data of the encryption level, the corresponding sensitive data is obtained according to the corresponding de-identification algorithm.
[0130] Step 630: When the target de-identified data is obtained from the low-encryption level sensitive data, the corresponding restored data is obtained according to the corresponding de-identification algorithm.
[0131] In one embodiment, since the de-identification process is also executed on the client side, prior knowledge can be used to determine the de-identification algorithm used for each target de-identified data and the relevant intermediate parameters. Therefore, for different target de-identified data, the corresponding de-identification algorithm can be used for restoration. It is important to note that only when the target de-identified data is obtained from sensitive data with a medium encryption level is the corresponding de-identification algorithm a reversible algorithm. Therefore, only when the target de-identified data is obtained from sensitive data with a medium encryption level can the corresponding sensitive data be restored from the target de-identified data using the corresponding de-identification algorithm. High-encryption-level and low-encryption-level target de-identified data are obtained through irreversible de-identification processes; therefore, they cannot be restored without loss of data to obtain the corresponding restored data.
[0132] As can be seen, the embodiments of this application introduce a "de-identification process" and a "restoration process" in the two key stages of the retrieval enhancement generation system, namely "knowledge construction" and "query response", respectively. This enables local de-identification processing of the original documents and local restoration processing of the query results, thereby ensuring data security and privacy compliance, while ensuring seamless querying for users.
[0133] Step 140: Obtain the question and answer results corresponding to the question data based on the de-identified search results, sensitive data, and restored data.
[0134] In one embodiment, generating question-and-answer results relies on a combination of anonymized search results, sensitive data, and restored data. First, the data in the anonymized search results, excluding the target sensitive data, is obtained as the basic search data. Next, the sensitive data and restored data are filled into the basic search data according to the positions of the corresponding target sensitive data to obtain the search results. Finally, the search results and the question data are input together into a large language model to generate the question-and-answer results.
[0135] For example, in a medical question-and-answer scenario, the anonymized search result is: "A patient's medical record number is 12345, and the diagnosis is pneumonia." This identifies sensitive data such as "a patient" and "medical record number 12345." Therefore, the basic search data is: "[Patient's name]'s medical record number is [medical record number], and the diagnosis is pneumonia." After restoring the sensitive data, the corresponding sensitive data or restored data is obtained. The sensitive data corresponding to "a patient" is "Patient Zhang San," and the restored data corresponding to "medical record number 12345" is "medical record number 54321." Therefore, the search result is: "Patient Zhang San's medical record number is 54321, and the diagnosis is pneumonia." The search results and the question data are then input into a large language model to generate the question-and-answer results.
[0136] In one embodiment, reference is made to Figure 7 , Figure 7 This is an overall flowchart of the retrieval enhancement method for question-and-answer scenarios provided in the embodiments of this application.
[0137] Reference Figure 7 First, the local user client sends the original document to the locally deployed end-to-end desensitization system. On the local device side, the original document, containing multiple types of sensitive data, undergoes comprehensive detection and desensitization to ensure that the knowledge base uploaded to the cloud completely isolates the original sensitive data, preventing the risk of sensitive data leakage. During the desensitization process, the original document is first converted into text data based on its document type. Then, the text data is input into a pre-trained document parsing model for data analysis to obtain at least one sensitive location. Based on this sensitive location, sensitive data is determined from the original document. The sensitive data includes corresponding security levels, with security levels including at least high, medium, and low encryption levels.
[0138] Then, the de-identification algorithm is adaptively determined based on the secrecy level and data type of the sensitive data. Specifically: for sensitive data with a high secrecy level, an irreversible de-identification algorithm is selected based on the data type; for sensitive data with a medium secrecy level, a reversible de-identification algorithm is selected based on the data type and target domain; and for sensitive data with a low secrecy level, a generalized de-identification algorithm is selected based on the data type.
[0139] Furthermore, when performing reversible desensitization on sensitive data with medium encryption levels, if the data type is numeric, the encryption key is obtained, and the sensitive data is encrypted while preserving its format using the encryption key and encryption parameters to obtain the target desensitized data. If the data type is string, the encryption key is obtained, the field characteristics of the sensitive data are obtained, the encrypted field and encryption format of the sensitive data are determined based on the field characteristics, the index data of the encrypted field in the preset dictionary is obtained, and the encrypted field is encrypted with pseudo-random numbers according to the user identifier, field characteristics, and encryption key to obtain the corresponding pseudo-random offset. The target desensitized data is obtained based on the pseudo-random offset and index data. If the data type is image, the encrypted area and semantic area are determined from the sensitive data according to the target domain, the encrypted area is erased, the encryption key is obtained, and the semantic area is encrypted while preserving its format based on the encryption key to obtain the target desensitized data.
[0140] Anonymous documents are generated based on the target anonymized data for building a cloud-based knowledge base and for use in cloud-based search enhancement generation services. After a user generates a question, the question data needs to be anonymized according to the anonymization method of the anonymized document to obtain an anonymized question. This anonymized question is then sent to the cloud-based search enhancement generation service, which performs the search and generation, returning the anonymized search results to the local user. The local user then locates the corresponding target anonymized data from the anonymized search results and restores the target anonymized data based on the corresponding anonymization algorithm, obtaining at least the sensitive data corresponding to the reversible anonymization algorithm and the restored data corresponding to other target anonymized data. Based on the anonymized search results, sensitive data, and restored data, the question-and-answer results corresponding to the question data are obtained.
[0141] In this process, the cloud-based search enhancement generation service does not directly process the original sensitive data, reducing the risk of cloud data leakage. Since the client-side is local to the user, it can perform de-identification and restoration without worrying about leakage issues, ensuring the security and privacy compliance of sensitive data throughout the entire knowledge base construction and query process.
[0142] The above process can be summarized as follows: 1) Local document preprocessing: The original document is converted into a de-identified document. After the user prepares the original document locally, the document content is first de-identified through the client-side de-identification system to generate a de-identified document. 2) Cloud-side knowledge base construction: The de-identified document is stored in the knowledge base. During this process, the cloud-based search enhancement generation service performs vectorization processing and index construction on the de-identified document to generate a de-identified knowledge base. 3) Local user questioning: The user submits a natural language question to the cloud-based search enhancement generation service, and the resulting question data is converted into a de-identified question. 4) Cloud-based knowledge base query and answering: Based on the de-identified question, the service performs retrieval and generation, returning the corresponding de-identified search results. Since the knowledge base is de-identified, the answer is also a de-identified answer. 5) Local result restoration: The client converts the de-identified search results to obtain the corresponding question-and-answer results. Thus, the user obtains the restored, complete, and readable original semantic answer, achieving secure and controllable access to private knowledge.
[0143] This application integrates the edge-side desensitization / restoration mechanism into the standard workflow of the cloud-based search enhancement generation service, forming a closed-loop sensitive information desensitization mechanism. Sensitive information is desensitized for documents and questions during the knowledge base construction phase and the query-response node, respectively. Furthermore, to maintain a seamless user experience, the answers are restored. This achieves end-to-end sensitive information protection in the cloud-based search enhancement generation scenario, ensuring that the cloud-based search enhancement generation service only accesses desensitized data, while guaranteeing that end users receive accurate, complete, and readable answers. In addition, differentiated but unified key-based desensitization strategies are adopted for different sensitive data types, simultaneously meeting the core requirements of format consistency, security, and reversibility. This is effectively applicable to the cloud-based search enhancement generation scenario and avoids semantic understanding issues.
[0144] The technical solution provided in this application involves acquiring user question data and sending it to a knowledge base for retrieval. The knowledge base stores anonymized documents in the target domain. These anonymized documents are obtained by anonymizing the original documents. During the anonymization process, after acquiring sensitive data from the original documents, an anonymization algorithm is adaptively determined based on the confidentiality level and data type of the sensitive data. The sensitive data is then anonymized using this algorithm to obtain corresponding target anonymized data. Anonymized documents are obtained from the target anonymized data, with at least one type of target anonymized data obtained using a reversible anonymization algorithm. The system receives anonymized retrieval results returned from the knowledge base and locates the corresponding target anonymized data from these results. The target anonymized data is then restored using the corresponding anonymization algorithm, yielding at least the sensitive data corresponding to the reversible anonymization algorithm and restored data corresponding to other target anonymized data. Finally, the question-and-answer results corresponding to the question data are obtained based on the anonymized retrieval results, the sensitive data, and the restored data. This application utilizes a knowledge base to store anonymized documents in a specific domain, protecting data privacy while introducing domain-specific knowledge. The anonymization process preserves non-sensitive context, maintains the vectorization quality of the original documents, and enhances retrieval effectiveness. Furthermore, the de-identification algorithm is dynamically selected based on the secrecy level and data type, preventing retrieval failures caused by full encryption and avoiding semantic loss due to static de-identification. Simultaneously, a reversible de-identification algorithm is used for sensitive data that needs to be used in retrieval, ensuring that the original semantics can be restored subsequently. This allows the knowledge base to participate in vectorized retrieval and support semantic understanding in generating large language models, avoiding semantic defects caused by irreversible de-identification in retrieval enhancement scenarios. This improves the accuracy of retrieval enhancement results, thereby improving the accuracy of question-answering results.
[0145] This application also provides a retrieval enhancement device for question-and-answer scenarios, which can implement the above-mentioned retrieval enhancement method for question-and-answer scenarios, as described above. Figure 8 The device includes:
[0146] Knowledge base construction module 810: This module is used to acquire user query data and send it to the knowledge base for retrieval. The knowledge base stores de-identified documents in the target domain. The de-identified documents are obtained by de-identifying the original documents. During the de-identification process, after acquiring the sensitive data of the original documents, the de-identification algorithm is adaptively determined based on the confidentiality level and data type of the sensitive data. The sensitive data is then de-identified according to the de-identification algorithm to obtain the corresponding target de-identified data. The de-identified documents are obtained from the target de-identified data. At least one type of target de-identified data is obtained by a reversible de-identification algorithm.
[0147] Restoration and positioning module 820: Used to receive the de-identified retrieval results returned based on the knowledge base, and locate the corresponding target de-identified data from the de-identified retrieval results.
[0148] Restoration module 830: used to restore the target desensitized data based on the corresponding desensitization algorithm, at least to obtain the sensitive data corresponding to the reversible desensitization algorithm and the restored data corresponding to other target desensitized data.
[0149] Question and answer result generation module 840: Used to obtain the question and answer results corresponding to the question data based on the de-identified search results, sensitive data, and restored data.
[0150] The specific implementation of the retrieval enhancement device for the question-and-answer scenario in this embodiment is basically the same as the specific implementation of the retrieval enhancement method for the question-and-answer scenario described above, and will not be repeated here.
[0151] This application also provides an electronic device, including:
[0152] At least one memory;
[0153] At least one processor;
[0154] At least one program;
[0155] The program is stored in a memory, and the processor executes the at least one program to implement the retrieval enhancement method for the question-and-answer scenario described above. The electronic device can be any smart terminal, including mobile phones, tablets, personal digital assistants (PDAs), and in-vehicle computers.
[0156] Please see Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:
[0157] The processor 901 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0158] The memory 902 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 to perform the retrieval enhancement method for the question-and-answer scenario in the embodiments of this application.
[0159] The input / output interface 903 is used to implement information input and output;
[0160] The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0161] Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904);
[0162] The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.
[0163] This application embodiment also provides a storage medium that stores a computer program. When the computer program is executed by a processor, it implements the retrieval enhancement method for the above-mentioned question-and-answer scenario.
[0164] Memory, as a non-transitory storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0165] The retrieval enhancement method, apparatus, device, and storage medium for question-and-answer scenarios proposed in this application acquire user question data and send it to a knowledge base for retrieval. The knowledge base stores de-identified documents in the target domain. These de-identified documents are obtained by de-identifying the original documents. During the de-identification process, after acquiring the sensitive data of the original documents, a de-identification algorithm is adaptively determined based on the confidentiality level and data type of the sensitive data. The sensitive data is then de-identified using the algorithm to obtain corresponding target de-identified data. De-identified documents are obtained from the target de-identified data, with at least one type of target de-identified data obtained using a reversible de-identification algorithm. The method receives de-identification retrieval results returned from the knowledge base and locates the corresponding target de-identified data from these results. The target de-identified data is then restored using the corresponding de-identification algorithm, yielding at least the sensitive data corresponding to the reversible de-identification algorithm and restored data corresponding to other target de-identified data. Finally, question-and-answer results corresponding to the question data are obtained based on the de-identification retrieval results, the sensitive data, and the restored data. This application utilizes a knowledge base to store anonymized documents for a specific domain, protecting data privacy while incorporating domain-specific knowledge. During the anonymization process, non-sensitive context is preserved, maintaining the vectorization quality of the original documents and improving retrieval enhancement. Furthermore, the anonymization algorithm is dynamically selected based on the secrecy level and data type, preventing retrieval failures caused by full encryption and avoiding semantic loss associated with static anonymization. Simultaneously, a reversible anonymization algorithm is employed for sensitive data required for retrieval, ensuring the original semantics can be restored subsequently. This allows the knowledge base to participate in vectorized retrieval and support semantic understanding in generating large language models, avoiding semantic defects caused by irreversible anonymization in retrieval enhancement scenarios. This improves the accuracy of retrieval enhancement results, thereby enhancing the accuracy of question-answering results.
[0166] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0167] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0168] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0169] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0170] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0171] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0172] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0173] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0174] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0175] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0176] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A retrieval enhancement method for question-and-answer scenarios, characterized in that, include: The system acquires user question data and sends it to a knowledge base for retrieval. The knowledge base stores de-identified documents in the target domain. These de-identified documents are obtained by de-identifying the original documents. During the de-identification process, after acquiring the sensitive data of the original documents, a de-identification algorithm is adaptively determined based on the confidentiality level and data type of the sensitive data. The sensitive data is then de-identified according to the algorithm to obtain the corresponding target de-identified data. The de-identified documents are obtained based on the target de-identified data. At least one type of target de-identified data is obtained using a reversible de-identification algorithm. Receive the de-identified search results returned by the knowledge base, and locate the corresponding target de-identified data from the de-identified search results; Based on the corresponding desensitization algorithm, the target desensitized data is restored to obtain at least the sensitive data corresponding to the reversible desensitization algorithm and the restored data corresponding to the other target desensitized data. The question and answer results corresponding to the question data are obtained based on the de-identified search results, the sensitive data, and the restored data.
2. The retrieval enhancement method for question-and-answer scenarios according to claim 1, characterized in that, Obtaining sensitive data from the original document includes the following steps: Based on the document type of the original document, the original document is converted into text data; The text data is input into a pre-trained document parsing model for data parsing to obtain at least one sensitive position. Based on the sensitive position, the sensitive data is determined from the original document. The sensitive data includes a corresponding secret level, and the secret level includes at least a high encryption level, a medium encryption level, and a low encryption level.
3. The retrieval enhancement method for question-and-answer scenarios according to claim 2, characterized in that, The step of adaptively determining the desensitization algorithm based on the confidentiality level and data type of the sensitive data includes: For the sensitive data with a high level of encryption, an irreversible desensitization algorithm is selected as the desensitization algorithm based on the data type. For the sensitive data with a security level of medium encryption, a reversible desensitization algorithm is selected as the desensitization algorithm based on the data type and the target domain. For the sensitive data with a low encryption level of secrecy, a generalized desensitization algorithm is selected as the desensitization algorithm based on the data type.
4. The retrieval enhancement method for question-and-answer scenarios according to claim 3, characterized in that, Based on the data type and the target domain, a reversible desensitization algorithm is selected as the desensitization algorithm. The sensitive data is then desensitized using the desensitization algorithm to obtain the corresponding target desensitized data, including: When the data type is a number, obtain the encryption key, and use the encryption key and encryption parameters to perform format-preserving encryption on the sensitive data to obtain the target desensitized data; When the data type is string, obtain the encryption key, obtain the field characteristics of the sensitive data, determine the encrypted field and encryption format of the sensitive data based on the field characteristics, obtain the index data of the encrypted field in the preset dictionary, encrypt the encrypted field with pseudo-random numbers according to the user identifier, the field characteristics, and the encryption key to obtain the corresponding pseudo-random offset, and obtain the target desensitized data based on the pseudo-random offset and the index data; When the data type is an image, the encrypted region and semantic region are determined from the sensitive data according to the target domain. The encrypted region is erased to obtain the encryption key. Based on the encryption key, the semantic region is encrypted to preserve the format, thereby obtaining the target desensitized data.
5. The retrieval enhancement method for question-and-answer scenarios according to claim 4, characterized in that, The process of obtaining the encryption key includes: The process that calls the key generation interface is determined to be a target process with only de-identification and restoration permissions. Based on the target process, the key generation interface is called to generate an initial key. Based on the data type, the initial key is specified for either format-preserving encryption or pseudo-random number encryption to obtain encryption purpose parameters; Determine the rotation strategy for the initial key, and generate the encryption key based on the rotation strategy, the encryption purpose parameter, and the initial key.
6. The retrieval enhancement method for question-and-answer scenarios according to claim 3, characterized in that, The process of restoring the target de-identified data based on the corresponding de-identification algorithm yields at least the sensitive data corresponding to the reversible de-identification algorithm and the restored data corresponding to the other target de-identified data, including: When the target de-identified data is obtained based on the sensitive data with the high encryption level, the corresponding restored data is obtained according to the corresponding de-identification algorithm; When the target de-identified data is obtained based on the sensitive data of the encryption level, the corresponding sensitive data is obtained according to the corresponding de-identification algorithm; When the target de-identified data is obtained based on the sensitive data with the low encryption level, the corresponding restored data is obtained according to the corresponding de-identification algorithm.
7. The retrieval enhancement method for question-and-answer scenarios according to claim 1, characterized in that, The step of obtaining the question-and-answer result corresponding to the question data based on the de-identified search result, the sensitive data, and the restored data includes: Obtain the data other than the target sensitive data from the de-identified search results as the basic search data; The sensitive data and the restored data are filled into the retrieval base data according to the positions of the corresponding target sensitive data to obtain the retrieval results; The search results and the question data are input together into a large language model to generate the question and answer, thus obtaining the question and answer results.
8. A retrieval enhancement device for a question-and-answer scenario, characterized in that, include: Knowledge base construction module: used to acquire user question data and send the question data to the knowledge base for retrieval and query. The knowledge base stores de-identified documents in the target domain. The de-identified documents are obtained by de-identifying the original documents. During the de-identification process, after acquiring the sensitive data of the original documents, the de-identification algorithm is adaptively determined according to the confidentiality level and data type of the sensitive data. The sensitive data is de-identified according to the de-identification algorithm to obtain the corresponding target de-identified data. The de-identified documents are obtained from the target de-identified data. At least one type of target de-identified data is obtained by a reversible de-identification algorithm. Restoration and positioning module: used to receive the de-identified retrieval results returned by the knowledge base, and locate the corresponding target de-identified data from the de-identified retrieval results; Restoration module: used to restore the target de-identified data based on the corresponding de-identification algorithm, to obtain at least the sensitive data corresponding to the reversible de-identification algorithm and the restored data corresponding to the other target de-identified data; Question and answer result generation module: used to obtain the question and answer results corresponding to the question data based on the de-identified search results, the sensitive data, and the restored data.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the retrieval enhancement method for the question-and-answer scenario as described in any one of claims 1 to 7.
10. A storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the retrieval enhancement method for the question-and-answer scenario as described in any one of claims 1 to 7.
Citation Information
Cited By
Role large model memory data privacy protection and authority level-to-level management method and system
CN121765752A
Role large model memory data privacy protection and permission hierarchical management method and system
CN121765752B
Method, device, medium and program product for processing sensitive data
CN121834870A
Data desensitization method based on large language model
CN122197055A