Data processing method and apparatus

WO2026175037A1PCT designated stage Publication Date: 2026-08-27CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/072472
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-19
Filing Date
2026-01-14
Publication Date
2026-08-27

Smart Images

  • Figure CN2026072472_27082026_PF_FP_ABST
    Figure CN2026072472_27082026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a data processing method and apparatus. The data processing method comprises: determining key question information of a target question, then performing information desensitization processing on the key question information to obtain desensitized key question information; performing a search using the desensitized key question information to obtain target reference data, and performing data processing on the target reference data and the target question using a language processing model to obtain an initial question processing result; and performing result desensitization processing on the initial question processing result to determine a target question processing result corresponding to the target question. In this way, data desensitization is performed throughout the entire data processing workflow to avoid the risk of sensitive information leakage and to ensure data security.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing method and apparatus Technical Field

[0001] This disclosure relates to the field of data security technology, and in particular to a data processing method. One or more embodiments of this disclosure also relate to a data processing apparatus, a computing device, an AI gateway, a Q&A system, a computer-readable storage medium, and a computer program product. Background Technology

[0002] With the continuous development of computer technology and artificial intelligence, using neural network models for data processing has become a common method to complete various types of data processing tasks; in this process, data security has also become a key concern.

[0003] Currently, when using neural network models to process data, the data may contain sensitive information. Therefore, how to avoid the leakage of sensitive information and ensure data security has become an urgent problem to be solved. Summary of the Invention

[0004] In view of the above, embodiments of this disclosure provide a data processing method. One or more embodiments of this disclosure also relate to a data processing apparatus, a computing device, an AI gateway, a question-answering system, a computer-readable storage medium, and a computer program product.

[0005] According to a first aspect of the present disclosure, a data processing method is provided, comprising: determining key information of a target problem; performing information desensitization processing on the key information to obtain desensitized key information; using the desensitized key information to perform retrieval to obtain target reference data; and using a language processing model to perform data processing on the target reference data and the target problem to obtain an initial problem processing result; and performing result desensitization processing on the initial problem processing result to determine the target problem processing result corresponding to the target problem.

[0006] According to a second aspect of the present disclosure, a data processing apparatus is provided, comprising: an information desensitization module configured to determine key information of a target problem, perform information desensitization processing on the key information to obtain desensitized key information; a result determination module configured to use the desensitized key information to perform retrieval, obtain target reference data, and use a language processing model to perform data processing on the target reference data and the target problem to obtain an initial problem processing result; and a result desensitization module configured to perform result desensitization processing on the initial problem processing result to determine the target problem processing result corresponding to the target problem.

[0007] According to a third aspect of the present disclosure, a computing device is provided, comprising: a memory and a processor; the memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, wherein the computer programs / instructions, when executed by the processor, implement the steps of the above-described data processing method.

[0008] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the data processing method described above.

[0009] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the data processing method described above.

[0010] According to a sixth aspect of the present disclosure, an AI gateway is provided, which is used to implement the following steps of a data processing method: determining key information of a target problem; performing information anonymization processing on the key information to obtain anonymized key information; using the anonymized key information to perform retrieval to obtain target reference data; and using a language processing model to perform data processing on the target reference data and the target problem to obtain an initial problem processing result; performing result anonymization processing on the initial problem processing result to determine the target problem processing result corresponding to the target problem.

[0011] According to a seventh aspect of the present disclosure, a question-answering system is provided, the system being used to implement the following steps of a data processing method: determining key information of a target question; performing information desensitization processing on the key information to obtain desensitized key information; using the desensitized key information to perform retrieval to obtain target reference data; and using a language processing model to perform data processing on the target reference data and the target question to obtain an initial question processing result; performing result desensitization processing on the initial question processing result to determine the target question processing result corresponding to the target question.

[0012] The data processing method provided in one or more embodiments of this disclosure can perform information anonymization processing on key information of the target problem before data processing using a language processing model, and use the anonymized key information to retrieve target reference data for processing the target problem, thereby reducing the possibility that the reference data contains sensitive information; and after processing the target reference data and the target problem using the language processing model to obtain the initial problem processing result, the initial problem processing result will be anonymized again, thereby avoiding the risk of sensitive information leakage and ensuring data security by anonymizing the entire data processing process. Attached Figure Description

[0013] Figure 1 is a schematic diagram illustrating the application of a data processing method provided in an embodiment of this disclosure.

[0014] Figure 2 is a flowchart of a data processing method provided in an embodiment of this disclosure.

[0015] Figure 3 is a flowchart of a data processing method provided in an embodiment of this disclosure.

[0016] Figure 4 is a schematic diagram of the structure of a data processing device provided in an embodiment of this disclosure.

[0017] Figure 5 is a structural block diagram of a computing device provided in an embodiment of this disclosure. Detailed Implementation

[0018] Numerous specific details are set forth in the following description to provide a full understanding of this disclosure. However, this disclosure can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this disclosure. Therefore, this disclosure is not limited to the specific implementations disclosed below.

[0019] The terminology used in one or more embodiments of this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this disclosure. The singular forms “a,” “the,” and “the” as used in one or more embodiments of this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this disclosure refers to and includes any or all possible combinations of one or more associated listed items.

[0020] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this disclosure, and similarly, second may also be referred to as first. Depending on the context, the word “if” as used herein may be interpreted as “when”, “in response to a determination”, or “when…”.

[0021] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this disclosure are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0022] In one or more embodiments of this disclosure, a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. A large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.

[0023] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0024] First, the terms and concepts involved in one or more embodiments of this disclosure will be explained.

[0025] Large Language Model (LLM): A large-scale neural network model built using deep learning techniques, capable of understanding and generating natural language text. LLMs are pre-trained on massive amounts of text data, possessing broad language understanding and generation capabilities, and are applied in various fields such as text generation, translation, and question answering.

[0026] ε-differential privacy: A system satisfies ε-differential privacy if, for all possible datasets D and D' (pairwise adjacent, i.e., differing by only one element), and for all possible outputs O, P[M(D)=O]≤e ε ×P[M(D')=O], where M(D) in the above formula represents the query operation performed on dataset D; M(D') represents the query operation performed on dataset D'; O represents the query result; P[M(D)=O] represents the probability that query M will get result O on database D; P[M(D')=O] represents the probability that query M will get result O on database D'; ε is a parameter of differential privacy, representing the degree of privacy protection; the smaller the value of ε, the better the privacy protection.

[0027] RAG (Retrieval-augmented Generation): This technique uses retrieval-enhanced generation to infuse relevant documents or knowledge into the generative model, thereby improving the accuracy and relevance of the generated results.

[0028] Laplace Mechanism: A method for implementing differential privacy, which achieves privacy protection by adding noise that follows a Laplace distribution to the query results.

[0029] PII (Personally Identifiable Information): Personally identifiable information refers to information that can be used to identify a specific individual, such as name, address, and telephone number.

[0030] Fine-tuning involves retraining a pre-trained large language model using a large amount of domain-specific text corpus to improve its performance in that domain. Fine-tuning aims to reduce "illusion" phenomena in generated content, making the model's output more accurate and professional.

[0031] logits: refers to the raw output value of the last layer in a deep learning model (such as a large language model). The logits are unnormalized prediction values.

[0032] With the continuous development of computer technology and artificial intelligence, using neural network models for data processing has become a common method to accomplish various types of data processing tasks; in this process, data security has also become a key concern. Currently, in the process of using neural network models to process data, the data may contain sensitive information, posing a risk of data leakage. For example, with the widespread application of Large Language Models (LLM) in various industries, enterprises and organizations are increasingly relying on private domain knowledge to improve service efficiency and user experience. However, private domain knowledge often contains sensitive corporate, product, or personal information; if applied directly to LLM without protection, it may lead to information leakage and privacy risks.

[0033] To address the aforementioned issues, this disclosure provides three solutions. The first solution is a data cleaning solution, which identifies and cleans all sensitive data, removing, replacing, or rewriting its content to achieve data protection. However, this solution has the following drawbacks: 1. High cost: The data cleaning process requires significant time and manpower, especially when processing massive amounts of data, resulting in extremely high costs. 2. Information loss: Over-cleaning may lead to the loss of critical information, affecting the retrieval and generation performance of the RAG system. 3. Maintenance difficulties: As the volume and types of data increase, continuous maintenance and updating of cleaning rules become difficult.

[0034] The second approach is fine-tuning, which further trains the LLM using a large amount of domain-specific text corpus, enabling it to identify and filter sensitive data. However, this approach has drawbacks: 1. High barrier to entry: Fine-tuning requires a large amount of corpus resources, professional AI engineering capabilities, and powerful computing resources (such as GPUs), making it difficult for general users to implement. 2. Risk of overfitting: Over-fine-tuning may lead to excellent model performance in a specific domain, but a decrease in generalizability and intelligence. 3. Inconvenient updates: After fine-tuning, if the model needs to adapt to new sensitive data or domains, it needs to be fine-tuned again, increasing maintenance difficulty.

[0035] The third approach is access control and permission management, which restricts user access to and manipulation of sensitive data through strict access control and permission management. However, this approach has the following drawbacks: 1. Limited user experience: Overly strict permission restrictions may affect the user's normal use and query experience. 2. Complex implementation: It requires meticulous permission division and dynamic management, increasing the system's complexity and maintenance costs. 3. Potential vulnerabilities: If the permission management mechanism has vulnerabilities, it may still lead to the leakage of sensitive information.

[0036] Based on this, a data processing method is provided in this disclosure. This disclosure also relates to a data processing apparatus, a computing device, and a computer-readable storage medium, which will be described in detail in the following embodiments.

[0037] Referring to Figure 1, which illustrates an application diagram of a data processing method according to an embodiment of this disclosure, as shown in Figure 1, a user can send a question (i.e., the target question) to a server 104 via a client 102. Upon receiving the question, the server 104 determines the corresponding question features (i.e., key question information) and adds noise data to these features to obtain de-identified question features (i.e., de-identified key question information). Then, based on the de-identified question post, a search is performed to obtain the corresponding reference text (i.e., target reference data). A large language model is used to process the question and reference text to obtain the initial answer (i.e., the initial question processing result). By de-identifying the initial answer, the target answer (i.e., the target question processing result) is obtained. After obtaining the target answer, the server 104 sends it to the client 102 for display, thereby ensuring data security while answering the user's question.

[0038] Referring to Figure 2, Figure 2 shows a flowchart of a data processing method provided according to an embodiment of the present disclosure, which specifically includes the following steps.

[0039] Step 202: Determine the key information of the target problem, and perform information desensitization processing on the key information to obtain desensitized key information.

[0040] Here, the target problem can be understood as a problem that needs to be processed using a language processing model, such as a user query sent by the client or a question raised by the user; the key information of the problem can be understood as information that represents the key content of the target problem, such as retrieval features (retrieval feature vectors) or matrices obtained by feature extraction of the target problem; or the key information of the problem can be keywords in the target problem.

[0041] Information desensitization can be understood as the process of removing sensitive information contained in key information of a problem; desensitizing key information of a problem can be understood as adding noisy data to key information of a problem; or, replacing or deleting sensitive information in key information of a problem.

[0042] In one or more embodiments provided in this disclosure, the key information of the problem is a retrieval feature; determining the key information of the target problem and performing information desensitization processing on the key information of the problem to obtain desensitized key information of the problem includes: extracting features from the target problem to obtain the retrieval feature of the target problem, and adding noise data to the retrieval feature to obtain desensitized retrieval feature.

[0043] Taking the application of the data processing method provided in this disclosure in the RAG process information protection scenario as an example, the data processing method is explained; specifically, the way this method desensitizes search features is to apply differential privacy to data retrieval, and the specific steps are as follows.

[0044] 1. Add noise to search features.

[0045] This method can perform feature transformation on user query to obtain a retrieval feature vector, and add noise data following a Laplace distribution to the retrieval feature vector to obtain a de-identified retrieval feature vector, thus protecting query privacy.

[0046] In one or more embodiments provided in this disclosure, before determining the key information of the target problem, the method further includes: receiving the target problem sent by the client, wherein the target problem is sent by the client when the user performs a problem input operation based on the problem processing interface.

[0047] The data processing method provided in this disclosure can be applied to the server side and can receive target questions sent by the client.

[0048] The problem handling interface can be understood as a human-computer interaction interface used to handle the target problem. For example, the problem handling interface can be a webpage, application interface, etc. The problem input operation can be understood as an operation performed based on the problem handling interface to input the target problem. The problem input operation can be an operation to input the target problem into the client through a data processing device (mouse, keyboard, touch screen, microphone, etc.).

[0049] Using the previous example, users can input their query into the client through the query processing interface displayed on the client. After receiving the user's query, the client will send it to the server where this data processing method is applied, thereby resolving the user's question through flexible human-computer interaction.

[0050] Step 204: Use the key information of the de-identified problem to retrieve target reference data, and use a language processing model to process the target reference data and the target problem to obtain the initial problem processing result.

[0051] Following the previous example, this method requires data collection and preprocessing before retrieving the target reference data. The specific execution method is as follows.

[0052] 1. Identify sensitive data: By combining automated tools with manual review, identify user data that needs protection (i.e., reference documents and reference data), including query content, document content, etc.

[0053] 2. Data Partitioning: Partition the dataset according to data sensitivity to ensure differentiated protection of data with different sensitivity levels in subsequent processing.

[0054] In one or more embodiments provided in this disclosure, the key information of the de-identification problem is de-identified keywords; the step of using the key information of the de-identification problem to retrieve target reference data includes: retrieving reference data similar to the de-identified keywords from multiple reference data in the data storage unit, and using the reference data similar to the de-identified keywords as target reference data.

[0055] In this context, the desensitized keywords can be understood as the operation of replacing, covering, or removing sensitive information (such as sensitive words) in keywords, or the operation of adding interfering words to keywords.

[0056] In one or more embodiments provided in this disclosure, the target reference data is target reference text; the step of using the key information of the de-identification problem to retrieve the target reference data includes: determining the text features of multiple reference texts in the data storage unit, and determining the feature similarity between each text feature and the retrieval feature; sorting the multiple reference texts in descending order based on the feature similarity to obtain a reference text sequence; and selecting a preset number of reference texts from the reference text sequence as the target reference text in a top-down manner.

[0057] The reference text can be understood as textual information used by the language processing model to process the target problem. For example, the reference text can be external knowledge text, reference documents (such as papers, books, etc.), web page content, etc. The target reference text can be understood as reference text associated with the target problem; for example, reference text similar to the target problem. The feature similarity can be understood as information used to represent the degree of similarity between text features and retrieval features. For example, the feature similarity can be cosine similarity or distance similarity. The preset number of texts can be set according to the actual application scenario; for example, the preset number of texts can be 10 or 5.

[0058] Following the previous example, this method can be applied to data retrieval using differential privacy. The specific steps are as follows.

[0059] 1. Utilize the de-identified search features to retrieve multiple reference documents from the database.

[0060] Determine the document features corresponding to multiple documents in the database, and calculate the feature similarity between the document features and the de-identified retrieval features; the feature similarity can be the cosine similarity or distance similarity between two feature vectors.

[0061] 2. Differential privacy implementation for top k selection.

[0062] Based on the above, when extracting the top k results (i.e., reference documents) from the database, this method first introduces noisy data into their retrieval features and performs document retrieval to obtain multiple reference documents. Then, to ensure the accuracy and effectiveness of the documents, this method can select target reference documents that are more relevant to the user's query from the multiple reference documents. Specifically, firstly, based on the document score (i.e., feature similarity), the multiple reference documents are sorted in descending order to obtain a document sequence; secondly, the top K results (i.e., a preset number of reference texts) are selected from the document sequence in a top-down manner, thereby ensuring the differential privacy nature of the selection process.

[0063] As can be seen from the above embodiments, this method can select target reference texts that are similar to the target problem from multiple reference texts, thereby facilitating the subsequent speech processing model to accurately process the target problem based on the target parameter text; avoiding the problem of inaccurate problem processing caused by insufficient knowledge learned by the speech processing model.

[0064] In one or more embodiments provided in this disclosure, the step of using a language processing model to process the target reference data and the target question to obtain an initial question processing result includes: using a feature extraction model to extract features from the target reference data and the target question to obtain initial model input features, and adding noise data to the initial model input features to obtain de-identified model input features; inputting the de-identified model input features into the language processing model, and performing character reasoning based on the de-identified model input features in the language processing model to obtain multiple candidate characters; determining any one of the multiple candidate characters as the target character, and continuing character reasoning based on the target character to obtain the initial question processing result.

[0065] The feature extraction model can be understood as a model used to extract features from the target reference data and the target problem. For example, the feature extraction model can be an embedding model, a data preprocessing model, an embedding network layer, etc.

[0066] A language processing model can be understood as a model capable of performing language reasoning. For example, a language processing model can be a deep learning model, a large model, a large language model (LLM), etc.

[0067] The initial problem processing result can be understood as the problem processing result output by the speech processing model after performing language reasoning based on the target reference data and the target problem; for example, the initial problem processing result can be feature vectors, logits, etc.

[0068] Following the previous example, in order to improve data security, this method requires input privacy protection before inputting data into a large model for processing, thereby protecting the private or sensitive data in the input data.

[0069] The privacy protection method can be privacy embedding perturbation. Specifically, this method can introduce noise into the embedding vector obtained in the retrieval stage by generating Gaussian noise with the same embedding dimension as the embedding vector and adding independent noise in each dimension, ensuring that the result of the generation process does not change significantly even if a single input is replaced. The specific implementation method is as follows.

[0070] 1. Preprocess or embed the input data to obtain the embedding vector.

[0071] For complex text tasks, word embedding layers, convolutional neural networks (CNNs), and recurrent neural networks (RNNs) can be used to extract features from the text. These high-level features are then passed to a large language model for processing, thereby improving the performance of the large language model. During this process, to ensure data security, noise can be added to the extracted features; the specific implementation is as follows.

[0072] First, the user query and the retrieved target reference documents are input into the text feature extraction module for preprocessing or embedding to obtain the embedding vector corresponding to the input data.

[0073] The text feature extraction module can be a word embedding layer, a contextualized embedding model, a convolutional neural network, a recurrent neural network, etc.

[0074] Secondly, for the embedded vector, independent noise data is added to each vector dimension to ensure that the result of the generation process does not change significantly even if a single input is replaced.

[0075] Based on the above, it can be seen that this method uses the query fuzzification operation to fuzzify the input data obtained from information retrieval, so that it does not fully reflect the user's true intention and prevents the leakage of sensitive information.

[0076] 2. Input the noise-added embedding vector into the large language model for processing to obtain the output of the large language model.

[0077] In this method, the large language model employs randomized output operations during the generation of output results (i.e., initial problem processing results). Specifically, this method introduces randomness into the decoding stage of the generative model (i.e., LLM). For multiple words (i.e. target characters) with different probabilities, a random sampling method is used to select the next word, rather than selecting the next word with the highest probability. Through continuous iterative reasoning, the final output result (text features) is obtained, thereby increasing the diversity of generated text and reducing the risk of sensitive information leakage.

[0078] Step 206: Perform result anonymization processing on the initial problem processing results to determine the target problem processing results corresponding to the target problem.

[0079] The result of the target question processing can be understood as the target answer or response content corresponding to the target question; the response content can be a prompt message or prompt content, such as a prompt that the question cannot be answered or the question processing failed.

[0080] The process of desensitizing the initial problem processing results can be understood as replacing, covering, removing, or adding interfering words to sensitive information in the initial problem processing results.

[0081] It should be noted that in one or more embodiments of this disclosure, the replacement of sensitive information can be carried out by replacing from back to front, thereby avoiding the problem of semantic inaccuracy caused by changes in position.

[0082] In one or more embodiments provided in this disclosure, the initial question processing result is an initial answer feature; the step of performing result desensitization processing on the initial question processing result to determine the target question processing result corresponding to the target question includes: adding noise data to the initial answer feature to obtain desensitized answer feature; and determining the target answer corresponding to the target question based on the desensitized answer feature.

[0083] The initial answer features can be understood as the text features output by the language processing model, such as logits or embedding vectors.

[0084] Following the previous example, corresponding noise data is added to the text features output by the large language model (i.e., the language processing model) to perform data anonymization on the text features, thereby obtaining the anonymized text features; based on the anonymized text features, the answer corresponding to the user query is obtained; thus ensuring data security and avoiding the leakage of sensitive information.

[0085] In one or more embodiments provided in this disclosure, determining the target answer corresponding to the target question based on the desensitized answer features includes: using a feature processing model to perform feature transformation on the desensitized answer features to obtain the target answer corresponding to the target question.

[0086] The feature processing model can be understood as a model that converts features into text. For example, the feature processing model can be a word embedding layer, a model output layer, etc.

[0087] Following the previous example, this method inputs the anonymized text features into the word embedding layer for feature transformation to obtain the answer corresponding to the user's query, thereby accurately obtaining the result required by the user.

[0088] In one or more embodiments of this disclosure, determining the target answer corresponding to the target question based on the desensitized answer features includes: determining the initial answer corresponding to the target question based on the desensitized answer features, and desensitizing sensitive information in the initial answer to obtain the target answer.

[0089] The initial answer can be understood as the initial answer text corresponding to the target question, which is obtained by feature transformation of the anonymized answer features. The target answer can be understood as the answer text obtained after anonymizing the initial answer.

[0090] Following the previous example, this method can protect the privacy of the output results during the output process. Specifically, this is achieved through output screening and correction operations: First, an algorithm is used to automatically detect and correct sensitive information in the generated text, such as using the PII detection algorithm; second, the detected sensitive information is modified, replaced, or removed to obtain the desensitized answer, thereby answering the user's questions while ensuring data security.

[0091] In one or more embodiments provided in this disclosure, after performing result anonymization processing on the initial problem processing result and determining the target problem processing result corresponding to the target problem, the method further includes: sending the target problem processing result to the client.

[0092] Specifically, this method, after obtaining the result of the target problem resolution, can send the result to the client, and display the result to the user through the client's problem resolution interface, thereby answering the user's target problem and meeting the customer's needs.

[0093] The data processing method provided in one or more embodiments of this disclosure can perform information anonymization processing on key information of the target problem before data processing using a language processing model, and use the anonymized key information to retrieve target reference data for processing the target problem, thereby reducing the possibility that the reference data contains sensitive information; and after processing the target reference data and the target problem using the language processing model to obtain the initial problem processing result, the initial problem processing result will be anonymized again, thereby avoiding the risk of sensitive information leakage and ensuring data security by anonymizing the entire data processing process.

[0094] The following description, in conjunction with Figure 3, uses the application of the data processing method provided in this disclosure in a RAG process information protection scenario as an example to further illustrate the data processing method. Figure 3 shows a flowchart of the processing procedure of a data processing method provided in an embodiment of this disclosure, specifically including the following steps.

[0095] Step 302: Differential privacy to data retrieval.

[0096] Specifically, in the data retrieval stage, this method can apply differential privacy to data retrieval. The specific execution steps may include: adding noise to the query results and implementing differential privacy for top k selection.

[0097] Adding noise to query results refers to introducing noisy data to disturb the search results during the process of retrieving reference documents. The specific implementation method is as follows.

[0098] 1. Add noise to search features.

[0099] The system can perform feature transformation on user queries to obtain retrieval features, and then add noisy data that follows a Laplace distribution to these retrieval features to obtain de-identified retrieval features.

[0100] 2. Utilize the de-identified search features to retrieve multiple reference documents from the database.

[0101] Determine the document features corresponding to multiple documents in the database, and calculate the feature similarity between the document features and the de-identified retrieval features; the feature similarity can be the cosine similarity between two feature vectors.

[0102] Among them, the differential privacy implementation of top k selection refers to selecting multiple target reference documents from multiple reference documents.

[0103] Specifically, based on the above, when extracting the top k results (i.e., reference documents) from the database, this method first introduces noisy data into their retrieval features and performs document retrieval to obtain multiple reference documents. Then, to ensure the accuracy and effectiveness of the documents, this method can select target reference documents that are more relevant to the user's query from among the multiple reference documents. The specific method is as follows: First, based on the document score (i.e., feature similarity), the multiple reference documents are sorted in descending order to obtain a document sequence; second, the top K results are selected from the document sequence in a top-down manner, thereby ensuring the differential privacy nature of the selection process.

[0104] Step 304: Enter privacy protection information.

[0105] This method can protect the privacy of input data during the input protection phase; specifically, it achieves this through privacy embedding perturbation operations, thereby achieving the effect of query fuzzification.

[0106] Specifically, in order to improve data security, this method requires input privacy protection before inputting data into a large model for processing, thereby protecting private or sensitive data in the input data.

[0107] This privacy protection method can be privacy-preserving embedding perturbation. This method introduces noise into the embedding vector obtained in the retrieval stage, adding independent noise in each dimension to ensure that even if a single input is replaced, the result of the generation process will not change significantly. The specific implementation method is as follows.

[0108] 1. Preprocess or embed the input data to obtain the embedding vector.

[0109] For complex text tasks, word embedding layers, convolutional neural networks (CNNs), and recurrent neural networks (RNNs) can be used to extract features from the text. These high-level features are then passed to a large language model for processing, thereby improving the performance of the large language model. During this process, to ensure data security, noise can be added to the extracted features; the specific implementation is as follows.

[0110] First, the user query and the retrieved target reference documents are input into the text feature extraction module for preprocessing or embedding to obtain the embedding vector corresponding to the input data.

[0111] The text feature extraction module can be a word embedding layer, a contextualized embedding model, a convolutional neural network, a recurrent neural network, etc.

[0112] Secondly, for the embedded vector, independent noise data is added to each vector dimension to ensure that the result of the generation process does not change significantly even if a single input is replaced.

[0113] Based on the above, it can be seen that this method achieves query fuzzification through the above operations, and performs fuzzification processing on the input data obtained from information retrieval so that it does not fully reflect the user's true intention and prevents the leakage of sensitive information.

[0114] 2. Input the noise-added embedding vector into the large language model for processing to obtain the output of the large language model.

[0115] In this method, the large language model employs randomized output operations during the generation of output results. Specifically, this method introduces randomness into the decoding stage of the generative model (i.e., LLM). For multiple words (i.e. target characters) with different probabilities, a random sampling method is used to select the next word, rather than selecting the next word with the highest probability. Through continuous iterative reasoning, the final output result (text features) is obtained, thereby increasing the diversity of generated text and reducing the risk of sensitive information leakage.

[0116] Step 306: Protect the privacy of the output results.

[0117] This method can protect the privacy of the output results during the output protection phase; the specific execution steps include: input screening and correction, and randomization of output operations.

[0118] The randomized output operation refers to the randomized output operation adopted by the large language model in the process of generating output results. The specific execution method is as follows: the embedding vector with added noise is input into the large language model, and randomness is introduced into the decoding stage of the generative model (i.e., LLM); for multiple words (i.e. target characters) with different probabilities, the next word is selected by random sampling, rather than selecting the next word with the highest probability; through continuous iterative reasoning, the final output result (text features) is obtained, thereby increasing the diversity of generated text and reducing the risk of sensitive information leakage.

[0119] Input screening and correction refers to the desensitization process performed on the output of the large language model.

[0120] First, the text features output by the large language model are input into the word embedding layer for feature transformation to obtain the answer corresponding to the user query.

[0121] Secondly, an algorithm is used to automatically detect and correct sensitive information in the generated text (i.e., the answer), and to modify or replace the detected sensitive information to obtain a desensitized answer. For example, this algorithm could be a PII detection algorithm.

[0122] It should be noted that for noisy data introduced in the RAG process, this method will perform differential privacy budget management. The specific steps of differential privacy budget management include: ε-budget allocation and privacy loss tracking.

[0123] Here, ε-budget allocation refers to the reasonable allocation of the ε budget between the retrieval and generation stages to achieve a balance between privacy protection and system performance. For example, a portion of ε can be allocated to the retrieval stage, and another portion to the generation stage. It should be noted that the noise data in one or more embodiments of this disclosure is determined based on the allocated budget; the budgets for different noise data can be different or the same.

[0124] Privacy loss tracking refers to continuously assessing and tracking privacy losses throughout the entire process to ensure that the total privacy loss does not exceed the preset ε budget and meets the requirements of differential privacy.

[0125] Based on the above steps, the data processing method in one or more embodiments of this disclosure provides a differential privacy-based RAG process information protection scheme. This scheme, through differential privacy-based RAG process information protection, ensures that while leveraging private domain knowledge to enhance generation capabilities, it effectively protects sensitive information and improves enterprise data security and user trust.

[0126] Furthermore, regarding the issue of "privacy protection of private domain knowledge," this method can effectively infuse private domain knowledge into the generative model during the RAG process without disclosing original sensitive data. Regarding the issue of "information leakage risk," this method can prevent users from reverse-engineering sensitive information in the private domain from the generated results. Regarding the issue of "balancing system performance and privacy protection," this method can ensure privacy protection without significantly reducing the system's retrieval and generation performance.

[0127] Compared to data cleaning solutions, this method avoids directly modifying the original data by introducing differential privacy in the retrieval and generation stages, reducing the risk of information loss. At the same time, it is highly automated, reducing labor and time costs.

[0128] Compared to model fine-tuning schemes, this method does not require fine-tuning of the LLM, which lowers the implementation threshold and computational resource requirements, avoids overfitting problems, and maintains the model's versatility and intelligence.

[0129] Compared to access control schemes, this method adds a privacy protection mechanism at the data processing level, does not rely on external permission management, and enhances the overall system security and robustness.

[0130] In summary, the technical effects achieved by this method can include: full-stage coverage of differential privacy in the RAG process, no additional token consumption required by LLM, privacy perturbation at the embedding level, dynamic privacy budget management, and multi-layered output filtering mechanism.

[0131] Differential privacy across all stages of the RAG workflow refers to integrating differential privacy technology into all aspects of data retrieval, input processing, and output generation to achieve end-to-end privacy protection. Obfuscating the input during processing can also help prevent reverse eavesdropping attacks to some extent.

[0132] No additional token consumption required by LLM: This means that the intervention of differential privacy technology does not require additional calls to LLM to generate token consumption, nor does it require LLM to generate multiple random results.

[0133] Privacy perturbation at the embedding level refers to introducing noise at the embedding vector level to ensure that even if sensitive information is retrieved, its high-dimensional representation is difficult to recover in reverse.

[0134] Dynamic privacy budget management refers to dynamically adjusting the allocation of the ε budget based on system needs and privacy requirements to achieve a balance between privacy protection and system performance.

[0135] Multi-layered output filtering mechanism: This refers to the combination of algorithm screening and randomization generation to ensure that the generated content does not contain sensitive information, thereby improving the reliability of privacy protection.

[0136] Corresponding to the above method embodiments, this disclosure also provides a data processing device embodiment. Figure 4 shows a schematic diagram of the structure of a data processing device provided in one embodiment of this disclosure. As shown in Figure 4, the device includes: an information desensitization module 402, configured to determine key information of a target problem, perform information desensitization processing on the key information to obtain desensitized key information; a result determination module 404, configured to use the desensitized key information to perform retrieval, obtain target reference data, and use a language processing model to perform data processing on the target reference data and the target problem to obtain an initial problem processing result; and a result desensitization module 406, configured to perform result desensitization processing on the initial problem processing result to determine the target problem processing result corresponding to the target problem.

[0137] Optionally, the key information of the problem is a retrieval feature; the information desensitization module 402 is further configured to: extract features from the target problem to obtain the retrieval feature of the target problem, and add noise data to the retrieval feature to obtain desensitized retrieval feature.

[0138] Optionally, the target reference data is target reference text; the result determination module 404 is further configured to: determine the text features of multiple reference texts in the data storage unit, and determine the feature similarity between each text feature and the retrieval feature; based on the feature similarity, sort the multiple reference texts in descending order to obtain a reference text sequence; and select a preset number of reference texts from the reference text sequence as the target reference text in a top-down manner.

[0139] Optionally, the result determination module 404 is further configured to: use a feature extraction model to extract features from the target reference data and the target question to obtain initial model input features, and add noise data to the initial model input features to obtain desensitized model input features; input the desensitized model input features into the language processing model, and perform character reasoning based on the desensitized model input features in the language processing model to obtain multiple candidate characters; determine any one of the multiple candidate characters as the target character, and continue character reasoning based on the target character to obtain the initial question processing result.

[0140] Optionally, the initial question processing result is an initial answer feature; the result desensitization module 406 is further configured to: add noise data to the initial answer feature to obtain desensitized answer feature; and determine the target answer corresponding to the target question based on the desensitized answer feature.

[0141] Optionally, the result desensitization module 406 is further configured to: use a feature processing model to perform feature transformation on the desensitized answer features to obtain the target answer corresponding to the target question.

[0142] Optionally, the result desensitization module 406 is further configured to: determine the initial answer corresponding to the target question based on the desensitized answer features, and perform desensitization processing on the sensitive information in the initial answer to obtain the target answer.

[0143] Optionally, the data processing device further includes a problem receiving module, configured to: receive the target problem sent by the client, wherein the target problem is sent by the client when the user performs a problem input operation based on the problem processing interface; the data processing device further includes a result sending module, configured to: send the processing result of the target problem to the client.

[0144] The data processing apparatus provided in one or more embodiments of this disclosure can perform information desensitization processing on key information of the target problem before data processing using a language processing model, and use the desensitized key information to retrieve target reference data for processing the target problem, thereby reducing the possibility that the reference data contains sensitive information; and after processing the target reference data and the target problem using the language processing model to obtain the initial problem processing result, the initial problem processing result will be desensitized again, thereby desensitizing the data throughout the entire data processing process, avoiding the risk of sensitive information leakage and ensuring data security.

[0145] The above is an illustrative scheme of a data processing apparatus according to this embodiment. It should be noted that the technical solution of this data processing apparatus and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the data processing apparatus, please refer to the description of the technical solution of the data processing method described above.

[0146] Figure 5 shows a structural block diagram of a computing device 500 according to an embodiment of the present disclosure. The components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and a database 550 is used to store data.

[0147] The computing device 500 also includes an access device 540, which enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 540 may include one or more of any type of wired or wireless network interface (e.g., a network interface controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.

[0148] In one embodiment of this disclosure, the aforementioned components of the computing device 500, as well as other components not shown in FIG. 5, may also be connected to each other, for example, via a bus. It should be understood that the computing device structural block diagram shown in FIG. 5 is merely for illustrative purposes and is not intended to limit the scope of this disclosure. Those skilled in the art can add or replace other components as needed.

[0149] Computing device 500 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). Computing device 500 can also be a mobile or stationary server.

[0150] The processor 520 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-described data processing method.

[0151] The various embodiments in this disclosure are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computing device embodiments are basically similar to the data processing method embodiments, so the description is relatively simple; relevant parts can be referred to the description of the data processing method embodiments.

[0152] This disclosure also provides an AI gateway in one or more embodiments, the AI ​​gateway being used to implement the following steps of a data processing method: determining key information of a target problem; performing information anonymization processing on the key information to obtain anonymized key information; using the anonymized key information to perform retrieval to obtain target reference data; and using a language processing model to process the target reference data and the target problem to obtain an initial problem processing result; performing result anonymization processing on the initial problem processing result to determine the target problem processing result corresponding to the target problem.

[0153] The AI ​​gateway can be understood as a RAG plugin. In data anonymization application scenarios, if customers need to reference knowledge bases (i.e., data storage units) or knowledge (i.e., reference text) that have privacy risks, the AI ​​gateway can be used as an auxiliary security enhancement method to improve the product's security experience.

[0154] The AI ​​gateway provided in one or more embodiments of this disclosure can perform information anonymization processing on key information of the target problem before using a language processing model for data processing, and use the anonymized key information to retrieve target reference data for processing the target problem, thereby reducing the possibility that the reference data contains sensitive information; and after using the language processing model to process the target reference data and the target problem to obtain the initial problem processing result, the initial problem processing result will be anonymized again, thereby avoiding the risk of sensitive information leakage and ensuring data security by anonymizing the data throughout the entire data processing process.

[0155] The above is an illustrative scheme of an AI gateway according to this embodiment. It should be noted that the technical solution of this AI gateway and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the AI ​​gateway, please refer to the description of the technical solution of the data processing method described above.

[0156] This disclosure also provides a question-answering system in one or more embodiments, which implements the following steps of a data processing method: determining key information of a target question; performing information desensitization processing on the key information to obtain desensitized key information; using the desensitized key information to perform retrieval to obtain target reference data; and using a language processing model to process the target reference data and the target question to obtain an initial question processing result; performing result desensitization processing on the initial question processing result to determine the target question processing result corresponding to the target question.

[0157] The Q&A system can be understood as a system for answering questions or making diagnoses. For example, it can be an internal expert Q&A system or a diagnostic system. Considering that much internal knowledge and experience data (i.e., reference text) contains certain sensitive information and is not suitable for direct external exposure, the Q&A system can be used to strengthen the protection of private knowledge when internal knowledge and experience data are referenced by external intelligent agents (i.e., language processing models).

[0158] The language processing model can be a speech processing model configured on devices other than the Q&A system.

[0159] The Q&A system provided in one or more embodiments of this disclosure can perform information anonymization processing on key information of the target question before data processing using a language processing model. It can then use this anonymized key information to retrieve target reference data for processing the target question, thereby reducing the possibility that the reference data contains sensitive information. Furthermore, after processing the target reference data and the target question using the language processing model to obtain the initial question processing result, the initial question processing result is anonymized again. This process of anonymizing data throughout the entire data processing flow avoids the risk of sensitive information leakage and ensures data security.

[0160] The above is an illustrative scheme of a question-answering system according to this embodiment. It should be noted that the technical solution of this question-answering system and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the question-answering system, please refer to the description of the technical solution of the data processing method described above.

[0161] An embodiment of this disclosure also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described data processing method.

[0162] The various embodiments in this disclosure are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computer-readable storage medium embodiments are basically similar to the data processing method embodiments, so the description is relatively simple; relevant parts can be referred to the description of the data processing method embodiments.

[0163] An embodiment of this disclosure also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described data processing method.

[0164] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the data processing method described above.

[0165] The foregoing has described specific embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0166] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0167] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this disclosure are not limited to the described order of actions, because according to the embodiments of this disclosure, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this disclosure.

[0168] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0169] The preferred embodiments disclosed above are merely illustrative of this disclosure. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments of this disclosure. These embodiments are selected and specifically described in this disclosure to better explain the principles and practical applications of the embodiments of this disclosure, thereby enabling those skilled in the art to better understand and utilize this disclosure. This disclosure is limited only by the claims and their full scope and equivalents.

Claims

1. A data processing method, comprising: Identify the key information of the target problem, and perform information desensitization processing on the key information to obtain desensitized key information of the problem; The target reference data is obtained by using the key information of the de-identified problem, and the target reference data and the target problem are processed by the language processing model to obtain the initial problem processing result. The initial problem processing results are anonymized to determine the target problem processing result corresponding to the target problem.

2. The data processing method according to claim 1, wherein the key information of the problem is a retrieval feature; The process of identifying key information about the target problem, and then performing information anonymization processing on that key information to obtain anonymized key information, includes: Feature extraction is performed on the target question to obtain the retrieval features of the target question, and noisy data is added to the retrieval features to obtain de-identified retrieval features.

3. The data processing method according to claim 2, wherein the target reference data is target reference text; The process of retrieving target reference data using key information related to the de-identification issue includes: The text features of multiple reference texts in the data storage unit are determined, and the feature similarity between each text feature and the retrieval feature is determined. Based on the feature similarity, the multiple reference texts are sorted in descending order to obtain a reference text sequence; Select a predetermined number of reference texts from the reference text sequence as the target reference texts in a top-down manner.

4. The data processing method according to any one of claims 1 to 3, wherein the step of using a language processing model to process the target reference data and the target problem to obtain an initial problem processing result includes: Using a feature extraction model, features are extracted from the target reference data and the target problem to obtain initial model input features, and noisy data is added to the initial model input features to obtain desensitized model input features; The de-identification model input features are input into the language processing model, and character reasoning is performed in the language processing model based on the de-identification model input features to obtain multiple candidate characters; The candidate character among the plurality of candidate characters is determined as the target character, and character reasoning is continued based on the target character to obtain the initial problem processing result.

5. The data processing method according to any one of claims 1 to 3, wherein the initial question processing result is an initial answer feature; The step of desensitizing the initial problem processing results to determine the target problem processing result corresponding to the target problem includes: Add noisy data to the initial answer features to obtain de-identified answer features; Based on the characteristics of the anonymized answer, the target answer corresponding to the target question is determined.

6. The data processing method according to claim 5, wherein determining the target answer corresponding to the target question based on the de-identified answer features includes: The desensitized answer features are transformed using a feature processing model to obtain the target answer corresponding to the target question.

7. The data processing method according to claim 5, wherein determining the target answer corresponding to the target question based on the de-identified answer features includes: Based on the desensitized answer features, the initial answer corresponding to the target question is determined, and the sensitive information in the initial answer is desensitized to obtain the target answer.

8. The data processing method according to any one of claims 1 to 3, further comprising, before determining the key information of the target problem: The client receives the target question sent by the client, wherein the target question is sent by the client when the user performs a question input operation based on the question processing interface; After performing result anonymization processing on the initial problem processing results to determine the target problem processing result corresponding to the target problem, the method further includes: The result of the target problem processing is sent to the client.

9. A data processing apparatus, comprising: The information desensitization module is configured to determine the key information of the target problem, perform information desensitization processing on the key information of the problem, and obtain desensitized key information of the problem. The result determination module is configured to retrieve target reference data by using the key information of the de-identified problem, and to process the target reference data and the target problem by using a language processing model to obtain the initial problem processing result. The result desensitization module is configured to perform result desensitization processing on the initial problem processing result to determine the target problem processing result corresponding to the target problem.

10. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which perform the following operations when executed by the processor: Identify the key information of the target problem, and perform information desensitization processing on the key information to obtain desensitized key information of the problem; The target reference data is obtained by using the key information of the de-identified problem, and the target reference data and the target problem are processed by the language processing model to obtain the initial problem processing result. The initial problem processing results are anonymized to determine the target problem processing result corresponding to the target problem.

11. The computing device according to claim 10, wherein the key information of the problem is a retrieval feature; The process of identifying key information about the target problem, and then performing information anonymization processing on that key information to obtain anonymized key information, includes: Feature extraction is performed on the target question to obtain the retrieval features of the target question, and noisy data is added to the retrieval features to obtain de-identified retrieval features.

12. The computing device according to claim 11, wherein the target reference data is target reference text; The process of retrieving target reference data using key information related to the de-identification issue includes: The text features of multiple reference texts in the data storage unit are determined, and the feature similarity between each text feature and the retrieval feature is determined. Based on the feature similarity, the multiple reference texts are sorted in descending order to obtain a reference text sequence; Select a predetermined number of reference texts from the reference text sequence as the target reference texts in a top-down manner.

13. The computing device according to any one of claims 10 to 12, wherein the step of using a language processing model to process the target reference data and the target problem to obtain an initial problem processing result includes: Using a feature extraction model, features are extracted from the target reference data and the target problem to obtain initial model input features, and noisy data is added to the initial model input features to obtain desensitized model input features; The de-identification model input features are input into the language processing model, and character reasoning is performed in the language processing model based on the de-identification model input features to obtain multiple candidate characters; The candidate character among the plurality of candidate characters is determined as the target character, and character reasoning is continued based on the target character to obtain the initial problem processing result.

14. The computing device according to any one of claims 10 to 12, wherein the initial problem processing result is an initial answer feature; The step of desensitizing the initial problem processing results to determine the target problem processing result corresponding to the target problem includes: Add noisy data to the initial answer features to obtain de-identified answer features; Based on the characteristics of the anonymized answer, the target answer corresponding to the target question is determined.

15. The computing device according to claim 14, wherein determining the target answer corresponding to the target question based on the de-identified answer features comprises: The desensitized answer features are transformed using a feature processing model. Obtain the target answer to the target question.

16. The computing device according to claim 14, wherein determining the target answer corresponding to the target question based on the de-identified answer features includes: Based on the desensitized answer features, the initial answer corresponding to the target question is determined, and the sensitive information in the initial answer is desensitized to obtain the target answer.

17. A computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 8.

18. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 8.

19. An AI gateway, the AI ​​gateway being used to implement the following steps of a data processing method: Identify the key information of the target problem, and perform information desensitization processing on the key information to obtain desensitized key information of the problem; The target reference data is obtained by using the key information of the de-identified problem, and the target reference data and the target problem are processed by the language processing model to obtain the initial problem processing result. The initial problem processing results are anonymized to determine the target problem processing result corresponding to the target problem.

20. A question-and-answer system, wherein the question-and-answer system is used to implement the following steps of a data processing method: Identify the key information of the target problem, and perform information desensitization processing on the key information to obtain desensitized key information of the problem; The target reference data is obtained by using the key information of the de-identified problem, and the target reference data and the target problem are processed by the language processing model to obtain the initial problem processing result. The initial problem processing results are anonymized to determine the target problem processing result corresponding to the target problem.