Private data question and answer method and system based on knowledge graph
By combining anonymization processing based on knowledge graphs and graph neural networks, the problem of insufficient knowledge adaptation of private data question-answering systems in vertical fields has been solved, efficient query and secure question-answering of private data have been achieved, and the accuracy and professionalism of the answers have been improved.
Patent Information
- Application Number
- CN202510661265.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-09-05
AI Technical Summary
Existing private data question-answering systems lack deep adaptation to vertical domain knowledge, have weak logical correlations between texts, have weak reasoning capabilities, and the generated answers lack security for data content.
A knowledge graph-based method is adopted to anonymize user nodes, construct anonymous distance matrices and clustering, perform attribute generalization, combine graph neural networks for optimization, generate anonymized knowledge graphs, and use graph structure relationship modeling and pre-trained large language models to generate answers.
While protecting user privacy, it improves the knowledge topology semantic integrity and answer security of the private data question-answering system, enhances the ability to understand complex semantic queries, and improves the accuracy and professionalism of questions and answers.
Smart Images

Figure CN120597314A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent question-answering technology, and in particular to a private data question-answering method and system based on a knowledge graph. Background Art
[0002] Private data scenarios typically refer to data resources owned by enterprises, institutions, or individuals and not made public. Compared to public data, this type of data is highly valuable, highly sensitive, and has strong domain characteristics (such as medical records, financial transactions, politics and military affairs). A private data question-answering system is an intelligent question-answering tool designed specifically for private data resources. It can efficiently manage and accurately query highly sensitive, high-value, and domain-specific data stored in the private data environments of enterprises, institutions, or individuals. Private data question-answering systems are committed to enabling the mining and efficient utilization of high-value internal enterprise data while ensuring privacy and security. They provide accurate and explainable decision support through semantic modeling and dynamic knowledge enhancement technologies. Compared with conventional general-domain question-answering systems, private data question-answering systems focus more on addressing data sensitivity, domain complexity, and real-time requirements. Therefore, balancing data security and knowledge sharing, and addressing the real-time reasoning needs in dynamic business scenarios, are important tasks that private data question-answering systems must address.
[0003] Existing private data question-answering systems often use large language models, which excel at understanding and generating language. However, they lack the ability to represent answers leveraging knowledge from knowledge bases. The current solution is to use a generative reader, which uses a decoder to independently encode each paragraph and uses the token embeddings of these paragraphs as input to the decoder.
[0004] The shortcomings of existing technologies are: First, the logical connections between texts are weak. Current methods encode each text data paragraph independently, thus ignoring the semantic connections between paragraphs, which are crucial for multi-hop reasoning on specific domain problems. Second, the generated answers lack security for the data content. Summary of the Invention
[0005] In view of the above-mentioned defects of the prior art, the present invention provides a private data question-answering method and system based on knowledge graph, which solves the problems of insufficient deep adaptation of traditional question-answering models to vertical field knowledge, weak logical correlation between texts, weak reasoning ability and insufficient security of generated answers for data content.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is:
[0007] A private data question answering method based on a knowledge graph includes the following steps:
[0008] S1, collect and preprocess private data;
[0009] S2. Build a knowledge graph based on the private data;
[0010] S3. Anonymizing the knowledge graph, including:
[0011] S31. Generate a similar user set for each user based on the knowledge graph and set an anonymization threshold; calculate and generate an anonymous distance between users; and construct an anonymous distance matrix based on the anonymous distance between users;
[0012] S32. Clustering similar user nodes in the similar user set according to the inter-user anonymous distance to generate initial clusters; the number of similar user nodes in each initial cluster is not less than the anonymization threshold;
[0013] S33: performing attribute generalization processing on the similar user nodes in each of the initial clusters to make the attribute values of the similar user nodes consistent, and generalizing the out-degree and in-degree of the similar user nodes.
[0014] Preferably, the similar user nodes are a set of nodes with the smallest attribute and degree loss between user nodes.
[0015] Preferably, the step S32 further includes merging or splitting the initial clusters whose number of similar user nodes is less than the anonymization threshold.
[0016] Preferably, the method further comprises step S4, wherein the triples of the anonymized knowledge graph are judged by a graph neural network.
[0017] The second aspect is a private data question-answering system based on knowledge graph, which includes a knowledge base module, a user dialogue module, a data retrieval and comparison module, a knowledge graph module, and a result call analysis module;
[0018] The knowledge base module is used to store knowledge texts; the user dialogue module accepts questions input by users and converts the questions into question vectors;
[0019] The data retrieval and comparison module receives the question vector, searches and compares the question vector in the vector library, and then outputs the vector library answer vector; the data retrieval and comparison module sends a retrieval instruction to the knowledge graph module;
[0020] The knowledge graph module receives the search instruction and performs data retrieval, and outputs a knowledge graph answer vector;
[0021] The result call analysis module is used to analyze and compare the vector library answer vector with the knowledge graph answer vector and generate a final answer;
[0022] The system is used to implement the steps of the private data question-answering method based on knowledge graph as described in the first aspect.
[0023] Compared with the prior art, the beneficial effects of the present invention are embodied in:
[0024] Different from the knowledge graph of traditional technology, this invention anonymizes the core user nodes in the knowledge graph, eliminating sensitive identity information while retaining the semantic integrity of the knowledge topology, which can protect personal information security to a certain extent. It desensitizes some highly sensitive private data (such as medical health records, financial transaction behaviors, and military and political information), providing a verifiable privacy security boundary for private data value mining. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a schematic flow chart of the method of Example 1;
[0026] Figure 2 This is a schematic diagram of the system structure framework of Example 2;
[0027] Figure 3 is a schematic diagram of the structure of the answer generator of Example 1;
[0028] Figure 4 A schematic diagram of a method for graph structure relationship modeling in the answer generator of Example 1;
[0029] Figure 5 Schematic diagram of the knowledge graph structure before and after anonymization in Example 1. DETAILED DESCRIPTION
[0030] In order to make the technical means, creative features, objectives and effects of the invention easier to understand, the present invention is further described with reference to specific figures. However, the present invention is not limited to the following implementation cases.
[0031] It should be noted that the structures, proportions, sizes, etc. illustrated in the drawings in this specification are only used to match the contents disclosed in the specification so that people familiar with this technology can understand and read them. They are not used to limit the conditions under which the present invention can be implemented. Therefore, they have no substantive technical significance. Any structural modification, change in proportional relationship or adjustment of size should still fall within the scope of the technical content disclosed in the present invention without affecting the efficacy and purpose that can be achieved by the present invention.
[0032] Example 1:
[0033] like Figure 1 、 Figure 2 The private data question answering method based on knowledge graph shown in FIG includes the following steps:
[0034] S1. Collect and preprocess private data.
[0035] Private data includes PDF, Word and other files uploaded by users. In this embodiment, users can create knowledge bases in different fields, such as finance, politics, military, medical and other specific fields. Users can add and delete private data in the knowledge base through front-end operations.
[0036] The preprocessing process involves storing the private data vectors generated by vectorizing private data in a vector database. This embodiment of the present invention utilizes a pretrained model to perform vector processing on user-uploaded text documents. This pretrained model first converts uploaded PDF files, such as PDFs, into DataFrames. Then, vector calculations are performed on the loaded text, and feature vectors are extracted for each row of data. The extracted feature vectors are stored in the vector database.
[0037] The user enters a question, and the system retrieves relevant literature based on the question. It then passes the question and literature to the decoder, and the vector database passes the relevant vectors based on the question. The relevant literature is obtained from the Internet and is external data.
[0038] In this embodiment, the question input by the user is represented as q, and the relevant literature is represented as D. q , the private data vector is denoted as v d .
[0039] S2. Construct a knowledge graph based on the private data.
[0040] The system constructs a knowledge graph based on the literature and private data in the vector database, and outputs knowledge graph triples. The specific operations for triple generation are as follows:
[0041] Generate question vector q through question q emb , for the relevant vector v passed into the vector database d , by calculating the problem vector q emb Compare the cosine similarity between the private data vectors in the vector database. The cosine similarity calculation formula is defined as:
[0042]
[0043] For question q, the purpose of the knowledge graph generator is to q Based on the above, a set of knowledge triples in the form of "(head entity, relation, tail entity)" are derived, and special symbols " <e>", the task is divided into two parts: entity recognition and relationship extraction; for entity recognition, the system uses the TAGME entity linking system to identify. q The entity identified in is denoted as E q ;
[0044] For the relationship extraction part, the system divides it into two submodules, namely contextual relationship extraction and indirect relationship extraction. Contextual relationship extraction is defined as:
[0045]
[0046] in with d i Indicates chapter d i Middle Entity E q and d i The in-context relation triples between .
[0047] For indirect relationship extraction, use API to extract.
[0048] For question q, the graph generated by the knowledge graph generator is represented as The generated graph will be passed to the decoder, such as Figure 3 As shown in the figure, the entity relationship extraction module (TripleRelationExtraction) first embeds the input entity pairs into spans. If the input entity dimension is different from the target dimension, the entity pairs are mapped. Finally, all span embeddings of each entity are weighted averaged, and the weight is determined by the entity mask (entity_mask). The final output dimension can be expressed as:
[0049] [batch_size,max_num_entity,ent_dim]
[0050] For the extraction of relation embedding, first for each entity pair (e i ,e j ), after concatenating the embeddings of the source entity and the target entity, the entity pair embedding is mapped to the relational space through the projection layer. The final output dimension can be expressed as:
[0051] [batch_size,max_num_entity,k,rel_dim]
[0052] Where k represents the maximum number of neighbors for each entity.
[0053] S3. Anonymize the knowledge graph, including:
[0054] S31. Generate a similar user set for each user based on the knowledge graph and set an anonymization threshold τ g ;Construct an anonymous distance matrix based on the anonymous distance between users;
[0055] There are multiple user points in the knowledge graph. For the knowledge graph G(v,e,r), the node V is divided into user sets v u With attribute value node v a , where v a ∪v u = V. Define user u in the graph, where u∈v u , for each user u in the knowledge graph, his attribute value and out-degree and in-degree are at least k in the graph u -1 other users cannot distinguish, the privacy parameter value of the user point u entering the question is k u .
[0056] First, get the k most similar points to user u u -1 user point, forming the nearest neighbor user set N(u). To measure the similarity between two users, define d adm The attribute and degree loss distance between two users. The greater the information loss, the lower the similarity. am , target user out-degree distance In-degree distance to target users Similar user nodes are the set of nodes with the smallest attribute and degree loss between user nodes.
[0057] Define the information loss of the generalized attribute value as AL(u,v), and define an attribute as r a , define the difference between the generalized attribute values of user u and user v as GV(G,u,v,r a ), define the original value of u and v as I α (G,r a ,u) and I α (G,r a ,u).
[0058]
[0059] The target user's in-degree information loss is expressed as the difference between the generalized in-degree and the original in-degree.
[0060]
[0061] Then the attribute and degree information loss metrics are:
[0062]
[0063] Then we can find k similar to u. u The cost of -1 user is defined as:
[0064]
[0065] Define k u With k v The k values defined for u and v respectively, then their anonymous distance is defined as:
[0066] d am (u,v)=max{cost(u),cost(v),d adm (u,v)}×max{k u ,k v }
[0067] Generate user anonymous distance matrix through anonymous distance between users
[0068] Specifically, the method for generating rules for user similarity sets is as follows. Figure 5 As shown, the attribute edge R of the user relationship UA ={age, occupation}, the out-degree of age is dom(age)={39,29}, and the out-degree of occupation is dom(occupation)={doctor, architect}. Assume that all attributes of the user are valued, let Li Si be u0 and Zhang San be u1, then I a (u0,age)=I a (u1, age) = {39, 29}, then information loss Calculate AL in the same way c (u1, occupation) = 0 and AL n (u1, age) = 0.5, then Then we can get the attribute distance For in-and-out and Calculate the target out-degree and in-degree distance The loss measure for attribute and degree information is Through the attribute and degree information loss of each other user point and u0, we can find k similar to u0 u For users with -1, the smaller the loss, the more similar the nodes are. The cost is
[0069] S32. First, sort all users by anonymization cost from low to high to form a set U, and enter a loop: as long as the number of users in the current user set is not less than the k of the first user u When the user has a value, we will extract the user and try to build a cluster for the user. In the process of building the cluster, we will start from the user with the smallest anonymous distance to the user and add them to the current cluster in turn, and dynamically update the minimum number of users required for the cluster k. c =max(k c ,k v ) until the anonymity level is met. Finally, the cluster is added to the result set and the next user is processed until the clustering conditions are no longer met. Finally, all generated initial clusters are returned. The clustering requirement is that the number of user nodes in each initial cluster C is not less than the anonymization threshold k u .
[0070] In order to ensure that the clustering results meet the personalized anonymity requirements set by each user and reduce the loss caused by information generalization, the original clusters are first divided according to whether they meet the anonymity constraint (i.e. whether the number of users in the cluster is not less than the maximum anonymity value k in the cluster). u ) is divided into effective clustering set C valid With the remaining user set U, by trying to reallocate each user not covered by the valid cluster to the existing cluster, by calculating the maximum anonymization distance between the user and all members of each cluster, if it meets the anonymity constraint, cluster size condition and tolerance distance threshold τ g , then assign it to the acceptable cluster with the lowest cost. After the assignment is completed, further check whether the size of each cluster exceeds twice the current maximum anonymity value. If it exceeds, split it into smaller sub-clusters with less information loss through sub-processes. Finally, all legal clusters are output.
[0071] S33: Perform attribute generalization processing on similar user nodes in each initial cluster C to make the attribute values of similar user nodes consistent, and generalize the out-degree and in-degree of similar user nodes.
[0072] For all users in the same cluster, the attributes of their users are generalized in the form of:
[0073]
[0074] First, traverse each cluster, define cluster C, merge the attribute sets in the cluster, and define the set as C A .
[0075] Add the missing attribute edge for each user in C and calculate the maximum out-degree (in-degree) d in the cluster max , for each user, if the out-degree is less than d max , add a pseudo-relationship edge, if the out-degree is greater than d max , delete the redundant edges
[0076] Store the anonymized knowledge graph data after generalization and complete privacy processing.
[0077] Through the anonymized graph, such as Figure 5 As shown, the user's identity is protected to a certain extent, and the user's sensitive information in the text can be protected, so that the user's privacy information will not be exposed through reverse inference through the system's answers.
[0078] S4. The anonymized triples are judged by graph neural network. Define temperature parameter τ to control the smoothness of probability distribution, and define z w is the unnormalized score, P τ (w) is the temperature parameter control probability distribution:
[0079]
[0080] Calculate the cumulative probability after sorting the probability distribution in descending order:
[0081] According to the threshold p, the minimum value set S that exceeds p is promoted, S is the triplet selected after arbitration, and the triplet after arbitration is passed to the decoder.
[0082] Methods based on graph structure relationship modeling such as Figure 4 As shown, the entity embedding and the relationship embedding are projected into a unified hidden dimension in the projection-layer, and the hidden dimension is represented as dim h , and perform weighted average calculation on the question embedding in the pooling layer. The question embedding is represented by e q .
[0083] Strengthen the calculation of relation embedding, and the calculation formula is expressed as:
[0084]
[0085] in Denoted as the initial relation embedding, r ev denoted as enhanced relation embedding, REM(·) denoted as relation embedding module, which is a two-layer feedforward neural network in the embodiment.
[0086] The system uses a multi-layer network to update entity embeddings. One layer in the network is defined as the Messagepassing layer. The entity embedding process of the system at layer l is defined as:
[0087]
[0088] where t e (l) represents the entity embedding of e at layer l, and t e (0) =t e ,FFN(·) represents the feedforward neural network layer, represents a set of adjacent triplets of edge e, q represents the average of the embeddings of question q, ⊙ represents element-wise multiplication, and w is a learnable parameter.
[0089] For each entity e, GNN fuses the aggregated messages from the entity triples to update the entity embedding.
[0090] After updating the entity embedding, the model calculates the similarity between the triple and the question q using the following formula:
[0091]
[0092] Based on this similarity measure, the system identifies and selects the top K triples with the highest relevance to the question, i.e.
[0093] The system concatenates this set of triplets in descending order of similarity as a paragraph, which is represented as
[0094] S5. The decoder embeds the question, paragraph, and triple and inputs them into the GNN model.
[0095] like Figure 3 As shown in Figure 3, the hidden state of the original input is concatenated with the triplet encoding in the fusion layer to form an enhanced input representation.
[0096] The model passes the output result to the encoder and outputs the answer; the formula for generating the answer is:
[0097]
[0098] The encoder and decoder both use the pre-trained large language model T5. A linear classifier is added to the GNN model to predict which entities are relevant to the question:
[0099] c q =Softmax(E q (L) w c )
[0100] E q (L) Expressed as the embedding matrix of all entities, c q Represents the probability that each entity is relevant to the question.
[0101] The GNN model is obtained by minimizing training:
[0102]
[0103] Expressed as ε q rel The calculated entity distribution relative to the true value, D KL Expressed as Kullback-Leibler divergence.
[0104] The T5 model is trained using the cross entropy loss between the predicted answer distribution and the true answer distribution, denoted as:
[0105]
[0106] The loss obtained from training the answer generator is:
[0107]
[0108] Where β is the trade-off hyperparameter between T5 loss and GNN loss.
[0109] Example 2:
[0110] A private data question-answering system based on knowledge graph, including a knowledge base module, a user dialogue module, a data retrieval and comparison module, a knowledge graph module, and a result call analysis module;
[0111] The knowledge base module is used to store knowledge text; the user dialogue module accepts questions input by users and converts the questions into question vectors;
[0112] The data retrieval and comparison module receives the question vector, searches and compares the question vector in the vector library, and then outputs the vector library answer vector; the data retrieval and comparison module sends a retrieval instruction to the knowledge graph module;
[0113] The knowledge graph module receives the search instruction and performs data retrieval, and outputs a knowledge graph answer vector;
[0114] The result call analysis module is used to analyze and compare the vector library answer vector with the knowledge graph answer vector and generate a final answer.
[0115] Specifically, the knowledge base module extracts the features of user uploaded files as follows:
[0116] A1: The system first converts the file into DataFrames format;
[0117] A2: After format conversion, perform vector calculations on the text loaded into memory and perform "feature vector extraction" on each row of data.
[0118] A3: Store the extracted feature vectors.
[0119] Specifically, the data retrieval and comparison module performs the following steps to retrieve and compare answers to user questions:
[0120] B1: The data retrieval and comparison module converts the user question into a vector, generates a retrieval instruction, and sends it to the knowledge graph;
[0121] B2: The data retrieval and comparison module searches for the closest vector in the knowledge base through vector retrieval, receives data from the knowledge graph, and compares it with the user's question;
[0122] B3: If the comparison result matches the answer to the user's question, the result data generated by the processing is set to T; if the comparison result does not match the answer to the user's question, the result data generated by the processing is set to F.
[0123] Specifically, the construction method of the knowledge graph construction module is as follows:
[0124] C1: The knowledge graph module first collects multi-source data and performs privacy processing;
[0125] C2: After data processing, domain knowledge extraction and alignment are performed. This task is divided into two parts: entity recognition (ER) and relationship extraction (RE). ER focuses on identifying entities in a paragraph, while RE focuses on inferring the relationship between entities. Finally, the identified entities are represented as E q ;
[0126] The result call analysis module construction method is as follows:
[0127] E1: Define an answer predictor that uses a graph neural network to select a set of anonymized graphs related to the question. Merge the triplets in the selected graphs into an additional paragraph for generating the answer.
[0128] E2: Temperature scale the similarity scores of the candidate triples and normalize them into probability distributions through the SoftMax function: where s i It is expressed as the original score of triple i, τ is the temperature parameter, and it is sorted in descending order of probability and the cumulative probability is calculated. The minimum triple set whose cumulative probability exceeds the threshold P for the first time is initially selected. For a given triple The answer is defined as
[0129] E3: By introducing a graph-structured relationship modeling method, we extract features from semantic units related to the question, and then use the encoder to obtain paragraph tag embeddings. To enhance the model's ability to capture key semantic elements, we implement a symbolic tagging strategy for specific semantic entities and their referential boundaries in the text sequence, inserting special symbols at the beginning and end of entities and mentions. <e>;
[0130] E4: After updating the entity embedding, calculate the triples again Similarity score to the question i The scaled probability distribution is obtained again by the temperature parameter τ, and the candidate triple set is dynamically selected by the threshold P.
[0131] E5: Pair of generated triples After padding and truncation, the generated embeddings are concatenated with the vectors extracted from the knowledge base and used as input to the T5 encoder to generate the answer.
[0132] Example 3:
[0133] This embodiment provides a storage medium, including: a storage medium storing a computer program for a private data question-answering system based on a knowledge graph, wherein the computer program enables a computer to execute a model of a private data question-answering system based on a knowledge graph described in Example 1.
[0134] Example 4:
[0135] This embodiment provides an electronic device, comprising one or more processors and memories and one or more programs, wherein the one or more programs are stored in the memories and are configured to be executed by the one or more processors, and the programs are used to execute a private data question-answering system based on a knowledge graph as described in Example 1.
[0136] In order to more fully verify and deeply understand the effectiveness of the embodiment of the present invention, it is demonstrated through experiments in this embodiment.
[0137] Specifically, this example uses the public datasets 2WikiMultiHopQA and MuSiQue to evaluate the knowledge graph generator and answer generator models used in a private data question-answering system based on a knowledge graph. The 2WikiMultiHopQA dataset contains approximately 570,000 question-answer pairs, split into training, validation, and test sets at an 8:1:1 ratio.
[0138] Specifically, each question is generated from Wikipedia data using templates or automated methods. Questions need to span multiple Wikipedia pages or information blocks, and comparing the attributes of two entities may require querying their respective pages separately. The data supports direct answers, including country names and selected domain experts.
[0139] Specifically, this example is based on the T5-base model. In the graph neural network (GNN) component, the triplet K is set to 10, and AdamW is used as the optimizer. The initial learning rate is set to 1e-4, the number of learning rate warmup steps is 3000, and the total number of training rounds is set to 5. The model is fully trained using the entire training set data. The training batch size is set to 4, the maximum length of the text sequence is 250, and the GNN message update is set to 3 rounds.
[0140] Specifically, the hardware configuration of the experiment in this embodiment is as shown in Table 1 below:
[0141] Table 1 Hardware Configuration
[0142] CPU Intel Xeon Gold 6130 GPU NVIDIA RTX A6000 48GB Memory 128GB Simulation Platform Python Modeling Platform Pytorch Platform
[0143] Table 1 defines the CPU, GPU, memory, simulation tools, and modeling platform used in this embodiment, providing a hardware foundation for experimentally verifying the response generation capability of the knowledge graph-based private data question-answering model.
[0144] In this example, four popular models are introduced to verify the improvement of the graph-based private data question-answering model in the embodiment of the present invention in terms of response strategy prediction accuracy and supportive response generation capabilities, including:
[0145] The DPR (Dense Passage Retrieval) model is a retrieval model based on dense vector representations, specifically designed for open-domain question answering. It uses a dual-encoder architecture to map questions and document paragraphs into the same dense vector space, and leverages vector similarity to quickly retrieve relevant paragraphs.
[0146] FiD: A generative model designed for open-domain question answering, it generates answers by fusing information from multiple retrieved documents into a decoder. FiD improves upon the classic RAG model, significantly enhancing its ability to leverage multi-document information.
[0147] GRAPE is a graph-based multi-hop question answering model that combines graph neural networks (GNNs) with retrieval-augmented generation (RAG) to enable multi-step reasoning for complex questions. The core idea of GRAPE is to use graph structures to model the relationships between entities and documents, enabling more efficient retrieval and reasoning of relevant information in multi-hop question answering tasks.
[0148] FiDO: is a model architecture for open-domain question answering (ODQA), which is based on the Retriever-Reader framework and answers open-domain questions by combining a retriever and a reader.
[0149] In this embodiment, performance (EM%) is used to verify the reliability of the model in the embodiment of the present invention.
[0150] Table 22 Evaluation indicators of the selected models on the WikiMultiHopQA dataset and MuSiQue dataset
[0151] Model 2WikiMultiHopQA MuSiQue DPR 54.8 16.9 FiD 74.1 29.9 GRAPE 73.4 28.3 FiDO 74.6 30.4 Graph-RAG (Ours) 75.3 30.1
[0152] Specifically, Table 2 shows a comparison of multiple indicators between this embodiment and other algorithm models. As shown in Table 2, this embodiment is not inferior to most of the most popular models on the MuSiQue dataset, and outperforms the above four models on the 2WikiMultiHopQA dataset, demonstrating the superiority of the present invention in predicting supportive response strategies.
[0153] The present invention improves the accuracy and responsiveness of private data question-and-answer systems in terms of domain knowledge analysis and intelligent question-and-answering. By combining knowledge graphs and vector retrieval technology, the system's ability to understand complex semantic queries is enhanced, overcoming the problem of insufficient support for professional domain knowledge in traditional question-and-answer systems. After the graph is generated, the graph content is anonymized, and through clustering calculations and attribute generalization, the data security of sensitive information and sensitive users is guaranteed to a certain extent. During the question-and-answer generation process, the system can construct semantic associations based on the knowledge graph, and combine retrieval optimization strategies to ensure the accuracy and professionalism of the answers. Compared with the existing technology, the present invention has achieved significant improvements in efficient query of private data, in-depth analysis of knowledge reasoning, and the relevance and interpretability of question-and-answer results, which will contribute to the widespread application of intelligent question-and-answer technology in professional fields such as finance, medicine, and law.< / e> < / e>
Claims
1. A private data question answering method based on knowledge graph, characterized in that: The following steps are involved: S1, collect and preprocess private data; S2. Build a knowledge graph based on the private data; S3. Anonymizing the knowledge graph, including: S31. Generate a similar user set for each user based on the knowledge graph and set an anonymization threshold; calculate and generate an anonymous distance between users; and construct an anonymous distance matrix based on the anonymous distance between users; S32. Clustering similar user nodes in the similar user set according to the inter-user anonymous distance to generate initial clusters; the number of similar user nodes in each initial cluster is not less than the anonymization threshold; S33: performing attribute generalization processing on the similar user nodes in each of the initial clusters to make the attribute values of the similar user nodes consistent, and generalizing the out-degree and in-degree of the similar user nodes.
2. The private data question-answering method based on knowledge graph according to claim 1 is characterized in that: The similar user nodes are a set of nodes with the smallest attribute and degree loss between user nodes.
3. The private data question-answering method based on knowledge graph according to claim 2 is characterized in that: The step S32 further includes merging or splitting the initial clusters whose number of similar user nodes is less than the anonymization threshold.
4. The private data question-answering method based on knowledge graph according to claim 1 is characterized in that: The method also includes step S4, wherein the triples of the anonymized knowledge graph are judged by a graph neural network.
5. A private data question-answering system based on knowledge graph, characterized in that: Including knowledge base module, user dialogue module, data retrieval and comparison module, knowledge graph module, and result call analysis module; The knowledge base module is used to store knowledge texts; the user dialogue module accepts questions input by users and converts the questions into question vectors; The data retrieval and comparison module receives the question vector, searches and compares the question vector in the vector library, and then outputs the vector library answer vector; the data retrieval and comparison module sends a retrieval instruction to the knowledge graph module; The knowledge graph module receives the search instruction and performs data retrieval, and outputs a knowledge graph answer vector; The result call analysis module is used to analyze and compare the vector library answer vector with the knowledge graph answer vector and generate a final answer; The system is used to implement the steps of the private data question-answering method based on knowledge graph as described in claim 1.