Neural map database system with privacy protection function and use method thereof
By introducing the P-NGDB system into the neural graph database, adversarial training technology and information classification are used to solve the privacy leakage risks in complex query answers in NGDB, and effective privacy protection and high-quality query answer performance are achieved.
Patent Information
- Application Number
- CN202411576484.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-10
- Filing Date
- 2024-11-06
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-11-06
AI Technical Summary
Neural graph database (NGDB) has the risk of privacy leakage in complex query responses. Attackers can infer sensitive information through a combination of multiple queries. The existing technology has not effectively solved this problem.
The Privacy Protection Neural Graph Database (P-NGDB) system was introduced. By dividing the information in the graph database into a privacy part and a public part, and classifying the answers of complex queries based on relevant privacy risks. Adversarial training technology is used to generate indistinguishable answers when querying using privacy information, thereby enhancing privacy protection.
It effectively reduces the privacy leakage risk of malicious queries, protects the privacy information in the graph database, and maintains high-quality complex query response performance. The experimental results show that P-NGDB shows effective privacy protection and query response capabilities on three data sets.
Smart Images

Figure CN119988464A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a privacy-preserving neural graph database (P-NGDB) system with a privacy protection function and a method for using the system, thereby reducing the risk of privacy leakage in a neural graph database (NGDB). Background Art
[0002] Graph databases play an important role in storing, organizing, and retrieving structured relational information and are widely used in downstream applications such as recommendation systems and fraud detection. However, traditional graph databases are often limited by incomplete graphs, which is a common problem in real-world knowledge graphs, such as the Freebase system. Because graph databases may not be able to capture all necessary relationships and connections between entities through traversal, this incompleteness can lead to the exclusion of relevant results.
[0003] To address these shortcomings, Neural Graph Database (NGDB) extends the concept of graph database by combining the flexibility of graph data models with the computational power of neural networks, thereby enabling efficient modeling, storage, and analysis of interconnected data. NGDB can provide unified storage for different entities in the embedding space, and achieve expressive and intelligent query results, reasoning results, and knowledge mining results through graph-based neural network technology. This capability is called complex query answering (CQA), which aims to identify answers that satisfy a given logical expression.
[0004] Logical expressions are usually defined in the form of predicate logic, with relational projection operations, existential quantifiers, Logical conjunction (∧), disjunction (∨), etc. The logical form of q includes terms like Win (V, Turing Award) and logical operators like conjunction (∧), which can express queries such as "find where Turing Award winners born in 1960 live".
[0005] Query encoding methods are usually used for complex query answering (CQA) and involve parameterizing entities, relations, and logical operators, encoding queries and entities into the same embedding space, and then retrieving answers based on the similarity between query embeddings and entity embeddings. For example, methods such as GQE and query2box can encode queries into vectors and hyperrectangles, respectively.
[0006] Although NGDB has made significant achievements in various fields, it does face unique privacy challenges compared to traditional graph databases. One of the risks worth noting is its generalization ability. Although generalization can effectively handle incomplete knowledge graphs and enrich retrieval information, it also enables attackers to use a combination of multiple complex queries to extract sensitive information from NGDB.
[0007] Figure 1 The structural diagram of NGDB is shown. With NGDB, the answer to a given complex logical query can be retrieved on an incomplete graph. However, since the graph information is stored in an embedding space and a query engine that is convenient for answering complex logical queries is used, privacy issues may arise. To explain the problem, assume that the scenario is to consider an example and an attacker tries to infer private information about Hinton's residence in NGDB. Through privacy risk detection, if the privacy is queried directly, it can be easily detected; however, the attacker can also use a combination of multiple queries to retrieve the desired privacy.
[0008] More specifically, if Figure 1 As shown in the figure, the residence information about an individual (Hinton in this case) is considered private information. Even if Hinton's residence is omitted when building the knowledge graph, or direct privacy queries are restricted, malicious attackers can still infer sensitive information without directly querying Hinton's residence. For example, by querying the residence of Turing Award winners born before 1950 and after 1940 in NGDB, the residence of LeCun's collaborators, etc. The intersection of these query results still has a high probability of exposing the residence of Turing Award winner Hinton.
[0009] In addition, graph privacy has always been an important concern in the field of graph research. Many studies have also shown that graph neural networks are vulnerable to various privacy leakage issues. For example, graph embeddings are vulnerable to attribute reasoning attacks, link prediction attacks, etc., and various protection methods have been proposed. However, these works focus on graph representation learning and assume that attackers have access to graph embeddings, which is not a typical scenario in NGDB. NGDB usually provides complex query answering services and does not publish learned embedding expressions, so that attackers can only use designed queries to infer privacy. In addition, these works only study node classification and link prediction problems, and do not explore query answering privacy in NGDB.
[0010] Therefore, since the privacy issues related to query answering in the NGDB field have not yet been resolved, an improved privacy protection method is needed to enhance the graph privacy of the neural graph database while providing high-quality public answers to queries. Summary of the invention
[0011] The object of the present invention is to provide a system and method to solve the above problems existing in the prior art.
[0012] In this disclosure, the potential privacy risks of neural graph databases (NGDBs) are revealed by formally defining and giving evaluations. Under the existing NGDB architecture, the present invention introduces a privacy protection goal and divides the answers of the knowledge graph to complex queries into private domains and public domains. In order to protect the privacy information in the knowledge graph, the NGDB should maintain high-quality retrieval of non-private answers to queries as much as possible, and at the same time deliberately obfuscate the correct but privacy-threatening answers involving privacy queries. For example, the corresponding answer set of the query raised by the privacy attacker can be obfuscated, so that the attacker will encounter greater challenges in predicting private information, thereby effectively protecting the privacy of the NGDB. In addition, corresponding privacy protection evaluation benchmarks are created on three datasets (such as Freebase, YAGO, and DBpedia) to evaluate the performance of complex query answering (CQA) in terms of public query answer quality and privacy information protection.
[0013] In order to alleviate the privacy leakage problem in NGDB, the present invention proposes a privacy-preserving neural graph database (P-NGDB) as one of the solutions. P-NGDB divides the information in the graph database into private and public parts, and classifies the answers to complex queries according to the relevant privacy risks, which depends on whether sensitive information is involved. P-NGDB can provide answers with different levels of accuracy for queries with different privacy risks. During the training phase of P-NGDB, adversarial techniques can be introduced to enable the system to generate indistinguishable answers when querying with private information, thereby enhancing the difficulty of inferring privacy through complex private queries. In the present invention, the main contributions are summarized as follows:
[0014] (1) The solution provided provides creative results in studying the privacy leakage problem of complex queries in the neural graph database (NGDB), and provides specific definitions on the privacy protection issues in the same field.
[0015] (2) Based on the three given public datasets, a benchmark is proposed to simultaneously evaluate the complex query ability (CQA) and privacy protection ability in NGDB.
[0016] (3) We introduce a framework called P-NGDB for protecting the privacy of NGDB. In addition, we conduct a large number of experiments on three datasets. The experimental results show that the adopted framework can effectively maintain high information retrieval performance while ensuring privacy protection.
[0017] According to a first aspect of the present invention, a system of a neural graph database with privacy protection function is provided, which is used to receive a query input from a user and provide an answer to the user in response to the query. The system includes: a user query input receiver, a graph builder module, a query encoder model, an answer retrieval module, a privacy risk identifier, a score calculation engine and an answer output module. The user query input receiver is used as a user interface to receive a structured query language from a user and use it as the input query. The graph builder module is used to receive a query from the user query input receiver and to build a computational graph to convert the query into a directed acyclic graph, where the edges represent operations on entities and attribute sets. The query encoder model is used to receive a query graph related to the computational graph from the graph builder module to process the query graph, wherein the query encoder model is configured to convert the query graph into a vector format. The answer retrieval module is used to receive the vector format of the query graph from the query encoder model and retrieve a candidate answer set, wherein the information output by the answer retrieval module represents the answer set of the entity or value retrieved by the computational graph. The privacy risk identifier is used to receive the information output by the answer retrieval module, determine whether the answer set has privacy risks, check sensitive information, and mark queries that cause privacy leakage, wherein the privacy risk identifier is also used to divide the answer set in the calculation graph into a public answer set and a private answer set. The score calculation engine is used to receive the public answer set and the private answer set, and use the corresponding loss function to calculate the probability of the public answer set and the private answer set, and calculate the score of each candidate answer, wherein the score calculation engine assigns a low threshold to the loss function of the public answer set, requiring the calculation result of the loss function of the public answer set to be lower than the threshold, and assigns a high threshold to the loss function of the private answer set, requiring the calculation result of the loss function of the private answer set to be higher than the threshold. The answer output module is used to refer to the result of the loss function and combine the threshold to select the final answer of the query of the original input.
[0018] According to a second aspect of the present invention, a method for a neural graph database with privacy protection function is provided, which is used to receive queries from user input and provide answers to users in response to the queries. The method includes the following steps: receiving a structured query language from user input as a query through a user query input receiver; building a computational graph through a graph builder module to convert the query into a directed acyclic graph, in which edges represent operations on entities and attribute sets; encoding the query graph related to the computational graph into a vector format through a query encoder model; retrieving the vector format of the query graph as a candidate answer set through an answer retrieval module, wherein the information output by the answer retrieval module represents an answer set of entities or values retrieved by the computational graph; determining whether the answer set has any privacy risks, checking sensitive information, and marking the query that causes privacy leakage through a privacy risk identifier; A risk identifier is provided to classify the answer set from the computational graph into a public answer set and a private answer set; a score calculation engine is used to calculate the probabilities of the public answer set and the private answer set using a corresponding loss function, and a score of each candidate answer is calculated, wherein the score calculation engine assigns a low threshold to the loss function of the public answer set, requiring that the calculation result of the loss function of the public answer set is lower than the threshold, and assigns a high threshold to the loss function of the private answer set, requiring that the calculation result of the loss function of the private answer set is higher than the threshold; and an answer output module is used to refer to the result of the loss function and combine the threshold to select the final answer to the query of the original input. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The embodiments of the present invention are described in more detail below with reference to the accompanying drawings, wherein:
[0020] Figure 1 The structural diagram of the Neural Graph Database (NGDB) is presented;
[0021] Figure 2 An example diagram of performing queries is presented, demonstrating that answers to privacy-threatening queries are retrieved;
[0022] Figure 3 An example diagram of calculating a privacy threat answer set in projection, intersection and union according to an embodiment of the present invention is presented;
[0023] Figure 4 Table 1 is provided, which provides the results of the comprehensive comparison;
[0024] Figure 5 A schematic diagram of eight common query types is presented;
[0025] Figure 6 Table 2 is provided, which provides the results of the comprehensive comparison;
[0026] Figure 7 Table 3 is provided, which provides the results of the comprehensive comparison;
[0027] Figure 8 Table 4 is provided, which provides the results of the comprehensive comparison;
[0028] Fig. 9 A schematic diagram of the GQE evaluation results under different privacy coefficients β is presented; and
[0029] Fig.10 A schematic diagram of the architecture of a privacy-preserving neural graph database (P-NGDB) system according to an embodiment of the present invention is presented. DETAILED DESCRIPTION
[0030] In the following description, systems and methods such as a privacy-preserving neural graph database (P-NGDB) are used as preferred examples. Those skilled in the art will appreciate that modifications, including additions and / or substitutions, may be made without departing from the scope and spirit of the invention. Specific details may be omitted so as not to obscure the invention; however, the present disclosure is written to enable those skilled in the art to practice the teachings herein without undue experimentation.
[0031] In the present invention, based on privacy protection in a graph embedding environment, a system using P-NGDB is provided to mitigate the risk of privacy leakage in a neural graph database (NGDB). In addition, adversarial training techniques are introduced during the training phase to force NGDB to generate indistinguishable answers when querying using private information, thereby increasing the difficulty of inferring sensitive information through the combination of multiple harmless queries. Extensive experimental results on three datasets show that P-NGDB can effectively protect private information in graph databases.
[0032] In this regard, there are several related studies / works targeting Complex Query Answering (CQA) and graph privacy, which are discussed below.
[0033] Technical Background - Complex Query Answering
[0034] A basic function of a neural graph database is to answer complex structured logical queries, or CQA, based on the data stored in the NGDB. Query embedding is one of the common methods for answering complex queries because it can demonstrate effectiveness and efficiency.
[0035] These query embedding methods utilize different structures to represent complex logical queries and effectively model queries in a variety of ranges. There are already some methods that encode query inputs into different geometric computational representations: the GQE model encodes query inputs into vectors in the embedding space, Query2box is encoded as hyperrectangular embedding, and HypeE is encoded as hyperbolic embedding to support separation logic operators. Query2Particle and ConE use multiple vector embeddings and cone embeddings to support negation operators, respectively. There are also studies that propose NRNs for encoding numerical values. In addition, some methods use probabilistic structures to encode logical queries: BetaE, Gamma, and PERM propose to use Beta distribution, Gamma distribution, and Gaussian distribution to encode logical knowledge graph queries, respectively. At the same time, neural structures are used to encode complex queries: BiQE, SQE, and QE-GNN utilize transformers, sequential encoders, and message passing graph neural networks, respectively. The query decomposition program proposes to decompose complex queries, while LMPNN proposes to use single-hop message passing on the query graph for CQA.
[0036] In past studies, although complex query answering has been extensively studied, privacy issues in NGDBs have been neglected. With the development of expressive power in NGDBs, the problem of privacy leakage becomes more and more important. Some studies have shown that graph representation learning is vulnerable to various privacy attacks. Some studies have shown that human reasoning attacks can be applied to identify the content of the source of training samples. Model extraction attacks and link stealing attacks attempt to infer information about the graph representation model and the original graph links, respectively. However, these works assume that the attacker has full access to the graph representation model. But in NGDB, sensitive information may be leaked at the query stage, which brings different privacy challenges. In this invention, the privacy leakage problem in NGDB is revealed, and a benchmark for comprehensive evaluation is proposed.
[0037] Technical Background - Graphics Privacy
[0038] In order to protect privacy in graph representation models, various privacy protection methods have been proposed. Anonymization techniques can be applied in graphs to reduce the probability of personal and link privacy leakage. Graph summarization technology aims to publish a set of anonymous graphs with private information removed. With the development of graph representation learning, learning-based methods have also been proposed, which regard graph representation learning as a combination of two subtasks: main target learning and privacy protection, and use adversarial training to remove sensitive information while maintaining the performance of the original task. At the same time, some methods separate sensitive information from the main learning goal. In addition, differential privacy introduces noise into the representation model and provides privacy guarantees for the individual privacy of the data set. Although the noise introduced by differential privacy can protect the model from privacy leakage, it will also significantly affect the performance of the original task. Federated learning utilizes differential privacy and encryption to prevent the transmission of participants' original data and is widely used in distributed graph representation learning. However, federated learning cannot be used to protect the intrinsic privacy information in the representation model.
[0039] Although there are various graph privacy protection methods, they only focus on simple node-level or edge-level privacy protection. However, in complex query answering tasks, the privacy risk leakage associated with neural graph databases has not been sufficiently researched and expanded. In NGDB, the development of CQA models introduces additional privacy leakage risks, and attackers can infer sensitive information through multiple combined queries. The P-NGDB proposed in this invention can effectively reduce the privacy leakage risk of malicious queries.
[0040] The following describes the preliminary question and related issues in formulating the question.
[0041] What we consider here is the privacy issue in multi-relational knowledge graphs with numerical attributes. The knowledge graph representation can be in Represents the vertex set corresponding to the entity in the knowledge graph. represents a set of relations, where r∈R and is a binary function defined as It is used to describe the relationship between entities in the knowledge graph. Among them, r(u,v)=1 means that there is a relationship r between entities u and v, otherwise r(u,v)=0. Represents a set of attributes, which can be divided into private attributes and public attributes: in and is a binary function defined as →{0,1}, used to describe the numerical attributes of an entity, where a(u,x)=1 means that entity u has attribute a, whose value is Otherwise a(u,x)=0. In addition, there is also the use of To prevent privacy leakage, entities The privacy attributes of the information cannot be exposed and should be protected to prevent inference.
[0042] Complex querying is a key task for NGDB, which can be defined in the form of regular first-order logic, with existential quantifiers It consists of various types of logical expressions such as logical conjunction (∧) and disjunction (∨). In a logical expression, each logical query has a unique variable V ? To represent the query target. Variables V1,…,V k and numerical variables X1,…,X l represents the existential quantified entity in complex queries. In addition, the anchor entity V a Sum value X a Has specific content in the query. Complex queries aim to identify the target entity V ? , so that there are knowledge graphs that can satisfy the given logical expression and Complex query expressions can be expressed in disjunctive normal form (DNF) as follows: q[V ? ]=V ? .V1,…,V k ,X1,…,X l :c1∨c2∨...∨c n ;c i =e i,1 ∧e i,2 ∧...∧e i,m
[0043] Atomic logical expression e i,j It can be used to represent: r(V,V′), which represents the relationship r between entities r and V′; a(V,X), which represents the attribute a of entity V, whose value is X; or f(X,X′), which represents the relationship between numerical attributes. i Represents several atomic logical expressions e i,j The symbols V, V′, X can be anchor entities, values, existential variables or numerical variables.
[0044] Regarding the statement of the problem, first assume that there is a graph Some attributes of the entity are considered private information. in It means if a p (u,x)=1, then entity u has sensitive attribute a p, whose value is x. Due to different privacy requirements, different entities handle attributes differently. NGDB stores graphs in an embedding space and can be queried using arbitrarily complex logical queries. Due to NGDB's generalization ability, attackers can easily infer sensitive attributes using complex queries. To protect private information, NGDB should retrieve ambiguous answers in case of queries with privacy risks, while maintaining high accuracy when querying public information.
[0045] In the present invention, a query program that is harmless and can maintain privacy is provided. Due to the presence of sensitive information in the neural graph database, attackers can use certain specific complex queries to infer privacy, which are defined as "harmless but privacy-related (IP) queries". Even if some answers to these queries can only be inferred when privacy information is involved, this poses a risk to privacy and should not be retrieved by NGDB. In order to meet this requirement, the program architecture and its mechanism for processing harmless but privacy-maintaining queries are described below.
[0046] In the following description, "query" / "query q" refers to the user input into the proposed system, which may consist of a structured query language / statement. The input query serves as a retrieval request for specific information or data in the neural graph database and is processed by various components to determine the most relevant answer while ensuring privacy protection.
[0047] The system's architecture consists of two parts: "computational graph (computational graph)" and "privacy-threatening query answers".
[0048] Computational graph (computational graph):
[0049] The query q can be parsed into a computation graph, i.e., a directed acyclic graph consisting of various types of directed edges representing operators on entities and attribute sets. Operations / operators in the computation graph include projection, intersection, and union.
[0050] (1) Projection: There are various types of projections. Given a set of entities S and the relationships between them Attribute type a, value set Through projection calculation, we can describe the entities or values that can be realized under the specified relationship and attribute type. For example, the attribute projection is expressed as P_a(S) = {x∈N|v∈S, a(v,x) = 1}, which describes the value that any entity in S can realize under attribute type a. In addition, the relational projection P is proposed r (S), reverse attribute projection and numerical projection P f (S), all these projections are denoted P(S).
[0051] (2) Intersection: Given a set of entities or value sets S1, S2, …, S n , use this logical operator to compute the intersection of these sets, i.e.
[0052] (3) Union: Given a set of entities or value sets S1, S2, …, S n , using this operator to compute the union of these sets, i.e. of.
[0053] Privacy Threat Query Answers:
[0054] After the complex query is parsed into the computational graph G, M = G(S) can represent the answer set of entities or values retrieved by the computational graph. To this end, the M retrieved by the complex query can be divided into the private answer set M according to whether the answer needs to be inferred with the participation of private information. private and the non-private answer set M public .
[0055] Figure 2 A query example is shown, showing that the answer to the privacy threat query is retrieved. Figure 2 In , node N represents privacy risk. Figure 2 In part (A) of , the logical knowledge graph query involves private information. Figure 2 In part (B) of the paper, there is an example of a complex query in the knowledge graph, where Toronto is considered a privacy threatening answer because it must be inferred from sensitive information.
[0056] Figure 2 In part (A), the listed query retrieves answers from the knowledge graph, and Toronto is considered a privacy threat because
[0057] Projection Operation / Operator: If the projection operation is used to infer privacy properties, the output of the projection operator will be considered as privacy-sensitive. Assume M = P(S) and the input S = S private ∪S public ; P(S)=P(S private )∪P(S public ). Without loss of generality, we can calculate the projection privacy answer set M in the two cases: private The formal definition of .
[0058] For projection on public input, it is:
[0059] For the projection on the private input, it is: M private =P(SPrivate ).
[0060] Figure 3 An example diagram of the calculation of the privacy threat answer set based on the projection, intersection and union of an embodiment of the present invention is shown. In the figure, node N1 represents a non-privacy answer, node N2 represents a privacy threat answer, and node N3 represents different privacy risks in the subset. The dotted arrows A1, A2 and A3 represent privacy projections. The range defined by the dotted lines D1, D2, D3 and D4 represents the privacy threat answer set. Figure 3 As shown in part (A) of the answer set, all nodes N2 in the answer set can be derived from the privacy attribute links alone and are classified as query answers with privacy risks. On the other hand, all nodes N3, although accessible through the privacy projection, can also be inferred from the public component, thus having a lower privacy risk.
[0061] Intersection operation / operator: The output of the intersection operator is considered private if the answer belongs to any private answer set. Given a set of answer sets M 1 ,M 2 ,…,M n , where each answer set After the intersection operator, like Figure 3 As shown in part (B) of , node N3 is classified as a privacy-threatening answer because its representation content involves the existence element within the answer subset.
[0062] Union Operation / Operator: If the answer is an element of the private answer set of the computation subgraph but not of the public answer set, the output of the union operator will be considered private. like Figure 3 As shown in part (C) of , all nodes N2 in the answer set can only be inferred from the privacy attribute links and are classified as privacy threatening answers.
[0063] Next, a privacy-preserving neural graph database according to an embodiment of the present invention is provided. The privacy-preserving neural graph database is used to protect sensitive information in the knowledge graph while also maintaining high-quality complex query answering performance. Here, two optimization goals are set for P-NGDB: maintaining the accuracy of retrieving non-privacy answers and obfuscating privacy-threatening answers.
[0064] Encoded representation of privacy-preserving neural graph database:
[0065] The query encoding method used can encode the query into the embedding space and compare it with the similarity of entities or attributes for answer retrieval. The encoding process can be expressed as: q=f(Query),
[0066] where f g is obtained by passing θ g Parameterized query encoding method, q∈R d Represents query embedding to simplify the representation content. Given a query, it iteratively calculates the query based on subquery embedding and logical operators. Assume that in the encoding process in step i, the subquery embedding is q i . Then the logical operators projection, intersection and union can be expressed as:
[0067] where f p 、f i and f u Respectively represent the projection, intersection, and union operators for parameterization. After query encoding, the score of each candidate can be calculated based on the query encoding and entity (attribute) embedding. Finally, the normalized probability can be calculated using the Softmax function as follows:
[0068] Where s is a scoring function, such as similarity, distance, etc. It is a candidate set.
[0069] Learning objectives of Privacy Preserving Neural Graph Databases:
[0070] In the privacy-preserving neural graph database, there are two learning objectives: given a query, P-NGDB should accurately retrieve non-private answers and obfuscate private answers. For public retrieval, given the query embedding q, the loss function can be expressed as:
[0071] Where N is the set of public open answers to a given query.
[0072] For privacy protection, the learning goal is opposite. The goal is not to retrieve the correct answer, but to provide obfuscated answers to solve the privacy problem. Therefore, for private answers, given the query embedding q, the goal can be expressed as:
[0073] in is the set of privacy risk answers for a given query.
[0074] In one embodiment, a query can be decomposed into subqueries and logical operators. For intersection and union, the logical operators will not generate new privacy-invading answers; while for projection, if the projection involves private information, the answer can be private. Therefore, the projection operator can be directly optimized to protect sensitive information. The privacy-preserving learning objective can be expressed as:
[0075] The solution provided by the present invention aims to achieve these two goals at the same time, and the final learning objective function can be expressed as: L=L u +βL p ,
[0076] Where β is the privacy coefficient that controls the protection strength in P-NGDB. The larger β is, the stronger the protection is.
[0077] The performance of the system on P-NGDB is demonstrated through experiments. In order to evaluate the privacy leakage problem in NGDB and the protection ability of the proposed P-NGDB, a benchmark is constructed on three real datasets and the performance of P-NGDB is evaluated based on it.
[0078] Experiment-Dataset
[0079] The benchmark is created based on a numerical complex query dataset, and the dataset used is built on three public knowledge graph datasets, namely FB15K-numerical, DB15K-numerical, and YAGO15K-numerica. In each knowledge graph, vertices describe entities and attributes, and edges describe entity relationships, entity attributes, and numerical relationships.
[0080] First, a set of edges representing entity attributes is randomly selected as private information. Then, the remaining edges are divided into 8:1:1 ratios to construct training, validation, and test edge sets. The training graph G_train, validation graph G_val, and test graph G_test are constructed on training edges, training+validation edges, and training+validation+test edges, respectively. The detailed statistical information of the three knowledge graphs is as follows: Figure 4 The data are shown in Table 1, where “#Nodes” represents the number of entities and attributes, “#Edges” represents the number of relationship triples and attribute triples, and “#Pri.Edges” represents the number of attribute triples that are considered private.
[0081] Experiments - Benchmark Build
[0082] like Figure 5 As shown, it evaluates the answering performance of complex queries on the following eight common query types, abbreviated as 1p, 2p, 2i, 3i, ip, ip, 2u and up.
[0083] Figure 5 Eight general query types are presented. The arrows represent projection, intersection, and union operators, respectively. For each general query type, each edge represents a projection or logical operator, and each node represents a set of entities or values. The sampling method is used to randomly sample complex queries from the knowledge graph. Training, validation, and test queries are randomly sampled from previously constructed graphs, respectively. For training queries, the corresponding training answers can be searched on the training graph. For validation queries, queries on the validation graph that have different answers from those on the training graph are used to evaluate the generalization ability. For test queries, a graph search is performed on the test graph with private edges to identify the test answers, and the answers are divided into private sets and public sets according to the definition of privacy risk query answers discussed earlier. A statistical analysis is performed on the number of privacy risk answers for different types of complex queries in the three knowledge graphs, and the statistical data are shown as follows. Figure 6 As shown in Table 2.
[0084] Experiment - Experiment Configuration
[0085] For the set benchmarks, the proposed P-NGDB can be applied to various complex query encoding methods to provide privacy protection.
[0086] Three commonly used complex query encoding methods can be selected to compare the performance with and without protection using P-NGDB. GQE: The Graph Query Encoder model encodes complex queries into vectors in an embedding space. Q2B: The Graph Query Encoder model encodes complex queries into a hyper-rectangular embedding space. Q2P: The graph query encoder model encodes complex queries into an embedding space with multiple vectors.
[0087] It is assumed here that no privacy-preserving methods are used in NGDB. Therefore, the comparison of these methods is based on noise perturbation, which is similar to differential privacy and is a common technique in database queries that introduces randomness to the answer by adding noise.
[0088] Regarding the evaluation indicators, the evaluation content includes two different parts: reasoning performance evaluation and privacy protection evaluation. In the performance evaluation, the generalization ability of the evaluation model is evaluated by calculating the ranking of answers that cannot be observed from the knowledge graph and can be directly retrieved. Given a test query q, the training, verification, and public test answers are represented as Mtrain, Mval, and Mtest, respectively. Then, the hit ratio (HR) and mean reciprocal ranking (MRR) are used to evaluate the quality of the retrieved answers. The HR@K indicator can evaluate the accuracy of the retrieval by measuring the percentage of correct hits among the first K retrieved items. The MRR indicator can evaluate the performance of the ranking model by calculating the average reciprocal ranking of the first relevant item in the result ranking list. The definition of the indicator can be:
[0089] If the metric is HR@K, then m(r) = 1[r≤K]; if the metric is MRR, then Here, a higher value indicates better inference performance.
[0090] In the privacy protection evaluation, the ranking metrics of answers that threaten privacy are calculated because these answers cannot be inferred from the observed graph, and smaller values indicate stronger protection. All models are trained using training queries and privacy attributes, and hyperparameters are tuned using validation queries. Then, the evaluation is performed on test queries, including performance evaluation on public answers and privacy evaluation on private answers. Finally, the experimental results for testing public and private queries are reported separately.
[0091] Regarding parameter settings, we first adjust the hyperparameters of the basic query encoding method on the validation query. NGDB uses the same parameters to ensure a fair comparison. For ease of illustration, the privacy penalty coefficient (β) is adjusted for each of the three datasets to achieve a balance between performance and privacy. For noise interference privacy protection, the noise intensity is also adjusted to ensure that the results are comparable to NGDB and maintain similar performance on public query answers.
[0092] In terms of performance evaluation, P-NGDB is applied to three different query encoding methods: GQE, Q2B, and Q2P, and the query answering performance with and without P-NGDB protection is compared. The experimental results are summarized in Figure 7 Table 3 and Figure 8 In Table 4.
[0093] Table 3 reports the average results of HR@3 and MRR, where higher values on the public answer set indicate better reasoning ability, while smaller values on the private answer set indicate stronger protection. The results show that without protection and if sensitive information exists in the knowledge graph, the basic encoding methods will suffer from privacy leakage. The proposed P-NGDB effectively protects private information with only a slight loss in complex query answering ability. For example, in FB15K-N, GQE has an HR@3 of 21.99 for public answers and 28.99 for private answers without privacy protection; while under the protection of P-NGDB, GQE has an HR@3 of 15.92 for public answers and 10.77 for private answers. After the private answers are protected, the HR@3 drops by 62.9% and the HR@3 of the public answers drops by 27.4%. Compared with noise interference protection, the proposed method can accurately protect sensitive information with much lower performance loss than noise interference.
[0094] The MRR results of P-NGDB for various query types are given in Table 4, and the percentages in brackets indicate the performance of the model protected by P-NGDB relative to the unprotected encoding model. The evaluation is performed on public and private answers separately to assess privacy protection. From the comparison of the MRR changes for public and private answers, it can be seen that P-NGDB provides protection for all query types. In addition, by comparing the average changes between these query types, it is obvious that P-NGDB exhibits different degrees of protection for various query types. For projection and union operations, the MRR of the protected model for private answers is significantly reduced, indicating stronger protection. For the intersection operator, the preservation provided by P-NGDB is as effective as the protection for other operators.
[0095] Regarding sensitivity studies, privacy protection requirements may vary between different knowledge graphs. The impact of the privacy penalty coefficient β is evaluated, where the larger the β, the stronger the privacy protection. The value of β is selected from {0.01, 0.05, 0.1, 0.5, 1}, and the retrieval accuracy of public and private answers is evaluated respectively. The retrieval performance of GQE on the FB15K-N dataset is evaluated. Compared with the unprotected GQE model, the changes in MRR under various β values are shown in Figure 2. Fig. 9 As shown in Figure 1, the evaluation results of GQE with different privacy coefficients β on FB15K-N are presented. The results show that the penalty mechanism can effectively control the level of privacy protection. In order to meet higher privacy requirements, P-NGDB must accept greater utility loss. In addition, since the privacy coefficient can be flexibly adjusted, P-NGDB can adapt to various scenarios with different privacy requirements.
[0096] The proposed system can be physically set up in a hardware device. Fig.10A schematic diagram of the architecture of a system 100 using P-NGDB according to an embodiment of the present invention is shown.
[0097] The system 100 can be used as a privacy-preserving neural graph database (i.e., P-NGDB) to receive queries from users, and then the privacy-preserving neural graph database responds to the input queries and provides answers to the users. The system 100 includes a user query input receiver 110, a graph builder module 120, a query encoder model 130, an answer retrieval module 140, a privacy risk identifier 150, a score calculation engine 160, and an answer output module 170.
[0098] The user query input receiver 110 may serve as a user interface. For example, through the user query input receiver 110, the user can input a query to the system 100, and the query form may be a question formed in a natural language or a structured query language.
[0099] The user query input receiver 110 can forward the input query to the graph builder module 120. The graph builder module 120 is responsible for building a computational graph, converting the input query into a directed acyclic graph, where the edges represent operations on entities and attribute sets.
[0100] The query encoder model 130 receives the calculated query graph from the graph builder module 120, and then processes the calculated query graph. The query encoder model 130 is configured to convert the query graph (i.e., the calculated query graph) into a vector format that can be processed by each component of the system 100 (i.e., it is readable by the components of the system 100). Through the encoding process provided by the query encoder model 130, the input query is converted into a vector representation (e.g., the input query is converted into a numerical representation).
[0101] The answer retrieval module 140 receives a vector / numerical representation of the query from the query encoder model 130. For ease of description, the received information is referred to as a query vector at this stage. The answer retrieval module 140 further uses the query vector processed by the query encoder model 130 to retrieve a set of candidate answers. In one embodiment, the system 100 also includes an embedding transformer module configured to encode the query vector into an embedding space. The processing result of the embedding transformer module can be an optimized query vector. In one embodiment, the answer retrieval module 140 includes a projection engine 142 and an operator setting model 144 for calculating a query graph.
[0102] In the computation graph, the projection engine 142 can perform various types of projection operations, such as attribute projection and relationship projection. The process performed by the projection engine 142 involves extracting related entities or values from the entity set and processing them according to the specified relationship and attribute types.
[0103] In answer retrieval, the operator set model 144 is responsible for processing set operations, such as intersection and union. The operator set model 144 performs set operations on multiple entities or value sets to obtain a result set. For example, the operator set model 144 can provide an intersection operation / operator on multiple entities or value sets to identify common elements shared between all sets; and the operator set model 144 can provide a union operation / operator on multiple entities or value sets to identify all elements in all sets.
[0104] After the query vector is processed by the answer retrieval module 140, the information output by the answer retrieval module 140 can represent the answer set of the entity or value retrieved by the computational graph.
[0105] The answer set may be passed to a privacy risk identifier 150. The privacy risk identifier 150 is configured to determine whether the answer set poses any privacy risk, check for sensitive information, and flag queries that may result in privacy breaches. In one embodiment, the privacy risk identifier 150 includes a privacy threat answer classification model 152 that collaboratively classifies answer sets from the computational graph into public answer sets (i.e., non-private answer sets) and private answer sets, the classification being based on whether the answer set requires inference in the context of private information.
[0106] The private answer set consists of answers that require the export of private information, while the public answer set (i.e., the non-private answer set) does not contain such information. When the projection operation result provided by the projection engine 142 involves a privacy attribute, the privacy risk identifier 150 uses the privacy threat answer classification model 152 to mark the obtained answer as a private answer. Similarly, if the result of the intersection operation provided by the operator setting model 144 belongs to the private answer set, even if it is derived from multiple answer sets, the privacy risk identifier 150 will still use the privacy threat answer classification model 152 to treat the result of the intersection operation as a private answer. For the union operation result provided by the operator setting model 144, if its result includes elements that belong to the private answer set but not to the public answer set, it is regarded as a private answer.
[0107] In one embodiment, the privacy risk identifier 150 includes a privacy assessor model 154 configured to assess whether the privacy answer set contains any privacy answers with lower privacy risk. Specifically, in the privacy answer set, if an answer can be accessed through a privacy projection, but can also be inferred from a common component using an intersection or union operator / operation, then the answer is determined to have a lower privacy risk.
[0108] Then, the public answer set (i.e., non-privacy answer set) and the privacy answer set are fed from the privacy risk identifier 150 to the score calculator engine 160. The score calculator engine 160 is configured to calculate the probability of the result set by using the corresponding loss function, thereby calculating the score of each candidate answer. In one embodiment, in the privacy answer set and the non-privacy answer set, the calculation goals are opposite.
[0109] Regarding the calculation loss function of the public answer set (i.e., the non-private answer set), it uses the public answer set of the given query as a calculation factor / parameter. The goal of public retrieval is to ensure that the system 100 can accurately retrieve public / non-private answers, and the loss function is used to measure the retrieval accuracy of public answers. For the loss function of the private answer set, it uses the privacy risk answer set of the given query as a calculation factor / parameter. The goal of privacy protection is to confuse private answers to prevent the leakage of private information, so the loss function is used to optimize the output information to effectively confuse private answers.
[0110] In one embodiment, the score calculation engine 160 can set relative thresholds for the two loss functions. For example, a low threshold is assigned to the loss function of the public answer set (i.e., the non-private answer set), requiring the loss function calculation result of the public answer set (i.e., the non-private answer set) to be lower than the threshold. On the contrary, a high threshold is assigned to the loss function of the private answer set, requiring the loss function calculation result of the private answer set to be higher than the threshold. This method is beneficial in that, under the condition of privacy protection, in addition to being able to effectively obfuscate private answers, it can also maintain a high accuracy rate for querying public information.
[0111] The answer output module 170 is configured to refer to the result of the loss function and consider the threshold to select the final answer. In one embodiment, the answer output module 170 can use an algorithm, an AI model, or a machine learning model, and then select the most appropriate answer from the processed answers that have been privacy-protected, sensitive information filtered, and scored. The algorithm, AI model, or machine learning model in the system 100 is well-trained. In one embodiment, adversarial techniques can be introduced during the training phase of the P-NGDB to generate answers that are difficult to distinguish or confuse when querying using private information, thereby making it more difficult to infer privacy through complex privacy queries while still accurately retrieving non-privacy answers.
[0112] Therefore, a system with P-NGDB can provide answers with different levels of precision for queries with different privacy risks. This approach simplifies the operation process of privacy protection while reducing computing power. Correspondingly, this process speeds up computing and maximizes efficiency while maintaining strict privacy protection.
[0113] As described above, in this disclosure, the privacy problem in neural graph databases is raised, showing that attackers can infer sensitive information through specified queries. It should be noted that some query answers may leak private information, which are called privacy risk answers. In order to systematically evaluate the proposed problems, a benchmark dataset is constructed based on FB15k-N, YAGO15k-N and DB15k-N. In addition, a new framework is proposed to protect the privacy of knowledge graphs from malicious queries by attackers. The experimental results constructed on the benchmark show that the proposed model (i.e., P-NGDB) can effectively protect privacy while sacrificing only a small amount of reasoning performance.
[0114] The functional units and modules of the processor and method according to the embodiments disclosed herein can be embodied in hardware or software. That is, the required processor can be fully implemented as a machine instruction or as a combination of a machine instruction and a hardware element. Hardware elements include, but are not limited to, computing devices, computer processors, or electronic circuits, including but not limited to application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), microcontrollers, and other programmable logic devices configured or programmed according to the teachings of this disclosure. A person skilled in the art of software or electronics can easily prepare computer instructions or software codes running in computing devices, computer processors, or programmable logic devices based on the teachings of this disclosure.
[0115] The system may include computer storage media, transient and non-transitory memory devices having computer instructions or software codes stored therein, which may be used to program or configure a computing device, a computer processor, or an electronic circuit to perform any of the processes of the present invention. The storage media, transient and non-transitory memory devices may include, but are not limited to, floppy disks, optical disks, Blu-ray disks, DVDs, CD-ROMs, magneto-optical disks, ROMs, RAMs, flash memory devices, or any type of medium or device suitable for storing instructions, codes, and / or data.
[0116] The system may also be configured as a distributed computing environment and / or cloud computing environment, where all or part of the machine instructions are executed in a distributed manner by one or more processing devices interconnected by a communication network, such as an intranet, a wide area network (WAN), a local area network (LAN), the Internet, and other forms of data transmission media.
[0117] The above description of the present invention is provided for the purpose of illustration and description. It is not intended to be exhaustive or to limit the present invention to the precise form disclosed. For those skilled in the art, many modifications and variations will be apparent.
[0118] The embodiments were chosen and described in order to best explain the principles of the invention and its practical application, thereby enabling others skilled in the art to understand the invention for various embodiments and with various modifications as are suited to the particular use contemplated.
Claims
1. A system for a privacy-preserving neural graph database, for receiving a query input from a user and providing an answer to the user in response to the query, characterized in that: include: A user query input receiver, used as a user interface to receive a structured query language from a user and use it as the query input; A graph builder module for receiving a query from a user query input receiver and for building a computational graph to transform the query into a directed acyclic graph, where edges represent operations on entities and attribute sets; a query encoder model for receiving a query graph related to the computation graph from the graph builder module to process the query graph, wherein the query encoder model is configured to convert the query graph into a vector format; an answer retrieval module, configured to receive the vector format of the query graph from the query encoder model and retrieve a set of candidate answers, wherein the information output by the answer retrieval module represents a set of answers to entities or values retrieved by the computational graph; A privacy risk identifier, used to receive the information output by the answer retrieval module, determine whether the answer set has a privacy risk, check sensitive information, and mark queries that cause privacy leakage, wherein the privacy risk identifier is also used to divide the answer set in the computational graph into a public answer set and a private answer set; a score calculation engine, configured to receive the public answer set and the private answer set, and calculate the probabilities of the public answer set and the private answer set using corresponding loss functions, and calculate the score of each candidate answer, wherein the score calculation engine assigns a low threshold to the loss function of the public answer set, requiring that the calculation result of the loss function of the public answer set be lower than the threshold, and assigns a high threshold to the loss function of the private answer set, requiring that the calculation result of the loss function of the private answer set be higher than the threshold; and The answer output module is used to refer to the result of the loss function and combine the threshold to select a final answer to the query originally input.
2. The system according to claim 1, characterized in that The query encoder model is also used to provide an encoding procedure that converts the query into a vector representation or a numerical representation.
3. The system according to claim 1, characterized in that The answer retrieval module includes a projection engine and an operator setting model for operating the computational graph.
4. The system according to claim 3, characterized in that The projection engine is used to perform various types of projection operations, including attribute projection operations and relationship projection operations.
5. The system according to claim 4, characterized in that The projection engine is used to perform various types of projection operations, including attribute projection and relationship projection, which involves extracting related entities or values from an entity set and processing them according to specified relationships and attribute types.
6. The system according to claim 5, characterized in that The operator set model is used to handle set operations, including intersection and union operations.
7. The system according to claim 6, characterized in that The operator set model provides an intersection operator on multiple entity or value sets to identify common elements shared between all sets, and the operator set model provides a union operator on multiple entity or value sets to identify all elements in all sets.
8. The system according to claim 7, characterized in that The privacy risk identifier includes a privacy threat answer classification model, which collaboratively classifies the answer set from the computational graph into the public answer set and the private answer set based on whether the answer set must be inferred in a case involving private information.
9. The system according to claim 8, characterized in that When the projection operation result provided by the projection engine involves privacy attributes, the privacy risk identifier uses the privacy threat answer classification model to mark the obtained answer as a private answer, and wherein, for the intersection operation result provided by the operator setting model, if it belongs to the private answer set, even if the result is from multiple answer sets, it is still regarded as a private answer by the privacy risk identifier using the privacy threat answer classification model, and wherein, for the union operation result provided by the operator setting model, if it includes elements that belong to the private answer set but not the public answer set, its result is regarded as a private answer.
10. The system according to claim 8, characterized in that The privacy risk identifier also includes a privacy assessor model for assessing whether the privacy answer set contains any privacy answer with a lower privacy risk.
11. The system according to claim 1, characterized in that The answer output module is used to select the most appropriate answer from the answer set using an algorithm, an AI model or a machine learning model.
12. According to the system of claim 11, the algorithm, the artificial intelligence model or the machine learning model has been trained, and adversarial technology is introduced in its training phase to generate indistinguishable or confusing answers when using private information queries, thereby making it more difficult to infer privacy through complex private queries while still being able to accurately retrieve non-private answers.
13. A method for a privacy-preserving neural graph database, for receiving a query input from a user and providing an answer to the user in response to the query, characterized in that: include: receiving, via a user query input receiver, a structured query language input from a user as a query; By means of a graph builder module, a computational graph is constructed to convert the query into a directed acyclic graph, wherein edges represent operations on entities and attribute sets; Encoding a query graph associated with the computation graph into a vector format through a query encoder model; Retrieving the vector format of the query graph as a candidate answer set through an answer retrieval module, wherein the information output by the answer retrieval module represents the answer set of the entity or value retrieved by the computational graph; Determining, by a privacy risk identifier, whether the answer set has any privacy risks, checking sensitive information, and marking the queries that cause privacy leakage; Classifying the answer set from the computational graph into a public answer set and a private answer set by the privacy risk identifier; By using a score calculation engine, the probabilities of the public answer set and the private answer set are calculated using a corresponding loss function, and the score of each candidate answer is calculated, wherein the score calculation engine assigns a low threshold to the loss function of the public answer set, requiring that the calculation result of the loss function of the public answer set is lower than the threshold, and assigns a high threshold to the loss function of the private answer set, requiring that the calculation result of the loss function of the private answer set is higher than the threshold; as well as The answer output module refers to the result of the loss function and combines the threshold to select the final answer of the query originally input.
14. The method according to claim 13, characterized in that The answer output module is used to select the most appropriate answer from the answer set using an algorithm, an AI model or a machine learning model.
15. The method according to claim 14, characterized in that The algorithm, the artificial intelligence model or the machine learning model has been trained, and adversarial technology is introduced during its training phase to generate indistinguishable or confusing answers when using private information queries, thereby making it more difficult to infer privacy through complex private queries while still being able to accurately retrieve non-private answers.
Citation Information
Patent Citations
Pre-training language model-oriented privacy disclosure risk assessment method and system
CN114676458A
Multi-hop question answering method based on multi-hop reasoning joint optimization
CN114780707A
Preserving privacy when statistically analyzing a large database
US20060161527A1
Structure-Preserving Subgraph Queries
US20160357799A1
Methods and systems for natural language processing of graph database queries
US20220414228A1