Evaluation method and system for knowledge graph
Through the incremental training and adversarial question-and-answer model based on TransE model, the security problem of knowledge graph evaluation under non-complete information conditions is solved, and efficient and safe knowledge graph quality evaluation is achieved.
Patent Information
- Application Number
- CN202510234140.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to safely evaluate knowledge graphs under non-complete information conditions, and there is a risk of information leakage.
The incremental training method based on the TransE model is used to generate a knowledge graph embedding model and train it through adversarial question-and-answer and answer models to evaluate the quality of the knowledge graph under non-complete information conditions while protecting information security.
The security evaluation of the knowledge graph under non-complete information conditions is realized, information leakage is avoided, and the best quality map among multiple knowledge graphs can be efficiently selected, and the evaluation process can be further improved through confrontation training.
Smart Images

Figure CN120146169A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of knowledge graph evaluation, and in particular to an evaluation method and system for knowledge graphs. Background Art
[0002] Knowledge graphs have strong reasoning capabilities and can be used to achieve intelligent scheduling. In some application scenarios of knowledge graphs, the quality of the knowledge graph determines the role it can play. For example, in the medical field, knowledge graphs can be used to assist diagnosis. If their quality is not high (such as lack of accurate or complete disease-drug relationships), it may lead to inaccurate diagnosis results and even affect the treatment effect of patients. Therefore, it is necessary to introduce a quality evaluation method for knowledge graphs.
[0003] In some existing solutions, the evaluation of knowledge graphs mostly depends on the details of the knowledge graphs, that is, the quality of the knowledge graphs is evaluated under complete information conditions. Such evaluation methods may lead to privacy exposure, resulting in loss of interests. In other existing solutions, there are also some methods for evaluating knowledge graphs under incomplete information conditions, but these methods still have the risk of key information leakage, such as leaking triple information in the knowledge graph during the evaluation process.
[0004] Therefore, how to evaluate the knowledge graph under incomplete information conditions and improve the security of the information in the knowledge graph during the evaluation process is a technical problem that needs to be solved urgently. Summary of the invention
[0005] In view of this, the present application discloses a method and system for evaluating a knowledge graph, which enables the evaluation of the knowledge graph under incomplete information conditions and improves the security of the information in the knowledge graph during the evaluation process.
[0006] In a first aspect, the present application discloses an evaluation method for a knowledge graph, including: based on the TransE model, training a first knowledge graph embedding model in an incremental manner using first triples of a first knowledge graph, and based on the TransE model, training a second knowledge graph embedding model in an incremental manner using second triples of a second knowledge graph; training a first question model and a first answer model based on the first knowledge graph embedding model, where the first answer model is configured to answer questions of the first question model, and the first question model is configured to adjust the difficulty of the questions according to the answers of the first answer model; training a second question model and a second answer model based on the second knowledge graph embedding model, where the second answer model is configured to answer questions of the second question model, and the second question model is configured to adjust the difficulty of the questions according to the answers of the second answer model; using the first answer model to answer questions of the second question model, recording a first score according to the answer results, and using the second answer model to answer questions of the first question model, recording a second score according to the answer results; evaluating the first knowledge graph and the second knowledge graph based on the first score and the second score.
[0007] Optionally, the evaluation method further includes: swapping the first knowledge graph embedding model and the second knowledge graph embedding model during training, so that the first triples train the second knowledge graph embedding model, and the second triples train the first knowledge graph embedding model.
[0008] Optionally, the first knowledge graph and the second knowledge graph are knowledge graphs in the same domain, and the entity embeddings and relationship embeddings in the first knowledge graph and the second knowledge graph have the same expressions in the corresponding first knowledge graph embedding model and second knowledge graph embedding model.
[0009] Optionally, the first question model being configured to adjust the difficulty of the questions according to the answers of the first answer model, and the second question model being configured to adjust the difficulty of the questions according to the answers of the second answer model both include the following processes: based on the answer results of the first answer model and the second answer model, respectively marking whether the answers to the questions of the first question model and the second question model are correct or incorrect, and repeating the marking process; using the marked questions to perform classification training on the question models corresponding to the first question model and the second question model; based on the classification training results, the question models perform question level classification on their own question sets; based on the question level classification, the first question model and the second question model adjust the question difficulty by selecting the question level classification.
[0010] Optionally, the first knowledge graph includes a first positive example set, and the first positive example set includes a number of first triples; training the first knowledge graph embedding model using the first triples of the first knowledge graph in an incremental manner includes: replacing the head entity or the tail entity in the first triple according to a random probability of a uniform distribution to obtain a third triple, and the first negative example set includes a number of third triples; reducing the first distance between the sum of the head entity and the first relation vector and the tail entity in each first triple of the first positive example set, and increasing the second distance between the sum of the head entity and the second relation vector and the tail entity in each third triple of the first negative example set; using the characteristics of the triples in the vector space, performing iterative calculations multiple times to make the first distance approach 0, or maximize the second distance, thereby determining the first relation vector or the second relation vector to obtain the first knowledge graph embedding model; the determination steps of the second knowledge graph embedding model are the same as those of the first knowledge graph.
[0011] Optionally, the evaluation method further includes: randomly extracting a subgraph from the knowledge graph, randomly deleting node information from the subgraph to construct a defective subgraph, using a graph convolutional network to construct a defective subgraph embedding model, training the first question model and the second question model with the defective subgraph embedding model, and using the vector representation of the defective subgraph embedding model as one of the input parameters of the answer model when training the answer model; wherein, the answer model includes a first answer model and a second answer model, the defective subgraph embedding model is used to represent the question, and the randomly deleted node information is the relationship information or entity information represented by any node in the subgraph, and the position information of the deleted node is saved in the defective subgraph embedding model.
[0012] Optionally, the defective subgraph embedding model further includes:
[0013] Using a graph convolutional network to learn the problem representation T of the defective subgraph, and the calculation formula is as follows:
[0014]
[0015] where W X is the feature matrix of the defective subgraph, W A is the adjacency matrix of the defective subgraph, W D is the degree matrix of the defective subgraph, W 1 is the weight matrix from the input layer to the hidden layer of the graph convolutional network, W 2 is the weight matrix from the hidden layer to the output layer of the graph convolutional network, N sg is the number of nodes in the defective subgraph.
[0016] Optionally, the evaluation method further includes:
[0017] Constructing an answer model based on a convolutional neural network and a fully connected neural network, and the calculation formula is as follows:
[0018] FA = LeakyReLU(∑ x∈fm x·k + b), fm ∈ QA, k ∈ K;
[0019]
[0020] P c = Sigmoid(W 0 [FA 1 ; FA 2 + b 0 );
[0021] Among them, FA is the feature extracted using the convolutional kernel K, fm is the feature map from QA, the answering model uses the first convolutional kernel and the second convolutional kernel. The first convolutional kernel is used to extract the local features FA of the defective subgraph problem representation and the candidate answer representation respectively 1 , and the second convolutional kernel is used to extract the global features FA of the defective subgraph problem representation and the candidate answer representation 2 , T is the problem representation of the defective subgraph, ca is the vector representation of the candidate answer, and P c is the probability that the candidate answer is the correct answer to the question, and b 0 is the bias term.
[0022] Optionally, the questioning model includes a first questioning model and a second questioning model, and the questioning model is a Naive Bayes model. Questions marked as answered correctly and answered incorrectly will enter the training of the questioning model at a certain proportion.
[0023] Second aspect, the present application discloses an evaluation system for a knowledge graph, including: a knowledge embedding module configured to receive a first knowledge graph and a second knowledge graph, and based on the TransE model, incrementally train to obtain a first knowledge graph embedding model and a second knowledge graph embedding model, where the first knowledge graph embedding model corresponds to the first knowledge graph and the second knowledge graph embedding model corresponds to the second knowledge graph; a first adversarial module configured to train a first question model and a first answer model based on the first knowledge graph embedding model, and train a second question model and a second answer model based on the second knowledge graph embedding model, where the first answer model is configured to answer questions of the first question model, the first question model is configured to adjust the difficulty of the questions according to the answers of the first answer model, the second answer model is configured to answer questions of the second question model, and the second question model is configured to adjust the difficulty of the questions according to the answers of the second answer model; a second adversarial module configured to use the first answer model to answer questions of the second question model, record a first score according to the answer result, and use the second answer model to answer questions of the first question model, and record a second score according to the answer result; a quality evaluation module configured to evaluate the first knowledge graph and the second knowledge graph based on the first score and the second score.
[0024] In summary, an evaluation method and system for a knowledge graph disclosed in the present application has at least the following beneficial effects:
[0025] (1) The present application realizes the quality evaluation of the knowledge graph under the condition of incomplete information, effectively protecting the privacy information of the knowledge graph;
[0026] (2) The present application compares the scores of two knowledge graphs and determines which knowledge graph has better quality according to the comparison result, so that the knowledge graph with the best quality can be efficiently selected from multiple knowledge graphs;
[0027] (3) The present application evaluates the quality of the knowledge graph in an adversarial form, enabling the subsequent improvement of the knowledge graph evaluation process according to the adversarial results;
[0028] (4) The present application adopts an incremental training method and exchanges the knowledge graph embedding model instead of the original data of the knowledge graph during the data exchange process, so that the data of the two knowledge graphs can be better protected during the training process;
[0029] (5) The question model and the answer model in the present application are equivalent to the relationship between a student and a teacher. The difficulty of the questions of the question model is adjusted according to the results of the answer model, and this property plays an important role in the training of the question model and the answer model. Description of the Drawings
[0030] The following is a brief introduction to the drawings used in the description of the embodiments of the present application:
[0031] Figure 1 It is a flowchart of an evaluation method for a knowledge graph provided by an embodiment of the present application;
[0032] Figure 2 It is a structural example diagram of an evaluation system for a knowledge graph provided by an embodiment of the present application. Detailed implementation manners
[0033] To more clearly illustrate the technical solutions of the embodiments of the present application, the specific implementation manners of the present application will be described below with reference to the accompanying drawings. The accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, and other implementation manners can be obtained. Adjustments and improvements made without departing from the concept of the present application fall within the protection scope of the present application.
[0034] To make the drawings concise, only the parts related to the corresponding embodiments are schematically shown in each drawing, and they do not represent the actual structure of the product. In addition, to make the drawings concise and easy to understand, in some drawings, parts with the same structure or function are only schematically shown partially, and there may actually be more or fewer parts with the same structure or function.
[0035] In the present application, unless otherwise clearly defined and limited, ordinal numbers, such as "first", "second", etc. are only used to distinguish and describe related objects, and cannot be understood as indicating or implying the relative importance or order between related objects; in addition, they do not represent the quantity of related objects. "Multiple" includes two or more, and other quantifiers are similar. " / " is used to describe the relationship between related objects, which means the "or" relationship between related objects. "And / or" is used to describe the relationship between related objects, which includes any combination relationship between related objects. For example, "a and / or b" includes: "a alone", "b alone", or "a and b". "One or more" or "at least one" among multiple objects refers to any object or any combination of multiple objects. For example, "one or more of a1, a2, a3" or "at least one of a1, a2, a3" includes: "a1 alone", "a2 alone", "a3 alone", "a1 and a2", "a1 and a3", "a2 and a3", or "a1, a2 and a3".
[0036] Knowledge graph is a knowledge management and reasoning technology based on graph data structure. Due to its strong reasoning ability, it is widely used in fields such as healthcare, finance, and industrial control, especially playing an important role in intelligent scheduling. For example, in the management of photovoltaic power plants, existing robots for installation, inspection, cleaning, etc. can well complete their own functions. These robots are like various parts of the human body, while the knowledge graph acts as the "brain" for unified scheduling, combining each module organically to achieve intelligent and efficient overall operation.
[0037] Benefiting from the data link method of graph, the knowledge graph can efficiently store and organize complex multi-dimensional data, and discover potential relationships and patterns among data through reasoning and querying. For example, in the healthcare field, the knowledge graph can associate patient symptoms, examination results with drug treatment plans, providing strong support for auxiliary diagnosis. However, the quality of the knowledge graph determines the role it can play. If the quality of the knowledge graph is not high (such as problems like incomplete disease-drug relationships and data redundancy), it will directly affect the accuracy of the diagnosis result and may even have an adverse impact on the treatment effect of patients. Therefore, how to evaluate the quality of the knowledge graph has become a key link in promoting its application.
[0038] Existing knowledge graph quality evaluation methods mainly follow the following two technical paths. Evaluation methods under complete information conditions usually require a comprehensive understanding of the detailed information of the knowledge graph (such as entities, relationships, and their triple representations), and design quality dimensions and related evaluation indicators based on this. By analyzing the construction process, data integrity, consistency, and semantic correctness of the knowledge graph, the quality evaluation result is obtained. However, these methods highly rely on the original data of the knowledge graph. Taking the photovoltaic maintenance knowledge graph as an example, if all equipment and operation data in the photovoltaic power plant need to be called during the evaluation process, this may lead to the leakage of internal sensitive information, thus causing privacy issues and loss of interests.
[0039] To address privacy and security issues, some solutions attempt to evaluate the quality of the knowledge graph under incomplete information conditions, such as sampling and analyzing partial data or anonymized data of the knowledge graph. However, these methods still have limitations: during the evaluation process, some key information (such as the content or structure of triples) may be indirectly leaked, thus unable to fully guarantee the privacy and security of the knowledge graph. In addition, existing methods mainly focus on the quality of the knowledge graph at the data level, while ignoring its performance at the ability level, that is, the actual ability of the knowledge graph in reasoning, answering questions, or performing tasks. This ability is the core of the knowledge graph quality evaluation.
[0040] In view of the above status, the current knowledge graph quality evaluation methods face two core challenges in practical applications: privacy protection and capability demonstration. How to evaluate the quality of knowledge graphs under incomplete information conditions and effectively avoid data leakage is a technical problem that needs to be solved urgently. For enterprises, knowledge graphs may contain a large amount of commercially sensitive data, such as technical formulas, customer relationships, key operating data, etc. How to ensure the security of this information during the evaluation process is crucial. At the same time, when designing evaluation dimensions, existing methods pay more attention to static indicators such as data integrity and consistency, and less attention to the performance of knowledge graphs in dynamic tasks, such as reasoning about complex relationships and the accuracy of answering user questions.
[0041] Photovoltaic power station management is one of the typical scenarios for knowledge graph application. Traditional power station inspection and maintenance rely on manual or single-function robots, which are inefficient and less intelligent. By introducing knowledge graphs, data sharing and intelligent linkage between robots can be achieved. For example, installation robots can adjust installation steps in real time according to the construction rules in the knowledge graph; inspection robots can quickly locate possible fault points and generate maintenance plans through knowledge graphs; cleaning robots prioritize cleaning key areas according to the task scheduling system in the knowledge graph. In these scenarios, the reasoning ability of knowledge graphs is the key to realizing intelligent scheduling, but the quality of their reasoning directly determines the effect of intelligent scheduling. Therefore, it is necessary to introduce a scientific, comprehensive and safe knowledge graph quality evaluation method.
[0042] The following is a description with reference to the accompanying drawings.
[0043] Figure 1 This is a flowchart of a method for evaluating a knowledge graph provided in an embodiment of the present application. Please refer to Figure 1 , an evaluation method for knowledge graph, comprising:
[0044] S100, based on the TransE model, incrementally train the first triple of the first knowledge graph to obtain a first knowledge graph embedding model, and based on the TransE model, incrementally train the second triple of the second knowledge graph to obtain a second knowledge graph embedding model;
[0045] S200, training a first questioning model and a first answering model based on the first knowledge graph embedding model, wherein the first answering model is configured to answer questions of the first questioning model, and the first questioning model is configured to adjust the difficulty of the question according to the answer of the first answering model;
[0046] S300, training a second questioning model and a second answering model based on a second knowledge graph embedding model, wherein the second answering model is configured to answer questions of the second questioning model, and the second questioning model is configured to adjust the difficulty of the question according to the answer of the second answering model;
[0047] S400 uses the first answering model to answer the questions of the second question model, records the first score according to the answering result, and uses the second answering model to answer the questions of the first question model, and records the second score according to the answering result;
[0048] S500 evaluates the first knowledge graph and the second knowledge graph based on the first score and the second score.
[0049] The first knowledge graph and the second knowledge graph are two different knowledge graphs. Among them, the first knowledge graph can be a newly developed knowledge graph by the user, and the second knowledge graph can be a knowledge graph widely used in the prior art. In this application, the quality of the two is evaluated through the confrontation between the first knowledge graph and the second knowledge graph. During the confrontation process, if the detailed information of the second knowledge graph is available, the information missing in the first knowledge graph compared with the second knowledge graph can be supplemented into the first knowledge graph manually or automatically, so as to improve the first knowledge graph.
[0050] In some embodiments of the present application, a large question set will be generated first. Each question model is based on the Naive Bayes classification method to predict which questions can be correctly answered by the current answering model. In this process, after each answer is completed, the model will be feedback annotated according to the accuracy of the answer, and the correct and wrong answers will be used as training data to gradually optimize the prediction ability of the question model. The trained question model can pre-classify other questions in the large question set and select the questions suitable for the current answering model to answer. These questions form a new question set. Through continuous iterative training and feedback, the question and answering models can more accurately evaluate and optimize the differences and confrontations between the first knowledge graph and the second knowledge graph, and further promote the improvement and optimization of the first knowledge graph.
[0051] In some embodiments of the present application, the first knowledge graph and the second knowledge graph are knowledge graphs in the same field, and the entity embeddings and relationship embeddings in the first knowledge graph and the second knowledge graph are expressed the same in the corresponding first knowledge graph embedding model and the second knowledge graph embedding model. Exemplarily, the first knowledge graph is a photovoltaic maintenance knowledge graph, and the second knowledge graph is a photovoltaic maintenance knowledge graph from a third party. The same entity or relationship is represented consistently in the two knowledge graphs, avoiding errors caused by embedding differences, so that the evaluation results are more objective and comparable. Moreover, the question model and the answering model can operate based on the consistent embedding representation when exchanging questions, without additional conversion or alignment, which helps to reduce the complexity of adversarial training. In addition, by sharing the same representation, it is easier to compare the quality of the two knowledge graphs, and it also provides a unified basis for subsequent knowledge fusion or completion.
[0052] Based on the TransE model, multiple groups of knowledge relationships in the knowledge graph are transformed into multiple triples. Among them, is the vector representation of the head entity, is the vector representation of the relation, is the vector representation of the tail entity. The triple has the characteristic in the vector space, that is, the head entity vector + the relation vector = the tail entity vector. In this way, the mapping of the knowledge in the knowledge graph to the vector space is realized.
[0053] A knowledge graph embedding model is a type of model that converts discrete symbols (such as entities and relations) represented in the knowledge graph into continuous vectors. Its goal is to enable the semantic information of the knowledge graph to be effectively represented in a low-dimensional vector space through this embedding, while retaining important structural information and semantic characteristics in the knowledge graph. In this application, the knowledge graph embedding model is trained in an incremental manner, that is, during the training process, instead of training with all the data at once, the model parameters are gradually updated in batches or steps. By means of incremental training, large-scale data can be efficiently processed, avoiding loading all the data at once, improving memory utilization; and, as new data is added, the embedding representation is gradually updated without having to retrain the entire model; in addition, during the training process, the data can be distributed to multiple nodes for parallel processing, further improving the training efficiency.
[0054] In some embodiments of this application, the evaluation method further includes: swapping the first knowledge graph embedding model and the second knowledge graph embedding model during training, so that the first triple trains the second knowledge graph embedding model, and the second triple trains the first knowledge graph embedding model.
[0055] During the training process, the first knowledge graph and the second knowledge graph continuously swap their knowledge graph embedding models and train them using their own triples. This operation is to enable the first knowledge graph and the second knowledge graph to understand each other's "language". If each model only trains its own knowledge graph embedding model using its own triples, it will be very difficult to understand the other party's questions during the subsequent answering and questioning process. Different from directly swapping the triple information in their respective knowledge graphs, this application can, under the premise of incomplete information, not expose its own data by swapping the knowledge graph embedding models, enabling the knowledge graphs of both parties to expand their reasoning capabilities in the semantic environment of the other party. In this way, through incremental training and swapping the knowledge graph embedding models instead of the original data of the knowledge graphs during the data exchange process, the data of the two knowledge graphs can be better protected during the training process.
[0056] In some embodiments of the present application, the first knowledge graph embedding model and the second knowledge graph embedding model are TransE models. The core of the TransE model lies in representing the semantic relationships of triples through a translation mechanism in the vector space. By mapping entities and relationships to vectors and learning their translation rules, TransE can efficiently handle tasks such as reasoning, completion, and quality evaluation in knowledge graphs.
[0057] In some embodiments of the present application, the first question model is configured to adjust the difficulty of questions according to the answers of the first answer model, and the second question model is configured to adjust the difficulty of questions according to the answers of the second answer model, both of which include the following processes: Based on the answer results of the first answer model and the second answer model, the questions of the first question model and the second question model are respectively marked as correct or incorrect answers, and the marking process is repeated; The marked questions are used to classify and train the question models corresponding to the first question model and the second question model; Based on the classification training results, the question models classify the question sets of themselves; Based on the question level classification, the first question model and the second question model adjust the question difficulty by selecting the question level classification.
[0058] The first question model and the second question model adjust the difficulty of questions through interaction with the answer model. Specifically, in each answering process, the answer model will answer based on the questions of the current question model, and then mark the questions as "correct" or "wrong" according to the answer results. This marking process will be repeated to accumulate marked data. Using these marked question data, the first question model and the second question model will respectively conduct classification training to learn how to adjust the difficulty of questions according to the feedback of the answer model. After multiple trainings, the question model can classify the questions by level. For example, in the question model, the questions are divided into three levels: "easy", "medium", and "difficult" according to the probability of the answer result approaching the correct answer, and more detailed level classification can also be made according to the probability values. Based on these classification levels, the question model will adjust the question difficulty according to the performance of the current answer model. If the answer model can accurately answer easy questions, the question model will increase the difficulty of the questions and select questions with a higher difficulty level; on the contrary, if the accuracy of the answer model is low, the question model will select questions with a low difficulty level to ask. This adjustment mechanism helps the question model achieve higher accuracy in question difficulty selection through continuous feedback and training, thereby promoting the continuous improvement of the quality of the knowledge graph.
[0059] In some embodiments of the present application, the first knowledge graph includes a first positive example set, and the first positive example set includes a number of first triples; training the first knowledge graph embedding model using the first triples of the first knowledge graph in an incremental manner includes: replacing the head entity or the tail entity in the first triple according to a random probability of a uniform distribution to obtain a third triple, and the first negative example set includes a number of third triples; reducing the first distance between the sum of the head entity and the first relation vector and the tail entity in each first triple of the first positive example set, and increasing the second distance between the sum of the head entity and the second relation vector and the tail entity in each third triple of the first negative example set; using the characteristics of the triples in the vector space, performing iterative calculations multiple times, and then determining the first relation vector or the second relation vector to obtain the first knowledge graph embedding model; the determination steps of the second knowledge graph embedding model are the same as those of the first knowledge graph embedding model.
[0060] Among them, the characteristics of the triples in the vector space are as described above, that is, head entity + relation vector = tail entity. The triples classified as positive examples can be represented as (h, r, t). Correspondingly, the triples classified as negative examples can be represented as (h', r, t) or (h, r, t'), etc. For each element in the knowledge graph triple set, the head entity or the tail entity of the triple can be replaced according to a random probability of a uniform distribution to form a negative example set. By minimizing the gap between and in the positive examples (making it as close to 0 as possible), and maximizing the gap between and in the negative examples, the embedding representations of the entities and relations are obtained through multiple rounds of iterative calculations, and the training steps of the knowledge graph embedding model are completed.
[0061] The training method of the first knowledge graph embedding model is given in the above embodiments. The second knowledge graph embedding model can also be obtained with reference to the above training method, and the present application will not elaborate on this.
[0062] During the training process, the first question model and the first answer model are also trained based on the first knowledge graph embedding model, and the second question model and the second answer model are trained based on the second knowledge graph embedding model. The question models include the above-mentioned first question model and the second question model, and the answer models include the above-mentioned first answer model and the second answer model. The answer model is configured to answer the questions raised by the question model, and the question model is configured to adjust the difficulty of the questions according to the answers of the answer model. In some embodiments of the present application, the question model is a Naive Bayes model, and the questions marked as answered correctly and answered incorrectly will enter the training of the question model at a certain ratio. In the adversarial training of the question model and the answer model, the answer model will divide the question set into two parts: correct answers and wrong answers. These questions are used as the training set to train the Naive Bayes model. The trained Naive Bayes model classifies the question bank and selects the same number of questions from the two categories to participate in the subsequent adversarial training. In the present application, the knowledge graphs training their respective question models and answer models can be called the first confrontation or internal confrontation; the knowledge graphs exchanging their respective questions and answering the questions of the other party can be called the second confrontation or external confrontation.
[0063] After the training of the question models and answer models of the first knowledge graph and the second knowledge graph is completed, exchange the questions generated by their respective question models, and have the answer models of the other party answer these questions, and record their respective scores. That is to say, use the first answer model to answer the questions raised by the second question model, record the first score according to the answer result, and use the second answer model to answer the questions raised by the first question model, record the second score according to the answer. The first score can be used to represent the quality of the first knowledge graph, and the second score can be used to represent the quality of the second knowledge graph. The higher the score, the higher the quality of the corresponding knowledge graph. For example, when the training of the answer model and the question model ends, each has 100 questions. The two parties use the answer model to answer 100 questions of the other party, and get one point for answering a question correctly. Finally, compare which score is higher. In this way, by comparing the scores of the two knowledge graphs and determining which knowledge graph has better quality according to the comparison result, the knowledge graph with the best quality among multiple knowledge graphs can be efficiently selected. In addition, the question model and the answer model are equivalent to the relationship between students and teachers. The difficulty of the questions of the question model is determined according to the results of the answer model. This property plays an important role in the training of the question model and the answer model.
[0064] In some embodiments of the present application, the evaluation method further includes: randomly extracting a sub-graph based on the knowledge graph, randomly deleting node information from the sub-graph to construct a defect sub-graph, using a graph convolutional network to construct a defect sub-graph embedding model, training a first question model and a second question model with the defect sub-graph embedding model, and using the vector representation of the defect sub-graph embedding model as one of the input parameters when training an answer model; wherein, the answer model includes a first answer model and a second answer model, the defect sub-graph embedding model is used to represent a problem, and the randomly deleted node information is the relationship information or entity information represented by any node in the sub-graph, and the position information of the deleted node is saved in the defect sub-graph embedding model.
[0065] The defect sub-graph includes multiple node information. Compared with the single relationship of the triple, the node relationships included in the defect sub-graph are more diversified, and more problem representations are constructed, which can enrich the questioning complexity of the question model. During training, the above-mentioned defect sub-graph embedding model and answer model can be jointly trained. The input of the answer model is the defect sub-graph represented by vectors. The defect sub-graph embedding model is to convert the defect sub-graph into a vector representation, and both models need to be trained. If the defect sub-graph embedding model is trained alone, there is no good evaluation index. When the two are jointly trained, the evaluation index is the accuracy of the answer model. In addition, joint training can optimize both models at the same time and save the model training time.
[0066] In some embodiments of the present application, the defect sub-graph embedding model further includes:
[0067] Using a graph convolutional network to learn the problem representation T of the defect sub-graph, the calculation formula is as follows:
[0068]
[0069] where, W X is the feature matrix of the defect sub-graph, W A is the adjacency matrix of the defect sub-graph, W D is the degree matrix of the defect sub-graph, W 1 is the weight matrix from the input layer to the hidden layer of the graph convolutional network, W 2 is the weight matrix from the hidden layer to the output layer of the graph convolutional network, N sg is the number of nodes in the defect sub-graph.
[0070] In some embodiments of the present application, the evaluation method further includes:
[0071] Constructing an answer model based on a convolutional neural network and a fully connected neural network, the calculation formula is as follows:
[0072] FA = LeakyReLU(∑ x∈fmx·k + b), fm ∈ QA, k ∈ K;
[0073]
[0074] P c = Sigmoid(W 0 [FA 1 ; FA 2 + b 0 );
[0075] Where, FA is the feature extracted using the convolution kernel K, fm is the feature map from QA, the solution model uses the first convolution kernel and the second convolution kernel, the first convolution kernel is used to extract the local features FA of the defect subgraph problem representation and the candidate answer representation respectively 1 , and the second convolution kernel is used to extract the global features FA of the defect subgraph problem representation and the candidate answer representation 2 , T is the problem representation of the defect subgraph, ca is the vector representation of the candidate answer, P c is the probability that the candidate answer is the correct answer to the question, b 0 is the bias term. The bias term b 0 will be continuously optimized during the training process.
[0076] In some embodiments of the present application, the evaluation method for the knowledge graph can be applied to the photovoltaic field, that is, to evaluate the quality of the knowledge graph in the photovoltaic field. Exemplarily, assume that there is already a photovoltaic maintenance knowledge graph, and an intelligent maintenance system has been constructed with this as the core. Initialize the knowledge graph embedding model, defect subgraph embedding model, solution model, and question model, and package them as sub-modules of a system. The input item of this sub-module is the knowledge graph, and the output item is the set of questions finally generated by the trained solution model and question model. All questions in this set are represented by vectors. Search for a third-party photovoltaic maintenance knowledge graph (free or paid). The third party uses the interface of this sub-module to train the model in its own data space. The applicant also uses the interface for model training. Neither party can see the detailed information of the other party's knowledge graph. The two parties exchange the set of questions, and use the solution model to answer. Finally, the two parties exchange the answers and have them scored by the other party. A higher score indicates better quality of the knowledge graph. If the third party can provide the original subgraph corresponding to the wrongly answered question, it can manually supplement the knowledge that may be missing in the applicant's knowledge graph, or it can be automatically supplemented in an intelligent way. The main function of this sub-module is to judge the quality of the knowledge graph. The wrong question set can be used as an intermediate result to supplement the missing knowledge, and can dynamically and continuously enrich the applicant's knowledge graph.
[0077] Based on a similar technical concept, the present application discloses an evaluation system for a knowledge graph. Figure 2 It is a structural schematic diagram of an evaluation system for a knowledge graph provided by an embodiment of the present application. Please refer toFigure 2 , an evaluation system 200 for a knowledge graph includes: a knowledge embedding module 210 configured to receive a first knowledge graph and a second knowledge graph, and based on the TransE model, incrementally train to obtain a first knowledge graph embedding model and a second knowledge graph embedding model, where the first knowledge graph embedding model corresponds to the first knowledge graph and the second knowledge graph embedding model corresponds to the second knowledge graph; a first adversarial module 220 configured to train a first question model and a first answer model based on the first knowledge graph embedding model, and train a second question model and a second answer model based on the second knowledge graph embedding model, where the first answer model is configured to answer questions of the first question model, the first question model is configured to adjust the difficulty of the questions according to the answers of the first answer model, the second answer model is configured to answer questions of the second question model, and the second question model is configured to adjust the difficulty of the questions according to the answers of the second answer model; a second adversarial module 230 configured to use the first answer model to answer questions of the second question model, record a first score according to the answer result, and use the second answer model to answer questions of the first question model, and record a second score according to the answer result; a quality evaluation module 240 configured to evaluate the first knowledge graph and the second knowledge graph based on the first score and the second score.
[0078] Related embodiments of this evaluation system can refer to the embodiments of the above evaluation method and will not be elaborated here.
[0079] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For parts not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. In addition, the above embodiments can be freely combined as needed.
Claims
1. A method for evaluating a knowledge graph, characterized in that: include: Based on the TransE model, the first triple of the first knowledge graph is incrementally used to train the first knowledge graph embedding model, and based on the TransE model, the second triple of the second knowledge graph is incrementally used to train the second knowledge graph embedding model; Training a first questioning model and a first answering model based on the first knowledge graph embedding model, wherein the first answering model is configured to answer questions of the first questioning model, and the first questioning model is configured to adjust the difficulty of the question according to the answer of the first answering model; Training a second questioning model and a second answering model based on the second knowledge graph embedding model, wherein the second answering model is configured to answer questions of the second questioning model, and the second questioning model is configured to adjust the difficulty of the question according to the answer of the second answering model; Using the first answer model to answer the questions of the second question model, recording a first score according to the answer result, and using the second answer model to answer the questions of the first question model, recording a second score according to the answer result; The first knowledge graph and the second knowledge graph are evaluated based on the first score and the second score.
2. The evaluation method according to claim 1, characterized in that: Also includes: During training, the first knowledge graph embedding model and the second knowledge graph embedding model are exchanged so that the first triplet trains the second knowledge graph embedding model, and the second triplet trains the first knowledge graph embedding model.
3. The evaluation method according to claim 1, characterized in that: The first knowledge graph and the second knowledge graph are knowledge graphs in the same field, and the entity embedding and relationship embedding in the first knowledge graph and the second knowledge graph are expressed the same in the corresponding first knowledge graph embedding model and the second knowledge graph embedding model.
4. The evaluation method according to claim 1, characterized in that: The first questioning model is configured to adjust the difficulty of the question according to the answer of the first answering model, and the second questioning model is configured to adjust the difficulty of the question according to the answer of the second answering model, both of which include the following process: Based on the answer results of the first answer model and the second answer model, mark the questions of the first question model and the second question model as correct or incorrect answers respectively, and repeat the marking process; Using the labeled questions, classify and train the question models corresponding to the first question model and the second question model; Based on the classification training result, the question model classifies its own question set by question level; Based on the question level classification, the first questioning model and the second questioning model adjust the question difficulty by selecting the question level classification.
5. The evaluation method according to claim 1, characterized in that: The first knowledge graph includes a first positive example set, and the first positive example set includes a plurality of the first triples; The method of incrementally training the first triplet of the first knowledge graph to obtain the first knowledge graph embedding model includes: Replacing the head entity or the tail entity in the first triplet according to a uniformly distributed random probability to obtain a third triplet, wherein the first negative example set includes a plurality of the third triplet; reducing a first distance between a sum of a head entity and a first relationship vector and a tail entity in each of the first triples of the first positive example set, and increasing a second distance between a sum of a head entity and a second relationship vector and a tail entity in each of the third triples of the first negative example set; Using the characteristics of triples in the vector space, performing multiple iterative calculations to make the first distance close to 0 or maximize the second distance, thereby determining the first relationship vector or the second relationship vector and obtaining the first knowledge graph embedding model; The steps for determining the second knowledge graph embedding model are the same as the steps for determining the first knowledge graph embedding model.
6. The evaluation method according to claim 1, characterized in that: Also includes: Randomly extracting a subgraph based on the knowledge graph, randomly deleting node information from the subgraph to construct a defect subgraph, using a graph convolutional network to construct a defect subgraph embedding model, training the first question model and the second question model with the defect subgraph embedding model, and using the vector representation of the defect subgraph embedding model as one of the input parameters of the answer model when training the answer model; Among them, the solution model includes the first solution model and the second solution model, the defect subgraph embedding model is used to characterize the problem, and the randomly deleted node information is the relationship information or entity information represented by any node in the subgraph, and the position information of the deleted node is saved in the defect subgraph embedding model.
7. The evaluation method according to claim 6, characterized in that: The defect subgraph embedding model also includes: Use graph convolutional network to learn the problem representation T of the defect subgraph, and the calculation formula is as follows: Among them, W X is the feature matrix of the defect subgraph, W A is the adjacency matrix of the defect subgraph, W D is the degree matrix of the defect subgraph, W1 is the weight matrix from the input layer to the hidden layer of the graph convolutional network, W2 is the weight matrix from the hidden layer to the output layer of the graph convolutional network, N sg is the number of nodes in the defect subgraph.
8. The evaluation method according to claim 7, characterized in that: Also includes: The solution model is constructed based on convolutional neural network and fully connected neural network, and the calculation formula is as follows: FA=LeakyReLU(∑ x∈fm x·k+b),fm∈QA,k∈K; P c =Sigmoid(W0[FA1;FA2]+b0); Wherein, FA is the feature extracted using convolution kernel K, fm is the feature map from QA, the answer model uses the first convolution kernel and the second convolution kernel, the first convolution kernel is used to extract the local features FA1 of the defect subgraph problem representation and the candidate answer representation, the second convolution kernel is used to extract the global features FA2 of the defect subgraph problem representation and the candidate answer representation, T is the problem representation of the defect subgraph, ca is the vector representation of the candidate answer, P c is the probability that the candidate answer is the correct answer to the question, and b0 is the bias term.
9. The evaluation method according to any one of claims 1 to 8, characterized in that: The question model includes the first question model and the second question model, and the question model is a naive Bayes model. Questions marked as correct answers and wrong answers will enter the training of the question model in a certain proportion.
10. An evaluation system for knowledge graphs, characterized in that: include: A knowledge embedding module is configured to receive a first knowledge graph and a second knowledge graph, and to incrementally train a first knowledge graph embedding model and a second knowledge graph embedding model based on a TransE model, wherein the first knowledge graph embedding model corresponds to the first knowledge graph, and the second knowledge graph embedding model corresponds to the second knowledge graph; A first adversarial module is configured to train a first question model and a first answer model based on the first knowledge graph embedding model, and to train a second question model and a second answer model based on the second knowledge graph embedding model, wherein the first answer model is configured to answer questions of the first question model, the first question model is configured to adjust the difficulty of the question according to the answer of the first answer model, the second answer model is configured to answer questions of the second question model, and the second question model is configured to adjust the difficulty of the question according to the answer of the second answer model; A second confrontation module is configured to use the first answer model to answer the question of the second question model, record a first score according to the answer result, and use the second answer model to answer the question of the first question model, and record a second score according to the answer result; A quality evaluation module is configured to evaluate the first knowledge graph and the second knowledge graph based on the first score and the second score.