A method and system for evaluating manufacturing map value

By using a corpus-based approach combined with sentence-level retrieval and question-answering models to evaluate manufacturing knowledge graphs, this approach solves the problems of single evaluation and manual intervention in existing technologies. It achieves efficient, flexible, and accurate knowledge graph value evaluation, making it suitable for downstream applications in the manufacturing industry.

CN117874257BActive Publication Date: 2025-12-16FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410062445.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-12-16
Estimated Expiration
2044-01-16

AI Technical Summary

Technical Problem

Existing methods for assessing the quality of knowledge graphs in the manufacturing industry are simplistic, neglect important factors, rely on manual intervention, are inefficient and inaccurate, cannot handle large-scale dynamic data, and suffer from inconsistencies and insufficient flexibility.

Method used

A corpus-based evaluation method is adopted, which generates a question set through a sentence-level retrieval model and a question-answering model, combines multiple machine reading comprehension models to evaluate the value of the knowledge graph, uses BERT and TF-IDF models to calculate sentence weights, dynamically generates question templates, integrates multiple question-answering models to eliminate bias, and achieves automated evaluation.

Benefits of technology

It achieves efficient and flexible knowledge graph value assessment, meets the needs of downstream applications in the manufacturing industry, has high interpretability and reusability, reduces the cost and time of manual intervention, and improves the accuracy and consistency of assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117874257B_ABST
    Figure CN117874257B_ABST
Patent Text Reader

Abstract

The application provides a manufacturing industry knowledge graph value evaluation method and system, wherein the method comprises the following steps: step 1, based on the information provided by the downstream application, the texts in the corpus are sorted, step 2, a question set is generated through the knowledge graph, and the score of each sentence answering the question set is obtained; step 3, based on the calculation results of step 1 and step 2, the value evaluation score of the knowledge graph for the downstream application is obtained. The application provides a manufacturing industry knowledge graph value evaluation method and system based on a corpus, which integrates value evaluation, question discovery and value improvement into one framework, is more in line with the actual needs of the downstream application of the manufacturing industry, has high interpretability and reusability, and is low in cost and high in efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application mainly relates to the field of knowledge graph and natural language processing, and particularly relates to a method and system for evaluating the value of a manufacturing industry knowledge graph. BACKGROUND

[0002] At present, the quality evaluation of a knowledge graph includes three stages of quality evaluation, problem discovery and quality improvement, and is mainly carried out from several dimensions such as accuracy, completeness, consistency, timeliness, credibility and usability. For the above dimensions, the existing methods proposed by researchers include manual methods, statistical learning-based methods and rule-based methods.

[0003] Among them, manual evaluation is usually the main method of quality evaluation. Using sampling and crowdsourcing technology, manual evaluation can be involved in the process of quality evaluation to give the quality of the knowledge graph, and the dimensions mainly focus on accuracy. Statistical learning-based methods are currently relatively mature. These methods detect outliers and predict missing types based on statistical distribution, train machine learning classifiers by manually extracting features such as type frequency and path, and use various representation learning techniques for link prediction to improve learning effect using external resources. These statistical learning methods mainly focus on the completeness dimension, and discover incomplete based on statistical learning and complete the knowledge graph. Rule-based methods use existing or mined predicate logic rules, RDF query language SPARQL, web ontology language (OWL) and graph pattern rules, and use symbolic reasoning techniques to reason about the knowledge graph and obtain parts that do not conform to the rules. These rule-based methods can be directly used to discover and improve error data and incompleteness in the knowledge graph, and have high interpretability and partial reusability.

[0004] Although these studies have made some progress, there are still some defects in the quality evaluation of manufacturing industry knowledge graphs. First, the current evaluation methods are usually single, only considering one or several specific indicators to measure the quality of the manufacturing industry knowledge graph, while ignoring other important factors. For example, some methods only focus on the coverage of the manufacturing industry knowledge graph, ignoring its accuracy and usability, which may lead to inaccurate evaluation results. In addition, each manufacturing industry knowledge graph has its specific purpose and goal, so the quality of the manufacturing industry knowledge graph needs to be evaluated according to different purposes and goals. However, single measurement indicators may not meet the different evaluation needs and may not be flexible enough in use.

[0005] Another problem is that manufacturing knowledge graph quality evaluation often requires a lot of manual intervention, some methods require manual annotation of data sets, some methods require manual definition of evaluation indicators, and some methods require manual judgment of the accuracy of evaluation results. This kind of manufacturing knowledge graph value evaluation method with manual intervention has some problems. First, they may be inefficient. Manual intervention requires a lot of time and effort, and manual intervention may also bring some errors. Therefore, the manufacturing knowledge graph value evaluation method using manual intervention may not meet the needs of manufacturing knowledge graph with large scale or higher demand scenarios. In addition, the manufacturing knowledge graph value evaluation method with manual intervention may not be reliable. Manual intervention depends on human subjective judgment, so it may be affected by personal bias or subjective factors. If manual intervention is not rigorous or experienced enough, the evaluation results may not be accurate.

[0006] Manufacturing knowledge graph quality evaluation also faces some technical challenges. First, manufacturing knowledge graph data sets are usually very large, containing a large number of entities and relationships. This means that efficient algorithms are needed to process these data, otherwise the evaluation process can be very slow. Second, the value of manufacturing knowledge graph is often dynamic, how to monitor and evaluate the value of manufacturing knowledge graph in real time is also a problem. In addition, manufacturing knowledge graph data sets are usually from different sources, such as enterprise databases, logs, text reports, etc. This means that there may be inconsistencies in the entities and relationships in the manufacturing knowledge graph, which will affect the consistency of the manufacturing knowledge graph and lead to biased evaluation results. Therefore, the manufacturing knowledge graph quality evaluation method needs to consider these inconsistencies. Finally, the manufacturing knowledge graph quality evaluation method may also face some practical application challenges. For example, the manufacturing knowledge graph quality evaluation method needs to be able to meet different application scenarios and evaluation goals. Therefore, the manufacturing knowledge graph quality evaluation method needs to have a certain flexibility, and can be customized according to different scenarios and goals. SUMMARY

[0007] The present application is to solve the above problems, and aims to provide a manufacturing knowledge graph value evaluation method and system based on corpus, which integrates value evaluation, problem discovery and value improvement into one framework.

[0008] The present application provides a manufacturing knowledge graph value evaluation method, which has the following characteristics, including the following steps:

[0009] Step 1, based on the information provided by the downstream application, the texts in the corpus are sorted to obtain the importance of each sentence in the text to the downstream application;

[0010] Step 2, generate a question set through the knowledge graph, and answer the question set by the text to obtain the score of each sentence answer question set;

[0011] Step 3, based on the calculation results of step 1 and step 2, obtain the value evaluation score of the knowledge graph for downstream applications.

[0012] In the method for evaluating the value of the manufacturing knowledge graph provided by the application, the information in step 1 can further have the following characteristics: the key application words / sentences correspond to the key words / sentences in the sentences, the sentences in the text are sorted according to the key application words / sentences, and a key word / sentence weight ordering model in each sentence in the text is obtained, which is used to represent the importance of each sentence to the downstream application.

[0013] In the method for evaluating the value of the manufacturing knowledge graph provided by the application, the specific process of step 1 can further have the following characteristics:

[0014] Obtain the corpus T and the set of key application words / sentences provided by the downstream application W, obtain the key word / sentence weight ordering model R in each sentence in the corpus T, and represent it as:

[0015]

[0016] Where S k is the set of key words / sentences in the kth sentence in the text of the corpus T, W k is the set of key words / sentences in the text of the corpus T, S k is the final similarity normalization score in the text, W k is larger, the sentence weight of the kth sentence is larger, the priority of S k is higher, W k is smaller, the sentence weight of the kth sentence is smaller, the priority of S k is lower, k=1,2....n, n represents the number of sentences in the text,

[0017] The weight ordering model includes a BERT fine-tuned sentence-level retrieval model and a TF-IDF similarity calculation model,

[0018] The sentence-level retrieval model is based on BERT, and is retrained using the MSMARCO retrieval task dataset to obtain the embedding of each sentence. The dot product result of the key word / sentence and the embedding of each sentence is taken as the similarity score, denoted as

[0019] The TF-IDF similarity calculation model converts the text into a term frequency representation to obtain the dot product result of the key word / sentence and each sentence as the similarity score, denoted as

[0020] Final similarity Wherein, alpha = 0.8, beta = 0.2, linear normalization is carried out to the final similarity, and finally

[0021] In the method for evaluating the value of the manufacturing knowledge graph provided by the application, the specific process of step 2 can have the following features:

[0022] The corpus T and the knowledge graph G are obtained, the knowledge graph G includes a series of triples <s, p, o>, a question set D is generated through the set of triples, and finally the score of each sentence in answering the questions in the question set D is obtained, which is represented as:

[0023]

[0024] Wherein S m is the mth sentence in the text of the corpus T, score m is the score of the question set in the question and answer of S m , m = 1, 2,..., n, and n represents the number of sentences in the text.

[0025] In the method for evaluating the value of the manufacturing knowledge graph provided by the application, the specific process of step 2 can have the following features: based on the entity in the triple, a question set is obtained by combining the artificially defined question template, and the question template can be dynamically replaced.

[0026] In the method for evaluating the value of the manufacturing knowledge graph provided by the application, the specific process of step 2 can have the following features: multiple question and answer models are integrated to eliminate the deviation of different models and improve the accuracy of question and answer.

[0027] In the method for evaluating the value of the manufacturing knowledge graph provided by the application, the specific process of step 2 can have the following features: the multiple question and answer models are three different machine reading comprehension question and answer models, which are ABCNN, MnemonicReader and QANet respectively, each word in the sentence is represented by a pre-trained word vector, and then the neural network is used to encode each question in the question set and each sentence in the text by combining the attention mechanism, and finally the candidate answer is predicted,

[0028] A multilayer perceptron is connected to the last layer of the three different machine reading comprehension question and answer models to calculate the similarity between the candidate answer and the actual answer, and the question and answer result is obtained through logistic regression, and the question and answer result of each machine reading comprehension question and answer model is 0 or 1, which is used to represent whether the sentence in the text can correctly answer the question generated by the triple.

[0029] If for the same question, if the i-th sentence S i If the question and answer results of two or more machine reading comprehension question and answer models are 1, S i The score of the question and answer question set score i is 1,

[0030] Otherwise, s i The score of the question and answer question set score i is 0.

[0031] In the method for evaluating the value of the manufacturing knowledge graph provided by the application, the method can have the following features: wherein,

[0032] In step 3, the value evaluation score is:

[0033]

[0034] Wherein, n represents the number of sentences in the text.

[0035] In the method for evaluating the value of the manufacturing knowledge graph provided by the application, the method can have the following features: wherein, in step 3, a plurality of knowledge graphs are selected, a plurality of values (G) corresponding to the plurality of knowledge graphs are calculated respectively, and the average value of the plurality of values (G) is calculated,

[0036] If the value (G) is higher than the average value, the questions generated by the triples of the knowledge graph can be correctly answered by the sentences in the text, wherein when the value (G) is 1, it means that the questions generated by each triple of the knowledge graph can be correctly answered by the sentences in the text,

[0037] If the value (G) is lower than the average value, the questions of the knowledge graph are analyzed and improved, when the sentences with larger sentence weights cannot correctly answer the questions of the question set, the knowledge graph needs to be supplemented or updated in time, when there are a large number of sentences with low sentence weights that can correctly answer the questions of the question set, the triples matched by the sentences can be preferentially deleted, when the value (G) is 0, it means that the questions generated by each triple cannot be correctly answered by the sentences in the text, the triples of the knowledge graph are completely irrelevant to the text, and the triples can be preferentially deleted.

[0038] The application provides a system for evaluating the value of a manufacturing knowledge graph, which uses the above-mentioned method for evaluating the value of a manufacturing knowledge graph and has the following features, comprising:

[0039] A sorting unit sorts the text in the corpus based on the information provided by the downstream application to obtain the importance of each sentence in the text to the downstream application;

[0040] The answering unit generates a question set through the knowledge graph and answers each sentence to obtain a score of each sentence answering question set;

[0041] The value evaluation unit obtains a value evaluation score of the knowledge graph for downstream applications based on the calculation results of the ranking unit and the answering unit.

[0042] Effects of the application

[0043] The application proposes a method and system for evaluating the value of a manufacturing knowledge graph based on a corpus, which integrates value evaluation, question discovery and value improvement into one framework.

[0044] The application is more in line with the actual needs of downstream applications in the manufacturing industry, has high interpretability and reusability, and is low-cost and efficient. BRIEF DESCRIPTION OF DRAWINGS

[0045] Figure 1 is a flowchart of the method for evaluating the value of a manufacturing knowledge graph in embodiment 1 of the application;

[0046] Figure 2 is a manufacturing knowledge graph value evaluation framework based on a corpus in embodiment 1 of the application; and

[0047] Figure 3 is an example of text and knowledge graph in the steel manufacturing industry in embodiment 1 of the application. DETAILED DESCRIPTION

[0048] In order to make the technical means, creative features, purposes and effects achieved by the application easy to understand, the following embodiments will be described in detail in combination with the drawings.

[0049] <Embodiment 1>

[0050] Figure 1 is a flowchart of the method for evaluating the value of a manufacturing knowledge graph in embodiment 1 of the application, Figure 2 is a manufacturing knowledge graph value evaluation framework based on a corpus in embodiment 1 of the application.

[0051] As shown in Figure 1 , 2 The present embodiment proposes a method for evaluating the value of a manufacturing knowledge graph, comprising the following steps:

[0052] Step 1: Based on the information provided by the downstream application, the texts in the corpus are ranked to obtain the importance of each sentence in the text to the downstream application.

[0053] The main purpose of Step 1 is to quantify the importance of the knowledge required by downstream applications. Considering the abstract nature of knowledge, text serves as a concrete carrier of knowledge, making it specific. Downstream applications can provide key application words / phrases to represent the knowledge they need. These key application words / phrases correspond to the keywords / phrases in the text. Finally, based on the key application words / phrases provided by the downstream applications, the sentences in the text are ranked to obtain a weighted ranking model of the keywords / phrases in each sentence. This model is used to represent the importance of each sentence to the downstream applications, i.e., the value of the knowledge.

[0054] Specifically, the input to step 1 is the set of key application words / phrases W from the corpus T and the information provided by the downstream application, and the final output is the weighted ranking model R of the keywords / phrases in each sentence of the corpus T, represented as:

[0055]

[0056] Among them, S k Let W be the set of keywords / sentences in the k-th sentence of the text in corpus T. k For S k The final similarity normalized score in the text, W k The larger S is, the greater the sentence weight of the k-th sentence. k The higher the priority, the better. k The smaller the value of S, the smaller the sentence weight of the k-th sentence. k The lower the priority, the more likely it is to be assigned a number of sentences in the text, k = 1, 2, ..., n.

[0057] The weighted ranking model mainly consists of a sentence-level retrieval model fine-tuned based on BERT and a TF-IDF similarity calculation model. The sentence-level retrieval model is retrained on top of BERT using the MS MARCO retrieval task dataset. This model outputs the embedding of each sentence, and the dot product of the keyword / sentence and the embedding of each sentence in the text is used as the similarity score, denoted as . The TF-IDF similarity calculation model converts text into a word frequency representation and obtains the similarity score as the dot product of keywords / sentences and each sentence, denoted as . Final similarity Based on our experimental results, α = 0.8 and β = 0.2 showed good performance. The final similarity was then linearly normalized to obtain...

[0058] Step 2: Generate a question set using a knowledge graph, and answer the questions with text to obtain a score for each sentence answering the question set.

[0059] The main purpose of step 2 is to compare the keywords / sentences of the text with the relations of the triples in the knowledge graph, i.e. whether the knowledge graph contains the keywords / sentences in the text. The advantage of using the question-answering method to evaluate the knowledge graph is that it can quantitatively evaluate the information in the knowledge graph and obtain accurate and complete evaluation results through the sentences in the text.

[0060] Specifically, step 2 inputs the corpus T and the knowledge graph G, where G consists of a series of triples <s, p, o>, and a question set D is generated from the set of triples in G. The final output is the score of each sentence answering the questions in the question set D, denoted as:

[0061]

[0062] where S m is the mth sentence in the text of the corpus T, score m is the score of the question in the question set for the sentence S m , m = 1, 2,..., n, and n represents the number of sentences in the text.

[0063] Each sentence only needs to correctly answer one question in the question set D, and the score of the question in the question set for the sentence is 1.

[0064] Specifically, first, based on the entities in the triples, a question set is generated by combining the artificially defined question templates, and questions related to the triple entities are proposed for the text.

[0065] For example, for the triple <car manufacturing, main material, steel>, the question template "s's p is ()" generates the question "the main material of car manufacturing is ()", and the correct answer to this question is the tail entity of the triple "steel". Users can also dynamically replace the question template according to the target to adjust the generated questions, which also improves the flexibility and universality of the evaluation method.

[0066] In this embodiment, step 2 integrates multiple question-answering models to eliminate the bias of different models and improve the accuracy of question-answering. Three different machine reading comprehension question-answering models are selected, namely ABCNN, MnemonicReader and QANet. These models represent each word in the sentence with a pre-trained word vector, and then use an attention mechanism to encode each question in the question set and each sentence in the text using a neural network. Finally, the candidate answers are predicted. A multi-layer perceptron is connected to the last layer of these models to calculate the similarity between the candidate answers and the actual answers, and the question-answering result is obtained through logistic regression. The question-answering result of each machine reading comprehension question-answering model is 0 or 1, indicating whether the sentence in the text can correctly answer the question generated from the triple. If for the same question, the i th sentence si In two or more machine reading comprehension question and answer models, if the output question and answer result is 1, S i The score of the question in the question and answer question set score i is 1, otherwise, S i The score of the question in the question and answer question set score i is 0.

[0067] Step 3, according to the calculation results of step 1 and step 2:

[0068] And

[0069]

[0070] The final evaluation result is obtained, that is, the value evaluation score of the knowledge graph to the downstream application:

[0071]

[0072] Wherein, n represents the number of sentences in the text.

[0073] Select multiple knowledge graphs, calculate multiple values (G) corresponding to multiple knowledge graphs, and calculate the average value of multiple values (G),

[0074] If the value (G) of the embodiment is higher than the average value, the question generated by the triples of the knowledge graph of the embodiment can be correctly answered by the sentences in the text, wherein when the value (G) of the embodiment is 1, it means that the question generated by each triple of the knowledge graph of the embodiment can be correctly answered by the sentences in the text,

[0075] If the value (G) of the embodiment is lower than the average value, analyze the problem and improvement of the knowledge graph of the embodiment, when the sentence with larger sentence weight cannot correctly answer the question set, the knowledge graph of the embodiment needs to be supplemented or updated in time; When there are a large number of sentences with low sentence weight that can correctly answer the question set, the triple matched by the sentence can be preferentially deleted, when the value (G) of the embodiment is 0, it means that the question generated by each triple cannot be correctly answered by the sentences in the text, the triples of the knowledge graph of the embodiment are completely irrelevant to the text, and the triples can be preferentially deleted.

[0076] Figure 3 The text and knowledge graph of the steel manufacturing industry in embodiment 1 of the application are examples.

[0077] For example, Figure 3The triples and sentences (S), sentence weights (W) and scores of the triple question answering related to the key sentence "the main material of automobile manufacturing" are shown. Some of the sentences are not associated with any triples, which means that the manufacturing knowledge graph involves less or no knowledge about these sentences. Using such a knowledge graph to support downstream applications related to key sentences may cause some problems, but the size of the impact depends on the sentence weight of these sentences. For example, the sentence weight of the sentence "Steel is the main material of automobile manufacturing, and its high strength and corrosion resistance make it the preferred material for automobile manufacturing." is greater than that of the sentence "Aluminum alloy is used in automobile manufacturing to manufacture body, engine and suspension system components to improve the fuel efficiency and performance of the automobile." Therefore, the priority of the keywords contained in the former is higher, and if there is no relevant content in the knowledge graph, it needs to be supplemented or updated in time. Triples not associated with any sentence are considered to be triples of little value. For example, the triple <car, main fuel, diesel> is not much related to the field of automobile manufacturing, because it is more about the use of cars rather than manufacturing, so when the manufacturing knowledge graph needs to be updated to remove redundancy, such triples can be deleted first.

[0078] <Embodiment 2>

[0079] The embodiment proposes a manufacturing knowledge graph value evaluation system using the above manufacturing knowledge graph value evaluation method, which comprises:

[0080] A sorting unit sorts the texts in the corpus based on the information provided by the downstream application to obtain the importance of each sentence in the text to the downstream application.

[0081] An answering unit generates a question set through the knowledge graph and answers the question set by the sentence to obtain the score of each sentence answering the question set.

[0082] A value evaluation unit obtains the value evaluation score of the knowledge graph for the downstream application based on the calculation results of the sorting unit and the answering unit.

[0083] It should be understood that each part of the manufacturing knowledge graph value evaluation system described in the embodiment can correspond to each step of the manufacturing knowledge graph value evaluation method described in Embodiment 1. Therefore, the operations, features and advantages described above for each step of the evaluation method also apply to each part of the evaluation system. For the sake of brevity, some operations, features and advantages are not described here.

[0084] Effects of the embodiment

[0085] The application provides a corpus-based manufacturing industry knowledge graph value evaluation method and system, which integrates value evaluation, problem discovery and value improvement into one framework.

[0086] The application is more in line with the actual needs of downstream applications in the manufacturing industry, has high interpretability and reusability, and is low-cost and efficient.

[0087] The application has certain flexibility and can be customized for evaluation according to different scenes and targets.

[0088] Those skilled in the art should understand that the application is not limited by the above embodiments, and the above embodiments and descriptions in the specification are only to illustrate the principles of the application, and various changes and improvements can be made without departing from the spirit and scope of the application, and these changes and improvements all fall within the scope of the claimed application. The scope of protection of the application is defined by the appended claims and their equivalents.

Claims

1. A method for evaluating the value of a manufacturing knowledge graph, characterized in that, The method comprises the following steps: Step 1, based on the information provided by the downstream application, the text in the corpus is sorted, obtaining the importance of each sentence in the text to the downstream application; In step 1, the information includes key application words / sentences, which correspond to the key words / sentences in the sentence, and the sentence in the text is sorted according to the key application words / sentences, and the weight ordering model of the key words / sentences in each sentence in the text is obtained, which is used to represent the importance of each sentence to the downstream application; The specific process of step 1 is: Obtain the corpus T and the set of key application words / sentences provided by the downstream application, and obtain the weight ordering model R of the key words / sentences in each sentence in the corpus T, which is represented as: wherein S k is the set of keywords / phrases in the kth sentence in the text of the corpus T, W k is the weight of the kth sentence in the text, S k is the final similarity normalized score in the text, W k is the weight of the kth sentence, S k is the priority of the kth sentence, W k is the weight of the kth sentence, S k is the priority of the kth sentence, k = 1, 2,... n, n represents the number of sentences in the text, The weight ordering model includes a BERT fine-tuned sentence-level retrieval model and a TF-IDF similarity calculation model, The sentence-level retrieval model is based on BERT, retrained using the MS MARCO retrieval task dataset, to obtain the embedding of each sentence, and the dot product result of the keyword / sentence and the embedding of each sentence is taken as the similarity score, denoted as The TF-IDF similarity calculation model converts the text into a word frequency representation, obtains the dot product result of the key words / sentences and each of the sentences as a similarity score, denoted as final similarity where a = 0.8, β = 0.2, the final similarity is linearly normalized, and finally Step 2, generate a question set through a knowledge graph, and answer the questions from the text to obtain the score of each sentence in answering the question set; The specific process of step 2 is: Obtain the corpus T and the knowledge graph G, the knowledge graph G includes a series of triples <s, p, o>, generate the question set D through the set of triples, and finally obtain the score of each sentence in answering the questions in the question set D, which is represented as: wherein S m is the mth sentence in the text of the corpus T, score m is the score of the mth sentence S m asks the question set of questions, m = 1, 2,... n, n represents the number of sentences in the text; Step 3, based on the calculation results of step 1 and step 2, obtain the value evaluation score of the knowledge graph for the downstream application.

2. The method for evaluating the value of the manufacturing knowledge graph according to claim 1, characterized in that: wherein Based on the entities in the triples, combined with the artificially defined question templates, the question set is obtained, and the question templates can be dynamically replaced.

3. The method for evaluating the value of the manufacturing knowledge graph according to claim 2, characterized in that: wherein Step 2 integrates multiple question and answer models to eliminate the bias of different models and improve the accuracy of question and answer.

4. The method for evaluating the value of the manufacturing knowledge graph according to claim 3, characterized in that: wherein, The multiple question and answer models are three different machine reading comprehension question and answer models, which are ABCNN, MnemonicReader and QANet, the machine reading comprehension question and answer model represents each word in the sentence with a pre-trained word vector, and then combines the attention mechanism to encode each question in the question set and each sentence in the text with a neural network, and finally predicts the candidate answer, In the last layer of the three different machine reading comprehension question and answer models, a multi-layer perceptron is connected to calculate the similarity between the candidate answer and the actual answer, and the question and answer result is obtained through logistic regression, and the question and answer result of each machine reading comprehension question and answer model is 0 or 1, which is used to represent whether the sentence in the text can correctly answer the question generated from the triples, If for the same question, if the i-th sentence S i In the case of two or more machine reading comprehension question and answer models, the question and answer results are 1, S i The score of the question set score i is 1, Otherwise, S i score of the question set i is 0.

5. The method for evaluating the value of the manufacturing knowledge graph according to claim 4, characterized in that: wherein, In step 3, the value evaluation score is: Where n represents the number of sentences in the text.

6. The method of claim 5, wherein: wherein In step 3, a plurality of knowledge graphs are selected, and a plurality of values (G) corresponding to the plurality of knowledge graphs are calculated, and an average value of the plurality of values (G) is calculated, If the value (G) is higher than the average value, the questions generated by the triples of the knowledge graph can be correctly answered by the sentences in the text, and when the value (G) is 1, it means that each question generated by the triples of the knowledge graph can be correctly answered by the sentences in the text, If the value (G) is lower than the average value, the questions of the knowledge graph are analyzed and improved, and when the sentences with larger sentence weights cannot correctly answer the questions in the question set, the knowledge graph needs to be supplemented or updated in time; when there are a large number of sentences with lower sentence weights that can correctly answer the questions in the question set, the triples matched by the sentences can be preferentially deleted, and when the value (G) is 0, it means that each question generated by the triples cannot be correctly answered by the sentences in the text, and the triples of the knowledge graph are completely irrelevant to the text, and the triples can be preferentially deleted.

7. A system for evaluating the value of a manufacturing knowledge graph, using the method for evaluating the value of a manufacturing knowledge graph according to any one of claims 1 to 6, characterized by, The method comprises: An ordering unit sorts the texts in the corpus based on information provided by a downstream application to obtain the importance of each sentence in the text to the downstream application; An answering unit generates a question set through a knowledge graph and answers the question set by the sentences to obtain a score of each sentence in answering the question set; A value evaluation unit obtains a value evaluation score of the knowledge graph for the downstream application based on the calculation results of the ordering unit and the answering unit.

Citation Information

Patent Citations

  • Knowledge graph-based knowledge question and answer verification code generation system and method

    CN111639187A

  • Vehicle maintenance scheme determination method and device, equipment and storage medium

    CN115034409A