Electric power knowledge graph question and answer method and system

By combining string matching technology and ERNIE 3.0 lightweight model, the knowledge graph is screened and matched, and the problems of low matching efficiency and low accuracy in knowledge graph questions and answers in the existing technology are solved, and a more efficient and accurate question and answer system is achieved.

CN120011486APending Publication Date: 2025-05-16XUCHANG XJ SOFTWARE TECHNOLOGIES LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311497145.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, when using deep learning models to perform knowledge graph questions and answers, the matching efficiency is reduced and the matching accuracy is not high.

Method used

By combining traditional string matching techniques (such as Jaccard algorithm and Word2Vec technology) and ERNIE 3.0 lightweight natural language model, the candidate knowledge set is first filtered, and then the deep learning model is used to calculate the statement similarity to obtain the matching results.

Benefits of technology

It improves the query speed and accuracy of the knowledge graph question and answer system, reduces the search for a large amount of knowledge graph data, saves time, and effectively solves the problems of matching efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011486A_ABST
    Figure CN120011486A_ABST
Patent Text Reader

Abstract

The invention relates to an electric power knowledge graph question answering method and system, and belongs to the technical field of knowledge graphs. The method comprises the following steps: firstly, acquiring an electric power knowledge question needing to be answered, preprocessing the electric power knowledge question, and then screening candidate knowledge sets by adopting a character string matching technology based on the preprocessed electric power knowledge question and the candidate knowledge sets in a knowledge graph; and finally, inputting the preprocessed electric power knowledge question and the screened candidate knowledge set into a trained deep learning model to obtain the similarity between each statement in the electric power knowledge question and the screened candidate knowledge set, and taking the statement with the highest similarity as a matching result. Compared with the prior art, the method solves the problems of low matching efficiency and low matching precision when a deep learning model is adopted for matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an electric power knowledge graph question-answering method and system, and belongs to the technical field of knowledge graphs. Background Art

[0002] The power grid safety operation management specification is a series of rules and regulations formulated by the State Grid Corporation of China to strengthen the safety production management of the power grid, standardize the safe operation behavior of the power grid, and ensure the safe and stable operation of the power grid. The power grid safety operation management specification involves many fields, including power equipment operation, maintenance, overhaul and construction. However, these specifications are mostly in the form of text, which is not conducive to users' rapid retrieval and standardized application. Therefore, it is very necessary to build a knowledge graph based on the power grid safety operation management specification.

[0003] One of the important applications of knowledge graph is as a knowledge base for automatic question-answering systems. Knowledge graph question-answering technology is an important part of the entire question-answering system. Currently, knowledge graph question-answering in existing technologies mostly uses keyword matching and template matching methods. Keyword matching can only match keywords and cannot consider semantic information. Therefore, it is impossible to accurately match some semantically similar words. Template matching requires manual construction of a large number of question templates with variables, selecting templates according to questions to form query expressions, and querying structured databases to generate answers. It takes a lot of manpower to proofread templates and maintain template libraries.

[0004] In order to solve the problems existing in knowledge graph question answering using keyword matching and template matching methods, the existing technology proposes to use deep learning models to match contextual semantics. However, when using deep learning models for matching, it is necessary to search data from a large amount of knowledge graph data, which will reduce matching efficiency and even affect matching accuracy. Summary of the invention

[0005] The purpose of the present invention is to provide an electric power knowledge graph question-answering method and system, which is used to solve the problems of reduced matching efficiency and low matching accuracy in the prior art when using deep learning models for matching.

[0006] To achieve the above purpose, the technical solution provided by the present invention is:

[0007] The present invention provides an electric power knowledge graph question-answering method, which comprises the following steps:

[0008] 1) Obtain the electricity knowledge questions that need to be answered and perform preprocessing;

[0009] 2) Based on the preprocessed power knowledge questions and candidate knowledge sets in the knowledge graph, the candidate knowledge sets are screened using string matching technology;

[0010] 3) The preprocessed electricity knowledge questions and the screened candidate knowledge set are input into the trained deep learning model to obtain the similarity between the electricity knowledge questions and the sentences in the screened candidate knowledge set, and the sentence with the highest similarity is taken as the matching result.

[0011] The present invention first preliminarily screens the input question, quickly excludes some irrelevant sentences and sentences with low correlation with the input sentence, reduces the number of sentences in the candidate knowledge set, and then uses the deep learning model to obtain the similarity between the sentences in the screened candidate knowledge set. Users do not need to search for data from a large amount of knowledge graph data, saving time in searching for data. Compared with the prior art, the present invention effectively solves the problems of reduced matching efficiency and low matching accuracy when using a deep learning model for matching.

[0012] Furthermore, the deep learning model adopts the ERNIE 3.0 lightweight model.

[0013] Furthermore, the ERNIE 3.0 lightweight model is obtained by compressing the ERNIE 3.0 natural language model through online distillation technology and quantization technology.

[0014] The present invention uses online distillation technology and quantization technology to compress the ERNIE 3.0 natural language model into an ERNIE 3.0 lightweight model. The online distillation technology and quantization technology can simplify a network model with a more complex structure into a model with a simple structure to save time for subsequent model training.

[0015] Furthermore, the training of the ERNIE 3.0 lightweight model includes a pre-training stage and a fine-tuning stage. The pre-training stage is to pre-train the ERNIE 3.0 lightweight model using the constructed training set; the fine-tuning stage is to fine-tune the pre-trained ERNIE 3.0 lightweight model using power knowledge questions as training data.

[0016] The present invention uses electricity knowledge questions as training data to fine-tune the pre-trained ERNIE 3.0 lightweight model, thereby obtaining a semantic model for the field of power failure, so that users' questions can be answered more accurately.

[0017] Furthermore, the ERNIE 3.0 lightweight model adopts the ERNIE 3.0-Base model.

[0018] Furthermore, in the step 2), the Jaccard algorithm and Word2Vec technology are used to screen the candidate knowledge set.

[0019] The present invention adopts Jaccard algorithm and Word2Vec technology to screen candidate knowledge sets. The Jaccard algorithm has high accuracy, simplicity and efficiency, and can greatly meet real-time requirements. The Word2Vec model is simple and has a fast training speed.

[0020] Furthermore, the preprocessing is to correct errors and remove redundancy for the electricity knowledge questions that need to be answered.

[0021] The present invention performs error correction and redundancy removal preprocessing operations on electric power knowledge questions that need to be answered, thereby improving the quality and accuracy of data and also improving the effect and reliability of data analysis.

[0022] In order to solve the above technical problems, the present invention also provides an electric power knowledge graph question and answer system, including a memory and a processor, and the processor is used to execute computer program instructions stored in the memory to implement the electric power knowledge graph question and answer method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is an overall implementation flow chart of the electric power knowledge graph question-answering method of the present invention;

[0024] Figure 2 It is a specific implementation flow chart of the electric power knowledge graph question-answering method of the present invention;

[0025] Figure 3 It is a model structure diagram of the ERNIE 3.0 natural language in the power knowledge graph question-answering method of the present invention. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0027] The electric power knowledge graph question-answering method proposed in the present invention first obtains the electric power knowledge question that needs to be answered and performs preprocessing, then uses string matching technology to screen the candidate knowledge set based on the preprocessed electric power knowledge question and the candidate knowledge set in the knowledge graph, and finally inputs the preprocessed electric power knowledge question and the screened candidate knowledge set into the trained deep learning model to obtain the similarity between the electric power knowledge question and each sentence in the screened candidate knowledge set, and takes the sentence with the highest similarity as the matching result. The present invention combines traditional string matching technology (Jaccard algorithm and Word2Vec technology) and ERNIE 3.0 natural language model, and can quickly and accurately answer users' questions in the knowledge graph, which not only improves the query speed and accuracy of the graph question-answering system, but also has good versatility and extensibility, and can also be applied to the construction of knowledge graph question-answering systems in other fields, effectively solving the problems of reduced matching efficiency and low matching accuracy in the prior art using deep learning models for matching.

[0028] Embodiment of the electric power knowledge graph question answering method:

[0029] The overall implementation process of the power knowledge graph question answering method of the present invention is as follows: Figure 1 As shown in the figure, first, the user inputs a question, and then the preprocessing module preprocesses the input question, and then performs rough knowledge matching. Through a fast matching strategy, the knowledge with low relevance to the input question is quickly eliminated, and then accurate knowledge matching is performed. According to the semantic similarity score of the question and the defective sentences in the knowledge set, they are sorted in descending order, and then the multi-hop query technology in the knowledge graph is used to gradually find the relevant knowledge nodes of the defect causes and corresponding solutions related to the question, and finally the obtained answers are organized into the form of knowledge cards for output.

[0030] The specific implementation process of the power knowledge graph question answering method of the present invention is as follows: Figure 2 As shown, the technical solution of the present invention is described as follows:

[0031] 1. Obtain the electricity knowledge questions that need to be answered and perform preprocessing

[0032] The present invention first obtains electric power knowledge questions that need to be answered, and then performs error correction and redundancy removal preprocessing operations on the electric power knowledge questions that need to be answered, so as to improve the quality and accuracy of data and improve the effect and reliability of data analysis.

[0033] 2. Based on the preprocessed power knowledge questions and candidate knowledge sets in the knowledge graph, the candidate knowledge sets are screened using string matching technology

[0034] The present invention adopts Jaccard algorithm and Word2Vec technology to compare the string similarity and coverage of the input sentence with the candidate knowledge set, and combines the Word2Vec similarity to obtain a score through comprehensive calculation. According to the preset threshold, some irrelevant candidate knowledge can be quickly excluded to complete the preliminary screening.

[0035] 3. Input the preprocessed power knowledge questions and the screened candidate knowledge set into the trained deep learning model to obtain the similarity between the power knowledge questions and the sentences in the screened candidate knowledge set, and take the sentence with the highest similarity as the matching result.

[0036] The deep learning model of the present invention adopts the ERNIE 3.0 lightweight model, which is formed by compressing the ERNIE 3.0 natural language model through online distillation technology and quantization technology. Online distillation technology uses a pre-trained large network to guide the training of a newly built small network, and quantization technology is based on data analysis to convert close values ​​into a number.

[0037] The training of the ERNIE 3.0 lightweight model of the present invention includes a pre-training stage and a fine-tuning stage. The pre-training stage is to pre-train the ERNIE 3.0 lightweight model using the constructed training set. The training data set uses the KBQA public data set, which is one of the questions in the NLPCC2018 competition. The KBQA public data set contains about 32,000 sentence pairs. Two labels of 0 and 1 are provided according to whether the sentence pairs are semantically the same. The model automatically learns the semantic similarity problem between the sentence pairs, ignoring different sentence expressions and question words; the fine-tuning stage is to use the power knowledge questions as training data to fine-tune the pre-trained ERNIE 3.0 lightweight model. Specifically, the ERNIE 3.0-Base model is used in this embodiment, and a structure of 12 layers, 768 hidden units and 64 heads is adopted.

[0038] The ERNIE 3.0 model uses the Encoder part of the Transformer as the model backbone for training. The model structure of the ERNIE 3.0 natural language in the power knowledge graph question answering method of the present invention is as follows: Figure 3As shown in the figure, the ERNIE 3.0 natural language model mainly consists of two parts: the text encoder T-Encoder and the knowledge encoder K-Encoder. For the T-Encoder, it is used to obtain the lexical and syntactic semantic information of the input token. It needs to sum the token embedding, segment embedding and positional embedding to obtain the input embedding, and then use the multi-layer bidirectional Transformer encoder to extract the semantic features. For K-Encode, it is used to integrate knowledge information into the text information from the bottom layer, so that the token information can be represented in a unified feature.

[0039] The ERNIE 3.0 model can be used to match electricity knowledge questions more accurately. The similarity between the electricity knowledge questions and the sentences in the screened candidate knowledge set is calculated, the similarities are sorted from high to low, and the sentence with the highest similarity is taken as the matching result.

[0040] Embodiment of the electric power knowledge graph question answering system:

[0041] The electric power knowledge graph question-answering system includes a memory and a processor, and the processor is used to execute computer program instructions stored in the memory to implement the electric power knowledge graph question-answering method of the present invention. The specific content of the electric power knowledge graph question-answering method has been described in detail in the electric power knowledge graph question-answering method embodiment, and will not be repeated in this embodiment.

Claims

1. A method for answering questions on an electric power knowledge graph, characterized in that: The electric power knowledge graph question answering method comprises the following steps: 1) Obtain the electricity knowledge questions that need to be answered and perform preprocessing; 2) Based on the preprocessed power knowledge questions and candidate knowledge sets in the knowledge graph, the candidate knowledge sets are screened using string matching technology; 3) The preprocessed electricity knowledge questions and the screened candidate knowledge set are input into the trained deep learning model to obtain the similarity between the electricity knowledge questions and the sentences in the screened candidate knowledge set, and the sentence with the highest similarity is taken as the matching result.

2. The electric power knowledge graph question-answering method according to claim 1, characterized in that: The deep learning model adopts the ERNIE 3.0 lightweight model.

3. The electric power knowledge graph question-answering method according to claim 2 is characterized in that: The ERNIE 3.0 lightweight model is obtained by compressing the ERNIE 3.0 natural language model through online distillation technology and quantization technology.

4. The electric power knowledge graph question-answering method according to claim 2, characterized in that: The training of the ERNIE 3.0 lightweight model includes a pre-training stage and a fine-tuning stage. The pre-training stage is to pre-train the ERNIE 3.0 lightweight model using the constructed training set; the fine-tuning stage is to fine-tune the pre-trained ERNIE3.0 lightweight model using power knowledge questions as training data.

5. The electric power knowledge graph question-answering method according to any one of claims 2 to 4, characterized in that: The ERNIE 3.0 lightweight model adopts the ERNIE 3.0-Base model.

6. The electric power knowledge graph question-answering method according to claim 1 or 2, characterized in that: In the step 2), the Jaccard algorithm and Word2Vec technology are used to screen the candidate knowledge set.

7. The electric power knowledge graph question-answering method according to claim 1, characterized in that: The preprocessing is to correct errors and remove redundancy for the electricity knowledge questions that need to be answered.

8. An electric power knowledge graph question-answering system, characterized in that: It includes a memory and a processor, and the processor is used to execute computer program instructions stored in the memory to implement the power knowledge graph question and answer method as described in any one of claims 1 to 7.