Intelligent energy knowledge question-answering method and system based on knowledge graph

By building an intelligent Q&A system based on knowledge graph, the internal knowledge asset management problems of energy enterprises are solved, rapid information acquisition and efficient customer service are achieved, and new business opportunities and more accurate market insights are brought to energy companies.

CN120011496APending Publication Date: 2025-05-16STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411937007.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Knowledge assets within energy companies are difficult to manage and share, resulting in inefficient information acquisition, long-term data transmission and error-prone.

Method used

Using an intelligent question-and-answer method based on knowledge graphs, collect energy knowledge text data through crawler tools, build a knowledge graph and integrate it with the BERT model, and establish an intelligent question-and-answer system to quickly respond to user queries and provide the required information.

Benefits of technology

It improves the efficiency of internal information acquisition and customer service of enterprises, promotes team collaboration and learning, provides energy companies with new business models and service methods, and helps enterprises more accurately grasp market dynamics and customer needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011496A_ABST
    Figure CN120011496A_ABST
Patent Text Reader

Abstract

The invention discloses an energy knowledge intelligent question-answering method and system based on a knowledge graph. The method comprises the following steps: S1, acquiring energy knowledge text data and generating a knowledge base; s2, constructing a knowledge graph and fusing the knowledge graph with a BERT model; and S3, constructing a question and answer system. According to the invention, enterprise employees and customer query can be quickly responded, required information is provided, and the working efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent question answering, and in particular to an energy knowledge intelligent question answering method and system based on a knowledge graph. Background Art

[0002] In life, intelligent question and answer is a method that uses natural language processing, machine language, data mining and other technologies to build a method that can understand users' natural language questions and provide accurate answers. It is based on massive data on the Internet and deep semantic understanding technology. At present, intelligent question and answer systems are widely used in customer service, business, law, finance, education, medical health and other fields. In the field of education, intelligent question and answer systems can provide students with personalized learning guidance. In the field of medical health, intelligent question and answer systems can help patients answer questions and provide preliminary diagnostic suggestions. Intelligent question and answer systems are playing an increasingly important role in daily work.

[0003] Energy companies usually involve a large amount of data and information. The knowledge assets within the company are difficult to manage and share. The traditional way for employees or customers to obtain information is through paper documents, emails, or access and search the internal database established by the company through specific query permissions; the traditional way of obtaining information within energy companies is more dependent on manual labor. It is inefficient, takes a long time to transmit data, and is prone to errors. Summary of the invention

[0004] The purpose of the present invention is to overcome the shortcomings of the prior art and provide an energy knowledge intelligent question-answering method and question-answering system based on knowledge graph, which can quickly respond to inquiries from enterprise employees and customers, provide the required information, and improve work efficiency.

[0005] A technical solution to achieve the above purpose is: an energy knowledge intelligent question answering method based on knowledge graph, comprising the following steps:

[0006] S1. Collect energy knowledge text data and generate knowledge base;

[0007] S2. Build a knowledge graph and integrate it with the BERT model;

[0008] S3. Build a question-answering system.

[0009] Furthermore, in S1, the specific method for collecting energy knowledge text data is to use crawler tools to obtain energy knowledge structured, semi-structured and unstructured text data from internal enterprise data, online databases and public open databases.

[0010] Furthermore, step S2 is specifically as follows:

[0011] S201. Using the OWL standard, perform entity extraction, relationship extraction and attribute extraction on the energy knowledge dataset, and perform semantic layer description, convert it into RDF triple format and store it in the Neo4j graph database;

[0012] S202, after embedding the entities in the knowledge graph into the input layer of BERT, a special tag [ENT] is created to identify the entity position in the text;

[0013] S203: Train and adjust the BERT model.

[0014] Furthermore, entity extraction is to identify and classify the contents of energy resources, energy facilities, energy companies, energy technologies, energy products and services, energy policies and regulations, and energy markets from the collected energy knowledge text data.

[0015] Furthermore, step S3 is specifically as follows:

[0016] S301, analyzing historical question and answer information to obtain a question classification set, a question template set, a keyword and hypernym classification set, and a query purpose set in the field of energy knowledge;

[0017] S302, using BiLSTM-CRF to process energy knowledge data;

[0018] The BiLSTM layer performs encoding and feature extraction. The CRF layer receives the output of the BiLSTM layer and encodes it into a fixed-length vector for sequence labeling.

[0019] S303, decoding the encoded vector into an answer sequence;

[0020] This is achieved through another LSTM layer, followed by a fully connected layer and a softmax activation function;

[0021] S304, using an open question-answering dataset to train a question-answering generation model, and then inputting data in the knowledge base into the model to obtain a corresponding question-answering dataset;

[0022] S305: Train a semantic analysis model based on the question-answering data set, use the semantic analysis model to analyze the question and find the answer to the question from the knowledge base.

[0023] Furthermore, the semantic analysis model includes an input layer, a word embedding layer, a multi-level fusion network, a fully connected layer, and an output layer;

[0024] The input layer receives raw text data and performs preprocessing;

[0025] The word vector layer maps the preprocessed words into word vectors in a high-dimensional space;

[0026] The multi-level fusion network transforms word vectors into multi-channel feature vectors;

[0027] The multi-channel feature vector is input into the fully connected layer for nonlinear transformation and extraction of high-level features;

[0028] The output layer is output through the Softmax layer.

[0029] Furthermore, the multi-level fusion network includes a word vector fusion layer, a syntactic structure fusion layer, a sentiment polarity fusion layer, and a multi-channel feature fusion layer;

[0030] The word vector fusion layer uses the attention mechanism to perform weighted fusion of word vectors; the syntactic structure fusion layer introduces the dependency syntactic analysis results to fuse the words in the sentence to capture the sentence structure information; the sentiment polarity fusion layer combines the sentiment dictionary to label the sentiment polarity of words and integrate the sentiment information into the semantic analysis; the multi-channel feature fusion layer splices the word vector fusion layer, the syntactic structure fusion layer and the sentiment polarity fusion layer to form a multi-channel feature vector.

[0031] Furthermore, entity recognition and attribute recognition are performed on the questions posed to the question-answering system, and the attributes corresponding to the entities are obtained in the knowledge graph. The cosine similarity of the encoded vector is calculated to represent the similarity between the question and the attribute, and multiple similarities are weighted and combined to obtain a comprehensive similarity. The candidate answers are sorted according to the comprehensive similarity, and the candidate answer sequence with the highest comprehensive similarity is selected as the answer sequence matching the user question.

[0032] Furthermore, a preset threshold is set for the total similarity of the candidate answers, and the candidate answers with similarity lower than the preset threshold are discarded and not displayed in the answer sequence matching the user question.

[0033] Another technical solution to achieve the above purpose is an energy knowledge intelligent question-answering system based on knowledge graph, including an energy knowledge text data collection module, a knowledge graph and BERT model fusion module and an intelligent question-answering module;

[0034] The crawler tool of the knowledge text data acquisition module obtains energy knowledge structured, semi-structured and unstructured text data from internal enterprise data, online databases and public databases to build a knowledge base;

[0035] The knowledge graph and BERT model fusion module extracts features from the energy knowledge dataset collected by the knowledge text data collection module, and forms a knowledge graph embedded with the BERT model in the knowledge base;

[0036] The intelligent question-answering module analyzes user questions and provides the closest answer sequence matching the user's question through the knowledge graph in the knowledge base based on the question.

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] 1. The intelligent question-answering method provided by the present invention integrates the knowledge assets of energy companies into a unified platform, which helps to promote team collaboration and learning, and can quickly retrieve relevant information from a large amount of data, thereby improving the efficiency of internal information acquisition and customer service;

[0039] 2. The intelligent question-answering method provided by the present invention can provide energy companies with new business models and service methods through large model technology and knowledge graphs, bring more diversified and personalized knowledge service products and solutions to enterprises, and promote the development of enterprises;

[0040] 3. The intelligent question-and-answer method provided by the present invention provides enterprises with information about customer needs, market trends, etc. through a large amount of question and answer data. By analyzing the data, it can more accurately grasp market dynamics and customer needs, which helps enterprises make decisions and optimize business strategies and service models. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 A schematic diagram of a process of an energy knowledge intelligent question-answering method based on a knowledge graph according to the present invention;

[0042] Figure 2 This is a schematic diagram of the architecture of an energy knowledge intelligent question-answering system based on a knowledge graph according to the present invention. DETAILED DESCRIPTION

[0043] In order to better understand the technical solution of the present invention, the following is a detailed description through specific embodiments:

[0044] See also Figure 1 The present invention provides an energy knowledge intelligent question-answering method based on knowledge graph, comprising the following steps:

[0045] S1. Collect energy knowledge text data and generate knowledge base.

[0046] The specific method for collecting energy knowledge text data is to use crawler tools to obtain energy knowledge structured, semi-structured and unstructured text data from internal enterprise data, online databases and public databases.

[0047] Structured data is data that has a clear, predefined data model and follows a consistent order. The most common structured data is data in a relational database. Structured data has three major characteristics, and data that meets all three characteristics can be called structured data. (1) It has a clear meaning (2) It has a strict, consistent order (3) It has a clear data type. Unstructured data is data that has no predefined data model and has an irregular or incomplete data structure. The most common unstructured data are documents, pictures, videos, etc. Semi-structured data refers to data that is between structured data and unstructured data, has certain structured characteristics, but does not fully meet structured characteristics. The most common semi-structured data include log files, XML documents, JSON documents, Email, HTML documents, etc.

[0048] S2. Build a knowledge graph and integrate it with the BERT model. Specifically include:

[0049] S201. Using the OWL standard, perform entity extraction, relationship extraction and attribute extraction on the energy knowledge dataset, and perform semantic layer description, convert it into RDF triple format and store it in the Neo4j graph database.

[0050] Entities are extracted from text corpus related to the question-answering knowledge domain to be constructed and the associations between entities are obtained. At the same time, the entity attribute names and entity attribute values ​​related to the entity are obtained during entity extraction. The extracted entities are used as nodes of the knowledge graph, and the associations between the extracted entities are used as edges connecting the nodes corresponding to the two entities. The entity attribute names are used as the attributes contained in the nodes corresponding to the entities, and the entity attribute values ​​are used as the attribute values ​​corresponding to the entity node attributes. A knowledge graph is constructed based on the nodes, edges, and the attributes and attribute values ​​of the nodes.

[0051] Entity extraction is to identify and classify the contents of energy resources, energy facilities, energy companies, energy technologies, energy products and services, energy policies and regulations, and energy markets from the collected energy knowledge text data.

[0052] S202. After embedding the entities in the knowledge graph into the input layer of BERT, a special tag [ENT] is created to identify the entity location in the text.

[0053] S203: Train and adjust the BERT model.

[0054] S3. Build a question-answering system. Specifically include:

[0055] S301. Analyze historical question and answer information to obtain a question classification set, a question template set, a keyword and hypernym classification set, and a query purpose set in the energy knowledge field.

[0056] By using past question-and-answer data, we analyze the templates of common questions and the query intent behind them. Through analysis and annotation by professionals, we replace the key information in the questions with their superordinate category words to form question templates. Based on the query intent embodied in these templates, we classify and summarize them, and then we can sort out a variety of question templates under different query intents.

[0057] S302. Use BiLSTM-CRF to process energy knowledge data.

[0058] The BiLSTM layer performs encoding and feature extraction. The CRF layer receives the output of the BiLSTM layer and encodes it into a fixed-length vector for sequence labeling.

[0059] S303: Decode the encoded vector into an answer sequence.

[0060] This is achieved with another LSTM layer, followed by a fully connected layer and a softmax activation function.

[0061] S304: Use an open question-answering dataset to train a question-answering generation model, and then input the data in the knowledge base into the model to obtain a corresponding question-answering dataset.

[0062] S305: Train a semantic analysis model based on the question-answering data set, use the semantic analysis model to analyze the question and find the answer to the question from the knowledge base.

[0063] The semantic model described by the semantic layer includes input layer, word vector layer, multi-level fusion network, fully connected layer and output layer;

[0064] The input layer receives raw text data and performs preprocessing;

[0065] The word vector layer maps the preprocessed words into word vectors in a high-dimensional space;

[0066] The multi-level fusion network transforms word vectors into multi-channel feature vectors;

[0067] The multi-channel feature vector is input into the fully connected layer for nonlinear transformation and extraction of high-level features;

[0068] The output layer is output through the Softmax layer.

[0069] The multi-level fusion network includes word vector fusion layer, syntactic structure fusion layer, sentiment polarity fusion layer and multi-channel feature fusion layer;

[0070] The word vector fusion layer uses the attention mechanism to perform weighted fusion of word vectors; the syntactic structure fusion layer introduces the dependency syntactic analysis results to fuse the words in the sentence to capture the sentence structure information; the sentiment polarity fusion layer combines the sentiment dictionary to label the sentiment polarity of words and integrate the sentiment information into the semantic analysis; the multi-channel feature fusion layer splices the word vector fusion layer, the syntactic structure fusion layer and the sentiment polarity fusion layer to form a multi-channel feature vector.

[0071] Perform entity recognition and attribute recognition on the questions asked to the question-answering system, obtain the attributes corresponding to the entities in the knowledge graph, calculate the cosine similarity of the encoded vector to represent the similarity between the question and the attribute, weightedly combine multiple similarities to obtain a comprehensive similarity, and sort the candidate answers according to the comprehensive similarity, and select the candidate answer sequence with the highest comprehensive similarity as the answer sequence matching the user question.

[0072] Furthermore, a preset threshold is set for the total similarity of the candidate answers, and the candidate answers with similarity lower than the preset threshold are discarded and not displayed in the answer sequence matching the user question.

[0073] See also Figure 2 The technical solution of the present invention also includes an energy knowledge intelligent question and answer system based on knowledge graph, including an energy knowledge text data acquisition module 41, a knowledge graph and BERT model fusion module 42 and an intelligent question and answer module 43.

[0074] The crawler tool of the knowledge text data acquisition module obtains energy knowledge structured, semi-structured and unstructured text data from internal enterprise data, online databases and public databases to build a knowledge base;

[0075] The knowledge graph and BERT model fusion module extracts features from the energy knowledge dataset collected by the knowledge text data collection module, and forms a knowledge graph embedded with the BERT model in the knowledge base;

[0076] The intelligent question-answering module analyzes user questions and provides the closest answer sequence that matches the user's question through the knowledge graph in the knowledge base based on the question.

[0077] Those skilled in the art should recognize that the above embodiments are only used to illustrate the present invention, and are not intended to limit the present invention. As long as they are within the spirit of the present invention, any changes or modifications to the above embodiments will fall within the scope of the claims of the present invention.

Claims

1. An energy knowledge intelligent question-answering method based on knowledge graph, characterized in that: The following steps are involved: S1. Collect energy knowledge text data and generate knowledge base; S2. Build a knowledge graph and integrate it with the BERT model; S3. Build a question-answering system.

2. According to claim 1, the energy knowledge intelligent question-answering method based on knowledge graph is characterized in that: In S1, the specific method for collecting energy knowledge text data is to use crawler tools to obtain energy knowledge structured, semi-structured and unstructured text data from internal enterprise data, online databases and public databases.

3. According to claim 1, the energy knowledge intelligent question-answering method based on knowledge graph is characterized in that: Step S2 is specifically as follows: S201. Using the OWL standard, perform entity extraction, relationship extraction and attribute extraction on the energy knowledge dataset, and perform semantic layer description, convert it into RDF triple format and store it in the Neo4j graph database; S202, after embedding the entities in the knowledge graph into the input layer of BERT, create a special tag [ENT] to identify the entity position in the text; S203: Train and adjust the BERT model.

4. According to claim 3, the energy knowledge intelligent question-answering method based on knowledge graph is characterized in that: Entity extraction is to identify and classify the contents of energy resources, energy facilities, energy companies, energy technologies, energy products and services, energy policies and regulations, and energy markets from the collected energy knowledge text data.

5. According to claim 1, the energy knowledge intelligent question-answering method based on knowledge graph is characterized in that: Step S3 is specifically as follows: S301, analyzing historical question and answer information to obtain a question classification set, a question template set, a keyword and hypernym classification set, and a query purpose set in the field of energy knowledge; S302, using BiLSTM-CRF to process energy knowledge data; The BiLSTM layer performs encoding and feature extraction. The CRF layer receives the output of the BiLSTM layer and encodes it into a fixed-length vector for sequence labeling. S303, decoding the encoded vector into an answer sequence; This is achieved through another LSTM layer, followed by a fully connected layer and a softmax activation function; S304, using an open question-answering dataset to train a question-answering generation model, and then inputting data in the knowledge base into the model to obtain a corresponding question-answering dataset; S305: Train a semantic analysis model based on the question-answering data set, use the semantic analysis model to analyze the question and find the answer to the question from the knowledge base.

6. The energy knowledge intelligent question-answering method based on knowledge graph according to claim 5 is characterized in that: The semantic analysis model includes input layer, word vector layer, multi-level fusion network, fully connected layer and output layer; The input layer receives raw text data and performs preprocessing; The word vector layer maps the preprocessed words into word vectors in a high-dimensional space; The multi-level fusion network transforms word vectors into multi-channel feature vectors; The multi-channel feature vector is input into the fully connected layer for nonlinear transformation and extraction of high-level features; The output layer is output through the Softmax layer.

7. The energy knowledge intelligent question-answering method based on knowledge graph according to claim 6 is characterized in that: The multi-level fusion network includes word vector fusion layer, syntactic structure fusion layer, sentiment polarity fusion layer and multi-channel feature fusion layer; The word vector fusion layer uses the attention mechanism to perform weighted fusion of word vectors; the syntactic structure fusion layer introduces the dependency syntactic analysis results to fuse the words in the sentence to capture the sentence structure information; the sentiment polarity fusion layer combines the sentiment dictionary to label the sentiment polarity of words and integrate the sentiment information into the semantic analysis; the multi-channel feature fusion layer splices the word vector fusion layer, the syntactic structure fusion layer and the sentiment polarity fusion layer to form a multi-channel feature vector.

8. The energy knowledge intelligent question-answering method based on knowledge graph according to claim 5 is characterized in that: Perform entity recognition and attribute recognition on the questions asked to the question-answering system, obtain the attributes corresponding to the entities in the knowledge graph, calculate the cosine similarity of the encoded vector to represent the similarity between the question and the attribute, weightedly combine multiple similarities to obtain a comprehensive similarity, and sort the candidate answers according to the comprehensive similarity, and select the candidate answer sequence with the highest comprehensive similarity as the answer sequence matching the user question.

9. The energy knowledge intelligent question-answering method based on knowledge graph according to claim 8 is characterized in that: A preset threshold is set for the total similarity of candidate answers, and candidate answers with similarity lower than the preset threshold are discarded and not displayed in the answer sequence matching the user's question.

10. A system for implementing any one of the energy knowledge intelligent question-answering methods based on knowledge graphs in claims 1 to 9, characterized in that: It includes energy knowledge text data collection module, knowledge graph and BERT model fusion module and intelligent question and answer module; The crawler tool of the knowledge text data acquisition module obtains energy knowledge structured, semi-structured and unstructured text data from internal enterprise data, online databases and public databases to build a knowledge base; The knowledge graph and BERT model fusion module extracts features from the energy knowledge dataset collected by the knowledge text data collection module, and forms a knowledge graph embedded with the BERT model in the knowledge base; The intelligent question-answering module analyzes user questions and provides the closest answer sequence that matches the user's question through the knowledge graph in the knowledge base based on the question.

Citation Information

Cited By

  • Power grid standard-oriented large model hallucination detection and trusted output control method

    CN122633825A