Knowledge question-answering method and system based on knowledge graph of hydropower station
By applying deep learning models in the field of hydropower stations for pre-processing, semantic analysis and knowledge retrieval of user problems, the shortcomings of existing systems in semantic understanding and knowledge retrieval are solved, and high-precision answer generation is achieved.
Patent Information
- Application Number
- CN202510051103.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-09
AI Technical Summary
The existing question and answer systems have insufficient semantic understanding capabilities and low knowledge retrieval accuracy in the field of hydropower stations, making it difficult to meet the actual needs of the engineering.
Using a knowledge question-and-answer method based on deep learning models, high-quality answers are generated by pre-processing of user questions, semantic analysis, SPARQL query and deep semantic matching sorting.
It significantly improves the accuracy and adaptability of the knowledge Q&A system in the hydropower station field, can accurately understand user intentions and provide accurate and complete answers.
Smart Images

Figure CN119961408A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge question answering, and in particular to a knowledge question answering method based on a hydropower station knowledge graph and a knowledge question answering system based on a hydropower station knowledge graph. Background Art
[0002] With the widespread application of knowledge graph technology, question-answering systems based on knowledge graphs have shown great potential in many fields. However, in the field of hydropower stations, existing question-answering systems still have technical bottlenecks when dealing with professional problems, especially in terms of semantic understanding and knowledge retrieval accuracy, which makes it difficult to meet the actual needs of engineering.
[0003] Insufficient semantic understanding is one of the main problems faced by existing systems. The field of hydropower stations involves a large number of professional terms and complex technical expressions, such as "hydraulic turbine dynamic and static balance test" or "transformer winding insulation impedance detection". These terms require deep semantic analysis to accurately extract core entities and relationships, while traditional question-answering systems usually rely on keyword matching and cannot correctly understand the user's technical intent, resulting in question-answering results that deviate from user needs.
[0004] In addition, the low accuracy of knowledge retrieval also limits the practicality of existing systems. Knowledge graphs in the field of hydropower stations usually contain multi-hop relationships and complex entity attributes. For example, "the maintenance cycle of a power station equipment" requires crossing multiple layers of associated information to obtain a complete answer. Existing question-answering systems mostly stay at shallow retrieval and cannot effectively handle multi-hop queries, which often leads to missing results or insufficient relevance, especially when the query conditions are unclear or the semantics are ambiguous.
[0005] Therefore, there is an urgent need to develop a knowledge question-answering method optimized for the hydropower station field, which can accurately understand user intentions through deep semantic analysis and use high-precision knowledge retrieval technology to provide accurate and complete answers, so as to solve the accuracy problems of existing solutions in technical question-answering in the field of hydropower station technology. Summary of the invention
[0006] The purpose of an embodiment of the present invention is to provide a knowledge question answering method based on a hydropower station knowledge graph, so as to at least solve the problem of insufficient accuracy of question answering in existing knowledge question answering solutions in the field of hydropower station technology.
[0007] In order to achieve the above-mentioned purpose, the first aspect of the present invention provides a knowledge question-answering method based on a hydropower station knowledge graph, the method comprising: in response to a user's question request signal, processing the user-entered question into natural language information, and performing preprocessing on the natural language information to obtain question information; performing semantic analysis on the question information based on a deep learning model, and extracting entities, relationships and attributes in the question information respectively to obtain user intent; performing SPARQL query in the knowledge graph based on the user intent to obtain multiple candidate answers; ranking the candidate answers by semantic matching based on a deep semantic matching model, and performing statistical conversion and natural language generation on the candidate answer with the highest semantic matching according to the question type to obtain and promote the corresponding question answer.
[0008] Optionally, the natural language information is preprocessed, including: performing noise filtering on questions input by the user, removing invalid characters and redundant symbols through regular expressions and / or pattern matching methods; using a word segmentation algorithm to segment the input text to generate a vocabulary sequence; and deleting meaningless words in the vocabulary sequence based on a stop word filter to obtain question information.
[0009] Optionally, a semantic analysis is performed on the question information based on a deep learning model, including: encoding the question information using a pre-trained model based on a Transformer structure to generate context-aware word vectors; obtaining dependencies between words in the context-aware word vectors through a multi-head self-attention mechanism, and extracting entities, relations, and attributes in the question as semantic features; and performing multi-task learning on the extracted semantic features, including entity recognition, relation extraction, and attribute matching, to obtain a semantic representation of the user's intent.
[0010] Optionally, performing a SPARQL query in the knowledge graph based on user intent includes: dynamically constructing a query template based on the entities and relationships obtained by parsing the user intent, and converting the constructed query template into a SPARQL query statement that conforms to the RDF format; based on the SPARQL query statement, using a recursive query mechanism for multi-hop relationships in the knowledge graph to obtain candidate answers related to the user's question; wherein, during the query process, the method also includes index preloading, caching strategy and / or query decomposition.
[0011] Optionally, the candidate answers are sorted by semantic matching based on a deep semantic matching model, including: encoding the question and each candidate answer into vector representations respectively, and constructing semantic matching features based on the entities, relations, types and context information extracted from the question; using a cross-attention mechanism in the model input layer to calculate the matching weight matrix between the question and the candidate answers; using a deep semantic network based on LSTM or Transformer to fuse the matching weight matrix between the question and the candidate answers, generate a matching score, and sort the candidate answers in descending order according to the score.
[0012] Optionally, the method also includes: when the semantic matching degree of the candidate answer is lower than a preset semantic matching degree threshold, triggering a rejection mechanism; wherein the rules of the rejection mechanism are: returning a prompt message that the answer was not found, and providing an analysis of the reasons for the query failure; retrieving a set of similar questions asked by the user, and recommending related questions using an algorithm based on edit distance or semantic similarity; calling a default FAQ library or a general knowledge base to provide the user with alternative answer suggestions.
[0013] Optionally, the candidate answers with the highest semantic match are statistically transformed according to the question type, including: for the question type, unit standardization conversion of the numerical data in the candidate answers; sorting or screening the candidate answers with time attributes according to the time series; for questions that require data calculation, calling the predefined operation rule module to perform the operation and generate the final answer.
[0014] Optionally, the rules for natural language generation are: based on the candidate answer sorting results, select the answer with the highest semantic match; use the template generation method to embed the attributes and statistical conversion results in the answer with the highest semantic match into natural language sentences to obtain candidate answer text; perform grammar checking and / or context optimization on the generated candidate answer text, and push the obtained answer to the question to the user end.
[0015] The second aspect of the present invention provides a knowledge question and answer system based on a hydropower station knowledge graph, the system comprising: a collection unit, for responding to a user's question request signal, processing the user-entered question into natural language information, and performing preprocessing on the natural language information to obtain question information; an intention recognition unit, for performing semantic analysis on the question information based on a deep learning model, and extracting entities, relationships and attributes in the question information respectively to obtain user intention; a query unit, for performing SPARQL query in the knowledge graph based on user intention to obtain multiple candidate answers; an output unit, for performing semantic matching ranking of candidate answers based on a deep semantic matching model, and performing statistical conversion and natural language generation on the candidate answer with the highest semantic matching degree according to the question type to obtain and promote the corresponding question answer.
[0016] On the other hand, the present invention provides a computer-readable storage medium having instructions stored thereon, which, when executed on a computer, enables the computer to execute the above-mentioned knowledge question-answering method based on the hydropower station knowledge graph.
[0017] Through the above technical scheme, the scheme of the present invention significantly improves the accuracy and adaptability of the knowledge question-answering system for the hydropower station field through a series of optimization steps. First, by preprocessing the natural language information of the user's question, noise and redundant information are effectively eliminated to ensure the clarity and parsability of the question. Secondly, the deep learning model is used to perform semantic analysis on the question information, which can accurately extract entities, relationships and attributes, thereby fully capturing the user's query intention. Based on the user's intention, the system obtains multiple candidate answers through SPARQL query in the knowledge graph to ensure the comprehensiveness of the answer coverage. Then, the candidate answers are ranked by matching degree with the help of a deep semantic matching model, which effectively solves the shortcomings of traditional keyword matching methods in complex semantic understanding and improves the relevance and accuracy of the answers. Finally, through statistical conversion and natural language generation, the system can output high-quality answers that are easy to understand and meet user needs according to the question type, further improving the user experience and the practicality of question-answering. This method is particularly suitable for problem solving in the field of hydropower station technology, overcomes the limitations of existing systems in semantic understanding and knowledge retrieval, and achieves high-precision answer generation.
[0018] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the embodiments of the present invention, but do not constitute a limitation on the embodiments of the present invention. In the accompanying drawings:
[0020] Figure 1 It is a flowchart of the steps of a knowledge question-answering method based on a hydropower station knowledge graph provided by an embodiment of the present invention;
[0021] Figure 2 It is a system structure diagram of a knowledge question and answer system based on a hydropower station knowledge graph provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0022] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.
[0023] Figure 1 1 is a method flow chart of a knowledge question answering method based on a hydropower station knowledge graph provided by an embodiment of the present invention. Figure 1 As shown in Figure 1 , an embodiment of the present invention provides a knowledge - based question - answering method based on a hydropower station knowledge graph. The method includes:
[0024] Step S10: In response to a user's question - asking request signal, process the question entered by the user into natural - language information, and perform pre - processing on the natural - language information to obtain question information.
[0025] Specifically, performing pre - processing on the natural - language information includes: filtering out noise from the question input by the user, removing invalid characters and redundant symbols through regular expressions and / or pattern - matching methods; using a word - segmentation algorithm to segment the input text to generate a sequence of words; and deleting meaningless words in the sequence of words based on a stop - word filter to obtain question information.
[0026] In an embodiment of the present invention, in response to a user's question - asking request signal, the knowledge - based question - answering method of the present invention processes the question entered by the user into natural - language information and performs multi - level pre - processing on this information to ensure the clarity of the question and the accuracy of subsequent parsing. Specifically, the pre - processing process includes the following steps:
[0027] 1) Perform noise filtering on the question input by the user. By using regular expressions and / or pattern - matching methods, the system can accurately identify and remove invalid characters, redundant symbols, and other irrelevant information in the input text, such as extra spaces, punctuation marks, or meaningless special characters. This step ensures the clear structure of the question text and reduces the impact of interference factors on subsequent analysis.
[0028] 2) Use a word - segmentation algorithm to perform word - segmentation processing on the input text to generate a sequence of words. To adapt to the technical requirements of the hydropower station field, the word - segmentation algorithm can adopt a Trie tree optimized based on the domain, a bidirectional maximum - matching algorithm, or a deep - learning word - segmentation model. These methods can accurately identify technical terms and compound nouns (such as "turbine governing system", "transformer winding insulation", etc.), thus disassembling complex technical descriptions into basic semantic units and providing high - quality input for subsequent semantic parsing.
[0029] 3) Apply a stop - word filter to delete meaningless words in the sequence of words. For example, eliminate insignificant function words such as "of", "is", "in", etc., while retaining core technical words. This process not only improves the density of the semantic information of the question but also effectively reduces the processing complexity, providing a more focused input for the subsequent model to identify the user's intention.
[0030] Through the above technical steps, the system can convert the questions input by the user into structured, concise, high-quality semantic information, thereby significantly improving the accuracy of problem parsing and semantic understanding. In the application scenario of hydropower stations, this method is particularly suitable for dealing with specialized problems such as equipment maintenance, operation and maintenance, and technical parameter inquiries. For complex questions such as "How to set the maintenance cycle of the turbine speed control system", the preprocessing module can identify the core keywords "turbine speed control system" and "maintenance cycle", filter out irrelevant characters, and form clear semantic units, laying a solid foundation for subsequent deep learning semantic parsing and knowledge retrieval.
[0031] Based on the solution of the present invention, through multi-level preprocessing, this method effectively solves the problem of low accuracy of question parsing in traditional question-answering systems caused by input text noise, inaccurate word segmentation, and interference from irrelevant information. Especially in the question-answering scenario of complex technical expressions in the field of hydropower stations, the preprocessing module can significantly improve the system's ability to capture the semantics of user questions, provide guarantees for the accuracy of subsequent knowledge retrieval and answer generation, and ultimately achieve efficient and reliable technical question-answering support.
[0032] Step S20: Perform semantic analysis on the question information based on the deep learning model, and extract entities, relationships, and attributes in the question information to obtain user intentions.
[0033] Specifically, a semantic analysis is performed on the question information based on a deep learning model, including: using a pre-trained model based on a Transformer structure to encode the question information and generate context-aware word vectors; obtaining the dependencies between words in the context-aware word vectors through a multi-head self-attention mechanism, and extracting entities, relations and attributes in the question as semantic features; performing multi-task learning on the extracted semantic features, including entity recognition, relation extraction and attribute matching, to obtain a semantic representation of the user's intention.
[0034] In an embodiment of the present invention, semantic analysis of question information based on a deep learning model is the core link to accurately understand user intent, especially in the field of hydropower stations, because its technical questions and answers involve complex equipment, relationships and attributes, and the accuracy of semantic analysis is crucial to the overall performance of the system. The present invention proposes a semantic analysis method based on a deep learning model, which can efficiently extract entities, relationships and attributes in questions and provide high-quality semantic representation for knowledge retrieval and answer generation.
[0035] Specifically, the scheme of the present invention adopts a pre-trained language model based on the Transformer structure (such as BERT, ALBERT) to encode the natural language questions input by the user. Through the encoding process, each word in the question information is converted into a high-dimensional word vector with context-awareness. These word vectors can capture the semantic changes of words in different contexts. For example, "water turbine" may form different semantic relationships with "speed control system" or "generator" in different questions. Context-aware word vectors lay the foundation for subsequent semantic analysis. On the basis of generating word vectors, the multi-head self-attention mechanism in the Transformer model is used to further explore the dependency relationship between words. This mechanism identifies the core vocabulary and its associated structure in the question by calculating the attention weight between each word and other words. For example, for the question "How to adjust the operating parameters of the water turbine speed control system?", the multi-head self-attention mechanism can clearly define "water turbine speed control system" as the core entity, forming a relationship with "operating parameters", while ignoring irrelevant functional vocabulary.
[0036] Furthermore, with the support of context awareness and dependencies, the system can extract key semantic features in the question, including entities, relations, and attributes:
[0037] 1) Entity Recognition: Through named entity recognition (NER) technology, the system extracts specific equipment or concepts from the question, such as "water turbine" and "speed control system".
[0038] 2) Relationship extraction: Through the relational classification model, the system extracts semantic relationships between entities, such as "belongs to", "adjusts", "runs", etc., to clarify the user's technical intent.
[0039] 3) Attribute matching: The system further extracts attributes related to the entity, such as "operating parameters", "speed control mode", etc., to enrich the integrity of the semantic expression.
[0040] In order to improve the comprehensive ability of semantic analysis, this method models entity recognition, relationship extraction and attribute matching as a multi-task learning problem. By sharing the encoding layer of the neural network, the system can efficiently extract multiple semantic features from the same input, avoid repeated calculations, and improve the overall performance through interaction between tasks. For example, in the field of hydropower stations, the relationship extraction results in the "equipment status query" task can assist the accuracy of attribute matching. Finally, the system combines the extracted entities, relationships, and attributes into a structured semantic representation. For example, for the question "How to adjust the operating parameters of the turbine speed control system?" The semantic representation may include: entity: turbine speed control system; relationship: adjustment; attribute: operating parameters. This semantic representation not only fully describes the user's question intention, but also serves as an input for subsequent knowledge retrieval and answer generation.
[0041] Step S30: Perform a SPARQL query in the knowledge graph based on the user's intention to obtain multiple candidate answers.
[0042] Specifically, a query template is dynamically constructed based on the entities and relationships obtained by parsing the user's intention, and the constructed query template is converted into a SPARQL query statement that conforms to the RDF format; based on the SPARQL query statement, a recursive query mechanism is used to obtain candidate answers related to the user's question for multi-hop relationships in the knowledge graph; wherein, during the query process, the method also includes index preloading, caching strategy and / or query decomposition.
[0043] In the embodiment of the present invention, the system dynamically generates a SPARQL query template based on the user intent parsing results (including extracted entities, relations and attributes). For example, the user question "How to adjust the operating parameters of the turbine speed control system?" The parsed entity is "turbine speed control system", the relation is "adjustment", and the attribute is "operating parameters". Based on this information, the system automatically constructs a query template. This dynamic template construction ensures the flexibility of the query and the adaptation to the semantics of different questions.
[0044] Furthermore, the constructed query template is converted into a SPARQL query statement that conforms to the RDF format, which is convenient for executing queries in the knowledge graph. For example, the knowledge graph may be represented by a triple: "turbine speed control system"-"hasParameter"-"operating parameters". The SPARQL query statement is consistent with this structure and can efficiently retrieve matching answers. The knowledge graph in the field of hydropower stations usually has multi-hop relationships. For example, from "equipment" to "operating parameters" may involve multiple intermediate entities (such as components or operating modes). The system gradually expands the query scope through a recursive query mechanism. For example, for the question "Maintenance cycle and precautions of a certain equipment", the system recursively traverses multi-level relationships such as "equipment"-"components"-"maintenance plan", and finally obtains a complete set of candidate answers.
[0045] Furthermore, the index data of the knowledge graph, such as entity types, relationship structures, and common attributes, is preloaded before the query to reduce the search scope during the query and improve efficiency. For high-frequency queries or queries with similar structures, the system caches the results to avoid repeated calculations. For example, the query results of "maintenance plan for turbine speed control system" can be cached for repeated use for similar questions. For complex queries, they are decomposed into multiple sub-queries and executed step by step. For example, "obtaining the operating parameters and adjustment methods of the equipment" is decomposed into two sub-queries, respectively obtaining the operating parameters and adjustment methods, and then merging the results.
[0046] Step S40: Sort the candidate answers by semantic matching based on the deep semantic matching model, and perform statistical conversion and natural language generation on the candidate answer with the highest semantic matching degree according to the question type to obtain and promote the corresponding answer to the question.
[0047] Specifically, the semantic matching degree of candidate answers is sorted based on the deep semantic matching model, including: encoding the question and each candidate answer into vector representations respectively, and constructing semantic matching features based on the entities, relations, types and context information extracted from the question; using the cross-attention mechanism in the model input layer to calculate the matching weight matrix between the question and the candidate answers; using a deep semantic network based on LSTM or Transformer to fuse the matching weight matrix between the question and the candidate answers, generate a matching score, and sort the candidate answers in descending order according to the score.
[0048] In the embodiment of the present invention, the semantic matching degree ranking of candidate answers based on the deep semantic matching model is an important step to achieve high-quality answer generation, especially in the field of hydropower stations, where questions involve complex technical relationships and multi-level semantics, and traditional methods are difficult to achieve accurate matching. Semantic vectorization of questions and candidate answers.
[0049] 1) Encoding into vector representation: The question and each candidate answer are encoded into high-dimensional vector representations through a Transformer-based pre-trained model (such as BERT, ALBERT). These vectors can capture contextual information at the semantic level. For example, "What maintenance measures are required for the turbine speed control system" can correctly associate "maintenance measures" with specific equipment attributes.
[0050] 2) Semantic feature construction: Combine the entities (such as "water turbine speed control system"), relations (such as "need") and types (such as "maintenance measures") extracted from the question to add domain-specific semantic features to the vector representation. The introduction of contextual information ensures that the matching of questions and answers takes into account the semantic integrity rather than just the literal level.
[0051] At the model input layer, the cross-attention mechanism is used to calculate the matching weight matrix between the question and each candidate answer. The cross-attention mechanism captures the semantic association between keywords by comparing the word vectors of the question and the answer word by word. For example, between "How to adjust the operating parameters of the turbine?" and "Operating parameters can be achieved by adjusting the frequency", the attention mechanism can focus on the two key concepts of "operating parameters" and "adjusting frequency", thereby improving the matching effect.
[0052] Furthermore, a deep semantic network based on LSTM or Transformer is used to fuse the matching weight matrix. By gradually integrating time series information, the matching information of questions and answers is processed in sequence, which is suitable for capturing the logical relationship in the answer. Through the global attention mechanism, the multi-level semantic association between questions and answers is captured, which is particularly suitable for handling the matching needs of multiple entities and multiple relationships in complex technical problems.
[0053] Furthermore, based on the output of the deep semantic network, a matching score is generated between the question and each candidate answer. The score reflects the similarity between the answer and the question in the semantic space. The matching scores of the candidate answers are sorted in descending order to ensure that the most relevant answers are ranked first. For example, for the question "What is the maintenance cycle of a hydropower station?", the system ranks "The maintenance cycle is every 12 months" first, rather than answers that are less semantically related to the question.
[0054] Preferably, the method also includes: when the semantic matching degree of the candidate answer is lower than a preset semantic matching degree threshold, triggering a rejection mechanism; wherein the rules of the rejection mechanism are: returning a prompt message that the answer was not found, and providing an analysis of the reasons for the query failure; retrieving a set of similar questions asked by the user, and recommending related questions using an algorithm based on edit distance or semantic similarity; calling a default FAQ library or a general knowledge base to provide the user with alternative answer suggestions.
[0055] In an embodiment of the present invention, in a hydropower station knowledge question and answer system, an optimization method based on a rejection mechanism is proposed for situations where the semantic matching degree of candidate answers is lower than a preset threshold. This mechanism significantly improves the system's fault tolerance and user experience through clear prompt information, failure cause analysis, and alternative answer suggestions. When the matching degree of the candidate answer calculated by the semantic matching model is lower than the preset threshold, the system determines that the current candidate answer cannot accurately meet user needs and triggers the rejection mechanism. The setting of the threshold is dynamically adjusted according to the model performance and the application scenario in the hydropower station field, and is usually determined by the trade-off between precision and recall during the model testing phase.
[0056] Furthermore, the system first returns a clear prompt message to the user, such as "No answer that fully matches your question was found. Please try to re-enter the question or view the following recommended content." This prompt avoids the user from being confused by the empty results and leaves room for subsequent interactions. Combining the intermediate results of question parsing and knowledge retrieval, the system analyzes the reasons for the query failure. For example, for the question "What are the design parameters of a certain device?", it may prompt "No design parameter information related to the device was found in the knowledge graph." This function helps users understand the reasons for the failure and may adjust the question statement or scope.
[0057] Furthermore, the system uses the edit distance algorithm or semantic similarity model to analyze the user's questions and retrieve a set of questions that are semantically similar to the current question. For example, for the question "What is the maintenance cycle of a turbine?", if there is no matching answer, the system may recommend "What is the maintenance cycle standard for turbines?" or "What are the regulations for the maintenance frequency of hydropower station equipment?". The edit distance algorithm works better for situations where the text expressions are similar, while the semantic similarity model can identify questions that are related at the semantic level. The two methods are used in combination to ensure the breadth and accuracy of the recommended questions.
[0058] Furthermore, if the recommended questions still cannot meet the user's needs, the system calls the FAQ or general knowledge base to provide alternative answers. For example, for technical detail questions that cannot be identified, the system may recommend common questions related to the field, such as "What are the key steps in the daily maintenance of hydropower station equipment?" or "What should be paid attention to when the transformer is running?".
[0059] Specifically, the candidate answers with the highest semantic matching degree are statistically transformed according to the question type, including: for the question type, the numerical data in the candidate answers are unit-standardized; for candidate answers with time attributes, they are sorted or screened according to the time series; for questions that require data calculation, the predefined operation rule module is called to perform the operation to generate the final answer.
[0060] In the embodiments of the present invention, in the field of hydropower stations, numerical data are usually expressed in different units. For example, flow rate may be expressed in cubic meters per second (m 3 / s) or liters per second (L / s). The system identifies the target unit when parsing questions and performs standardized conversions on the numerical data involved in the candidate answers. For example, for the question "What is the maximum flow rate of a turbine?", if the candidate answer is in L / s, but the user's question implicitly requires m 3 / s, the system uses predefined conversion rules (such as 1m 3 / s=1000L / s) to unify the units. Ensure the consistency of the answer units to facilitate user understanding and application.
[0061] Furthermore, time-related questions are more common in the maintenance and operation of hydropower station equipment, such as "When was the last equipment maintenance?" The system identifies the time attribute requirements of the question and sorts or filters the candidate answers according to the time series. For example, for the question "What is the annual power generation in the past five years?", the system retrieves the annual power generation data from the knowledge graph, arranges it in ascending order of time, and filters the records of the past five years. Through time sorting and filtering, an intuitive time series answer that meets the requirements of the question is generated.
[0062] Furthermore, many technical questions require answers through calculations, such as "What is the average monthly power generation of a certain device?" The system calls predefined operation rule modules for such questions. For example, after retrieving the monthly power generation data, the average value formula is used for calculation to generate the final answer. For more complex calculations (such as integral calculations of load curves), the system completes high-precision calculations by integrating professional calculation modules (such as Python's NumPy library). The system can automatically handle the calculation requirements in user questions and output directly usable results, avoiding additional calculations by users.
[0063] Specifically, the rules for natural language generation are: based on the candidate answer sorting results, select the answer with the highest semantic match; use the template generation method to embed the attributes and statistical conversion results in the answer with the highest semantic match into the natural language sentence pattern to obtain the candidate answer text; perform grammar checking and / or context optimization on the generated candidate answer text, and push the obtained answer to the question to the user end.
[0064] In an embodiment of the present invention, natural language generation (NLG) is a key link in the hydropower station knowledge question and answer system to convert complex technical data into user-readable answers. After the candidate answers are deeply semantically matched and sorted, the system automatically selects the answer with the highest semantic match as the basis for the final output. For example, for the question "What is the power generation of a certain hydropower station?", the candidate answer with the highest matching degree may be "The annual power generation is 12 billion kWh", and the system gives priority to this answer to enter the generation process. When matching multiple answers, the system gives priority to answers with higher accuracy or more detailed attributes to ensure the reliability of the output content.
[0065] Furthermore, through predefined natural language templates, the entities, attributes, relationships, and statistical conversion results in the candidate answers are embedded in the sentence structure. For example, the question "What is the maintenance cycle of a certain equipment?" may match the answer "Maintenance cycle: every 6 months", and the corresponding template generation rule is "The maintenance cycle of the equipment is {maintenance cycle}." The generated answer is "The maintenance cycle of the equipment is every 6 months." The template is dynamically adjusted according to the terminology library in the field of hydropower stations to ensure the professionalism of the generated content. For example, the description of "power generation" may use a variety of terms such as "unit power output" or "annual cumulative power generation" to match specific scenario requirements.
[0066] Furthermore, by integrating language correction tools (such as GPT or rule-based grammar parsers), the system checks the grammatical structure and spelling of the generated natural language answers to avoid incorrect expressions caused by template problems. Combined with the semantic context of the user's question, the generated answer is reconstructed to ensure language fluency. For example, the answer generated for "What is the operating status of a certain turbine?" may be optimized to "The turbine is currently operating normally and all parameters meet the design requirements." The final generated natural language answer is pushed to the user end through the system interface and presented to the user in a clear text form, while supporting further interaction or extended queries.
[0067] Figure 2 : is a system structure diagram of a knowledge question answering system based on a hydropower station knowledge graph provided by an embodiment of the present invention. Figure 2 As shown, an embodiment of the present invention provides a knowledge question-answering system based on a knowledge graph of a hydropower station, the system comprising: a collection unit, for responding to a user's question request signal, processing a user-entered question into natural language information, and performing preprocessing on the natural language information to obtain question information; an intention recognition unit, for performing semantic analysis on the question information based on a deep learning model, and extracting entities, relationships, and attributes in the question information respectively to obtain user intention; a query unit, for performing SPARQL query in the knowledge graph based on user intention to obtain multiple candidate answers; an output unit, for performing semantic matching ranking of candidate answers based on a deep semantic matching model, and performing statistical conversion and natural language generation on the candidate answer with the highest semantic matching degree according to the question type to obtain and promote the corresponding question answer.
[0068] An embodiment of the present invention also provides a computer-readable storage medium, which stores instructions, which, when executed on a computer, enable the computer to execute the above-mentioned knowledge question-answering method based on the hydropower station knowledge graph.
[0069] Those skilled in the art will understand that all or part of the steps in the method for implementing the above-mentioned embodiments can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including several instructions for making a single-chip microcomputer, a chip or a processor (processor) perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0070] The optional embodiments of the present invention are described in detail above in conjunction with the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above embodiments. Within the technical concept of the embodiments of the present invention, the technical scheme of the embodiments of the present invention can be subjected to a variety of simple modifications, and these simple modifications all belong to the protection scope of the embodiments of the present invention. It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the embodiments of the present invention will not further describe various possible combinations.
[0071] In addition, various embodiments of the present invention may be arbitrarily combined, and as long as they do not violate the concept of the embodiments of the present invention, they should also be regarded as the contents disclosed in the embodiments of the present invention.
Claims
1. A knowledge question answering method based on a hydropower station knowledge graph, characterized in that: The method comprises: In response to a user's question request signal, processing the user's input question into natural language information, and performing preprocessing on the natural language information to obtain question information; Performing semantic analysis on the question information based on a deep learning model, respectively extracting entities, relationships, and attributes in the question information to obtain user intent; Perform SPARQL queries in the knowledge graph based on user intent to obtain multiple candidate answers; The candidate answers are sorted by semantic matching based on the deep semantic matching model, and the candidate answers with the highest semantic matching degree are statistically converted and natural language generated according to the question type to obtain and promote the corresponding answer to the question.
2. The method according to claim 1, characterized in that Preprocessing the natural language information includes: Perform noise filtering on user input questions and remove invalid characters and redundant symbols through regular expressions and / or pattern matching methods; Use the word segmentation algorithm to segment the input text and generate a vocabulary sequence; Based on the stop word filter, meaningless words in the word sequence are removed to obtain problem information.
3. The method according to claim 1, characterized in that Performing semantic analysis on the question information based on a deep learning model includes: Use a pre-trained model based on the Transformer structure to encode question information and generate context-aware word vectors; The multi-head self-attention mechanism is used to obtain the dependencies between words in the context-aware word vector, and to extract the entities, relations, and attributes in the question as semantic features. Multi-task learning is performed on the extracted semantic features, including entity recognition, relationship extraction, and attribute matching, to obtain the semantic representation of user intent.
4. The method according to claim 1, characterized in that: Perform SPARQL queries in the knowledge graph based on user intent, including: Dynamically construct query templates based on the entities and relationships obtained by parsing user intent, and convert the constructed query templates into SPARQL query statements that conform to the RDF format; Based on SPARQL query statements, a recursive query mechanism is used to obtain candidate answers related to user questions for multi-hop relationships in the knowledge graph; During the query process, the method further includes index preloading, caching strategy and / or query decomposition.
5. The method according to claim 1, characterized in that The semantic matching degree of candidate answers is sorted based on the deep semantic matching model, including: Encode the question and each candidate answer into vector representations, and construct semantic matching features based on the entities, relations, types, and context information extracted from the question. Use the cross-attention mechanism at the model input layer to calculate the matching weight matrix between the question and the candidate answer; Use a LSTM or Transformer-based deep semantic network to fuse the matching weight matrix between the question and the candidate answer, generate a matching score, and sort the candidate answers in descending order according to the score.
6. The method according to claim 1, characterized in that The method further comprises: When the semantic matching degree of the candidate answer is lower than the preset semantic matching degree threshold, the rejection mechanism is triggered; The rules of the rejection mechanism are: Returns a prompt message indicating that the answer was not found, and provides an analysis of the reasons for the query failure; Retrieve a set of similar questions asked by users and recommend related questions using algorithms based on edit distance or semantic similarity; Call the default FAQ library or general knowledge base to provide users with alternative answer suggestions.
7. The method according to claim 1, characterized in that The candidate answers with the highest semantic matching degree are statistically transformed according to the question type, including: According to the question type, the numerical data in the candidate answers are standardized and converted to units; For candidate answers with time attributes, sort or filter them according to the time series; For problems that require data calculation, the predefined operation rule module is called to perform the operation and generate the final answer.
8. The method according to claim 1, characterized in that The rules for natural language generation are: According to the sorting results of candidate answers, select the answer with the highest semantic match; Using the template generation method, the attributes and statistical conversion results of the answer with the highest semantic matching degree are embedded into the natural language sentence pattern to obtain the candidate answer text; Perform grammar checking and / or context optimization on the generated candidate answer text, and push the obtained question answer to the user end.
9. A knowledge question answering system based on a hydropower station knowledge graph, characterized in that: The system comprises: A collection unit, configured to respond to a user's question request signal, process the user's input question into natural language information, and perform preprocessing on the natural language information to obtain question information; An intention recognition unit, used to perform semantic analysis on the question information based on a deep learning model, and extract entities, relationships and attributes in the question information to obtain user intentions; The query unit is used to perform SPARQL queries in the knowledge graph based on user intent and obtain multiple candidate answers; The output unit is used to sort the candidate answers by semantic matching degree based on the deep semantic matching model, and perform statistical conversion and natural language generation on the candidate answers with the highest semantic matching degree according to the question type, so as to obtain and promote the corresponding answer to the question.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the knowledge question-answering method based on the hydropower station knowledge graph as described in any one of claims 1 to 8.
Citation Information
Cited By
Intelligent question and answer method and device for water conservancy industry and medium
CN120596602A
Intelligent file searching method and system based on natural language processing and knowledge graph
CN120821699A
Question answering system construction method and system based on large language model
CN120910197A
Huangmuo paddling inheriting and innovating method based on large model
CN120930743A
Question answering method and device for beacon information, question answering equipment and storage medium
CN121166855A