User intention recognition method and system for knowledge questions and answers of hydropower station

By structuring the natural language problems input by users in the hydropower plant field and deep semantic feature extraction, combined with entity recognition model and knowledge graph linking, the limitations of the existing system in user intention recognition are solved, and accurate analysis and efficient response to complex problems are achieved.

CN119988626APending Publication Date: 2025-05-13GUODIAN DADU RIVER POWER ENG
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510056345.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing hydropower station knowledge Q&A system has shown significant limitations in user intention recognition, and cannot accurately analyze complex user problems, resulting in a degradation of the overall performance of the Q&A system.

Method used

By structuring the natural language problems input by users, deep semantic features are extracted using the BERT model, and combined with the entity recognition model of Bi-LSTM and CRF models, key entities are identified and linked to the knowledge graph, and semantic structures are constructed for verification, and finally output user intentions.

Benefits of technology

It realizes accurate intention identification of complex user problems in the hydropower plant field, and improves the analytical ability and response accuracy of the question-and-answer system in complex field scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988626A_ABST
    Figure CN119988626A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a user intention recognition method and system for knowledge questions and answers of a hydropower station, and belongs to the technical field of knowledge questions and answers. The method comprises the following steps: performing structured processing on the natural language question to obtain structured question information; the problems are classified, intention information in the structured problem information is extracted based on a classification structure and a BERT model, limiting conditions of the intention information are recognized, and problem analysis information is obtained; key entities in the question analysis information are recognized based on a pre-trained entity recognition model, and the key entities are linked to a knowledge graph through dictionary matching and / or context analysis; and the problem analysis information and the information of the key entity linked to the knowledge graph are constructed into a semantic structure, and a user intention is output. According to the scheme, it is ensured that the finally-output user intention has high accuracy and consistency, and therefore the analysis ability and response accuracy of the question answering system for the problems in the hydropower station field are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge question answering, and in particular to a method for identifying user intentions in a hydropower station knowledge question answering and a system for identifying user intentions in a hydropower station knowledge question answering. Background Art

[0002] With the development of natural language processing technology, question-answering systems based on knowledge graphs have been widely used in many fields. However, in the knowledge question-answering scenario in the field of hydropower stations, existing question-answering technologies have shown significant limitations in user intent recognition, mainly in the inability to accurately parse complex user questions, thus affecting the overall performance of the question-answering system.

[0003] Questions in the field of hydropower stations usually involve complex technical terms, professional equipment names, and a variety of qualifications. For example, users may ask about the specific maintenance status or operation records of a certain generator set. Such questions often contain multiple entities and conditions (such as time range, equipment number, etc.). Existing question-answering systems lack the ability to deeply understand domain knowledge and accurately parse intent when dealing with these problems. On the one hand, the intent recognition models of traditional question-answering systems are mostly based on general corpus training, which makes it difficult to capture the semantic features unique to the field of hydropower stations, resulting in insufficient semantic understanding of the questions; on the other hand, the questions entered by users are expressed in a variety of ways and may be ambiguous. Existing systems cannot accurately identify and associate key entities in questions and their contextual semantics, resulting in entity linking errors or incomplete parsing of qualification conditions.

[0004] In addition, the questions asked by users in the hydropower station question-and-answer scenario are often context-dependent. For example, in a multi-round conversation, a user may ask a vague question at the beginning, and then gradually clarify his or her intention by adding additional information. However, traditional question-and-answer systems lack the ability to dynamically model the semantics of multi-round conversations, and cannot effectively integrate contextual information to adjust and optimize user intentions, resulting in a further decrease in the accuracy and completeness of intent recognition results in complex scenarios.

[0005] Therefore, in response to the above-mentioned problems in the field of hydropower stations, there is an urgent need for a question-answering technology that can deeply understand the domain semantics and accurately identify user intentions, so as to realize the efficient application of question-answering systems in complex domain scenarios. Summary of the invention

[0006] The purpose of the embodiments of the present invention is to provide a method for identifying user intentions in hydropower station knowledge question and answering, so as to at least solve the problem of insufficient accuracy of user intention identification in the existing knowledge question and answer process in the field of hydropower station knowledge.

[0007] In order to achieve the above-mentioned purpose, the first aspect of the present invention provides a user intention recognition method for hydropower station knowledge question and answer, the method comprising: recovering natural language questions input by users, and performing structured processing on the natural language questions to obtain structured question information; classifying questions based on the structured question information, and extracting intent information in the structured question information based on the classification structure and the BERT model and identifying the limiting conditions of each intent information to obtain question parsing information; recognizing key entities in the question parsing information based on a pre-trained entity recognition model, and linking the key entities to a knowledge graph through dictionary matching and / or context analysis; wherein the entity recognition model is obtained based on the coupled training of the BERT model, the Bi-LSTM model and the CRF model; constructing the question parsing information and the information linking the key entities to the knowledge graph into a semantic structure, and verifying it through semantic pattern matching with the knowledge graph, and outputting the user intention based on the verification result.

[0008] Optionally, the classification of questions is implemented based on a combination of rule matching and deep learning methods, including: the rule matching recognizes question words in the question and their contextual semantics based on predefined grammatical patterns and question word category rules, and determines the category of the question based on the recognition result as a candidate category; the deep learning method uses a bidirectional recurrent neural network or a recurrent convolutional neural network to verify the candidate category by learning the contextual information and local features of the question text.

[0009] Optionally, the identification of limiting conditions for each intent information is achieved through a slot labeling model, including: the slot labeling model labels the candidate information in the user input question as a specific slot type based on the Bi-LSTM-CRF model; wherein, the labeling process of labeling the candidate information in the user input question as a specific slot type is combined with the context embedding vector provided by the BERT pre-trained model, and the long-distance dependencies between slots are captured through the Bi-LSTM network; based on the slot labeling results, the overall labeling sequence is optimized through the conditional random field layer to obtain the limiting conditions for each intent information.

[0010] Optionally, the rules for the context analysis are: based on a bidirectional gated loop, the context information of multiple rounds of dialogues in the process of recycling natural language questions input by the user is organized, and the semantic features of the question currently input by the user and the historical questions and answers are integrated; the historical semantics are weighted using an attention mechanism to extract the information most relevant to the current question as the result of the context analysis.

[0011] Optionally, linking key entities to the knowledge graph includes: constructing an entity name dictionary for the key entities based on dictionary matching and / or context analysis results, wherein: the entity name dictionary contains the full name, abbreviations, synonyms, spelling variations and / or common nicknames of the entity; generating a set of candidate entities based on the entity name dictionary, and the dictionary's full match and / or partial match strategy, and calculating the matching confidence of the candidate entities in combination with contextual semantic information; verifying the relationship consistency between the candidate entity with the highest matching confidence and the semantic structure in the knowledge graph, and executing the corresponding candidate entity link with the highest matching confidence after completing the verification.

[0012] Optionally, the question resolution information and the information linking key entities to the knowledge graph are constructed into a semantic structure and verified through semantic pattern matching with the knowledge graph, including: mapping the question resolution information to the entity and relationship nodes in the knowledge graph, retrieving the matching status of the resolution results and the relationship chains in the knowledge graph through a path search algorithm; calculating the global consistency score for each key entity and its associated relationship in the resolution result using a graph embedding algorithm; when the global consistency score deviates from the preset standard consistency score range, reselecting the candidate relationship chain based on the weight-adjusted optimization algorithm, and updating the semantic structure to re-execute verification until the global consistency score is within the preset standard consistency score range.

[0013] Optionally, the training rules of the entity recognition model are: extracting contextual features of the text based on the BERT model, and sharing the underlying feature representation for the entity recognition task and the intent classification task; during the training process, optimizing the annotation accuracy of the named entity boundaries and the generalization performance of the intent classification by jointly optimizing the multi-objective loss function; combining the self-attention mechanism to perform boundary recognition ability training for entity recognition, and obtaining an entity recognition model.

[0014] Optionally, outputting user intent based on the verification result includes: calling a structured data interface for generating question resolution to execute user intent output; wherein the user intent includes question type, knowledge graph mapping of key entities, and parameter information of limiting conditions; and enhancing the intent result during the output process by combining the contextual relationship of entities in the knowledge graph and the verified semantic consistency score.

[0015] The second aspect of the present invention provides a user intention recognition system for hydropower station knowledge questions and answers, the system comprising: a collection unit, used to recover natural language questions input by users, and perform structured processing on the natural language questions to obtain structured question information; a parsing unit, used to classify questions based on the structured question information, and extract intent information in the structured question information based on the classification structure and the BERT model and identify the limiting conditions of each intent information to obtain question parsing information; a linking unit, used to identify key entities in the question parsing information based on a pre-trained entity recognition model, and link the key entities to a knowledge graph through dictionary matching and / or context analysis; wherein the entity recognition model is obtained based on the coupled training of the BERT model, the Bi-LSTM model and the CRF model; an output unit, used to construct the information linking the question parsing information and the key entities to the knowledge graph into a semantic structure, and verify it through semantic pattern matching with the knowledge graph, and output the user intention based on the verification result.

[0016] On the other hand, the present invention provides a computer-readable storage medium having instructions stored thereon, which, when executed on a computer, enables the computer to execute the above-mentioned method for identifying user intentions in a hydropower station knowledge question and answer session.

[0017] Through the above technical scheme, the method provided by the scheme of the present invention realizes accurate intention recognition of natural language questions of users through multi-step processing, and effectively solves the problem of inaccurate intention recognition in knowledge question and answer in the field of hydropower stations. First, by structurally processing the natural language questions input by the user, grammatical and semantic information is extracted to form structured question information, laying the foundation for subsequent semantic analysis. Subsequently, the questions are classified based on the structured information, and the deep semantic features are extracted using the BERT model, while the limiting conditions are identified to generate accurate question analysis information, thereby improving the depth of semantic understanding. By combining the pre-trained entity recognition model of BERT, Bi-LSTM and CRF models, the key entities in the questions are accurately identified, and the key entities are accurately linked to the entities in the knowledge graph through dictionary matching and context analysis methods. Furthermore, by constructing the semantic structure of question analysis and entity linking, and combining the semantic pattern verification and optimization of the knowledge graph, it is ensured that the user intention of the final output has high accuracy and consistency, thereby greatly improving the analysis ability and response accuracy of the question and answer system for complex field problems.

[0018] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the embodiments of the present invention, but do not constitute a limitation on the embodiments of the present invention. In the accompanying drawings:

[0020] Figure 1 is a flowchart of the steps of a method for identifying user intentions in a hydropower station knowledge question and answer session provided by an embodiment of the present invention;

[0021] Figure 2 It is a system structure diagram of a user intention recognition system for hydropower station knowledge question and answer provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0022] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.

[0023] Figure 1 1 is a method flow chart of a method for identifying user intentions in a hydropower station knowledge question and answer system provided by an embodiment of the present invention. Figure 1 As shown, an embodiment of the present invention provides a method for identifying user intentions in a hydropower station knowledge question and answer session, the method comprising:

[0024] Step S10: Recover the natural language question input by the user, and perform structured processing on the natural language question to obtain structured question information.

[0025] Specifically, in the application scenario of hydropower stations, users usually input complex and diverse questions through natural language, such as equipment operating status, maintenance records, power generation efficiency, etc. These questions may contain professional terms, multiple entities, multiple limiting conditions (such as time and place), and semantic ambiguity. The natural language questions input by users are recovered, and irrelevant information such as stop words and punctuation marks are removed through preprocessing steps, and the word segmentation algorithm is used to divide the questions into smaller semantic units. Then, a context-aware embedding vector is generated through the BERT model to capture the semantic associations of each word in the question, thereby constructing the preliminary semantic structure of the question.

[0026] Furthermore, at the grammatical level, the system combines syntactic analysis tools to parse the syntax tree of the question, identify the subject, predicate, object and modifiers of the sentence, and extract the core semantic elements in the question. At the semantic level, the semantic features generated by BERT are combined with pre-trained models in specific fields to identify the key content in the question, such as device name, operating status, and time range. Through these processing steps, the system converts natural language questions into formatted data with clear semantics and structure, that is, structured question information.

[0027] Based on the solution of the present invention, through the combination of word segmentation and semantic embedding, the problem of parsing professional terms and diversified expressions of natural language is solved, providing an accurate semantic basis for subsequent intent recognition and entity linking. Secondly, through comprehensive analysis of syntax and semantics, the system can extract multi-dimensional information in the question to ensure that the limiting conditions and contextual semantics are not omitted. Finally, the generation of structured question information lays the foundation for subsequent knowledge graph linking and question and answer generation, enabling the system to quickly and accurately understand user needs in complex domain scenarios and provide accurate answers. This method not only improves the system's ability to process natural language, but also significantly improves the intelligence level of the hydropower station knowledge question and answer system.

[0028] Step S20: Based on the structured question information, the questions are classified, and based on the classification structure and the BERT model, the intent information in the structured question information is extracted and the limiting conditions of each intent information are identified to obtain the question analysis information.

[0029] Specifically, the classification of questions is implemented based on a combination of rule matching and deep learning methods, including: the rule matching recognizes question words and their contextual semantics in the questions based on predefined grammatical patterns and question word category rules, and determines the category of the questions based on the recognition results as candidate categories; the deep learning method uses a bidirectional recurrent neural network or a recurrent convolutional neural network to verify the candidate categories by learning the contextual information and local features of the question text.

[0030] In the embodiment of the present invention, in the application scenario of a hydropower station, the questions raised by users usually involve multiple dimensions such as equipment status, maintenance operations, and power generation statistics. These questions are diverse in form and may contain multiple semantic levels. In order to accurately resolve user questions, it is necessary to classify the questions and extract the intent information and limiting conditions therein. The method of the present invention achieves efficient classification and semantic analysis of questions by combining rule matching and deep learning technology.

[0031] Furthermore, based on the structured problem information, the problems are preliminarily classified, combining rule matching and deep learning methods, which are:

[0032] 1) Rule matching: The system uses grammatical patterns and question word category rules predefined by domain experts to identify question words (such as "how", "how much", "whether", etc.) and their contextual semantics in the question. The position of the question word in the sentence and its dependencies are extracted through the grammatical parse tree, and the category of the question is analyzed in combination with the context, such as statistical questions (such as "how much electricity is generated this month") and factual questions (such as "the status of a certain device"). Candidate categories are generated based on the recognition results, which provide preliminary references for subsequent deep learning models.

[0033] 2) Deep learning verification: Candidate categories are verified by a bidirectional recurrent neural network (Bi-LSTM) or a recurrent convolutional neural network (RCNN). Bi-LSTM learns long-distance dependencies by capturing the contextual information of the question text, and is suitable for analyzing long questions with complex semantic relationships; RCNN extracts local features in combination with convolution operations, and is suitable for identifying key phrases in questions. After the system inputs the question text into the deep learning model, it outputs the final question category through the classification layer, compares it with the candidate categories that match the rules, and selects the category with the highest confidence as the final classification result.

[0034] Furthermore, based on the classification structure and semantic embedding of the BERT model, the core intent of the question is further analyzed. For example, for the question "What was the operating status of a certain device last week?", the intent can be parsed as "querying the device status." The system captures the relationship between important words in the question through BERT's contextual semantic modeling capabilities, and combines the contextual characteristics of the question category to extract explicit and implicit intent information.

[0035] Furthermore, based on intent analysis, slot labeling technology is used to identify limiting conditions in the question (such as time range, specific equipment, etc.). The system uses the Bi-LSTM+CRF model to classify conditions through sequence labeling. For example, "last week" is labeled as a time slot, and "a certain device" is labeled as a device slot. The labeled limiting conditions and intent information together constitute the question analysis information, providing an accurate semantic basis for subsequent knowledge graph links.

[0036] Step S30: Identify key entities in question parsing information based on a pre-trained entity recognition model, and link the key entities to the knowledge graph through dictionary matching and / or context analysis.

[0037] Specifically, the entity recognition model is obtained based on the coupled training of the BERT model, the Bi-LSTM model and the CRF model.

[0038] Furthermore, the identification of limiting conditions of each intent information is achieved through a slot labeling model, including: the slot labeling model labels the candidate information in the user input question as a specific slot type based on the Bi-LSTM-CRF model; wherein, the labeling process of labeling the candidate information in the user input question as a specific slot type is combined with the context embedding vector provided by the BERT pre-training model, and the long-distance dependency relationship between the slots is captured through the Bi-LSTM network; based on the slot labeling results, the overall labeling sequence is optimized through the conditional random field layer to obtain the limiting conditions of each intent information.

[0039] Furthermore, the rules of the context analysis are: based on a bidirectional gated loop, the context information of multiple rounds of dialogues in the process of natural language questions input by the user is recycled, and the semantic features of the question currently input by the user are integrated with the historical questions and answers; the historical semantics are weighted using an attention mechanism to extract the information most relevant to the current question as the result of the context analysis.

[0040] Furthermore, the key entities are linked to the knowledge graph, including: constructing an entity name dictionary of the key entities based on dictionary matching and / or context analysis results, wherein: the entity name dictionary contains the full name, abbreviations, synonyms, spelling variations and / or common nicknames of the entity; generating a set of candidate entities based on the entity name dictionary, and the dictionary's full match and / or partial match strategy, and calculating the matching confidence of the candidate entities in combination with contextual semantic information; verifying the relationship consistency between the candidate entity with the highest matching confidence and the semantic structure in the knowledge graph, and executing the corresponding candidate entity link with the highest matching confidence after completing the verification.

[0041] Furthermore, the training rules of the entity recognition model are as follows: extracting contextual features of the text based on the BERT model, and sharing the underlying feature representation for the entity recognition task and the intent classification task; during the training process, optimizing the annotation accuracy of the named entity boundaries and the generalization performance of the intent classification by jointly optimizing the multi-objective loss function; combining the self-attention mechanism to perform boundary recognition ability training for entity recognition, and obtaining an entity recognition model.

[0042] In the embodiment of the present invention, in the application scenario of a hydropower station, user questions often contain a large amount of complex entity information (such as equipment name, operating status, time range, etc.) and limiting conditions (such as "maintenance record of a certain unit" or "power generation in the past week"). In order to ensure that the question-answering system can accurately analyze these problems, the present invention proposes a technical method based on a pre-trained entity recognition model, which solves the core technical problem of complex field problem analysis through entity recognition, slot labeling, context analysis and knowledge graph linking.

[0043] Specifically, the present invention adopts an entity recognition model based on the coupled training of the BERT model, the Bi-LSTM model and the CRF model to perform high-precision recognition of key entities in user questions. Model design:

[0044] 1) BERT model: responsible for generating contextual embedding vectors of question texts, capturing long-distance dependency information between semantics, and providing deep semantic representation for subsequent processing.

[0045] 2) Bi-LSTM model: Through the bidirectional recurrent neural network (Bi-LSTM), the sequential dependencies in the embedding vector are further learned, which is particularly suitable for processing complex entity descriptions in the field of hydropower stations.

[0046] 3) CRF model: Use conditional random fields (CRF) to globally optimize the annotation results of entity boundaries and types to ensure the consistency of the recognition sequence. For example, "Unit No. 1" is accurately labeled as the equipment name entity in "Maintenance Unit No. 1".

[0047] The model optimizes both entity recognition and intent classification tasks, using shared underlying features to improve the performance of both. Through weighted optimization of the loss function, it takes into account both the accuracy of entity boundary annotation and the generalization ability of intent classification. Introducing the self-attention mechanism in the entity recognition process strengthens the model's perception of entity contextual relationships, which is particularly effective when dealing with polysemous words or long sentences.

[0048] Furthermore, in the hydropower station scenario, questions often contain multiple constraints, such as time, location, equipment name, etc. The slot labeling model is based on the Bi-LSTM-CRF architecture and achieves accurate extraction of constraints through the following steps:

[0049] 1) Context embedding vector generation: The BERT model is used to generate context embedding vectors to capture the semantic information of the candidate slots.

[0050] 2) Long-distance dependency learning: The Bi-LSTM network learns the long-distance dependencies between slots through forward and backward recurrent neural networks, such as extracting time slots from “the past week”.

[0051] 3) Global optimization: The CRF layer performs global optimization on the annotation sequence to eliminate potential conflicts and ensure the consistency of slot annotations.

[0052] Through the slot labeling model, the system can accurately divide the limiting conditions in the user's question into specific slot types, such as time slots, device slots, etc., which constitute an important part of the problem resolution information.

[0053] In the hydropower station question-and-answer scenario, the user's questions may be gradually refined across multiple rounds of dialogue. To solve this problem, the present invention uses a bidirectional gated recurrent unit (BiGRU) and an attention mechanism to dynamically integrate the contextual semantics in multiple rounds of dialogue:

[0054] 1) Contextual semantic fusion: BiGRU is used to capture the association between the current question and historical questions and answers, and to fuse the question currently input by the user with historical semantic features.

[0055] 2) Attention mechanism: Weight the semantic information in multiple rounds of dialogue and extract the semantic information most relevant to the current question, thereby enhancing the accuracy of question parsing.

[0056] The context analysis results not only optimize the semantic understanding of the current question, but also provide contextual guidance for entity recognition and slot labeling, ensuring parsing consistency in multi-round dialogue scenarios.

[0057] Furthermore, after entity recognition is completed, the system links key entities to the knowledge graph to ensure that the entity information in the question can be associated with the structured data in the knowledge graph. The specific implementation includes the following steps:

[0058] 1) Entity name dictionary construction: Build an entity name dictionary containing the full name, abbreviations, synonyms, spelling variants and common nicknames of entities to cover various expressions of entities in the field of hydropower stations.

[0059] 2) Candidate entity generation: Based on the dictionary matching strategy, a set of candidate entities is generated using full matching and partial matching methods. Combined with contextual semantic information, the most relevant entities are screened by calculating the matching confidence of the candidate entities.

[0060] 3) Verification and linking: Verify the consistency of the relationship between the candidate entity with the highest confidence and the semantic structure in the knowledge graph to ensure that the selected entity accurately matches the question semantics.

[0061] After verification, the matching entities are linked to the knowledge graph to form a complete entity relationship mapping.

[0062] Step S40: Construct the question resolution information and the information linking the key entities to the knowledge graph into a semantic structure, verify it through semantic pattern matching with the knowledge graph, and output the user intention based on the verification result.

[0063] Specifically, the question resolution information and the information linking the key entities to the knowledge graph are constructed into a semantic structure and verified through semantic pattern matching with the knowledge graph, including: mapping the question resolution information to the entities and relationship nodes in the knowledge graph, and retrieving the matching status of the resolution results and the relationship chains in the knowledge graph through a path search algorithm; for each key entity and its associated relationship in the resolution result, a global consistency score is calculated using a graph embedding algorithm; when the global consistency score deviates from the preset standard consistency score range, a candidate relationship chain is reselected based on a weight-adjusted optimization algorithm, and the semantic structure is updated to re-execute verification until the global consistency score is within the preset standard consistency score range.

[0064] Furthermore, outputting the user intent based on the verification result includes: calling a structured data interface for generating question analysis, and executing the user intent output; wherein the user intent includes the question type, the knowledge graph mapping of key entities, and parameter information of the limiting conditions; and in the output process, enhancing the intent result by combining the contextual relationship of the entities in the knowledge graph and the verified semantic consistency score.

[0065] In the embodiment of the present invention, in the hydropower station application scenario, the questions raised by users usually contain complex semantic information and specialized entity relationships. In order to accurately understand the user's intention, the present invention links the question analysis information and key entities to the knowledge graph, builds a semantic structure, verifies and optimizes the semantic model of the knowledge graph, and finally outputs the user's intention based on the verification result.

[0066] First, the key entities and their limiting conditions (such as time, location, etc.) in the problem resolution information are mapped to the entities and relationship nodes in the knowledge graph. For example, the user question "query the power generation of a certain unit last week" will be parsed to "a certain unit" as the key entity and "last week" as the time limiting condition. The semantic structure will include the unit entity and its relationship chain related to time and power generation. The path search algorithm is used to retrieve the key entities and their associations in the semantic structure and match the relationship chain in the knowledge graph. The path search algorithm finds the shortest path or the optimal path based on the topological structure of the knowledge graph to verify whether the semantic relationship in the problem resolution information is consistent with the actual relationship in the knowledge graph. For example, in the above example, the association relationship of "unit-power generation-time" in the semantic structure needs to find a qualified relationship chain in the knowledge graph to ensure semantic matching. For each matched key entity and its association relationship, the system uses the graph embedding algorithm to calculate the global consistency score. The graph embedding algorithm measures the semantic similarity between the semantic structure and the entity relationship in the knowledge graph by mapping the entities and relationships in the knowledge graph to the vector space, and generates a global consistency score. When the global consistency score deviates from the preset standard range, the system reselects the candidate relationship chain based on the weight-adjusted optimization algorithm and updates the semantic structure. The updated semantic structure will be re-verified until the global consistency score falls within the standard range to ensure that the semantic structure matches the knowledge graph accurately.

[0067] Furthermore, based on the verified semantic structure, the system calls to generate a structured data interface to output the user's intent in a structured form. User intent includes question type (such as statistical, factual), knowledge graph mapping of key entities (such as the ID or attributes of a certain device), and parameter information of limiting conditions (such as time range). In the process of intent output, the intent results are enhanced by combining the contextual relationship of the entities in the knowledge graph and the verified semantic consistency score. For example, through the contextual relationship, the system can further clarify ambiguous intent descriptions, such as "querying the status of the device" can be supplemented as "querying the operating status of a certain unit" according to the semantic structure.

[0068] The output user intent can be directly used as the input of the question-answer generation module to provide users with accurate answers. In multi-round dialogue scenarios, the system also supports dynamic updating of user intent to adapt to the gradual refinement of user needs.

[0069] Figure 2 1 is a system structure diagram of a user intention recognition system for hydropower station knowledge question and answer provided by an embodiment of the present invention. Figure 2 As shown, an embodiment of the present invention provides a user intention recognition system for hydropower station knowledge question and answer, and the system includes: a collection unit, which is used to recover natural language questions input by users, and perform structured processing on the natural language questions to obtain structured question information; a parsing unit, which is used to classify questions based on the structured question information, and extract intent information in the structured question information based on the classification structure and the BERT model and identify the limiting conditions of each intent information to obtain question parsing information; a linking unit, which is used to identify key entities in the question parsing information based on a pre-trained entity recognition model, and link the key entities to the knowledge graph through dictionary matching and / or context analysis; wherein the entity recognition model is obtained based on the coupled training of the BERT model, the Bi-LSTM model and the CRF model; an output unit, which is used to construct the information linking the question parsing information and the key entities to the knowledge graph into a semantic structure, and verify it through semantic pattern matching with the knowledge graph, and output the user intention based on the verification result.

[0070] An embodiment of the present invention further provides a computer-readable storage medium, on which instructions are stored, which, when executed on a computer, enable the computer to execute the above-mentioned method for identifying user intentions in hydropower station knowledge questions and answers.

[0071] Those skilled in the art will understand that all or part of the steps in the method for implementing the above-mentioned embodiments can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including several instructions for making a single-chip microcomputer, a chip or a processor (processor) perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0072] The optional embodiments of the present invention are described in detail above in conjunction with the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above embodiments. Within the technical concept of the embodiments of the present invention, the technical scheme of the embodiments of the present invention can be subjected to a variety of simple modifications, and these simple modifications all belong to the protection scope of the embodiments of the present invention. It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the embodiments of the present invention will not further describe various possible combinations.

[0073] In addition, various embodiments of the present invention may be arbitrarily combined, and as long as they do not violate the concept of the embodiments of the present invention, they should also be regarded as the contents disclosed in the embodiments of the present invention.

Claims

1. A method for identifying user intentions in a hydropower station knowledge question and answer session, characterized in that: The method comprises: Retrieving the natural language question input by the user, and performing structured processing on the natural language question to obtain structured question information; Based on the structured question information, the questions are classified, and based on the classification structure and BERT model, the intent information in the structured question information is extracted and the limiting conditions of each intent information are identified to obtain the question analysis information; Based on the pre-trained entity recognition model, key entities in question parsing information are identified, and key entities are linked to the knowledge graph through dictionary matching and / or context analysis; The entity recognition model is obtained based on the coupled training of the BERT model, the Bi-LSTM model and the CRF model; The question parsing information and key entity links to the information of the knowledge graph are constructed into a semantic structure, verified by semantic pattern matching with the knowledge graph, and the user intention is output based on the verification results.

2. The method according to claim 1, characterized in that: The problem classification is based on the combination of rule matching and deep learning methods, including: The rule matching recognizes the interrogative words in the question and their contextual semantics based on the predefined grammatical patterns and interrogative word category rules, and classifies the question based on the recognition results as candidate categories; The deep learning method adopts a bidirectional recurrent neural network or a recurrent convolutional neural network to verify candidate categories by learning context information and local features of the question text.

3. The method according to claim 1, characterized in that: The identification of the limiting conditions of each intent information is realized through the slot annotation model, including: The slot labeling model labels the candidate information in the user input question as a specific slot type based on the Bi-LSTM-CRF model; wherein, The process of labeling the candidate information in the user input question as a specific slot type is combined with the context embedding vector provided by the BERT pre-training model to capture the long-distance dependency between slots through the Bi-LSTM network; Based on the slot labeling results, the overall labeling sequence is optimized through the conditional random field layer to obtain the limiting conditions of each intent information.

4. The method according to claim 1, characterized in that The rules of the context analysis are: Based on the bidirectional gated loop, the context information of multiple rounds of dialogue in the process of natural language questions input by users is collected, and the semantic features of the current question input by the user are integrated with the historical questions and answers. The attention mechanism is used to weight the historical semantics and extract the most relevant information to the current problem as the result of context analysis.

5. The method according to claim 1, characterized in that Link key entities to the knowledge graph, including: Build an entity name dictionary for key entities based on dictionary matching and / or context analysis results, where: The entity name dictionary contains the full name, abbreviations, synonyms, spelling variations and / or common nicknames of the entity; Generate a candidate entity set based on an entity name dictionary and a complete match and / or partial match strategy of the dictionary, and calculate the matching confidence of the candidate entity in combination with contextual semantic information; In the knowledge graph, the consistency between the candidate entity with the highest matching confidence and the semantic structure is verified, and after the verification is completed, the corresponding candidate entity link with the highest matching confidence is executed.

6. The method according to claim 1, characterized in that The information of question resolution and key entity links to the knowledge graph is constructed into a semantic structure and verified by semantic pattern matching with the knowledge graph, including: Map the problem analysis information to the entity and relationship nodes in the knowledge graph, and use the path search algorithm to retrieve the matching between the analysis results and the relationship chain in the knowledge graph; For each key entity and its associated relationship in the parsing result, a global consistency score is calculated using a graph embedding algorithm; When the global consistency score deviates from the preset standard consistency score range, the optimization algorithm based on weight adjustment reselects the candidate relationship chain, updates the semantic structure and re-executes verification until the global consistency score is within the preset standard consistency score range.

7. The method according to claim 1, characterized in that The training rules of the entity recognition model are: Extract contextual features of text based on the BERT model and share the underlying feature representation for entity recognition and intent classification tasks; During the training process, the annotation accuracy of named entity boundaries and the generalization performance of intent classification are optimized simultaneously by jointly optimizing the multi-objective loss function; Combined with the self-attention mechanism, the boundary recognition ability training of entity recognition is performed to obtain the entity recognition model.

8. The method according to claim 1, characterized in that Output user intent based on the verification results, including: Call the structured data interface for generating question parsing and execute the user intention output; The user intent includes question type, knowledge graph mapping of key entities and parameter information of limiting conditions; During the output process, the contextual relationships of entities in the knowledge graph and the verified semantic consistency scores are combined to enhance the intent results.

9. A user intention recognition system for hydropower station knowledge question and answer, characterized in that: The system comprises: A collection unit, used for collecting natural language questions input by users, and performing structured processing on the natural language questions to obtain structured question information; A parsing unit is used to classify questions based on structured question information, extract intent information from the structured question information based on the classification structure and the BERT model, identify the limiting conditions of each intent information, and obtain question parsing information; A linking unit is used to identify key entities in question parsing information based on a pre-trained entity recognition model, and link the key entities to the knowledge graph through dictionary matching and / or context analysis; wherein, The entity recognition model is obtained based on the coupled training of the BERT model, the Bi-LSTM model and the CRF model; The output unit is used to construct the question resolution information and the information linking the key entities to the knowledge graph into a semantic structure, and verify it through semantic pattern matching with the knowledge graph, and output the user intention based on the verification result.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the method for identifying user intentions in a hydropower station knowledge question and answer session as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Electric knowledge base accurate retrieval and intelligent question-answering system and method based on AI large model and knowledge graph

    CN121029937A

  • A Precise Retrieval and Intelligent Question Answering System and Method for Electrical Knowledge Base Based on AI Large Model and Knowledge Graph

    CN121029937B