Method and system for natural language query of SCD files of smart substation
By importing SCD files into a graph database and converting them into Cypher query statements using knowledge graphs and semantic triples, the flexibility and inefficiency of SCD file query tools are solved, and the friendliness and efficiency of natural language queries are achieved.
Patent Information
- Application Number
- CN202111496011.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2041-12-08
AI Technical Summary
The query mode of existing SCD file query tools is rigid, the query conditions are not flexible enough, and it is difficult to meet teaching needs. In addition, the SCD file code is complex, the query efficiency is low, and it is difficult for professionals to read and understand it directly.
Import the SCD file into the graph database, use the knowledge graph to perform natural language queries, and convert semantic triples into Cypher query statements to achieve flexible queries.
It reduces the difficulty for beginners to query SCD files, provides more natural and friendly exploratory learning support, and improves query efficiency and flexibility.
Smart Images

Figure CN114168615B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power systems, and in particular to a method and system for querying SCD files of smart substations using natural language. Background Art
[0002] Smart substations are critical infrastructure for achieving a robust smart grid. Their design must adhere to the IEC 61850 standard, proposed by IEC TC57. This standard enables interoperability between intelligent electronic devices (IEDs) from different manufacturers, improving the scalability and maintainability of substation automation systems. The Substation Configuration Language (SCL), defined in Part 6 of the IEC 61850 standard, is the foundation for IED interoperability. The most critical SCD (Substation Configuration Description) file within the SCL file system describes key configuration information for the smart substation's secondary system, primarily including IED configuration information and communication parameters.
[0003] Thomas et al. proposed that experimental design for substation automation should incorporate relevant IEC 61850 protocols and analysis software. However, the SCL language standard system is complex, making it difficult for general engineers to directly read SCD files. Furthermore, SCD file codes are very long. A typical SCD file for a 220 kV smart substation may contain nearly 10 million lines of code. Searching for data using XML text editing tools such as XMLSpy and XML Notepad is complex and inefficient.
[0004] Currently, multiple equipment manufacturers and software companies have developed SCD configuration software for secondary systems. These software generally provide SCD information browsing capabilities based on predetermined design and debugging requirements. These software features rigid query modes, inflexible query input, and generally lack a mapping between query results and SCD codes, resulting in significant limitations for learning SCD files. Current SCD research focuses primarily on SCD file version comparison and verification, as well as virtual circuit construction. However, significant room for improvement remains in enhancing the flexibility and convenience of SCD file queries. Consequently, current secondary system configuration software struggles to meet the teaching needs of exploratory SCD experiments.
[0005] In recent years, some researchers have proposed natural language-based data query techniques, exploring the use of natural language questions to query BIM design files or databases. More recently, a study has explored the use of natural language combined with knowledge graphs to query fire information from building IFC files. However, because SCD files and IFC files adhere to different standards and target different disciplines, the natural language-based IFC query techniques proposed by previous researchers are not suitable for querying SCD files. Therefore, it is necessary to develop natural language interactive tools for SCD file queries, tailored to the design and management requirements of smart substation secondary logic systems. Summary of the Invention
[0006] In view of the defects in the prior art, the purpose of the present invention is to provide a method and system for natural language query of SCD files of smart substations.
[0007] According to the present invention, a method for querying an SCD file of a smart substation using natural language is provided, comprising:
[0008] Import step: import the SCD file into the graph database;
[0009] Information supplementation step: obtain the natural language query statement input into the graph database, modify it through the knowledge graph to obtain the suggested question sentence, supplement the omitted attributes in the suggested question sentence, and replace the synonyms of professional terms with standard reference terms;
[0010] Semantic information extraction step: extracting multi-relational semantic information from the modified natural language query statement, and expressing the obtained multi-relational semantic information as semantic triples;
[0011] Conversion step: obtaining a general question sentence based on the semantic triple, querying the corresponding assembly template from an existing assembly template database based on the general question sentence, converting the semantic triple into a Cypher code segment, and assembling the Cypher code segment into a Cypher query statement for the graph database using the assembly template obtained from the semantic triple query;
[0012] Query step: Use the Cypher query statement to query the content of the corresponding SCD file in the graph database.
[0013] Preferably, the information supplementing step includes:
[0014] The knowledge graph is used to add SCL engineering attributes to natural language query statements and replace synonyms of professional terms with standard reference terms.
[0015] Preferably, the information extraction step includes:
[0016] Calculating character vectors of the characters in the modified natural language query statement;
[0017] Evaluate the contextual features of each character to obtain the linguistic or semantic relationship between characters;
[0018] All subjects are identified based on contextual features, and the objects and predicates associated with each subject are identified to obtain semantic triples.
[0019] Preferably, the conversion step comprises:
[0020] The assembly template for assembling the Cypher code segment is selected according to a set of semantic triples obtained from the Cypher query statement.
[0021] Preferably, the suggested question obtained after the knowledge graph correction includes: searching for attribute K [1] is the value V [1] and / or ... attribute K [n] is the value V [n] The attribute L of node m related to node n [1] ...and the property L [n] .
[0022] According to the present invention, a system for querying SCD files of smart substations using natural language is provided, comprising:
[0023] Import module: import SCD files into graph database;
[0024] Information supplementation module: obtains natural language query statements input into the graph database, modifies them through the knowledge graph to obtain suggested questions, supplements omitted attributes in the suggested questions, and replaces synonyms of professional terms with standard reference terms;
[0025] Semantic information extraction module: extracting semantic information from the modified natural language query statement, and expressing the obtained semantic information as semantic triples;
[0026] Conversion module: obtains a general question sentence based on the semantic triple, searches for the corresponding assembly template from the existing assembly template database based on the general question sentence, converts the semantic triple into a Cypher code segment, and uses the assembly template obtained from the semantic triple query to assemble the Cypher code segment into a Cypher query statement for the graph database;
[0027] Query module: uses the Cypher query statement to query the content of the corresponding SCD file in the graph database.
[0028] Preferably, the information supplement module includes:
[0029] The knowledge graph is used to add SCL engineering attributes to natural language query statements and replace synonyms of professional terms with standard reference terms.
[0030] Preferably, the information extraction module includes:
[0031] Calculating character vectors of the characters in the modified natural language query statement;
[0032] Evaluate the contextual features of each character to obtain the linguistic or semantic relationship between characters;
[0033] All subjects are identified based on contextual features, and the objects and predicates associated with each subject are identified to obtain semantic triples.
[0034] Preferably, the conversion module comprises:
[0035] The assembly template for assembling the Cypher code segment is selected according to a set of semantic triples obtained from the Cypher query statement.
[0036] Preferably, the suggested question obtained after the knowledge graph correction includes: searching for attribute K [1] is the value V [1] and / or ... attribute K [n] is the value V [n] The attribute L of node m related to node n [1] ...and the property L [n] .
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] This paper uses natural language to query SCD files. It first converts SCD files into graph data storage and proposes a method for automatically generating Cypher query statements from natural language. Experimental results based on SCD files from real-world engineering examples demonstrate that this method can reduce the difficulty of querying SCD data for beginners and provide a more natural and user-friendly support for exploratory learning of the SCL language and the structure of SCD files. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:
[0040] Figure 1 It is a workflow diagram of the present invention;
[0041] Figure 2 The extraction flow chart of triples;
[0042] Figure 3 Generate a flowchart for the query statement;
[0043] Figure 4 Schematic diagram of the query result of the present invention;
[0044] Figure 5 A schematic diagram of generating a Cypher query statement from a triple set;
[0045] Figure 6 This is a structural diagram of the subject extractor. DETAILED DESCRIPTION
[0046] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.
[0047] like Figure 1 As shown, this embodiment provides a method for querying the SCD file of a smart substation using natural language, which includes the following steps.
[0048] 1. Import steps: Import the SCD file into the graph database.
[0049] Although the hierarchical relationships between data nodes in SCD files stored in XML format are very clear, many secondary system business queries involve complex relationships between nodes at multiple levels. Querying directly in the form of XML files is uneconomical in terms of programming cost and efficiency. If a relational database is used to store SCD files, a large number of Join operations will be generated during queries, resulting in low query efficiency and high computing power consumption. Therefore, considering using a NoSQL database to store SCD files. Compared with key-value storage databases and document databases, graph databases have the advantages of high efficiency and simple operation for path-based queries between nodes.
[0050] Graph databases use nodes as data elements and edges between two nodes to represent binary relationships between data elements. Both nodes and edges allow for the definition of attributes, making them more suitable for complex relationships between data objects. Compared to other types of databases, graphs are more suitable for searching multi-level, mesh-connected data based on path conditions. Their search engines also optimize for relational queries, effectively preventing the need to read all data caused by localized data queries. Commonly used graph databases include Neo4j, JanusGraph, and HugeGraph. Neo4j's graph analysis capabilities are relatively comprehensive, supporting billions of data points and providing a rich set of graph theory algorithms, including path search, similarity, centrality, community detection, and link prediction. It also offers query result visualization, enabling the import and storage of large SCD files. Therefore, SCD files can be imported into the Neo4J graph database programmatically.
[0051] Graph databases feature a unified query language. Neo4J's query language, Cypher, is a declarative language that eliminates the need to describe specific computational steps like in traditional procedural programming. Instead, it uses a formal language to describe the results and constraints of data filtering, making it easier to map to natural language questions. Table 1 lists commonly used Cypher commands, including those for creating, deleting, and querying data nodes, attributes, and paths. In addition to the basic data manipulation commands mentioned above, Cypher query statements can also utilize an extended function library. For example, the APOC string concatenation function can be used to extract the reference address of an output virtual terminal from the six attributes of an ExtRef node.
[0052] Table 1. Basic Cypher operation commands
[0053]
[0054]
[0055] In particular, Cypher allows queries along non-deterministic paths, eliminating the need to explicitly define each node in the path, and requiring only some key path information. For example, in the third query example in Table 1, we need to query the IEDs associated with a virtual circuit described as "breaker location." The connection path between the two IEDs is very long, and the virtual circuit association that satisfies the "breaker location" condition in the desc (description) attribute is just a "segment" on the path. This is very similar to the way humans think about queries in natural language, and was a major reason for choosing Neo4J.
[0056] Second, the information supplementation step: A natural language query is input into the graph database, which is then modified using the knowledge graph to generate suggested questions. Omitted attributes are added to the suggested questions, and synonyms for specialized terms are replaced with standard reference terms. Specifically, the knowledge graph is used to add SCL engineering attributes to the natural language query and replace synonyms for specialized terms with standard reference terms.
[0057] Many existing NLP training corpora primarily draw from news, social networks, e-commerce websites, and scientific literature. Unfortunately, terminology related to substation automation is relatively limited, especially in SCD documents. Many existing NLP tools still struggle to recognize that "CYG Sunri" is the name of the equipment manufacturer. Training existing NLP models to recognize such specialized vocabulary or named entities is relatively costly, so a knowledge graph was developed to assist in extracting semantic information using the information extraction model described in the next paragraph. Another reason for incorporating a knowledge graph is to replace synonyms in interrogative sentences with "standard" terms. For example, "electronic intelligent device," "intelligent device," and "smart device" all refer to the same concept of intelligent electronic devices defined in the IEC 61850 standard, the abbreviation "IED" officially used in SCD documents. In the knowledge graph database, all of these synonyms are equivalent to the single word "IED" in the SCD document.
[0058] Furthermore, considering that the application background of the present invention is that professional technicians query SCD, in extreme cases, the characteristic words indicating semantic relationships are even directly omitted, which brings great difficulties to computer recognition semantic management. Therefore, it is necessary to supplement information before performing semantic analysis. For example, the query user asks the query "Give me the IED of NARI Relay Protection", in which "manufacturer" is omitted, because professionals in this context know that "NARI Relay Protection" is the manufacturer of relay protection intelligent equipment. To solve this problem, the values of SCL attributes can be obtained from the collected engineering design SCD files to obtain a list of SCL keywords corresponding to named entities. Each named entity, such as "NARI Relay Protection" or "NRR", corresponds to an SCL language engineering attribute "LED.manufacture", where IED is an XML node tag and manufacture is the attribute key name of the IED node. With this knowledge graph, when engineering attributes are missing in natural language questions, the system can automatically fill in the blanks. For example, the system can automatically fill in the above question with "Give me the IED whose manufacturer is NARI Relay Protection". The automatically filled "manufacturer" is a standard reference term in the synonym library, and its SCL engineering attributes can be inferred by querying the database through the entity name in the knowledge graph.
[0059] like Figure 2 As shown, the information-supplemented natural language question is input into the information extraction model, which is used to extract semantic triples from the corrected natural language query statement. Figure 2 The semantic information extraction pipeline from input question to generated semantic triples is shown, which mainly consists of four functions.
[0060] The Character Embedding module computes a character vector for each Chinese character in the query. The Shared Semantic Encoder module then evaluates the contextual features of each character to capture the linguistic or semantic relationships between them. Here, context refers to the characters before and after the character being identified, and a Bi-LSTM model is used accordingly. The Subject Extractor module then uses the computed contextual features to identify all candidate subjects. Subsequently, the Object and Predicate Extractor module uses the contextual features computed in the previous step to identify the associated object entities and predicates for each identified subject entity. This allows for the extraction of multiple semantic triples from natural language questions.
[0061] A single Chinese character alone cannot store effective semantic information, whereas a specific word conveys more explicit and comprehensive semantic information than a single Chinese character. Therefore, using only word-level vectorization will not fully utilize the semantic information carried by words for text analysis. In particular, for the present invention, a specification often contains a large number of domain-specific terms, and the semantic information of these proprietary terms is crucial for analyzing standard text data.
[0062] After obtaining the word vector and word vector representation of the standard text, the word vector representation and word vector representation are fused. In general, given a standard sentence S = {c1, c2, c3, ..., c n}, where c i For the i-th character of the sentence, this paper defines the word mixture vector x of the i-th character i as follows:
[0063] x i =[e c (c i 10W T e w (word(c i ))],
[0064] where e c and e w Respectively represent the query pre-trained word vector model and query pre-trained word vector model operations; word(c i ) represents the i-th character c i The words it belongs to are realized by the Chinese word segmentation tool; Is a trainable transformation matrix used to transform the word vector dimension d w With word vector dimension d c Align.
[0065] The word vector of mixed words is to transform the word vector sequence to the same dimension as the word vector through a matrix, and add the two together. During the model training process, the fixed pre-trained word vector model and the pre-trained word vector model remain unchanged, so the transformation matrix needs to be continuously optimized. From another perspective, it can also be considered that the prior word semantic information in the word vector model is integrated into the word vector through the transformation matrix, which not only makes full use of the prior semantic information of the word, but also retains the flexibility of the word vector. The word vector representation of mixed words only provides the prior semantic information of each word in the standard text. It is difficult to directly express the actual semantic information expressed by the sentence by relying solely on the prior semantic information of the word. It is necessary to further model the word sequence and mine the semantic information carried in the sequence. Therefore, the present invention designs a semantic encoder based on a bidirectional long short-term memory network to perform sequence modeling on the standard text, fully mine the context information, thereby extracting potential semantic features, and sharing the extracted features as underlying parameters for downstream tasks, so it is also called a shared semantic encoder.
[0066] Information Extraction Step 3: Extract multi-relational semantic information from the modified natural language query and represent the resulting multi-relational semantic information as semantic triples. Specifically, character vectors are calculated for the characters in the modified natural language query, and the contextual features of each character are evaluated to determine the linguistic or semantic relationships between the characters. Based on the contextual features, all subjects are identified, along with the objects and predicates associated with each subject, to generate multi-relational semantic triples.
[0067] The triple structure is a binary relational model consisting of three parts: subject (Subject Entity), predicate (Predicate), and object (ObjectEntity), expressed as "[Subject, Predicate, Object]", also known as Subject-Predicate-Object (SPO) triple. Among them, the predicate defines the semantic relationship between the subject and the object (from subject to object). In the field of information extraction, the subject is also called the head entity, the object is called the tail entity, and the predicate is also called the semantic relationship. The extraction of semantic relationships is divided into two categories: the extraction of semantic relationships involving only two named entities is called binary relationships; the extraction of semantic relationships involving three or more entities is called multi-relation extraction. Generally, multi-relationships can be constructed by extracting multiple binary semantic relationships.
[0068] The core of the shared semantic encoder utilizes a bidirectional long short-term memory (BiLSTM) model to model the context of text sentences. Traditional recurrent neural networks struggle with the loss of early memory (gradient vanishing) and the inability to write later memory (gradient explosion) when processing long time series data. Long short-term memory networks, by introducing a gating mechanism to continuously correct memory, address this issue of long-term dependency learning.
[0069] An LSTM model includes gate mechanisms for "remembering" (acquiring new information) and "forgetting" (discarding unnecessary information). However, a unidirectional LSTM can only predict the output of the next moment based on the temporal information of the previous moment, but in some problems, the output of the current moment is not only related to the previous state, but may also be related to the future state. For example, in this article, to predict whether certain consecutive characters in a standard text are named entities, it is necessary not only to judge based on the previous text, but also to consider the content after it in order to make a correct judgment. Therefore, this article adopts a bidirectional LSTM model, that is, extracting contextual features from both the forward and backward directions. Specifically, in the forward (forward reading) direction, the Bi-LSTM network calculates from left to right along the input vector sequence, while in the reverse (reverse reading) direction, it calculates from right to left. In this way, the association between a character and its surrounding characters (on the left and right) is encoded by connecting the forward and backward LSTM hidden states. In general, the output of the BiLSTM model can be expressed as follows:
[0070]
[0071] in, and Represent the hidden states of the forward (from left to right) and backward (from right to left) directions at time step t, Represents the matrix concatenation operation. With the help of BiLSTM model, the context feature h of the standard text can be extracted t In order to further explore the global features of text sentences, this paper performs max pooling on the hidden state output by the BiLSTM model to obtain the sentence feature g. Finally, the context feature h of the standard text is t Combined with the sentence feature g to form the task shared feature [h t ; g]. Task-shared features will be used jointly by downstream models, thereby realizing underlying parameter sharing and achieving the goal of joint learning.
[0072] For each sentence to be processed, all candidate subjects in the standard text are first extracted, and then the objects and semantic relations corresponding to the candidate subjects are extracted based on the semantic information of each candidate subject. The function of the subject extractor (SubjectExtractor) is to identify all possible candidate subjects in the standard text. In Chinese query sentences, candidate subjects are often composed of multiple consecutive characters. In order to identify these named entities and further discover overlapping semantic relations on this basis, this paper introduces a pointer annotation structure, which models the extraction problems of candidate subjects, objects and semantic relations as sequence annotation tasks. The overlapping semantic relations here refer to the different types of semantic relations between a subject and multiple objects in a sentence.
[0073] The pointer annotation structure is used here to semantically encode the query question and, based on the encoding results, to find the corresponding answer in the corresponding text. The pointer annotation structure provides a starting and ending position pointer to segment the text into segments to form the answer. The present invention uses the pointer annotation structure to extract triples from the input text.
[0074] After receiving the task-shared features output by the shared semantic encoder, the subject extractor first uses a BiLSTM model to extract task-specific features from the shared features. It then uses two multi-head self-attention models to learn the starting and ending position features of the subject, respectively. The self-attention model helps identify which words are most important for each subtask in multi-task learning.
[0075] This paper selects the scaled dot product model as the attention scoring function, is the task-specific feature output by the BiLSTM layer, then in the self-attention mechanism, the relationship between the query matrix Q, the key matrix K and the value matrix V is Q = K = V = h se The attention function can be expressed as follows:
[0076]
[0077] in d represents the hidden state dimension of the BiLSTM layer output, which is equal to 2d h Furthermore, the multi-head self-attention mechanism connects multiple self-attentions. Operationally, it transforms the query matrix Q, key matrix K, and value matrix V into m Q, K, and V through a linear mapping. Through the multi-head attention mechanism, the model can process sequence data using representation information from different subspaces at different sequence positions, achieving better recognition results.
[0078] Assuming that the multi-head self-attention mechanism contains m heads, the i-th attention head can be expressed as follows:
[0079]
[0080] in, is the projection parameter to be trained, d k =2d h / m. The final result of the multi-head self-attention mechanism is composed of the concatenation of each attention head:
[0081]
[0082] in are the parameters to be trained.
[0083] Afterwards, the output of the multi-head self-attention mechanism is input into a fully connected layer with a Softmax activation function to generate a label probability distribution on each character. t The label calculation is as follows:
[0084]
[0085]
[0086] in is a training parameter, |T| is the number of output label types, and the value of |T| is the named entity category plus 1.
[0087] Similarly, the multi-head self-attention mechanism can also be used to learn the terminal position features of the subject. Considering that the information of the subject's starting position can provide useful help for predicting the terminal position, this paper converts the output h of the first multi-head self-attention mechanism into se-sta With task-specific features h se After splicing, the Q, K, and V parameters of the second multi-head self-attention mechanism are initialized. Intuitively, this operation enables the model to fully utilize the information of the starting position when predicting the ending position of the subject, and strengthen the potential correlation between the starting position and the ending position. The output of the second multi-head self-attention mechanism is recorded as Then in the process of marking the end position of the subject, the character c t The label calculation is as follows:
[0088]
[0089]
[0090] in is the training parameter.
[0091] After the above process, the subject extractor decomposes the identification of candidate subjects into two sequence labeling tasks. The first sequence labeling task is responsible for identifying the starting position of the candidate subject. If a character is identified as the starting character of the candidate subject, the corresponding position is annotated with the named entity type label of the subject. The second sequence labeling task is responsible for identifying the ending position of the candidate subject. The labeling process is the same as the starting position labeling process.
[0092] Finally, cross entropy is used to measure the loss between the probability distribution predicted by the subject extractor and the true distribution. The training loss function of the subject extractor can be written as:
[0093]
[0094] in, and are the actual starting and ending position labels of the i-th character, and n is the length of the design specification text.
[0095] After obtaining all candidate subjects, the object and semantic relationship extractor is responsible for extracting all corresponding objects and the semantic relationships between the subject and the object for each candidate subject.
[0096] For a candidate subject, its own semantic information is crucial for predicting its semantically associated objects and semantic relations. Therefore, in the object and semantic relationship extractor, the given subject must first be semantically encoded. This paper inputs a fragment of the task-shared features corresponding to the given subject into an LSTM model. The last hidden state output by the LSTM model serves as the semantic encoding of the given subject. The semantic encoding of the given subject is concatenated with each feature vector of the task-shared features to produce a new feature matrix. This new feature matrix can be considered as a semantic encoding result that carries both contextual features and the given subject's features.
[0097] The subsequent computational process of the object and semantic relationship extractor is similar to that of the subject extractor. In the first step, the feature matrix carrying the semantic information of a specific subject is input into a BiLSTM model to extract task-specific features. In the second step, two sequence labeling tasks are constructed using two multi-head self-attention mechanisms and fully connected layers. The first sequence labeling task is responsible for labeling the starting position of all candidate objects, and the second sequence labeling task is responsible for labeling the ending position of all candidate objects. Unlike the subject extractor, if a character is labeled as the starting position or ending position of a candidate object, its label is not the type of named entity, but the type of semantic relationship between the given subject and the candidate object. In this way, given a subject, its objects and corresponding semantic relationships can be extracted simultaneously.
[0098] Generally, for a given subject, the start and end position of the object are marked, and the character c t The label calculations are as follows:
[0099]
[0100]
[0101]
[0102]
[0103] in, are the outputs of the first and second multi-head self-attention mechanisms at timestamp t, respectively; b ope-sta ,as well as are all training parameters; |T′| is the number of output label types, and the value of |T′| is the number of semantic relations plus 1. Finally, cross entropy is used to measure the loss between the probability distribution predicted by the object and semantic relationship extractor and the true distribution. The training loss function of the object and semantic relationship extractor is:
[0104]
[0105] in, are the actual starting and ending position labels of the i-th character, and n is the length of the standard text.
[0106] During the model training phase, the subject extractor and the object and semantic relationship extractor are jointly trained using the task-shared features provided by the shared semantic encoder. In each training instance, a subject is randomly selected from a standard dataset of canonical text as input to the object and semantic relationship extractor. The loss function values for the subject extractor and the object and semantic relationship extractor are calculated by measuring the difference between the model's predictions and the standard results. Finally, the two losses are summed to form the final loss function of the joint model:
[0107]
[0108] This paper uses the Adam algorithm to optimize the final loss function of the model, so that the errors generated in the subject extraction, object and semantic relationship extraction processes affect each other. The errors generated by each subtask are constrained by other tasks, thereby strengthening the potential interaction between named entities and semantic relationships. After the model training is completed, the pseudo code of the inference algorithm using the model to predict triples in the standard text is shown in Table 3: The pseudo code lines 1-3 mainly show the preparatory work before using the model for inference: setting the standard text length parameter n, initializing the candidate subject set and triple sets Lines 4 - 12 of the pseudocode describe the process of extracting all candidate subjects based on the annotation results of the subject extractor; lines 13 - 24 of the pseudocode describe the process of extracting objects and semantic relations.
[0109] Table 3. Pseudocode of the inference algorithm
[0110]
[0111]
[0112] As can be seen from the above inference algorithm, generally for a specification statement containing k subjects, the entire entity relation joint extraction task is decomposed into 2 + 2k sequence labeling tasks. The first 2 sequence labeling tasks are completed in the subject extractor, mainly aiming to identify all candidate subjects in the statement; the latter 2k sequence labeling tasks are completed in the object and semantic relation extractor, mainly aiming to identify the objects and corresponding semantic relations for each subject.
[0113] IV. Transformation steps: Obtain a general question based on the semantic triple, query the corresponding assembly template from the existing assembly template database according to the general question, convert the semantic triple into a Cypher code segment, and assemble the Cypher code segment into a Cypher query statement for the graph database using the assembly template queried according to the semantic triple.
[0114] Figure 3 It illustrates the process of converting a natural language question into a query statement. The Chinese question "Find the MAC address associated with *IL1101:RPIT / CBXCBR1.Pos.stVal*" is first supplemented and perfected. Specifically, there are often some omissions in the original query question, such as "iedname", "FCDA", and "value". Figure 3 It illustrates that these omitted words are automatically supplemented with the help of the knowledge graph specially developed for SCD. At the same time, the synonym "give" of the Chinese word is replaced by the standard word "find" and stored in the knowledge graph database.
[0115] Subsequently, the system performs semantic information extraction, generates a set of semantic triples, and then converts these semantic triples into Cypher code segments, such as Figure 3For example, according to the predefined conversion rules, the two semantic triples "Find(Address, Value)" and "Has_Attribute(Mac, Address)" are converted into the Cypher code segment "MATCH aMac:Mac RETURN aMac.Address". There are four main conversion rules for the main predicate types, including Find, Has_Attribute, Has_Connection, and Has_Value.
[0116] A code snippet is only a component of a complete Cypher query statement, and the template of a typical query sentence is needed to assemble the converted Cypher code snippet into a complete Cypher query statement. The Cypher code snippet derived from the semantic triple should be further assembled into a complete query statement according to the code assembly template. The choice of this assembly template can be inferred from the set of semantic triples. After interviewing 15 engineers and designers from the State Grid Corporation of China and researchers from a university specializing in substation automation, more than 400 questions were collected and some complex questions were removed from the sample library because they were too difficult for computers to understand. Many of the eliminated questions were then broken down into several simple questions that can be understood by NLP tools. Of all human questions, more than 52% follow Figure 3 Finally, the Cypher code snippet is automatically assembled into a complete query statement according to the assembly template, such as Figure 3 shown.
[0117] Converting a semantic triple into a Cypher query statement is divided into three steps:
[0118] First, the semantic triples are converted into corresponding Cypher code segments. Table 4 shows the template for converting semantic triples into Cypher code segments.
[0119] Table 4. Template for semantic triple conversion Cypher code snippet
[0120]
[0121]
[0122] At the same time, a corresponding general question is obtained based on the semantic triple. Based on the general question, a corresponding assembly template is searched from an existing assembly template database. The Cypher code segment is then assembled into a complete Cypher query statement based on the assembly template obtained from the query. Based on the relationship between the subject, predicate, and object in the semantic triple, a general question is obtained. Different subjects, predicates, and objects can each generate a corresponding general question. The assembly template database stores different assembly templates, each corresponding to a general question. The present invention does not limit the number or form of assembly templates and general questions.
[0123] A Cypher code snippet is only part of a complete Cypher statement. Assembly templates corresponding to common questions are needed to assemble the converted Cypher code snippets into complete Cypher query statements. Common questions were obtained through survey interviews. Through interviews with engineering technicians, professional course teachers, and students majoring in power systems and automation, 400 typical questions were identified. After summarization and organization, 95 complex questions that were difficult for computers to understand were eliminated. These eliminated questions can basically be broken down into multiple simple questions. The remaining 305 questions were corrected using the knowledge graph to obtain the common question pattern shown in the following example:
[0124] "Find attribute K [1] is the value V [1] and / or ... attribute K [n] is the value V [n] The attribute L of node m related to node n [1] ...and the property L [n] ”.
[0125] Corresponding assembly templates are provided for the various general questions described above, which are used to assemble and connect Cypher code segments converted from triples to generate Cypher query statements. The general questions described above are merely examples, and those skilled in the art will appreciate that different general questions can be generated based on different subjects, predicates, and objects. Figure 5 The following figure shows the correspondence between the above example general question and Cypher query. The above example general question is a relatively complex sentence structure, and some language components can be deleted. For example, the question "Find the type of the logical node of the intelligent device named IL1101" corresponds to the Cypher query statement:
[0126] MATCH(A:IED)WHERE A.name="IL1101"MATCH(A)-[*1..10]->(B)RETURNB.lnClass.
[0127] 5. Query step: Use the graph database query statement to query the content of the corresponding SCD file in the graph database.
[0128] Figure 4 Five typical natural language queries were presented, and the system provided correct results for all five. Further analysis of the software's log files revealed that the automatically generated Cypher statements were also correct. Experimental results demonstrate that the prototype software system correctly translates natural language into Cypher statements and outputs correct query results.
[0129] The natural language query method of the present invention can greatly reduce the difficulty for novice engineers / designers to query SCD files, and provide more natural and friendly support for exploring and learning the structure of SCD files.
[0130] The present invention also provides a system for querying the SCD file of a smart substation using natural language, comprising:
[0131] Import module: import SCD files into the graph database.
[0132] Information supplement module: obtains the natural language query statement input into the graph database, modifies it through the knowledge graph to obtain the suggested question sentence, supplements the omitted attributes in the suggested question sentence, and replaces the synonyms of professional terms with standard reference words.
[0133] Information extraction module: extracts semantic information from the modified natural language query statement and represents the obtained semantic information as semantic triples.
[0134] Conversion module: obtains general questions based on semantic triples, queries corresponding assembly templates from an existing assembly template database based on the general questions, converts semantic triples into Cypher code segments, and uses the assembly templates obtained from the semantic triple query to assemble the Cypher code segments into Cypher query statements for the graph database.
[0135] Query module: uses the Cypher query statement to query the content of the corresponding SCD file in the graph database.
[0136] Those skilled in the art will appreciate that, in addition to implementing the system and its various devices, modules, and units provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same functions of the system and its various devices, modules, and units provided by the present invention in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; the devices, modules, and units for implementing various functions can also be considered as both software modules implementing the method and structures within the hardware component.
[0137] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.
Claims
1. A method for querying the SCD file of a smart substation using natural language, characterized in that: include: Import step: import the SCD file into the graph database; Information supplementation step: obtain the natural language query statement input into the graph database, modify it through the knowledge graph to obtain the suggested question sentence, supplement the omitted attributes in the suggested question sentence, and replace the synonyms of professional terms with standard reference terms; Semantic information extraction step: extracting multi-relational semantic information from the modified natural language query statement, and expressing the obtained multi-relational semantic information as semantic triples; Conversion step: obtaining a general question sentence based on the semantic triple, querying the corresponding assembly template from an existing assembly template database based on the general question sentence, converting the semantic triple into a Cypher code segment, and at the same time, obtaining a corresponding general question sentence based on the relationship between the subject, predicate, and object in the semantic triple, querying the corresponding assembly template from an existing assembly template database based on the general question sentence, and assembling the Cypher code segment into a Cypher query statement for the graph database using the assembly template obtained based on the semantic triple query; Query step: using the Cypher query statement to query the content of the corresponding SCD file in the graph database; The information extraction step comprises: Calculating character vectors of the characters in the modified natural language query statement; Evaluate the contextual features of each character to obtain the linguistic or semantic relationship between characters; Identify all subjects based on contextual features, and identify the objects and predicates associated with each subject to obtain semantic triples; The conversion step comprises: Selecting the assembly template for assembling the Cypher code segment according to a set of semantic triples obtained from the Cypher query statement; After the knowledge graph is modified, the suggested questions include: Find attribute K [1] is the value V [1] and / or ... attribute K [n] is the value V [n] The attribute L of node m related to node n [1] ...and the property L [n] ; The semantic information extraction step uses the BiLSTM model to extract the contextual features h of the standard text t , perform maximum pooling on the hidden state output by the BiLSTM model to obtain the sentence feature g, and finally, the context feature h of the standard text t Combined with the sentence feature g to form the task shared feature [h t ;g]; For each sentence to be processed, all candidate subjects in the standard text are first extracted. Then, based on the semantic information of each candidate subject, the corresponding object and semantic relationship are extracted. The extraction of candidate subjects, objects and semantic relationships are modeled as sequence labeling tasks. The pointer labeling structure segments the text by providing a start position pointer and an end position pointer. After receiving the task-shared features, the subject extractor first uses the BiLSTM model to extract task-specific features from the task-shared features. Then, two multi-head self-attention models are used to learn the starting position features and ending position features of the subject respectively. Select the scaled dot product model as the attention scoring function, record is the task-specific feature output by the BiLSTM layer, then in the self-attention mechanism, the relationship between the query matrix Q, the key matrix K and the value matrix V is Q = K = V = h se , the attention function is expressed as follows: in d represents the hidden state dimension of the BiLSTM layer output, which is equal to 2d h ; Assuming that the multi-head self-attention mechanism contains m heads, the i-th attention head is expressed as follows: in, is the projection parameter to be trained, d k =2d h / m, the final result of the multi-head self-attention mechanism is spliced together by each attention head: in are the parameters to be trained; Afterwards, the output of the multi-head self-attention mechanism It is input into a fully connected layer with a Softmax activation function to generate a label probability distribution on each character. In the process of marking the starting position of the subject, the character c t The label calculation is as follows: in is a training parameter, |T| is the number of output label types, and the value of |T| is the named entity category plus 1; The output h of the first multi-head self-attention mechanism se-sta With task-specific features h se After splicing, initialize the Q, K, and V parameters of the second multi-head self-attention mechanism; The output of the second multi-head self-attention mechanism is recorded as Then in the process of marking the end position of the subject, the character c t The label calculation is as follows: in is the training parameter; After the above process, the subject extractor decomposes the recognition of candidate subjects into two sequence labeling tasks. The first sequence labeling task is responsible for identifying the starting position of the candidate subject. If a character is recognized as the starting character of the candidate subject, the position corresponding to the character will be labeled with the named entity type label of the subject. The second sequence labeling task is responsible for identifying the ending position of the candidate subject. Its labeling process is the same as the starting position labeling process. Finally, cross entropy is used to measure the loss between the probability distribution predicted by the subject extractor and the true distribution. The training loss function of the subject extractor is written as: in, and are the actual starting and ending position labels of the i-th character, respectively, and n is the length of the design specification text; After obtaining all candidate subjects, the object and semantic relationship extractor is responsible for extracting all the objects corresponding to each candidate subject and the semantic relationship between the subject and the object; In the object and semantic relationship extractor, the given subject needs to be semantically encoded first. The fragment of the task-shared features corresponding to the given subject is input into an LSTM model. The last hidden state output by the LSTM model is used as the semantic encoding of the given subject. The semantic encoding of the given subject is concatenated with each feature vector of the task-shared features to obtain a new feature matrix. This new feature matrix is considered to carry both contextual features and the semantic encoding results of the given subject features. The subsequent computational process of the object and semantic relationship extractor includes the following steps: First, a feature matrix carrying subject-specific semantic information is input into a BiLSTM model to extract task-specific features. Second, two sequence labeling tasks are constructed using two multi-head self-attention mechanisms and fully connected layers. The first sequence labeling task is responsible for labeling the starting position of all candidate objects, and the second sequence labeling task is responsible for labeling the ending position of all candidate objects. Unlike the subject extractor, if a character is labeled as the starting position or ending position of a candidate object, its label is not the type of named entity, but the type of semantic relationship between the given subject and the candidate object. For a given subject, the starting position and ending position of the object are marked in the process of character c t The label calculations are as follows: in, are the outputs of the first and second multi-head self-attention mechanisms at timestamp t, respectively; b ope-sta ,as well as are all training parameters; |T'| is the number of output label types, and the value of |T'| is the number of semantic relations plus 1; finally, the cross entropy is used to measure the loss between the probability distribution predicted by the object and semantic relationship extractor and the true distribution. The training loss function of the object and semantic relationship extractor is: in, are the actual starting and ending position labels of the i-th character, respectively, and n is the length of the standard text; During the model training phase, the subject extractor and the object and semantic relationship extractor are jointly trained by sharing the task-shared features provided by the semantic encoder. In each training instance, a subject is randomly selected from the standard dataset of canonical text as the input of the object and semantic relationship extractor. By measuring the difference between the model prediction result and the standard result, the loss function values of the subject extractor and the object and semantic relationship extractor are calculated. Finally, the two losses are added together to form the final loss function of the joint model: The Adam algorithm is used to optimize the final loss function of the model, so that the errors generated in the subject extraction, object and semantic relationship extraction processes affect each other. The error generated by each subtask is constrained by other tasks, thereby strengthening the potential interaction between named entities and semantic relationships.
2. The method for querying the SCD file of a smart substation using natural language according to claim 1, characterized in that: The information supplementation step includes: The knowledge graph is used to add SCL engineering attributes to natural language query statements and replace synonyms of professional terms with standard reference terms.
3. A natural language query system for smart substation SCD files, characterized by: include: Import module: import SCD files into graph database; Information supplementation module: obtains natural language query statements input into the graph database, modifies them through the knowledge graph to obtain suggested questions, supplements omitted attributes in the suggested questions, and replaces synonyms of professional terms with standard reference terms; Semantic information extraction module: extracting semantic information from the modified natural language query statement, and expressing the obtained semantic information as semantic triples; Conversion module: obtains a general question sentence based on the semantic triple, queries the corresponding assembly template from the existing assembly template database based on the general question, converts the semantic triple into a Cypher code segment, and simultaneously obtains a corresponding general question sentence based on the relationship between the subject, predicate, and object in the semantic triple, queries the corresponding assembly template from the existing assembly template database based on the general question, and assembles the Cypher code segment into a Cypher query statement for the graph database using the assembly template obtained from the semantic triple query; Query module: uses the Cypher query statement to query the content of the corresponding SCD file in the graph database; The information extraction module includes: Calculating character vectors of the characters in the modified natural language query statement; Evaluate the contextual features of each character to obtain the linguistic or semantic relationship between characters; Identify all subjects based on contextual features, and identify the objects and predicates associated with each subject to obtain semantic triples; The conversion module comprises: Selecting the assembly template for assembling the Cypher code segment according to a set of semantic triples obtained from the Cypher query statement; After the knowledge graph is modified, the suggested questions include: Find attribute K [1] is the value V [1] and / or ... attribute K [n] is the value V [n] The attribute L of node m related to node n [1] ...and the property L [n] ; The semantic information extraction module uses the BiLSTM model to extract the contextual features h of the standard text t , perform maximum pooling on the hidden state output by the BiLSTM model to obtain the sentence feature g, and finally, the context feature h of the standard text t Combined with the sentence feature g to form the task shared feature [h t ;g]; For each sentence to be processed, all candidate subjects in the standard text are first extracted. Then, based on the semantic information of each candidate subject, the corresponding object and semantic relationship are extracted. The extraction of candidate subjects, objects and semantic relationships are modeled as sequence labeling tasks. The pointer labeling structure segments the text by providing a start position pointer and an end position pointer. After receiving the task-shared features, the subject extractor first uses the BiLSTM model to extract task-specific features from the task-shared features. Then, two multi-head self-attention models are used to learn the starting position features and ending position features of the subject respectively. Select the scaled dot product model as the attention scoring function, record is the task-specific feature output by the BiLSTM layer, then in the self-attention mechanism, the relationship between the query matrix Q, the key matrix K and the value matrix V is Q = K = V = h se , the attention function is expressed as follows: in d represents the hidden state dimension of the BiLSTM layer output, which is equal to 2d h ; Assuming that the multi-head self-attention mechanism contains m heads, the i-th attention head is expressed as follows: in, is the projection parameter to be trained, d k =2d h / m, the final result of the multi-head self-attention mechanism is spliced together by each attention head: in are the parameters to be trained; Afterwards, the output of the multi-head self-attention mechanism It is input into a fully connected layer with a Softmax activation function to generate a label probability distribution on each character. In the process of marking the starting position of the subject, the character c t The label calculation is as follows: in is a training parameter, |T| is the number of output label types, and the value of |T| is the named entity category plus 1; The output h of the first multi-head self-attention mechanism se-sta With task-specific features h se After splicing, initialize the Q, K, and V parameters of the second multi-head self-attention mechanism; The output of the second multi-head self-attention mechanism is recorded as Then in the process of marking the end position of the subject, the character c t The label calculation is as follows: in is the training parameter; After the above process, the subject extractor decomposes the recognition of candidate subjects into two sequence labeling tasks. The first sequence labeling task is responsible for identifying the starting position of the candidate subject. If a character is recognized as the starting character of the candidate subject, the position corresponding to the character will be labeled with the named entity type label of the subject. The second sequence labeling task is responsible for identifying the ending position of the candidate subject. Its labeling process is the same as the starting position labeling process. Finally, cross entropy is used to measure the loss between the probability distribution predicted by the subject extractor and the true distribution. The training loss function of the subject extractor is written as: in, and are the actual starting and ending position labels of the i-th character, respectively, and n is the length of the design specification text; After obtaining all candidate subjects, the object and semantic relationship extractor is responsible for extracting all the objects corresponding to each candidate subject and the semantic relationship between the subject and the object; In the object and semantic relationship extractor, the given subject needs to be semantically encoded first. The fragment of the task-shared features corresponding to the given subject is input into an LSTM model. The last hidden state output by the LSTM model is used as the semantic encoding of the given subject. The semantic encoding of the given subject is concatenated with each feature vector of the task-shared features to obtain a new feature matrix. This new feature matrix is considered to carry both contextual features and the semantic encoding results of the given subject features. The subsequent computational process of the object and semantic relationship extractor includes the following steps: First, a feature matrix carrying subject-specific semantic information is input into a BiLSTM model to extract task-specific features. Second, two sequence labeling tasks are constructed using two multi-head self-attention mechanisms and fully connected layers. The first sequence labeling task is responsible for labeling the starting position of all candidate objects, and the second sequence labeling task is responsible for labeling the ending position of all candidate objects. Unlike the subject extractor, if a character is labeled as the starting position or ending position of a candidate object, its label is not the type of named entity, but the type of semantic relationship between the given subject and the candidate object. For a given subject, the starting position and ending position of the object are marked in the process of character c t The label calculations are as follows: in, are the outputs of the first and second multi-head self-attention mechanisms at timestamp t, respectively; b ope-sta ,as well as are all training parameters; |T'| is the number of output label types, and the value of |T'| is the number of semantic relations plus 1; finally, the cross entropy is used to measure the loss between the probability distribution predicted by the object and semantic relationship extractor and the true distribution. The training loss function of the object and semantic relationship extractor is: in, are the actual starting and ending position labels of the i-th character, respectively, and n is the length of the standard text; During the model training phase, the subject extractor and the object and semantic relationship extractor are jointly trained by sharing the task-shared features provided by the semantic encoder. In each training instance, a subject is randomly selected from the standard dataset of canonical text as the input of the object and semantic relationship extractor. By measuring the difference between the model prediction result and the standard result, the loss function values of the subject extractor and the object and semantic relationship extractor are calculated. Finally, the two losses are added together to form the final loss function of the joint model: The Adam algorithm is used to optimize the final loss function of the model, so that the errors generated in the subject extraction, object and semantic relationship extraction processes affect each other. The error generated by each subtask is constrained by other tasks, thereby strengthening the potential interaction between named entities and semantic relationships.
4. The system for natural language query of smart substation SCD files according to claim 3 is characterized in that: The information supplement module includes: The knowledge graph is used to add SCL engineering attributes to natural language query statements and replace synonyms of professional terms with standard reference terms.
Citation Information
Patent Citations
Intelligent question-answering method and system based on power grid field scheduling scene knowledge graph
CN112527997A
Query statement generation method, device and system and computer readable storage medium
CN112989145A