A power standard knowledge graph construction method, a knowledge question and answer system and device
Patent Information
- Application Number
- CN202211320954.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-26
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-10-26
AI Technical Summary
究其原因,是由于构建电力标准知识图谱过程中,相关知识抽取难度较高,导致知识图谱构建困难
[0060]本发明有益效果为:针对文本信息的知识抽取通过设计的模型实现电力标准知识中实体和属性的联合抽取,不仅可以保证知识抽取的可靠性,还能够保证抽取的效率;针对图像信息的知识抽取通过设计的模型有效克服了电力标准知识抽取困难(由于存在公式图像用于表征数值限定、计算方式等相关信息的数据,现有技术无法实现有效的知识抽取)的问题,不仅能够保证公式图像中相关知识的抽取,还可以保证此类知识抽取的可靠性。并且,针对公式文本采用设计的WordBert子模型,不涉及分词操作,不仅可以减少处理过程,还能够有效保留信息,避免传统Bert模型中分词导致的公式信息提取错误的问题。而将抽取实体和属性的向量序列处理后再输入至关系抽取子模型,实现对实体间关系的抽取,可以利用Bert子模型已经处理得到的向量序列来进行后续的关系处理,且能够进行相应处理后再进行关系的抽取,不仅可以有效减少知识抽取的工作量(因为不需要再进行重复的实体抽取过程),还由于已经确定了实体,在关系抽取的过程中能够事半功倍。
Smart Images

Figure CN115934955B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power technology, and in particular to a method for constructing a power standard knowledge graph, a knowledge question-and-answer system, and a device. Background Technology
[0002] Knowledge graphs are a modern theory that integrates theories and methods from applied mathematics, computer graphics, information visualization, and information science with bibliometric methods such as citation analysis and co-occurrence analysis. They utilize visualized graphs to vividly display the core structure, development history, cutting-edge fields, and overall knowledge architecture of a discipline, achieving multidisciplinary integration. Through data mining, information processing, knowledge measurement, and graphical representation, complex knowledge domains can be displayed, revealing their dynamic development patterns and providing practical and valuable references for disciplinary research.
[0003] There are many applications based on knowledge graphs, such as intelligent question answering, personalized recommendations, knowledge reasoning, and visualization. Knowledge question answering systems, similar to search engines, are information retrieval tools. However, unlike search engines, knowledge question answering systems can understand and process natural language questions at the semantic level and directly return the answers, achieving semantic retrieval. If a knowledge graph is used as the knowledge source for a question answering system, it constitutes a knowledge-based question answering system. This system can accept questions in natural language form, understand the meaning of the questions through semantic analysis, and then query and return the answers from the knowledge base.
[0004] Existing knowledge related to the power industry typically relies on search engines, with no intelligent question-answering systems found in this specific field. This is because extracting relevant knowledge during the construction of a power standards knowledge graph is challenging, making the graph construction difficult. Summary of the Invention
[0005] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.
[0006] In view of the problems existing in the above and / or existing knowledge in the power industry, the present invention is proposed.
[0007] Therefore, the problem that this invention aims to solve is how to extract relevant knowledge during the construction of a power standard knowledge graph.
[0008] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0009] In a first aspect, embodiments of the present invention provide a method for constructing a power standard knowledge graph, which includes,
[0010] By collecting power standard data, an ontology structure for a power standard knowledge graph is constructed, which includes entities, attributes, and relationships between entities.
[0011] Acquire basic data containing knowledge of power standards, and extract knowledge from the basic data to extract entities, attributes, and relationships between entities;
[0012] Based on the extracted knowledge, knowledge is integrated and stored to construct a power standard knowledge graph.
[0013] As a preferred embodiment of the power standard knowledge graph construction method of the present invention, the step of obtaining basic data containing power standard knowledge and extracting knowledge from the basic data includes,
[0014] The basic data is preprocessed to obtain multiple text information, or multiple text information and at least one image information;
[0015] For each piece of text information, the text information is segmented and input into the Bert sub-model to obtain the corresponding vector sequence. Then, the vector sequence is input into the BGRU sub-model to output a state matrix that reveals the label scores of each word in the text information. The state matrix is then input into the CRF sub-model to calculate the optimal label sequence, thereby realizing the extraction of entities and attributes.
[0016] For each image information, the image information is input into an externally called formula recognition sub-tool to obtain the converted text information. The converted text information is processed to obtain at least one formula text. Each formula text is input into the WordBert sub-model to obtain the corresponding vector sequence. Then, the vector sequence is input into the BGRU sub-model to output a state matrix that reveals the label scores corresponding to each formula text in the converted text information. The state matrix is then input into the CRF sub-model to calculate the optimal label sequence and achieve attribute extraction.
[0017] The vector sequences of extracted entities and attributes are processed and then input into the relation extraction sub-model to extract the relations between entities.
[0018] As a preferred embodiment of the power standard knowledge graph construction method of the present invention, the knowledge extraction method for each text information includes:
[0019] After segmenting the text information into words, we obtain a segmented text w of length n;
[0020] The segmented text w = ([CLS], w1, w2, ..., w n The input of [SEP] is fed into the Bert sub-model to obtain the vector sequence l = (l0, l1, l2, ..., l) corresponding to the segmented text w. n ,l n+1 ), l i ∈R n×L Where i∈[0,n+1], and the vector sequence l=(l0,l1,l2,…,l n ,l n+1 ) represents the hidden state corresponding to the segmented text w in the last layer of the Bert sub-model, [CLS] is the start symbol, [SEP] is the end symbol, and L is the dimension of the hidden state of the Bert sub-model;
[0021] Let the vector sequence l = (l0, l1, l2, ..., l n ,l n+1 Each word vector sequence l in ) i As input to each time step in the BGRU sub-model;
[0022] The hidden state sequence output by the forward GRU in the BGRU sub-model and the hidden state sequence of the reverse GRU output Calculations are performed to obtain the hidden state sequence h corresponding to the vector sequence l. n+1 h n+1 ∈R n×H H is the dimension of the hidden state of the BGRU sub-model;
[0023] The hidden state sequence h n+1 Map from H dimensions to k dimensions, where k is the number of labels;
[0024] Calculate the tag score for each word segmentation to be classified into k tags, and obtain the state matrix E = (e0, e1, e2, ..., e n ,e n+1 ), where e i ∈R k is a column vector;
[0025] The state matrix is input into the CRF sub-model to calculate the optimal label sequence.
[0026] As a preferred embodiment of the power standard knowledge graph construction method of the present invention, the step of inputting the state matrix into the CRF sub-model and calculating the optimal label sequence includes,
[0027] Let the state matrix E = (e0, e1, e2, ..., e n ,e n+1 Input into the CRF sub-model;
[0028] Based on the constraint matrix F and the input state matrix E introduced in the CRF sub-model, each label sequence is calculated. Total score:
[0029]
[0030] Where F∈R (k+2)×(k+2) , Represents the label sequence The total score, where α is the adjustment factor. This represents the probability that the i-th word segment in the state matrix E is classified into the j-th tag. Indicates a sequence of labels The probability that the j-th label will be transferred to the (j+1)-th label;
[0031] Based on each label sequence Total score Calculate the optimal label sequence
[0032]
[0033] in, It is the set of all possible label sequences.
[0034] As a preferred embodiment of the power standard knowledge graph construction method of the present invention, the knowledge extraction method for each image information includes:
[0035] The converted text information is identified to determine whether the target symbol "=" exists;
[0036] If the target symbol "=" does not exist, then the converted text information is determined to be a formula text;
[0037] If the target symbol "=" exists, the converted text information will be split using the target symbol "=" to obtain multiple formula texts;
[0038] Combine the formula text v = ([CLS], v1, v2, ..., v m The input of [SEP] is fed into the WordBert sub-model to obtain the vector sequence l = (l0, l1, l2, ..., l) corresponding to the formula text combination v. m ,l m+1 ), l i ∈R m×L Where i∈[0,m+1], and the vector sequence l=(l0,l1,l2,…,l m ,l m+1) represents the hidden state corresponding to the formula text combination v in the last layer of the WordBert sub-model, [CLS] is the start symbol, [SEP] is the end symbol, and L is the dimension of the hidden state of the WordBert sub-model;
[0039] Let the vector sequence l = (l0, l1, l2, ..., l m ,l m+1 The formula vector sequence l in ) i As input to each time step in the BGRU sub-model; the hidden state sequence output by the forward GRU in the BGRU sub-model. and the hidden state sequence of the reverse GRU output Calculations are performed to obtain the hidden state sequence h corresponding to the vector sequence l. m+1 h m+1 ∈R m×H H is the dimension of the hidden states in the BGRU sub-model; the hidden state sequence h m+1 Map from H dimensions to k dimensions, where k is the number of labels; calculate the label score for each formula classified into k labels, resulting in the state matrix E = (e0, e1, e2, ..., e m ,e m+1 ), e i ∈R k is a column vector;
[0040] The state matrix is input into the CRF sub-model to calculate the optimal label sequence.
[0041] As a preferred embodiment of the power standard knowledge graph construction method of the present invention, the step of inputting the state matrix into the CRF sub-model and calculating the optimal label sequence includes,
[0042] Let the state matrix E = (e0, e1, e2, ..., e m ,e m+1 Input into the CRF sub-model;
[0043] Based on the input state matrix E, calculate each label sequence Total score:
[0044]
[0045] in, Represents the label sequence Total score, This represents the probability that the i-th component in the state matrix E is classified into the j-th label;
[0046] Based on each label sequence Total score Calculate the optimal label sequence in, It is the set of all possible label sequences.
[0047] As a preferred embodiment of the power standard knowledge graph construction method of the present invention, the step of processing the vector sequence of extracted entities and attributes and then inputting it into the relation extraction sub-model to extract the relations between entities includes:
[0048] Based on the extracted entities, the vector sequence l = (l0, l1, l2, ..., l) corresponding to the segmented text w is... n ,l n+1 Mark the corresponding vector in );
[0049] The labeled vector sequence l′ is input into the relation extraction sub-model;
[0050] For the labeled vectors in the vector sequence l′, perform binary cross-grouping on all labeled vectors so that each labeled vector has a pairing relationship with other labeled vectors;
[0051] For each pair of labeled vectors that have a combination relationship, the two labeled vectors in the pair are concatenated to obtain the combined vector;
[0052] Calculate the score for each combined vector under each relation category;
[0053] The optimal score corresponding to each combined vector is obtained and sorted. The lowest optimal score is eliminated. For each remaining optimal score, the entity relationships with corresponding relationship categories between the entities corresponding to its combined vector are determined, thereby realizing the extraction of entity relationships.
[0054] Secondly, embodiments of the present invention provide a power standard knowledge graph construction system, which includes:
[0055] The data layer includes a pre-built knowledge graph of power standards, and a word segmentation dictionary built based on entities and attributes in the knowledge graph of power standards;
[0056] The Web layer is used to receive user queries and generate and display answer information based on the query results from the query layer, wherein the query information is in natural language form.
[0057] The query layer is used to convert the query information into a Cypher query statement, send it to the Neo4j graph database for querying, and obtain the query results.
[0058] Thirdly, embodiments of the present invention provide a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any step of the above-described method.
[0059] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements any step of the above-described method.
[0060] The beneficial effects of this invention are as follows: For knowledge extraction from textual information, the designed model achieves joint extraction of entities and attributes from power standard knowledge, ensuring both reliability and efficiency in knowledge extraction. For knowledge extraction from image information, the designed model effectively overcomes the difficulty of extracting power standard knowledge (due to the existence of formula images used to represent data related to numerical constraints, calculation methods, etc., existing technologies cannot achieve effective knowledge extraction), ensuring both the extraction of relevant knowledge from formula images and the reliability of such knowledge extraction. Furthermore, the designed WordBert sub-model for formula text does not involve word segmentation, reducing processing steps and effectively preserving information, avoiding the errors in formula information extraction caused by word segmentation in traditional BERT models. By processing the vector sequences of extracted entities and attributes and then inputting them into the relation extraction sub-model, the relations between entities can be extracted. The vector sequences already processed by the BERT sub-model can be used for subsequent relation processing. This not only effectively reduces the workload of knowledge extraction (because there is no need to repeat the entity extraction process), but also makes the relation extraction process more efficient since the entities have already been determined. Attached Figure Description
[0061] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:
[0062] Figure 1 A flowchart for the method of constructing a knowledge graph for power standards.
[0063] Figure 2 A schematic diagram of a model for building a knowledge graph of power standards.
[0064] Figure 3 This is a schematic diagram of an electrical standards knowledge Q&A system. Detailed Implementation
[0065] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0066] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0067] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0068] This invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of this invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not adhering to the usual scale. Furthermore, the schematic diagrams are merely examples and should not be construed as limiting the scope of protection of this invention. In actual fabrication, the three-dimensional spatial dimensions of length, width, and depth should be included.
[0069] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0070] Unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" in this invention should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; similarly, they can refer to mechanical connections, electrical connections, or direct connections, or indirect connections through an intermediate medium, or internal connections between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0071] Example 1
[0072] Reference Figure 1 and Figure 2This is the first embodiment of the present invention, which provides a method for constructing a power standard knowledge graph, including:
[0073] S100: Construct the ontology structure of the power standard knowledge graph by collecting power standard data. The ontology structure includes entities, attributes and relationships between entities.
[0074] It should be noted that, considering the domain of power standard knowledge graphs, the ontology structure can be constructed using top-down and bottom-up approaches. A portion of the ontology structure can be pre-designed, such as: power standard name (e.g., building lightning protection design code) - indicator (e.g., lightning protection device), indicator (lightning protection device) - lower-level indicator (lightning protection wire), etc., and new ontology structures can be discovered and added during the subsequent knowledge extraction process.
[0075] S200: Obtain basic data containing power standard knowledge, and extract knowledge from the basic data to extract entities, attributes, and relationships between entities.
[0076] It should be noted that obtaining basic data containing knowledge of power standards can be achieved through methods such as collecting documents and crawling web pages. For example, one can crawl information about power standards from web pages, or obtain it from a constructed dataset (since power standards knowledge belongs to a very vertical field and the knowledge is relatively stable). The basic data containing power standards knowledge can be plain text data (such as Word documents, PDF documents, TXT documents, etc.), or a combination of text data and formulas / images (such as PDF documents containing formulas, Word documents containing formula images, etc.). The basic data can also be documents obtained by processing and organizing data crawled from web pages.
[0077] It should be noted that for the text data in the basic data, the text data in the basic data can be split into multiple text information based on sentence delimiters.
[0078] It should be noted that the acquired basic data can be preprocessed to obtain multiple text information, or multiple text information and at least one image information.
[0079] Furthermore, if the underlying data contains formula images, then for each formula image in the underlying data, the image can be processed to obtain the corresponding image information. For example, the formula image can be input into Mathpix to obtain the output formula. The output LaTeX format can be converted to TeX, and then MathType can be used to convert the LaTeX to MathML format, which is a plain text format, and can be used to obtain a Word document.
[0080] Furthermore, for each image information, a unique identifier can be assigned. Similarly, the same identifier can be assigned to all text information corresponding to the paragraph containing the formula image and its adjacent paragraphs, establishing a relationship between image and text information. This method allows for the establishment of associations between text and image information, facilitating the subsequent determination of the entity to which an attribute belongs, thus ensuring the accuracy and reliability of the knowledge graph.
[0081] It should be noted that, in order to achieve knowledge extraction from text information (joint extraction of entities and attributes), for each piece of text information, the text information can be segmented and input into the Bert sub-model to obtain the corresponding vector sequence.
[0082] Specifically, after segmenting the text information into words, a segmented text w of length n is obtained. Then, the segmented text w = ([CLS], w1, w2, ..., w n The input of [SEP] is fed into the Bert sub-model to obtain the vector sequence l = (l0, l1, l2, ..., l) corresponding to the segmented text w. n ,l n+1 ), l i ∈R n×L ;
[0083] Where i∈[0,n+1], the vector sequence l=(l0,l1,l2,…,l n ,l n+1 ) represents the hidden state corresponding to the segmented text w in the last layer of the Bert sub-model, [CLS] is the start symbol, [SEP] is the end symbol, and L is the dimension of the hidden state of the Bert sub-model (e.g., 100-dimensional, 200-dimensional, etc.).
[0084] Furthermore, after obtaining the vector sequence l output by the Bert sub-model, it can be input into the BGRU sub-model, which then outputs a state matrix that reveals the corresponding label scores for each word in the text information.
[0085] Specifically, the vector sequence l = (l0, l1, l2, ..., l) n ,l n+1 Each word vector sequence l in ) i These are used as inputs to each time step (requiring n+2 time steps) in the BGRU sub-model, and then the hidden state sequence output by the forward GRU in the BGRU sub-model is used. and the hidden state sequence of the reverse GRU output Calculations are performed to obtain the hidden state sequence h corresponding to the vector sequence l. n+1 h n+1 ∈R n×H , where H is the dimension of the hidden state of the BGRU sub-model.
[0086] It should be noted that the hidden state sequence output by the forward GRU The hidden state sequence of the reverse GRU output The hidden state sequence h is obtained by summing the bits and then averaging them (to further improve accuracy, a weighted summation can also be used). n+1 .
[0087] Furthermore, the hidden state sequence h n+1 Mapping from H dimensions to k dimensions, where k is the number of tags, and then calculating the tag score for each word segmentation to k tags, yields the state matrix E = (e0, e1, e2, ..., e n ,e n+1 ), where e i ∈R k e i It is a column vector;
[0088] The state matrix is input into the CRF sub-model to calculate the optimal label sequence, thereby enabling the extraction of entities and attributes.
[0089] Specifically, the state matrix E = (e0, e1, e2, ..., e n ,e n+1 The input is fed into the CRF sub-model, based on the constraint matrix F introduced in the CRF sub-model and the input state matrix E, where F∈R (k+2)×(k+2) Calculate each tag sequence using the following formula. Total score:
[0090]
[0091] in, Represents the label sequence The total score, where α is the adjustment factor. This represents the probability that the i-th word segment in the state matrix E is classified into the j-th tag. Indicates a sequence of labels The probability of the j-th label being transferred to the (j+1)-th label.
[0092] Then, it can be based on each label sequence Total score Substitute into the following formula to calculate the optimal label sequence
[0093]
[0094] in, It is the set of all possible label sequences.
[0095] In addition, to ensure the applicability of the introduced constraint matrix F, the following loss function can be added to the CRF sub-model. During the training phase, the constraint matrix F is learned by minimizing this loss function.
[0096]
[0097] Where y represents the correct label sequence. It is the set of all possible label sequences.
[0098] The designed model achieves joint extraction of entities and attributes from power standard knowledge, ensuring both reliability and efficiency in knowledge extraction. By employing a BERT+BGRU+CRF model, word segmentation can be performed before processing with the BERT model to achieve joint extraction of entities and attributes, improving the accuracy of entity and attribute extraction and reducing the complexity of model design. Furthermore, the introduced constraint matrix F constrains the state matrix E, preventing the output of invalid label sequences. Additionally, during the calculation of each label sequence... The total score introduces an adjustment factor α, which has stronger applicability in the joint extraction of entities and attributes, ensures the accuracy of entity and attribute extraction, and overcomes the problem caused by the difference in the constraint matrix F required in the entity and attribute extraction process (using the same standard constraint matrix for attributes and entities will lead to high entity extraction accuracy but low attribute extraction accuracy, or low entity extraction accuracy but high attribute extraction accuracy).
[0099] To achieve knowledge extraction (attribute extraction) from image information, for each image, the image information can be input into an externally invoked formula recognition sub-tool to obtain converted text information. The converted text information can then be processed to obtain at least one formula text.
[0100] It should be noted that the converted text information can be identified to determine whether the target symbol "=" exists. If the target symbol "=" does not exist, the converted text information is determined to be a formula text; if the target symbol "=" exists, the converted text information is split into multiple formula texts by the target symbol "=" (if there are 4 target symbols "=", it can be split into 5 formula texts).
[0101] The "=" sign can be used to break down the properties involved in a formula into property identifiers (such as the symbolic representation of the property) and property constraints (such as the numerical constraints of the property, the range of parameter values, etc.), and some even include the intermediate derivation process of the property.
[0102] For each formula text, all formula texts can be input together into the WordBert sub-model to obtain the corresponding vector sequence. Here, each formula text represents the formula text obtained by splitting the text information converted from the same image data.
[0103] Specifically, the formula text combination v = ([CLS], v1, v2, ..., v m The input of [SEP] is fed into the WordBert sub-model to obtain the vector sequence l = (l0, l1, l2, ..., l) corresponding to the formula text combination v. m ,l m+1 ), l i ∈R m×L Where i∈[0,m+1], and the vector sequence l=(l0,l1,l2,…,l m ,l m+1 ) represents the hidden state corresponding to the formula text combination v in the last layer of the WordBert sub-model. [CLS] is the start symbol, [SEP] is the end symbol, and L is the dimension of the hidden state of the WordBert sub-model (e.g., 100, 200, consistent with the dimension of the hidden state of the Bert sub-model).
[0104] The resulting vector sequence can then be input into the BGRU sub-model, which outputs a state matrix that reveals the corresponding label scores for each formula text in the text information of the transformation.
[0105] Specifically, the vector sequence l = (l0, l1, l2, ..., l) m ,l m+1 The formula vector sequence l in ) i The hidden state sequence output by the forward GRU in the BGRU sub-model is then used as the input for each time step in the BGRU sub-model. and the hidden state sequence of the reverse GRU output Calculations are performed to obtain the hidden state sequence h corresponding to the vector sequence l. m+1 h m+1 ∈R m×H H is the dimension of the hidden states in the BGRU sub-model. Then, the hidden state sequence h... m+1 Mapping from H dimensions to k dimensions, where k is the number of labels, calculate the label score for each formula classified into k labels, resulting in the state matrix E = (e0, e1, e2, ..., e m ,e m+1 ), e i ∈R k e i It is a column vector. The process here is similar to the operation of the BGRU sub-model introduced earlier, so it will not be repeated here.
[0106] The state matrix E = (e0, e1, e2, ..., e) is obtained. m ,e m+1 After that, the state matrix can be input into the CRF sub-model to calculate the optimal label sequence.
[0107] Specifically, the state matrix E = (e0, e1, e2, ..., e m ,e m+1 The input is fed into the CRF sub-model; based on the input state matrix E, each label sequence is calculated. Total score:
[0108]
[0109] in, Represents the label sequence Total score, This represents the probability that the i-th component in the state matrix E is classified into the j-th label.
[0110] Based on each label sequence Total score Calculate the optimal label sequence
[0111]
[0112] in, It is the set of all possible label sequences.
[0113] It should be noted that in this embodiment, the vector sequence l = (l0, l1, l2, ..., l) output based on the WordBert sub-model is considered. m ,l m+1 The determined state matrix E = (e0, e1, e2, ..., e) m ,e m+1 The vector sequence l = (l0, l1, l2, ..., l) based on the output of the Bert sub-model was adopted. n ,l n+1 The determined state matrix E = (e0, e1, e2, ..., e) n ,e n+1 Different calculation methods are used to calculate the tag sequence. The overall score is higher because the differentiated calculation method for the state matrices obtained in these two cases yields better results. Of course, the vector sequence l = (l0, l1, l2, ..., l) output by the WordBert sub-model is also more effective. m ,l m+1 The determined state matrix E = (e0, e1, e2, ..., e) m ,e m+1 ), in calculating the label sequence When calculating the total score, the calculation method of formula (1) can also be used, because the method of formula (1) also considers the difference between entities and attributes (especially text information formulas) and introduces the adjustment factor α. However, relatively speaking, when used only for attribute extraction, the effect of formula (1) is slightly inferior to that of formula (4). However, the method of formula (1) is better than not performing differentiation processing and labeling entities and attributes. The performance will be much better when the total score is calculated.
[0114] Therefore, it is possible to extract attributes based on formula images.
[0115] This approach, through a designed model, effectively extracts relevant knowledge (all attributes) from formula images within power standard knowledge. It overcomes the difficulty of extracting power standard knowledge (current technologies cannot effectively extract knowledge from formula images used to represent numerical constraints, calculation methods, and other related information). This not only ensures the extraction of relevant knowledge from formula images but also guarantees the reliability of such knowledge extraction. Furthermore, it eliminates the need for word segmentation, and the WordBert sub-model does not require training on segmented text but rather on complete sentences (especially formulas, characters, operators, etc.), significantly improving the accuracy of formula attribute extraction. The designed WordBert sub-model for formula text, without word segmentation, reduces processing steps and effectively preserves information, avoiding errors in formula information extraction caused by word segmentation in traditional BERT models.
[0116] Furthermore, after extracting entities and attributes, the extracted entity and attribute vector sequences can be processed and then input into the relation extraction sub-model to extract the relationships between entities.
[0117] By processing the vector sequences of extracted entities and attributes and then inputting them into the relation extraction sub-model, the relations between entities can be extracted. The vector sequences already processed by the BERT sub-model can be used for subsequent relation processing. This not only effectively reduces the workload of knowledge extraction (because there is no need to repeat the entity extraction process), but also makes the relation extraction process more efficient since the entities have already been determined.
[0118] Specifically, based on the extracted entities, the vector sequence l = (l0, l1, l2, ..., l) corresponding to the segmented text w is... n ,l n+1 The corresponding vector in ) is marked.
[0119] For example, the vector sequence l = (l0, l1, l2, ..., l) corresponding to the segmented text w.n ,l n+1 In the vector sequence l = (l0, l1, l2, ..., l5), the tokens corresponding to l1, l3, and l5 are extracted as entities. n ,l n+1 The corresponding vectors in the sequence are labeled to obtain the labeled vector sequence l′=(l0,l′1,l2,l′3,l4,l′5,…,l n ,l n+1 ).
[0120] The labeled vector sequence l′ can then be input into the relation extraction sub-model. This relation extraction sub-model is a BERT-based relation extraction model.
[0121] For the labeled vectors in the vector sequence l′, perform binary cross-grouping on all labeled vectors so that each labeled vector has a paired combination relationship with other labeled vectors, and represent them as labeled vector pairs.
[0122] Following the previous example, regarding the labeled vector sequence l′=(l0,l′1,l2,l′3,l4,l′5,…,l n ,l n+1 ), perform binary cross-grouping on all the label vectors (l′1,l′3,l′5) to obtain three types of label vector pairs: (l′1,l′3), (l′1,l′5) and (l′3,l′5).
[0123] For each pair of labeled vectors that have a combinatorial relationship, the two labeled vectors of the pair are concatenated to obtain the combined vector. The concatenation method can be as follows: the two labeled vectors of the pair are concatenated end-to-end to obtain the corresponding combined vector. For example, the labeled vector pair (l′1, l′3) is concatenated to obtain the combined vector l. c1 The combined vector l is obtained by concatenating the marked vector pairs (l′1, l′5). c2 The combined vector l is obtained by concatenating the marked vector pair (l′3, l′5). c3 .
[0124] Then we can calculate the score of each combined vector under each relation category, and obtain the corresponding score vector. Among them, P x Let q represent the score vector formed by the scores of the x-th combination vector under each relation category, where q is the number of relation categories.
[0125] Then it can be based on the vector sequence p xOnce the optimal score is determined, the relationship category corresponding to the optimal score represents the relationship category of the combined vectors. Then, the optimal scores are sorted, and the lowest-ranked optimal score is removed.
[0126] For each remaining optimal score, the entity relationships with corresponding relationship categories between the entities corresponding to its combined vector can be determined, thus enabling the extraction of entity relationships. This allows for the rapid, efficient, and accurate extraction of entity relationships through multi-entity binary mutual combination, taking into account the relationships between individual entities.
[0127] The correspondence between attributes and entities can be determined during the joint extraction of entities and attributes; or after the entities and attributes are identified, the attributes and entities can be classified; or the classification between entities and attributes can be extracted from the webpage using a wrapper (for example, by entering the URL, using a tool to crawl the webpage, using a wrapper to extract the corresponding attributes of the entities provided by the webpage, and then classifying the extracted attributes).
[0128] It should be noted that, regarding the data sources for power standard knowledge, for each document (especially those whose content belongs to normative documents, such as: Overvoltage Protection Design Code for Industrial and Civil Power Installations, Grounding Design Code for Industrial and Civil Power Installations, Lightning Protection Design Code for Buildings, Design Code for Power Installations in Explosion and Fire Hazardous Locations, etc.), the title can be extracted separately to extract a basic entity object, and key attributes such as compilation time, application scenario, and publishing unit can be extracted as important factors in subsequent applications of power standard knowledge graphs such as intelligent question answering and personalized recommendations.
[0129] S300: Based on the extracted knowledge, knowledge fusion is performed, and the fused knowledge is stored in the Neo4j graph database to construct the power standard knowledge graph.
[0130] It should be noted that there are many ways to integrate knowledge, but the main ones require entity alignment and entity disambiguation. For example, the Jaccard algorithm based on string similarity can be used to achieve entity alignment and entity disambiguation.
[0131] Furthermore, a strategy of extracting and storing simultaneously can be adopted: the results of knowledge extraction are temporarily stored in memory as JSON data, and then submitted to the Neo4j graph database for persistent storage through Python's py2neo library.
[0132] In this way, a knowledge graph of power standards can be constructed.
[0133] Furthermore, this embodiment also provides an electricity standard knowledge question-and-answer system, including:
[0134] The data layer includes a pre-built knowledge graph of power standards, and a word segmentation dictionary built based on entities and attributes in the knowledge graph of power standards;
[0135] The Web layer is used to receive user queries and generate and display answer information based on the query results of the Web layer, wherein the query information is in natural language form;
[0136] The query layer is used to convert the query information into a Cypher query statement, send it to the Neo4j graph database for querying, and obtain the query results.
[0137] This embodiment also provides a computer device applicable to the power standard knowledge graph construction method, including:
[0138] The system includes a memory and a processor. The memory stores computer-executable instructions, and the processor executes these instructions to implement the power distribution area household-transformer relationship identification method proposed in the above embodiments.
[0139] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0140] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the method for constructing a power standard knowledge graph as proposed in the above embodiments.
[0141] The storage medium proposed in this embodiment and the data storage method proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0142] Example 2
[0143] Reference Figures 1-3This is the second embodiment of the present invention, which provides a method for constructing a knowledge graph of power standards. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through experiments.
[0144] In this embodiment, a dataset was constructed using 984 labeled base data points. This dataset was divided into a training set (689 data points), a validation set (197 data points), and a test set (98 data points) in a 7:2:1 ratio. The model was then trained, validated, and tested. Precision, recall, and F1 score were used as evaluation metrics to verify the model's performance.
[0145] (1) Precision P represents the accuracy of the model's predictions, and the calculation formula is as follows:
[0146]
[0147] Where M represents the sample set that the model predicts to be positive, and T represents the sample set that is actually positive.
[0148] (2) Recall R represents the comprehensiveness of the model's predictions, and is calculated using the following formula:
[0149]
[0150] (3) The F1 value is a combination of precision P and recall R, and the calculation formula is as follows:
[0151]
[0152] Based on the model's performance validation, the relevant evaluation data are: precision P≈0.84, recall R≈0.90, and F1≈0.87. It is evident that the model performs very well and is effective in extracting knowledge from power standards.
[0153] Based on the constructed power standard knowledge graph, a power standard knowledge question-and-answer system can be further built.
[0154] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for constructing a knowledge graph of power standards, characterized in that: include, By collecting power standard data, an ontology structure for a power standard knowledge graph is constructed, which includes entities, attributes, and relationships between entities. Acquire basic data containing knowledge of power standards, and extract knowledge from the basic data to extract entities, attributes, and relationships between entities; The process of acquiring basic data containing electricity standard knowledge and extracting knowledge from the basic data includes, The basic data is preprocessed to obtain multiple text information, or multiple text information and at least one image information; For each piece of text information, the text information is segmented and input into the Bert sub-model to obtain the corresponding vector sequence. Then, the vector sequence is input into the BGRU sub-model to output a state matrix that reveals the label scores of each word in the text information. The state matrix is then input into the CRF sub-model to calculate the optimal label sequence, thereby realizing the extraction of entities and attributes. For each image information, the image information is input into an externally called formula recognition sub-tool to obtain the converted text information. The converted text information is processed to obtain at least one formula text. Each formula text is input into the WordBert sub-model to obtain the corresponding vector sequence. Then, the vector sequence is input into the BGRU sub-model to output a state matrix that reveals the label scores corresponding to each formula text in the converted text information. The state matrix is then input into the CRF sub-model to calculate the optimal label sequence and achieve attribute extraction. The vector sequences of extracted entities and attributes are processed and then input into the relation extraction sub-model to extract the relations between entities. Based on the extracted knowledge, knowledge is integrated and stored to construct a power standard knowledge graph.
2. The method for constructing a power standard knowledge graph as described in claim 1, characterized in that: The knowledge extraction method for each text information includes: After segmenting the text information into words, the length is obtained as follows: Word segmentation text ; Word segmentation text Inputting the data into the Bert sub-model yields the segmented text. The corresponding vector sequence Vector sequence For the word segmentation text in the last layer of the Bert sub-model The corresponding hidden state, As the start symbol, This is the end-of-line character; vector sequence Each word vector sequence is used as the input to each time step in the BGRU sub-model; The hidden state sequences output by the forward GRU and the backward GRU in the BGRU sub-model are calculated to obtain the vector sequence l and the corresponding hidden state sequence. Let be the dimension of the hidden state of the BGRU sub-model; The hidden state sequence from Dimension mapping to dimension, For the number of tags; Calculate the classification of each word into The state matrix is obtained by calculating the label scores of each label. ,in is a column vector; The state matrix is input into the CRF sub-model to calculate the optimal label sequence.
3. The method for constructing a power standard knowledge graph as described in claim 2, characterized in that: The step of inputting the state matrix into the CRF sub-model and calculating the optimal label sequence includes: The state matrix Input into the CRF sub-model; Constraint matrix introduced in the CRF sub-model and the input state matrix Calculate each label sequence Total score; Based on each label sequence Total score Calculate the optimal label sequence : , in, It is the set of all possible label sequences.
4. The method for constructing a power standard knowledge graph as described in claim 3, characterized in that: The knowledge extraction method for each image information includes: The converted text information is identified to determine whether the target symbol "=" is present. If the target symbol "=" does not exist, then the converted text information is determined to be a formula text; If the target symbol "=" exists, the converted text information will be split using the target symbol "=" to obtain multiple formula texts; Combine formula text Inputting into the WordBert sub-model yields a combination of formula text. The corresponding vector sequence Vector sequence The final layer of the WordBert submodel contains the formula text combination. The corresponding hidden state, As the start symbol, The terminator is ; vector sequence The vector sequences of each formula in the model are used as inputs to each time step in the BGRU sub-model. The hidden state sequences output by the forward GRU and the backward GRU in the BGRU sub-model are calculated to obtain vector sequence l, and the corresponding hidden state sequence. The hidden state dimension of the BGRU sub-model; the hidden state sequence is derived from... Dimension mapping to dimension, For the number of tags; calculate the category of each formula. The state matrix is obtained by calculating the label scores of each label. , is a column vector; The state matrix is input into the CRF sub-model to calculate the optimal label sequence.
5. The method for constructing a power standard knowledge graph as described in claim 4, characterized in that: The step of inputting the state matrix into the CRF sub-model and calculating the optimal label sequence includes: The state matrix Input into the CRF sub-model; Based on the input state matrix Calculate each label sequence Total score; Based on each label sequence Total score Calculate the optimal label sequence ,in, It is the set of all possible label sequences.
6. The method for constructing a power standard knowledge graph as described in claim 5, characterized in that: The process of processing the vector sequences of extracted entities and attributes and then inputting them into the relation extraction sub-model to extract the relationships between entities includes... Based on the extracted entities, the segmented text is processed. The corresponding vector sequence The corresponding vectors in the data are marked; The labeled vector sequence Input into the relation extraction sub-model; For vector sequences The labeled vectors are used to perform binary cross-grouping on all labeled vectors so that each labeled vector has a pairing relationship with other labeled vectors. For each pair of labeled vectors that have a combination relationship, the two labeled vectors in the pair are concatenated to obtain the combined vector; Calculate the score for each combined vector under each relation category; The optimal score corresponding to each combined vector is obtained and sorted. The lowest optimal score is eliminated. For each remaining optimal score, the entity relationships with corresponding relationship categories between the entities corresponding to its combined vector are determined, thereby realizing the extraction of entity relationships.
7. A power standard knowledge question-and-answer system, based on the power standard knowledge graph construction method described in claims 1-6, characterized in that: include, The data layer includes a pre-built knowledge graph of power standards, and a word segmentation dictionary built based on entities and attributes in the knowledge graph of power standards; The Web layer is used to receive user queries and generate and display answer information based on the query results from the query layer, wherein the query information is in natural language form. The query layer is used to convert the query information into a Cypher query statement, send it to the Neo4j graph database for querying, and obtain the query results.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-6.
Citation Information
Patent Citations
Electric power standard knowledge graph construction method and device, computer equipment and medium
CN112231418A