Knowledge graph construction method and device of electric power facility, equipment and storage medium

By building a knowledge graph of power facilities and using digital twin models for entity identification and information integration, the problem of dispersed power facilities is solved, and management efficiency and information retrieval are improved.

CN120106202APending Publication Date: 2025-06-06BEIJING CHINA POWER INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510003151.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The information of power facilities is scattered and difficult to integrate, resulting in low management efficiency.

Method used

By obtaining structured and unstructured data of power facilities, a digital twin model is built, entity recognition is carried out, and the target knowledge graph is obtained through completion.

Benefits of technology

The systematized integration of power facility information has been realized, the efficiency and accuracy of information retrieval has been improved, and the management and application capabilities of power facilities have been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106202A_ABST
    Figure CN120106202A_ABST
Patent Text Reader

Abstract

The invention provides an electric power facility knowledge graph construction method and device, equipment and a storage medium. The method comprises the following steps: acquiring structured data and unstructured data of an electric power facility; preprocessing the structured data and the unstructured data to obtain electric power facility data, and constructing a digital twinborn model of the electric power facility based on the electric power facility data; performing entity identification on the power facility data by using the digital twin model to obtain an entity classification result; and constructing an initial knowledge graph based on the entity classification result, and complementing the initial knowledge graph to obtain a target knowledge graph. In this way, the context information and the local features of the power facility data can be identified by using the digital twin model, so that the obtained entity classification result is more accurate. The target knowledge graph is obtained by complementing the initial knowledge graph, so that a relatively systematic and comprehensive electric power facility knowledge graph is obtained, and the performance of the target knowledge graph is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a method, device, equipment and storage medium for constructing a knowledge graph of electric power facilities. Background Art

[0002] With the development of the power industry, the data generated by power facilities has skyrocketed. However, the current data management methods often lead to scattered power facility information and difficulty in integration, resulting in low efficiency in power facility management.

[0003] In view of this, how to improve the efficiency of power facility management by managing power facility information has become a technical problem that needs to be solved urgently. Summary of the invention

[0004] In view of this, the purpose of the present disclosure is to propose a knowledge graph construction method, device, equipment and storage medium for power facilities to solve or partially solve the above-mentioned technical problems.

[0005] Based on the above objectives, the first aspect of the present disclosure proposes a method for constructing a knowledge graph of electric power facilities, the method comprising:

[0006] Obtain structured and unstructured data of power facilities;

[0007] Preprocessing the structured data and the unstructured data to obtain power facility data, and constructing a digital twin model of the power facility based on the power facility data;

[0008] Using the digital twin model to perform entity recognition on the power facility data to obtain entity classification results;

[0009] An initial knowledge graph is constructed based on the entity classification results, and the initial knowledge graph is completed to obtain a target knowledge graph.

[0010] Based on the same inventive concept, the second aspect of the present disclosure proposes a knowledge graph construction device for electric power facilities, comprising:

[0011] An acquisition module configured to acquire structured data and unstructured data of the electric power facility;

[0012] A model building module is configured to preprocess the structured data and the unstructured data to obtain power facility data, and build a digital twin model of the power facility based on the power facility data;

[0013] An entity recognition module is configured to perform entity recognition on the power facility data using the digital twin model to obtain an entity classification result;

[0014] The knowledge graph construction module is configured to construct an initial knowledge graph based on the entity classification results, and complete the initial knowledge graph to obtain a target knowledge graph.

[0015] Based on the same inventive concept, the third aspect of the present disclosure proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the method as described above when executing the computer program.

[0016] Based on the same inventive concept, a fourth aspect of the present disclosure proposes a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method as described above.

[0017] From the above, it can be seen that the knowledge graph construction method, device, equipment and storage medium for power facilities provided by the present disclosure. Obtain structured data and unstructured data of the power facility. Preprocess the structured data and unstructured data to obtain power facility data, and build a digital twin model of the power facility based on the power facility data. Use the digital twin model to perform entity recognition on the power facility data to obtain entity classification results. Construct an initial knowledge graph based on the entity classification results, and complete the initial knowledge graph to obtain a target knowledge graph. In this way, the digital twin model can be used to identify the contextual information and local features of the power facility data, so that the entity classification results are more accurate. By constructing an initial knowledge graph, power facility applications can be supported. By completing the initial knowledge graph to obtain a target knowledge graph, a relatively systematic and comprehensive knowledge graph of power facilities is obtained, thereby improving the performance of the target knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are only embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0019] Figure 1 A flowchart of a method for constructing a knowledge graph of electric power facilities according to an embodiment of the present disclosure;

[0020] Figure 2 An overall flow chart for constructing a knowledge graph of electric power facilities according to an embodiment of the present disclosure;

[0021] Figure 3 It is a schematic diagram of the structure of the BiLSTM feature extraction module of the embodiment of the present disclosure;

[0022] Figure 4 A flowchart for constructing a knowledge graph for an embodiment of the present disclosure;

[0023] Figure 5 A schematic diagram of the structure of the distribution network knowledge graph according to an embodiment of the present disclosure;

[0024] Figure 6 A schematic diagram of the structure of a knowledge graph construction device for electric power facilities according to an embodiment of the present disclosure;

[0025] Figure 7 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0026] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.

[0027] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0028] Based on the description of background technology, it is a challenging task to retrieve digital twin power facilities using knowledge graphs. First, a digital grid based on digital twins was proposed, a closed-loop grid data empowerment system was introduced, a digital grid based on digital twins was defined, and three characteristics and attributes of digital twins were elaborated, enriching the connotation of digital grids. Later, a new digital twin framework was proposed to provide high-fidelity power system simulation. The research conclusions show that the digital twin framework significantly improves prediction accuracy, speeds up analysis, improves system reliability, and contributes to effective risk management of power systems. The researchers also proposed a method to establish a hybrid digital twin model for electrical components of smart grid end users. This method simulates the component specifications of the reference smart grid system. This setting improves energy management, energy efficiency, and the use of renewable energy.

[0029] The researchers proposed a top-down and bottom-up combined knowledge graph construction method for power grid fault handling, which condensed a large amount of unstructured text content in the power grid dispatching operation link into an expressible, operable, and reasonable structured knowledge network, effectively improving the power grid's emergency handling capabilities and dispatching intelligence level. The researchers also proposed a bolt defect classification method based on knowledge graph and feature fusion, adaptively combining the knowledge graph of bolts and nuts with the normalization of the original regional features, and then obtaining the classification score vector based on the fusion features and the bolt-nut pair from the classifier, and then fused these score vectors with the score vector of the original regional features, and obtained good experimental results. The researchers also proposed a logistics resource allocation method based on dynamic spatiotemporal knowledge graph, using deep neural networks to analyze the signal data of large-scale IoT devices, establishing a spatiotemporal consistent digital twin model DSTKG, and performing relationship reasoning and completion based on PL task information. The efficient allocation of logistics resources is achieved through graph algorithms, and the effectiveness of the method is verified by examples.

[0030] With the continuous development and maturity of technologies such as the Internet of Things, artificial intelligence, and deep learning, the data generated by intelligent devices used in power grid operation, control, and monitoring has skyrocketed in a geometric form. The data generated by virtual devices in the construction of digital twin power grids has also increased synchronously. Therefore, it is necessary to use intelligent means to convert a large amount of heterogeneous data in the digital twin power model into knowledge, so as to assist relevant personnel in making decisions quickly and clear the fault when a power grid fault occurs. As a semantic network, knowledge graph can model entities, concepts, attributes, and their relationships in the real world. It has strong expression ability and modeling flexibility, and can be widely used in simple query services such as equipment query, line query, and fault query, as well as intelligent question and answer and knowledge reasoning services. Digital twin power model retrieval based on knowledge graph can solve the problems of fragmented facility information, low retrieval accuracy, and complex management in the current power facility management. Traditional management methods often lead to scattered facility information that is difficult to integrate, making it difficult for decision makers to grasp the overall situation of the facility; at the same time, traditional retrieval methods may be based on simple rules or manual judgment, which are prone to retrieval errors, affecting resource allocation and decision-making effects; as the scale of facilities expands and the complexity increases, traditional knowledge management methods can no longer meet the needs, resulting in low management efficiency.

[0031] As mentioned above, how to improve the efficiency of power facility management by managing power facility information has become an important research issue.

[0032] Based on the above description, if Figure 1 As shown, the knowledge graph construction method of electric power facilities proposed in this embodiment includes:

[0033] Step 101: Acquire structured data and unstructured data of power facilities.

[0034] Step 102: preprocess the structured data and the unstructured data to obtain power facility data, and construct a digital twin model of the power facility based on the power facility data.

[0035] Step 103: Use the digital twin model to perform entity recognition on the power facility data to obtain entity classification results.

[0036] Step 104: construct an initial knowledge graph based on the entity classification results, and complete the initial knowledge graph to obtain a target knowledge graph.

[0037] In specific implementation, the disclosed embodiment proposes a method for constructing a knowledge graph of electric power facilities. By integrating fragmented electric power facility information, using knowledge graphs for accurate management, and building a digital twin model to achieve comprehensive monitoring and intelligent management, the efficiency and accuracy of electric power facility information retrieval are improved.

[0038] The present disclosure proposes a method for constructing a knowledge graph of digital twin power facilities based on deep learning. First, various information and attributes of the power facilities are digitized to form a digital twin power model. Secondly, with the help of deep learning technology, the digital twin power model is subjected to entity recognition to identify the power facility entities with power attributes and values. Thirdly, a knowledge graph is constructed for the identified power facility entities, and the knowledge graph is completed through the TransE algorithm, which systematically integrates multiple information such as the attributes, functions and locations of the power facilities, and clearly expresses and manages the complex relationships between digital twin models, so that staff can better understand and analyze the dependencies of power facilities. After the knowledge graph of power facilities is established, the entities and relationships in the knowledge graph can be used to carry out intelligent applications of power facilities, so that business personnel can better understand and analyze the operating status, dependencies and potential risks of power facilities.

[0039] Specifically, a digital twin model of power facilities is constructed, and deep learning is used to identify entities in the digital twin power facility model. A knowledge graph is created based on the identified entities, and the missing parts in the knowledge graph are completed to form a relatively comprehensive and systematic knowledge graph of power facilities. Through the entities and relationships in the knowledge graph, the retrieval and management functions of digital twin power facilities can be realized.

[0040] Figure 2 The overall flow chart for constructing the knowledge graph of electric power facilities in the embodiment of the present disclosure is as follows. Figure 2As shown in the figure, in order to construct the knowledge graph of power facilities, data preparation is required first. Data preparation is to collect structured data and unstructured data related to power facilities as the basis for knowledge graph construction. After data preparation is completed, the corresponding digital twin model is generated for the power facilities through data cleaning and attribute fusion. After that, based on the digital twin model of power facilities, the entity recognition operation of the knowledge graph is performed to identify all entities, relationships and attributes related to power facilities. Then the initial knowledge graph is constructed. After the initial knowledge graph is constructed, due to the lack of data and algorithm limitations, there will be some omissions in the constructed initial knowledge graph, so it is necessary to go through the initial knowledge graph completion stage. After the initial knowledge graph is completed, a relatively systematic and comprehensive target knowledge graph of power facilities can be obtained. After that, the relevant power business applications can be expanded based on the completed target knowledge graph.

[0041] Through the above embodiments, structured data and unstructured data of power facilities are obtained. The structured data and unstructured data are preprocessed to obtain power facility data, and a digital twin model of the power facility is constructed based on the power facility data. The digital twin model is used to perform entity recognition on the power facility data to obtain entity classification results. An initial knowledge graph is constructed based on the entity classification results, and the initial knowledge graph is completed to obtain a target knowledge graph. In this way, the digital twin model can be used to identify the contextual information and local features of the power facility data, so that the entity classification results are more accurate. By constructing an initial knowledge graph, power facility applications can be supported. By completing the initial knowledge graph to obtain the target knowledge graph, a relatively systematic and comprehensive power facility knowledge graph is obtained, thereby improving the performance of the target knowledge graph.

[0042] In some embodiments, the digital twin model includes: a BERT vectorization module, a BiLSTM-CNN feature extraction module and a CRF sequence labeling module; step 103 includes:

[0043] Step 1031: Use the BERT vectorization module to vectorize the power facility data to obtain a power facility vector.

[0044] Step 1032: Use the BiLSTM-CNN feature extraction module to perform power facility feature extraction processing on the power facility vector to obtain power facility features.

[0045] Step 1033: Use the CRF sequence labeling module to perform sequence labeling processing on the power facility features to obtain a power facility labeling sequence, and use the power facility labeling sequence as an entity classification result.

[0046] In the specific implementation, based on the power facility equipment inspection report, power facility documents and other materials, character-level features are integrated according to the characteristics of the power grid data to extract the power features in the power grid data. The entity classification is completed while the power grid equipment entity extraction task is performed. A BERT-BiLSTM-CNN-CRF hybrid model that integrates entity keyword features is proposed. The hybrid model consists of the following three parts:

[0047] 1. BERT vectorization module: Perform vectorization operations by extracting keyword features of power facilities and performing vectorized representation.

[0048] 2. BiLSTM-CNN feature extraction module. This module is used for feature extraction. The word vector obtained after processing by the BERT vectorization module is first input into the BILSTM feature extraction module to extract the features of the power facilities, and then the prediction results are input into the CNN feature extraction module for further feature extraction.

[0049] 3. CRF sequence labeling module: The results obtained after processing by the CNN feature extraction module are input into the CRF sequence labeling module for sequence labeling to obtain the corresponding labeling sequence, thereby realizing the entity recognition of power facilities.

[0050] Keywords, also known as reserved words, are often the central theme of a text, the focus of a sentence, or representative key information in a vocabulary. Keywords or key words are often used in natural language processing tasks. Focusing on data in the power field, we study the task of identifying entities in power facilities. By observing and analyzing data in the power field, we can find that the entities to be extracted in the power field data often contain keyword information. Therefore, keyword information is used to assist the task of extracting entities in power facilities.

[0051] First, the Term Frequency-Inverse Document Frequency (TF-IDF) algorithm is used to assist in building keywords. The TF-IDF algorithm is an algorithm used for information retrieval and text analysis, mainly used to evaluate the importance of a word to a document. The TF-IDF algorithm mainly consists of two parts: term frequency (TF) and inverse document frequency (IDF).

[0052] Considering that keyword information generally appears in keywords, TF-IDF is used to first extract entity keywords in the power field, and then construct keywords through keywords. TF-IDF evaluates the importance of a word based on the number of times a word appears in the text.

[0053] When using TF-IDF to extract keywords, you first need to build a stop word list to remove irrelevant punctuation and words in the corpus. Then you need to calculate the word frequency (that is, the number of times a word appears in the file).

[0054]

[0055] Among them, TF k is the frequency of word k in the text, N k is the number of times word k appears in the text, and N is the number of words in the text.

[0056] Then, based on a large amount of data in the field of power facilities, the inverse document frequency (IDF) is calculated.

[0057]

[0058] Among them, IDF k is the inverse document frequency of word k in the text, Y is the total number of documents in the corpus, and Y k is the number of documents containing word k.

[0059] Finally, calculate the TF-IDF value.

[0060] TF-IDF k =TF k *IDF k

[0061] The obtained TF-IDF values ​​were arranged in descending order, and the first 800 were selected as the extracted keywords. Next, the keywords were manually screened and extracted from the extracted keywords, and finally a keyword table in the field of power facilities was constructed by professionals. The keywords were then used as features and input into the model together with word vectors for entity recognition.

[0062] Through the above scheme, a hybrid model of BERT-BiLSTM-CNN-CRF that integrates entity keyword features is used to capture the context sequence information in the power facility report through the BILSTM feature extraction module, and the local features in the power facility report are captured through the CNN feature extraction module, thereby achieving full mining of the power facility text report and improving the effect of entity recognition.

[0063] In some embodiments, step 1031 includes:

[0064] Step 1031A: convert the electric power facility data into an initial electric power facility sequence.

[0065] Step 1031B, adding a start character to the start position of the initial electric facility sequence, and adding an end character to the end position of the initial electric facility sequence, to obtain an updated electric facility sequence.

[0066] Step 1031C, using the BERT vectorization module, vectorizes the updated power facility sequence to obtain a word embedding vector.

[0067] Step 1031D, using the BERT vectorization module, divide the updated power facility sequence according to the location of the sentence to obtain a sentence embedding vector.

[0068] Step 1031E, using the BERT vectorization module, mark the sequence position of the updated power facility sequence to obtain a position embedding vector.

[0069] Step 1031F, combining the word embedding vector, the sentence embedding vector and the position embedding vector to obtain a power facility vector.

[0070] In specific implementation, the BERT model can recognize and extract semantic information in sentences, so the word vectors generated by the BERT model contain semantic information, which will be of great help to entity recognition in the field of power facilities. Therefore, this method uses the BERT vectorization module to generate word vectors, which improves the quality of the generated word vectors, thereby improving the accuracy of power facility entity recognition.

[0071] In the specific operation process, input the initial power facility sequence X = (x 1 ,x 2 ,x 3 ,…,x n ) Add "[CLS]" characters to the beginning of each sequence to store the semantic information of the entire input sequence, and then use "[SEP]" to separate and distinguish sentences, and add special characters "[SEP]" at the end of each sentence. To train word vectors using the BERT vectorization module, three embedding operations are required, namely word embedding, sentence embedding, and position embedding.

[0072] 1) Word embedding: The input part is character-level data. First, word embedding is performed, that is, the BERT vectorization module performs an embedding operation on the experimental data in the field of power facilities, and converts the input characters into word embedding vectors.

[0073] 2) Sentence embedding: It is necessary to distinguish whether the input text is sentence A or sentence B, and divide the s sentences in the corpus according to the parity of the sentence position, such as N = {n 1 ,n 2 ,…,n k ,…,n n}If the position k of the sentence is an odd number, define the Segment Embeddings of each word vector in this sentence as E A, if the sentence position k is an even number, then define E B .

[0074] 3) Position embedding: Select position information for position embedding to mark the position of the character in the input data.

[0075] 4) Combine the results generated by the three embeddings to obtain the power facility vector Y generated by the BERT model = (T 1 ,T 2 ,T 3 ,…,T n ). The characters in the corpus and the input keyword features are trained through the BERT vectorization module to generate word vectors as the input part of the BiLSTM-CNN feature extraction model.

[0076] Through the above scheme, the BERT vectorization module is used to vectorize the power facility data, so that the obtained power facility vector contains the semantic information of the power facility, thereby improving the accuracy of power facility entity recognition. In addition, the BERT vectorization module is used to determine the word embedding vector, sentence embedding vector and position embedding vector, and then the word embedding vector, sentence embedding vector and position embedding vector are combined to obtain the power facility vector, so that the obtained power facility vector is more comprehensive and accurate.

[0077] In some embodiments, step 1032 includes:

[0078] Step 1032A, using the BiLSTM feature extraction module, extracting the above information of the power facility vector to obtain a forward sequence, extracting the below information of the power facility vector to obtain a reverse sequence, and concatenating the forward sequence and the reverse sequence to obtain a target sequence.

[0079] Step 1032B, using the CNN feature extraction module, performs local feature extraction on the target sequence to obtain an output matrix, and uses the output matrix as the power facility feature.

[0080] In specific implementation, the BiLSTM-CNN feature extraction module combines the advantages of the BiLSTM model and the CNN model. The BiLSTM model is good at processing long texts, and the CNN model is good at extracting local features. In the process of entity recognition, contextual information is often needed to improve the extraction effect.

[0081] Figure 3 Schematic diagram of the structure of the BiLSTM feature extraction module of the embodiment of the present disclosure. Figure 3As shown in the figure, the BiLSTM feature extraction module is composed of a forward LSTM feature extraction module and a backward LSTM feature extraction module. The forward LSTM feature extraction module can memorize the previous information, while the backward LSTM feature extraction module can memorize the following information. The BiLSTM feature extraction module can make full use of the context information. Therefore, the bidirectional LSTM feature extraction module is used for feature extraction to convert the forward sequence output at each position into and reverse sequence Splicing at the merging layer to get the complete target sequence

[0082] The BiLSTM feature extraction module consists of five layers in total, the first is the input layer, the second is the forward feedback layer, the third is the backward feedback layer, the fourth is the activation layer, and the fifth is the output layer.

[0083] The BiLSTM feature extraction module is integrated with the CNN feature extraction module for feature extraction. A 4-layer CNN is added after the BiLSTM feature extraction module, and the convolution kernel size is (5,256) dimensions. The power facility vector T generated by the BERT vectorization module is (T 1 ,T 2 ,T 3 ,…,T n ) is subjected to feature extraction by the BiLSTM feature extraction module. The target sequence h output by the BiLSTM feature extraction module is h = (h 1 ,h 2 ,h 3 ,…,h n ) contains rich context information, initial tag information, keyword feature information and output tag information. Then the feature vector output by the BiLSTM feature extraction module is input into the CNN feature extraction module to extract local features and obtain the output matrix P n*k =(p 1 ,p 2 ,…,p n ), where k is the number of defined labels. n*k =(p 1 ,p 2 ,…,p n ), P n It is the score value of each label of the input word, such as P ij The i-th word is the score of the j-th label. However, the prediction based on the score alone is not accurate. This requires the introduction of the CRF sequence labeling module to improve the accuracy of classification.

[0084] Through the above scheme, the BiLSTM feature extraction module can be used to extract the context information of the power facility vector, and the CNN feature extraction module can be used to extract the local features of the power facility vector. In this way, the obtained power facility features integrate the context information and the local features, so that entity recognition can be performed accurately.

[0085] In some embodiments, step 1033 includes:

[0086] Step 1033A: using the CRF sequence labeling module, predicting the power facility features to obtain a predicted power facility sequence, and determining the score of the predicted power facility sequence.

[0087]

[0088] Wherein, S(X,Y) is the score of the predicted power facility sequence, is the transition probability of the i-th label, The output for the i-th position is y i The output probability of .

[0089] Step 1033B, determining a target tag sequence according to the scores of the predicted power facility sequence by using the Viterbi algorithm in the dynamic programming algorithm.

[0090] Step 1033C: label the electric facility features according to the target label sequence to obtain an electric facility label sequence.

[0091] Y * = argmaxScore(X,Y r )

[0092] Among them, Y * is the power facility labeling sequence, X is the initial power facility sequence, Y r Characteristics of power facilities.

[0093] In specific implementation, Conditional Random Fields (CRF) is a statistical model used for sequence labeling tasks. It can more accurately label sequence data by considering the dependency between input sequence and output tag. CRF model is often used in natural language processing tasks such as part-of-speech tagging, word segmentation, named entity recognition, etc., and is integrated with LSTM model to solve entity extraction problems.

[0094] The CRF sequence labeling module is introduced to perform sequence labeling. The CRF sequence labeling module can add constraints to the predicted labels and use the existing labeling information in the labeling process. For example, the label of the next word of the word labeled "B-" in the entity should be "I-" or "O". The CRF sequence labeling module can also learn certain constraints from the data set during the training process, such as the label of the first word in the entity should be "B-" or "O".

[0095] Through the BILSTM-CNN feature extraction module, we can know that the initial power facility sequence X = (x 1 ,x 2 ,x 3 ,…,x b ) After feature extraction, the output matrix P is obtained n*k =(p 1 ,p 2 ,…,p n ), for the forecast power facility sequence Y=(y 1 ,y 2 ,y 3 ,…,y n ), define the score of the predicted power facility sequence, the function is as follows:

[0096]

[0097] The score S of each position consists of two parts, one is the transfer score matrix R of the CRF sequence labeling module, and the other is the output matrix Q of the BILSTM-CNN feature extraction module. is the transition probability from label i to label j, The softmax output for the i-th position is y i The sum of the scores at each position is the score of the entire sequence.

[0098] When performing entity recognition, the scores S corresponding to all possible y are calculated based on the trained parameters. When making predictions, the CRF sequence labeling module uses the Viterbi algorithm in the dynamic programming algorithm to obtain the optimal label sequence (i.e., the target label sequence), labels it according to the optimal label sequence, and then outputs it as the prediction result.

[0099] Y * = argmaxScore(X,Y r )

[0100] Among them, Y * is the power facility labeling sequence, Y r It is the characteristic of power facilities (i.e. the real labeled data sequence).

[0101] Through the above scheme, the CRF sequence labeling module can consider the dependency between the input sequence and the output tag, so as to label the sequence data more accurately. Therefore, the CRF sequence labeling module can accurately label the power facility features, making the obtained power facility labeling sequence more accurate, so as to accurately perform entity recognition.

[0102] In some embodiments, step 104 includes:

[0103] Step 1041 , determining entity concepts and entity relationships according to the entity classification results, and constructing an ontology model based on the entity concepts and the entity relationships.

[0104] Step 1042: construct a mapping relationship between power facility entities and power facility relationships based on the ontology model, and store the mapping relationship in a database.

[0105] Step 1043: Establish a power facility data structure based on the ontology model and the database, and establish a corresponding relationship between the ontology model and the database.

[0106] Step 1044, manage the power facility concept, power facility entity and power facility data separately, and integrate the power facility concept, power facility entity and power facility data separately to obtain an initial knowledge graph.

[0107] Step 1045, performing key path analysis, entity distribution analysis and knowledge exploration analysis on the initial knowledge graph to obtain analysis results.

[0108] When implemented specifically, a knowledge graph is a graph-based data structure that emphasizes contextual understanding of data by linking metadata. This is ideal for applications in scenarios that require large-scale integration, management, and extraction of value from different sources, and dynamic general knowledge graphs are well suited for digital twin scenarios. Compared to traditional data models, knowledge graphs have multiple advantages in modeling, structuring, managing, and analyzing heterogeneous and complex data with dynamic relationships, allowing complex knowledge abstractions to be represented in specific or cross-domains. Unlike relational or NoSQL models, knowledge graphs allow data to evolve flexibly because no predefined schema is required. In addition, these structures can interoperate by mapping entities to current ontologies, thereby enabling transactions of data and its context between software applications.

[0109] After identifying the power facility entities, it is necessary to build the corresponding digital twin power facility knowledge graph. Figure 4 Flow chart of knowledge graph construction of the embodiment of the present disclosure. Figure 4 As shown, through Figure 4The paper introduces the construction process of the knowledge graph of digital twin power facilities, which includes ontology model construction, knowledge mapping, data extraction, knowledge management and graph analysis.

[0110] (1) Ontology model: used to construct the components of power facilities, establish the concepts of entities in power facilities, and establish the relationships between entities in power facilities.

[0111] (2) Knowledge mapping: Based on the ontology model and entity relationships, the knowledge mapping of power facilities is constructed, and the relationship mapping between different entities is established. The corresponding structure can be stored in a variety of databases, such as MySQL, SQL Server, Hive, and Oracle. The specific storage database selected needs to be determined according to the application environment.

[0112] (3) Data extraction: With the help of the ontology model and the corresponding database, the data structure of the power facilities is established, and a mapping relationship is established between the ontology model and the database. In the process of data extraction, full extraction and incremental extraction can be performed. Full extraction means extracting all the data in the corresponding data system, and incremental extraction means extracting only the newly added and changed data when extracting the data from the corresponding data system.

[0113] (4) Knowledge management: Management of concepts, entities and attributes of power facilities, including concept management, concept fusion, power facility entity management, power facility entity fusion, and attribute management and fusion of power facility entities.

[0114] (5) Graph analysis: After the above steps, the construction process of the initial knowledge graph of power equipment is basically completed. Then, based on the obtained initial knowledge graph, key path analysis, entity distribution analysis, knowledge exploration analysis and other power facility applications can be carried out.

[0115] Figure 5 Schematic diagram of the structure of the distribution network knowledge graph of the embodiment of the present disclosure. Figure 5 As shown in the figure, the knowledge graph structure of the distribution network forms three sub-areas of the knowledge graph according to the three concepts of power supply area, power supply voltage and power supply mode. Among them, the power supply area is divided into urban power supply area, rural power supply area and factory power supply area. The power supply mode is divided into direct current power supply mode and alternating current power supply mode. The power supply voltage is divided into high voltage, medium voltage and low voltage.

[0116] Through the above scheme, the method for constructing a digital twin knowledge graph based on the power facility entity realizes the construction of the initial knowledge graph of the digital twin power facility through the processes of ontology model construction, knowledge mapping, data extraction, knowledge management, and graph analysis, so as to support the application of power facilities.

[0117] In some embodiments, step 104 includes:

[0118] Step 104A, using the TransE algorithm, determine the triple vector of the initial knowledge graph and determine the objective function of the triple vector.

[0119] Step 104B, determining the triple distance according to the objective function through the L2 norm, taking the triple distance as a hyperparameter, and determining the loss function according to the hyperparameter.

[0120] Step 104C: train the initial knowledge graph according to the loss function to obtain an updated knowledge graph.

[0121] Step 104D, determining the triplet parameters of the updated knowledge graph, determining the similarity between each candidate entity among multiple candidate entities and the triplet vector parameters, and determining the target entity with the greatest similarity from the multiple candidate entities.

[0122] Step 104E, using the target entity, complete the triple parameters of the updated knowledge graph to obtain the target knowledge graph.

[0123] In specific implementation, through the construction of knowledge graph, a relatively complete initial knowledge graph of power facilities can be basically formed. However, the initial knowledge graph formed by relying on the data extracted from the system and the collected power facility inspection reports often has some omissions, so it is necessary to complete the facility knowledge graph.

[0124] Use the TransE algorithm for knowledge graph completion. The TransE algorithm is a knowledge graph completion method based on translation embedding. It represents both entities and relations as vectors, and simulates the semantic relationship between entities and relations through translation operations in the vector space. The core idea of ​​the TransE algorithm is to regard the relationship as a translation operation from the head entity to the tail entity. Specifically, if a triple (h, r, t) holds (where h represents the head entity, r represents the relationship, and t represents the tail entity), then in the vector space, the head entity vector h plus the relationship vector r should be close to or equal to the tail entity vector t, that is, h+r≈t. The TransE algorithm uses the maximum margin method, and the objective function is as follows:

[0125]

[0126] Among them, S is the correct triple, S ′ is the triplet of errors obtained by replacing h or t, γ is the interval distance hyperparameter, [x] + represents a positive function, that is, when x>0, [x] + =x; when x≤0, [x] + =0.

[0127] The L2 norm is used to calculate the triple distance. The result is a hyperparameter of the TransE algorithm. First, the L2 norm is derived, and the positive and negative values ​​are determined element by element. If it is positive, it is assigned 1, and if it is negative, it is -1. The calculation formula is as follows:

[0128]

[0129] The loss function is:

[0130]

[0131] During the model training process, the loss function will gradually decrease, which reflects the improvement of model performance. Regarding the distance difference between positive triples and negative triples, when the distance of positive triples (dpositive triples) decreases and the distance of negative triples (dnegative triples) increases, the difference between the two (dpositive triples - dnegative triples) will appear as a smaller and smaller negative number. As a positive number, margin is used to set a threshold to ensure that the distance difference between positive triples and negative triples does not exceed this threshold. As the training progresses, when the absolute value of a negative number exceeds margin, although the entire expression will become negative in theory, in actual calculations, only the positive part is taken, that is, when the expression is negative, the expression is set to zero. In this way, the maximum distance between positive and negative triples is limited to the range set by margin. Therefore, margin actually represents the maximum distance allowed between positive and negative triples. The introduction of margin effectively avoids the situation where the distance of negative samples increases infinitely, thereby ensuring the stability and performance of the model.

[0132] After the model training is completed, the obtained embedding vector is used to complete the knowledge graph. For a given head entity and relationship, the tail entity is predicted by searching the vector space for the closest vector to the head entity vector plus the relationship vector; and vice versa. In addition, the prediction ability of the TransE algorithm can be used to complete incomplete triples. The process of predicting missing elements is completed by searching the third element that best matches the known elements (head entity, relationship, or two of the tail entities) in the trained embedding vector space. Specifically, for an incomplete triple, the model will traverse all possible candidate entities, calculate the vector relationship between these candidate entities and the known elements (that is, for a given head entity and relationship, calculate the similarity between the candidate tail entity vector and the "head entity vector + relationship vector", or vice versa), and select the candidate entity with the highest similarity (that is, the smallest distance) as the prediction result to complete the triple. This method takes advantage of the linear relationship characteristics between entity and relationship vectors learned by the TransE model.

[0133] Through the above scheme, the initial knowledge graph is completed based on the TransE algorithm to obtain the target knowledge graph, and the incomplete triples are completed using the prediction ability of the TransE algorithm. The missing elements are predicted by searching for the third element that best matches the known elements (two of the head entity, relationship or tail entity) in the trained embedding vector space, thereby improving the performance of the generated target knowledge graph of power facilities.

[0134] Through the above embodiments, structured data and unstructured data of power facilities are obtained. The structured data and unstructured data are preprocessed to obtain power facility data, and a digital twin model of the power facility is constructed based on the power facility data. The digital twin model is used to perform entity recognition on the power facility data to obtain entity classification results. An initial knowledge graph is constructed based on the entity classification results, and the initial knowledge graph is completed to obtain a target knowledge graph. In this way, the digital twin model can be used to identify the contextual information and local features of the power facility data, so that the entity classification results are more accurate. By constructing an initial knowledge graph, power facility applications can be supported. By completing the initial knowledge graph to obtain the target knowledge graph, a relatively systematic and comprehensive power facility knowledge graph is obtained, thereby improving the performance of the target knowledge graph.

[0135] It should be noted that the method of the embodiment of the present disclosure can be performed by a single device, such as a computer or a server. The method of the present embodiment can also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present disclosure, and the multiple devices will interact with each other to complete the described method.

[0136] It should be noted that the above describes some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0137] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a knowledge graph construction device for power facilities.

[0138] refer to Figure 6 , the knowledge graph construction device of the electric power facility comprises:

[0139] An acquisition module 301 is configured to acquire structured data and unstructured data of a power facility;

[0140] A model building module 302 is configured to pre-process the structured data and the unstructured data to obtain power facility data, and build a digital twin model of the power facility based on the power facility data;

[0141] An entity identification module 303 is configured to perform entity identification on the power facility data using the digital twin model to obtain an entity classification result;

[0142] The knowledge graph construction module 304 is configured to construct an initial knowledge graph based on the entity classification result, and complete the initial knowledge graph to obtain a target knowledge graph.

[0143] In some embodiments, the digital twin model includes: a BERT vectorization module, a BiLSTM-CNN feature extraction module and a CRF sequence labeling module; the entity recognition module 303 includes:

[0144] A vectorization processing unit, configured to use the BERT vectorization module to perform vectorization processing on the power facility data to obtain a power facility vector;

[0145] A feature extraction unit is configured to use the BiLSTM-CNN feature extraction module to perform power facility feature extraction processing on the power facility vector to obtain power facility features;

[0146] The sequence labeling unit is configured to use the CRF sequence labeling module to perform sequence labeling processing on the power facility features to obtain a power facility labeling sequence, and use the power facility labeling sequence as an entity classification result.

[0147] In some embodiments, the vectorized processing unit includes:

[0148] a conversion subunit, configured to convert the electric power facility data into an initial electric power facility sequence;

[0149] a character adding subunit configured to add a start character at the start position of the initial power facility sequence and to add an end character at the end position of the initial power facility sequence to obtain an updated power facility sequence;

[0150] A word embedding vector determination subunit is configured to use a BERT vectorization module to perform vectorization processing on the updated power facility sequence to obtain a word embedding vector;

[0151] The sentence embedding vector determination subunit is configured to use the BERT vectorization module to divide the updated power facility sequence according to the location of the sentence to obtain the sentence embedding vector;

[0152] A position embedding vector determination subunit is configured to use a BERT vectorization module to mark the sequence position of the updated power facility sequence to obtain a position embedding vector;

[0153] The combination processing subunit is configured to perform combination processing on the word embedding vector, the sentence embedding vector and the position embedding vector to obtain a power facility vector.

[0154] In some embodiments, the feature extraction unit comprises:

[0155] The context information extraction subunit is configured to use a BiLSTM feature extraction module to extract context information from the power facility vector to obtain a forward sequence, extract context information from the power facility vector to obtain a reverse sequence, and concatenate the forward sequence and the reverse sequence to obtain a target sequence;

[0156] The local feature extraction subunit is configured to use the CNN feature extraction module to perform local feature extraction on the target sequence to obtain an output matrix, and use the output matrix as the power facility feature.

[0157] In some embodiments, the sequence labeling unit includes:

[0158] The score determination subunit is configured to use the CRF sequence labeling module to perform prediction processing on the power facility features to obtain a predicted power facility sequence, and determine the score of the predicted power facility sequence,

[0159]

[0160] Wherein, S(X,Y) is the score of the predicted power facility sequence, is the transition probability of the i-th label, The output for the i-th position is y i The output probability of

[0161] A target tag sequence determination subunit is configured to determine a target tag sequence according to the scores of the predicted power facility sequence by using a Viterbi algorithm in a dynamic programming algorithm;

[0162] The labeling processing subunit is configured to label the electric power facility features according to the target label sequence to obtain an electric power facility labeling sequence.

[0163] Y * = argmaxScore(X,Y r )

[0164] Among them, Y *is the power facility labeling sequence, X is the initial power facility sequence, Y r Characteristics of power facilities.

[0165] In some embodiments, the knowledge graph construction module 304 includes:

[0166] An ontology model building unit, configured to determine entity concepts and entity relationships according to the entity classification result, and build an ontology model based on the entity concepts and the entity relationships;

[0167] A mapping relationship construction unit, configured to construct a mapping relationship between power facility entities and power facility relationships based on the ontology model, and store the mapping relationship in a database;

[0168] A data extraction unit, configured to establish a power facility data structure based on the ontology model and the database, and to establish a corresponding relationship between the ontology model and the database;

[0169] The knowledge management unit is configured to manage the power facility concept, the power facility entity and the power facility data respectively, and to fuse the power facility concept, the power facility entity and the power facility data respectively to obtain an initial knowledge graph;

[0170] The graph analysis unit is configured to perform key path analysis, entity distribution analysis and knowledge exploration analysis on the initial knowledge graph to obtain analysis results.

[0171] In some embodiments, the knowledge graph construction module 304 includes:

[0172] An objective function determination unit is configured to determine a triple vector of the initial knowledge graph using a TransE algorithm, and determine an objective function of the triple vector;

[0173] a loss function determination unit, configured to determine a triplet distance according to the objective function through an L2 norm, use the triplet distance as a hyperparameter, and determine a loss function according to the hyperparameter;

[0174] A knowledge graph training unit is configured to train the initial knowledge graph according to the loss function to obtain an updated knowledge graph;

[0175] a target entity determination unit configured to determine a triplet parameter of the updated knowledge graph, determine a similarity between each candidate entity among a plurality of candidate entities and the triplet vector parameter, and determine a target entity with the greatest similarity from among the plurality of candidate entities;

[0176] The knowledge graph completion unit is configured to use the target entity to complete the triple parameters of the updated knowledge graph to obtain the target knowledge graph.

[0177] For the convenience of description, the above device is described by dividing it into various modules according to its functions. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0178] The device of the above embodiment is used to implement the knowledge graph construction method of the corresponding power facilities in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0179] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the knowledge graph construction method of the power facilities described in any of the above embodiments is implemented.

[0180] Figure 7 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.

[0181] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0182] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0183] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0184] The communication interface 1040 is used to connect a communication module (not shown in the figure) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB (Universal Serial Bus), network cable, etc.), or through a wireless mode (such as mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).

[0185] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0186] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.

[0187] The electronic device of the above-mentioned embodiment is used to implement the knowledge graph construction method of the corresponding power facilities in any of the above-mentioned embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0188] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the knowledge graph construction method of the power facility as described in any of the above embodiments.

[0189] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0190] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the knowledge graph construction method of the power facilities as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0191] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a computer program product, including computer program instructions. When the computer program instructions are run on a computer, the computer executes the knowledge graph construction method of the power facilities described in any of the above embodiments, which has the beneficial effects of the corresponding method embodiments and will not be repeated here.

[0192] It is understandable that before using the technical solutions of each embodiment of the present disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.

[0193] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can independently choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.

[0194] As an optional but non-limiting implementation, in response to receiving the user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0195] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0196] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure is limited to these examples. Based on the concept of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.

[0197] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the known power / ground connections to the integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure will be implemented (that is, these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it is apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with changes in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0198] Although the present disclosure has been described in conjunction with specific embodiments of the present disclosure, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0199] The embodiments of the present disclosure are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the present disclosure. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.

Claims

1. A method for constructing a knowledge graph of electric power facilities, characterized in that: The method comprises: Obtain structured and unstructured data of power facilities; Preprocessing the structured data and the unstructured data to obtain power facility data, and constructing a digital twin model of the power facility based on the power facility data; Using the digital twin model to perform entity recognition on the power facility data to obtain entity classification results; An initial knowledge graph is constructed based on the entity classification results, and the initial knowledge graph is completed to obtain a target knowledge graph.

2. The method according to claim 1, characterized in that The digital twin model includes: a BERT vectorization module, a BiLSTM-CNN feature extraction module and a CRF sequence labeling module; The using the digital twin model to perform entity recognition on the power facility data to obtain entity classification results includes: Using the BERT vectorization module, vectorize the power facility data to obtain a power facility vector; Using the BiLSTM-CNN feature extraction module, performing power facility feature extraction processing on the power facility vector to obtain power facility features; The CRF sequence labeling module is used to perform sequence labeling processing on the power facility features to obtain a power facility labeling sequence, and the power facility labeling sequence is used as an entity classification result.

3. The method according to claim 2, characterized in that The using the BERT vectorization module to vectorize the power facility data to obtain a power facility vector includes: converting the electric utility data into an initial electric utility sequence; Adding a start character at the start position of the initial power facility sequence, and adding an end character at the end position of the initial power facility sequence, to obtain an updated power facility sequence; Using the BERT vectorization module, vectorize the updated power facility sequence to obtain a word embedding vector; Using the BERT vectorization module, the updated power facility sequence is divided according to the location of the sentence to obtain a sentence embedding vector; Using the BERT vectorization module, marking the sequence position of the updated power facility sequence to obtain a position embedding vector; The word embedding vector, the sentence embedding vector and the position embedding vector are combined to obtain a power facility vector.

4. The method according to claim 2, characterized in that: The using the BiLSTM-CNN feature extraction module to perform power facility feature extraction processing on the power facility vector to obtain power facility features includes: Using the BiLSTM feature extraction module, extracting the above information of the power facility vector to obtain a forward sequence, extracting the below information of the power facility vector to obtain a reverse sequence, and concatenating the forward sequence and the reverse sequence to obtain a target sequence; The CNN feature extraction module is used to perform local feature extraction on the target sequence to obtain an output matrix, and the output matrix is ​​used as the power facility feature.

5. The method according to claim 2, characterized in that: The using the CRF sequence labeling module to perform sequence labeling processing on the electric facility features to obtain an electric facility labeling sequence includes: Using the CRF sequence labeling module, the power facility features are predicted to obtain a predicted power facility sequence, and the score of the predicted power facility sequence is determined. Wherein, S(X,Y) is the score of the predicted power facility sequence, is the transition probability of the i-th label, The output for the i-th position is y i The output probability of Determine a target tag sequence according to the scores of the predicted power facility sequence by using a Viterbi algorithm in a dynamic programming algorithm; The electric power facility features are labeled according to the target label sequence to obtain an electric power facility label sequence, Y * =argmaxScore(X,Y r ) Among them, Y * is the power facility labeling sequence, X is the initial power facility sequence, Y r Characteristics of power facilities.

6. The method according to claim 1, characterized in that The constructing an initial knowledge graph based on the entity classification result includes: Determine entity concepts and entity relationships according to the entity classification results, and construct an ontology model based on the entity concepts and the entity relationships; Based on the ontology model, construct a mapping relationship between power facility entities and power facility relationships, and store the mapping relationship in a database; Establishing a power facility data structure based on the ontology model and the database, and establishing a corresponding relationship between the ontology model and the database; The power facility concepts, power facility entities and power facility data are managed separately, and the power facility concepts, power facility entities and power facility data are integrated separately to obtain an initial knowledge graph; The initial knowledge graph is subjected to key path analysis, entity distribution analysis and knowledge exploration analysis to obtain analysis results.

7. The method according to claim 1, characterized in that The step of completing the initial knowledge graph to obtain a target knowledge graph includes: Using the TransE algorithm, determine the triple vector of the initial knowledge graph, and determine the objective function of the triple vector; Determine a triplet distance according to the objective function by using an L2 norm, use the triplet distance as a hyperparameter, and determine a loss function according to the hyperparameter; According to the loss function, the initial knowledge graph is trained to obtain an updated knowledge graph; Determine the triplet parameters of the updated knowledge graph, determine the similarity between each candidate entity in a plurality of candidate entities and the triplet vector parameters, and determine the target entity with the greatest similarity from the plurality of candidate entities; The target entity is used to complete the triple parameters of the updated knowledge graph to obtain the target knowledge graph.

8. A knowledge graph construction device for electric power facilities, characterized in that: include: An acquisition module configured to acquire structured data and unstructured data of the electric power facility; A model building module is configured to preprocess the structured data and the unstructured data to obtain power facility data, and build a digital twin model of the power facility based on the power facility data; An entity recognition module is configured to perform entity recognition on the power facility data using the digital twin model to obtain an entity classification result; The knowledge graph construction module is configured to construct an initial knowledge graph based on the entity classification results, and complete the initial knowledge graph to obtain a target knowledge graph.

9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and running on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 7.