A knowledge graph construction method and device for a power scenario
By using BiLSTM-CRF technology for entity recognition and relation extraction in power scenarios, and combining entity and image semantic matching, a universal and applicable power knowledge graph is constructed. This solves the problems of large data volume, complex data types, and low data quality in the construction of power knowledge graphs in existing technologies, and improves data processing efficiency and accuracy.
Patent Information
- Application Number
- CN202411603165.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2026-03-20
- Estimated Expiration
- 2044-11-11
AI Technical Summary
In existing technologies, the construction of power knowledge graphs suffers from problems such as massive data volume, complex data types, diverse knowledge content, low value density, and low data quality, making it difficult to form a universal and applicable knowledge graph for the power field.
Bidirectional Long Short-Term Memory (BiLSTM) network and Conditional Random Field (CRF) are used for entity recognition and relation extraction. Combined with entity and image semantic matching methods, a knowledge graph of the power scenario is constructed, including entity normalization, entity alignment and relation fusion.
It improves the efficiency and accuracy of power scenario data processing, reduces reliance on domain-specific expert knowledge, lowers the cost of manual intervention, and forms a universal and applicable knowledge graph, supporting the power industry to better understand development trends and electricity demand.
Smart Images

Figure CN119539050B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and apparatus for constructing a knowledge graph for a power scenario. Background Technology
[0002] In recent years, the power industry has developed rapidly, becoming one of the most important sectors of the economy. According to statistics, the power industry has shown a steady growth trend. With the development of internet technology, business data in the power sector is also constantly increasing, encompassing numerous power entities such as power equipment, power plants, and transmission lines. To understand the future development trends and electricity demand in the power sector, examining the correlations between power elements on a power knowledge graph can provide a systematic understanding of power equipment and electricity consumption, laying a solid foundation for stable power supply and energy management. The construction of a power knowledge graph needs to consider the characteristics of multimodal, unstructured, and unlabeled data, as well as the high costs of manpower and time. Therefore, building a high-quality power knowledge graph system is one of the significant challenges currently facing the power industry.
[0003] In the existing technology, entity recognition methods based on deep learning have been proposed one after another, and methods for constructing power knowledge graphs based on power system data (such as dispatch management systems) have also been proposed. However, challenges such as huge data volume, complex data types, diverse knowledge content, low value density, and low data quality still exist.
[0004] How to construct a knowledge graph in the power sector that is universal and applicable is a technical problem that needs to be solved. Summary of the Invention
[0005] This invention provides a method and apparatus for constructing a knowledge graph for a power scenario, in order to address the deficiencies in the existing technology.
[0006] This invention provides a method for constructing a knowledge graph in a power scenario, comprising the following steps:
[0007] Text and image data of the target power scenario are acquired, and entity recognition extraction is performed on the text and image data of the target power scenario based on Bidirectional Long Short Time Memory Network (BiLSTM) and Conditional Random Field (CRF) to obtain multimodal data entity recognition results.
[0008] Relationships are extracted based on the multimodal data entity recognition results to determine the corresponding target triples; wherein, the target triples include: entities and the association relationships between entities;
[0009] Based on the target triples, knowledge fusion and knowledge graph construction are performed to obtain the knowledge graph of the target power scenario.
[0010] According to a knowledge graph construction method for a power scenario provided by the present invention, the method involves acquiring text data and image data of a target power scenario, and performing entity recognition extraction on the text data and image data of the target power scenario based on a bidirectional long short-term memory network (BiLSTM) and a conditional random field (CRF) to obtain multimodal data entity recognition results, including:
[0011] The text data of the target power scenario is acquired, and the first entity recognition result corresponding to the text data is extracted based on the bidirectional long short-term memory network BiLSTM and the conditional random field CRF.
[0012] Image data of the target power scene is acquired, and the second entity recognition result corresponding to the image data is extracted based on the entity and image semantic matching method.
[0013] The multimodal data entity recognition result is constituted based on the first entity recognition result and the second entity recognition result.
[0014] According to the knowledge graph construction method for a power scenario provided by the present invention, the step of acquiring text data of the target power scenario and extracting the first entity recognition result corresponding to the text data based on the Bidirectional Long Short-Term Memory Network (BiLSTM) and the Conditional Random Field (CRF) includes:
[0015] The text data of the target power scenario is acquired, and the text data is preprocessed to generate a first embedding vector corresponding to each word; wherein, the first embedding vector is obtained by concatenating the high-dimensional vector of the word generated based on the first pre-training algorithm and the unique encoding identifier of the word;
[0016] Based on the first embedding vector corresponding to each word, the contextual sentence-level features in the text data are extracted through the bidirectional long short-term memory network BiLSTM, and a complete labeled prediction sequence is obtained.
[0017] The neural unit dropout rate is set to obtain the predicted probability value of the label corresponding to each word; wherein, the neural unit dropout rate is used to prevent overfitting;
[0018] A Conditional Random Field (CRF) is constructed, and based on the complete labeled prediction sequence and the predicted probability value of the label corresponding to each word, the first entity recognition result corresponding to the text data is obtained.
[0019] According to the knowledge graph construction method for a power scenario provided by the present invention, the step of acquiring image data of the target power scenario and extracting a second entity recognition result corresponding to the image data based on an entity and image semantic matching method includes:
[0020] Image data of the target power scene is acquired, and image preprocessing is performed on the image data to obtain an overall image representation;
[0021] Based on the entity and image semantic matching method, target information is extracted from the overall image representation, and a second entity recognition result corresponding to the image data is obtained based on the target information.
[0022] According to the knowledge graph construction method for a power scenario provided by the present invention, the relationship extraction includes: an embedding layer, a bidirectional long short-term memory network (BiLSTM) layer, and a conditional random field (CRF) layer;
[0023] The step of extracting relationships based on the multimodal data entity recognition results and determining the corresponding target triples includes:
[0024] For the embedding layer, a second embedding vector is generated for each word in the target power scene data, and a second embedding vector sequence is generated; wherein, the second embedding vector is obtained by concatenating the high-dimensional vector generated by the first pre-training algorithm and the high-dimensional vector generated by the second pre-training algorithm for each word;
[0025] The second embedded vector sequence is input into the Bidirectional Long Short-Term Memory (BiLSTM) layer to obtain a vector sequence of the target dimension; wherein, the target dimension is the number of relation labels;
[0026] The vector sequence of the target dimension is input into the Conditional Random Field (CRF) layer for sequence labeling to obtain the output label sequence, and the target triplet is identified based on the output label sequence.
[0027] According to a knowledge graph construction method for a power scenario provided by the present invention, the knowledge fusion includes: entity normalization, entity alignment, and relationship fusion;
[0028] The process of knowledge fusion and knowledge graph construction based on the target triples to obtain the knowledge graph of the target power scenario includes:
[0029] Based on the target triples, entity normalization, entity alignment, and relation fusion are performed respectively to obtain the knowledge fusion result; wherein, entity normalization is: removing duplicate similar entities and simplifying entities according to semantic similarity; entity alignment is: fusing concepts in the entity set; relation fusion is: removing redundant relations based on the result of entity normalization;
[0030] A knowledge graph is constructed from the knowledge fusion results to obtain the knowledge graph of the target power scenario.
[0031] The present invention also provides a knowledge graph construction device for power scenarios, comprising the following modules:
[0032] The entity recognition module is used to acquire text data and image data of the target power scene, and to perform entity recognition extraction on the text data and image data of the target power scene based on the bidirectional long short time memory network BiLSTM and the conditional random field CRF to obtain multimodal data entity recognition results.
[0033] The relationship extraction module is used to extract relationships based on the multimodal data entity recognition results and determine the corresponding target triples; wherein, the target triples include: entities and the association relationships between entities;
[0034] The knowledge graph construction module is used to perform knowledge fusion and knowledge graph construction based on the target triples to obtain the knowledge graph of the target power scenario.
[0035] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a knowledge graph construction method for any of the above-described power scenarios.
[0036] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the knowledge graph construction method for the power scenario as described above.
[0037] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements a knowledge graph construction method for any of the above-described power scenarios.
[0038] This invention provides a method and apparatus for constructing a knowledge graph for a power scenario. It acquires text and image data of a target power scenario and performs entity recognition extraction on the text and image data based on a Bidirectional Long Short-Term Memory (BiLSTM) network and a Conditional Random Field (CRF) to obtain multimodal entity recognition results. Based on the multimodal entity recognition results, it extracts relationships to determine corresponding target triples. Each target triple includes an entity and the relationships between entities. Based on the target triples, it performs knowledge fusion and knowledge graph construction to obtain the knowledge graph of the target power scenario. Therefore, this invention utilizes BiLSTM-CRF deep learning technology to achieve named entity recognition and simultaneously forms a universal and applicable knowledge graph construction method for better management and utilization of relevant knowledge in the power field. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0040] Figure 1 This is a flowchart illustrating the knowledge graph construction method for power scenarios provided by the present invention.
[0041] Figure 2 This is a complete flowchart of the knowledge graph construction method for power scenarios provided by the present invention.
[0042] Figure 3 This is a schematic diagram of the knowledge graph construction device for power scenarios provided by the present invention.
[0043] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0045] The following is combined with Figures 1-4 This invention describes a method and apparatus for constructing a knowledge graph for a power scenario.
[0046] It's important to note that entities are the basic units in knowledge graphs, including attributes, attribute values, and the correspondences between related entities. Entity recognition methods typically require designing features for text from different domains, lacking general applicability. With the increasing scale of data, the design of multi-layered architectures and the introduction of deep learning have become crucial. To reduce the human effort required for manually formulating rules and improve the generalization ability of models, deep learning-based entity recognition methods have been proposed in recent years. To address the issue that existing methods are designed for English data and cannot be applied to Chinese data, a method for extracting Chinese relationships based on attention models using entity character features has been proposed.
[0047] The power sector encompasses a wide variety of scenarios, with complex relationships between them. Furthermore, the lack of large-scale publicly available data in the power sector has resulted in relatively limited research on knowledge graphs related to power scenarios. Existing power knowledge graphs built upon power system data (such as dispatch management systems) face challenges including massive data volume, complex data types, diverse knowledge content, low value density, and low data quality. Therefore, this invention provides a method for constructing a knowledge graph for power scenarios to address at least one of the aforementioned problems.
[0048] Figure 1 This is a flowchart illustrating the knowledge graph construction method for power scenarios provided by the present invention, as shown below. Figure 1 As shown, the method includes the following:
[0049] Step 100: Obtain text data and image data of the target power scene, and perform entity recognition extraction on the text data and image data of the target power scene based on Bidirectional Long Short-Term Memory Network (BiLSTM) and Conditional Random Field (CRF) to obtain multimodal data entity recognition results.
[0050] Figure 2 This is a complete flowchart of the knowledge graph construction method for power scenarios provided by the present invention. The following is a combination of... Figure 2 The method for constructing a knowledge graph for power scenarios provided by this invention is described.
[0051] It should be noted that step 100, multimodal data entity recognition, specifically includes entity recognition in text data and entity recognition in image data.
[0052] Specifically, step 100 includes:
[0053] Step 110: Obtain the text data of the target power scenario, and extract the first entity recognition result corresponding to the text data based on the Bidirectional Long Short-Term Memory Network (BiLSTM) and the Conditional Random Field (CRF).
[0054] Step 110 specifically includes:
[0055] Step 111: Obtain the text data of the target power scenario, perform text preprocessing on the text data, and generate a first embedding vector corresponding to each word; wherein, the first embedding vector is obtained by concatenating the high-dimensional vector of the word generated based on the first pre-training algorithm and the unique encoding identifier of the word.
[0056] Step 112: Based on the first embedding vector corresponding to each word, extract the contextual sentence-level features in the text data through the bidirectional long short-term memory network BiLSTM, and obtain the complete labeled prediction sequence.
[0057] Step 113: Set the neural unit dropout rate and obtain the predicted probability value of the label corresponding to each word; wherein, the neural unit dropout rate is used to prevent overfitting.
[0058] Step 114: Construct a Conditional Random Field (CRF) and, based on the complete labeled prediction sequence and the predicted probability value of the label corresponding to each word, obtain the first entity recognition result corresponding to the text data.
[0059] In one embodiment, steps 111 to 114 are described.
[0060] 1. Processing Text Data. Text sentences in the power industry context can be viewed as sequences of multiple words. The words in the power industry context are sorted according to their Chinese Pinyin alphabetical order to form a power industry vocabulary dictionary. A one-hot encoding is generated for each word according to the order in the dictionary. A high-dimensional vector for each word is generated using a pre-trained BERT algorithm. The one-hot encoding and the high-dimensional vector are concatenated to form the first embedding vector for each word.
[0061] It's important to note that BERT is a language model developed through unsupervised training on a large amount of unlabeled text. It uses a Transformer-based encoder for bidirectional encoding. By building a tokenized language model, BERT can randomly overwrite or replace any word in a sentence, enabling the model to predict the randomly overwritten parts of the context and obtain distributed contextual representations of words. Furthermore, during the pre-training phase, BERT simultaneously performs the next sentence prediction task, allowing the model to understand the relationship between two sentences.
[0062] 2. Long text segments are deconstructed using BiLSTM to extract sentence-level features of vocabulary and context in natural text, and the embedding vector of each sentence is optimized. Then, different nodes are input into the BiLSTM network layer. Finally, the observation sequences output by the forward and backward propagation layers of the corresponding nodes are paired to obtain the complete labeled prediction sequence.
[0063] 3. Set the neural unit dropout rate to prevent overfitting. This involves mapping and converting the output hidden state vector to the number of labels in the detection label set, and obtaining matrix P by setting equal dimensions to obtain contextual features and the probability value of each label prediction.
[0064] 4. Construct a CRF layer, based on the Chi-square Markov hypothesis and the independent observation hypothesis, to reintegrate the sequence labeling from word-level to sentence-level. For label prediction of the original text sequence, add start and end frames corresponding to the first and last parts of the sentence to define the sentence boundaries. Then, the model calculates the degree of matching between the predicted label of sentence x and the true label y:
[0065]
[0066] in, This represents the transition score matrix between labels. The matching degree is influenced by the output probability values of the LSTM network layers. and the matrix of CRF layer prediction output The effect of this is further obtained using SoftMax normalization:
[0067]
[0068] in, This represents all possible predicted label sequences.
[0069] The model optimization training method is to maximize the above formula. After training, the label category is directly output based on the highest score of all labels for each word, and the label category is the result of entity recognition.
[0070] Step 120: Obtain image data of the target power scene, and extract the second entity recognition result corresponding to the image data based on entity and image semantic matching method.
[0071] Step 120 specifically includes:
[0072] Step 121: Obtain image data of the target power scene and perform image preprocessing on the image data to obtain an overall image representation.
[0073] Step 122: Based on the entity and image semantic matching method, extract the target information in the overall image representation, and obtain the second entity recognition result corresponding to the image data based on the target information.
[0074] In one embodiment, steps 121 to 122 are described.
[0075] 5. Image Data Processing. The power sector also contains actual image information under different scenarios. A VGG19 network pre-trained on the ImageNet dataset was used, with the output of the last convolutional layer (layer 16) as the image feature. First, each image... The image was resized to 256×256 and randomly cropped five times. Each randomly cropped image was then mirrored horizontally to obtain ten different preprocessed image copies. To extract entity information from the images, a pre-trained Mask R-CNN model was used. The pre-trained Mask R-CNN model is based on a feature pyramid network and a ResNet101 backbone, generating bounding boxes and segmentation masks for each object in the image. Input image The segmentation mask generated by the pre-trained Mask R-CNN model is used to obtain image copies containing only entities from the original image. Subsequently, 10 different image copies and the entity-only image copy are input into the VGG19 feature extractor to obtain the output of the model's last convolutional layer. Then, these feature vectors are subjected to average pooling to obtain the overall image representation. .
[0076] 6. To further extract specific target information such as scene information and label information from the image, an entity and image semantic matching method is adopted. A multi-label image classification model is used. This is used to detect whether an entity appears in an image, where the entity is the result of entity extraction in step four. (Classification model) The goal is to generate a vector. As the detected entities in the image, where the vector This refers to the vocabulary length extracted in step 4. For the... An entity, in a vector The Each position uses 1 or 0 to indicate presence or absence. For example, if the vocabulary of entities extracted in step 4 is {person, transformer, high-voltage line, ladder}, and the image... of If the range is [0, 1, 1, 0], then the image contains two entities: a transformer and a high-voltage line. Then, directly... As a result of image entity extraction.
[0077] Multi-label image classification model The architecture is based on a spatial transformation layer and an LSTM network. The spatial transformation layer locates attention regions from the convolutional feature maps. The LSTM network predicts semantic label scores on the located regions sequentially and captures the global dependencies of these regions.
[0078] Step 130: Based on the first entity recognition result and the second entity recognition result, construct the multimodal data entity recognition result.
[0079] Specifically, the entities extracted in steps 4 and 6 are combined to form multimodal data entity recognition results.
[0080] Step 200: Based on the multimodal data entity recognition results, perform relationship extraction to determine the corresponding target triples; wherein, the target triples include: entities and the association relationships between entities.
[0081] It should be noted that the relation extraction method consists of three layers: an embedding layer, a bidirectional LSTM (BiLSTM) layer, and a conditional random field (CRF) layer.
[0082] Specifically, step 200 includes:
[0083] Step 210: For the embedding layer, generate a second embedding vector corresponding to each word in the target power scene data, and generate a second embedding vector sequence; wherein, the second embedding vector is obtained by concatenating the high-dimensional vector generated by the first pre-training algorithm and the high-dimensional vector generated by the second pre-training algorithm for each word.
[0084] Step 220: Input the second embedded vector sequence into the Bidirectional Long Short-Term Memory (BiLSTM) layer to obtain a vector sequence of the target dimension; wherein, the target dimension is the number of relation labels.
[0085] Step 230: Input the vector sequence of the target dimension into the Conditional Random Field (CRF) layer for sequence labeling to obtain the output label sequence, and identify the target triple based on the output label sequence.
[0086] In one embodiment, steps 210 to 230 are described.
[0087] 1. For the embedding layer, the pre-trained BERT algorithm and GloVe are used to generate a high-dimensional vector for each word. The high-dimensional vectors obtained by the two algorithms are concatenated to form the second embedding vector of the word.
[0088] 2. Word embedding vector sequence It is fed into a BiLSTM layer consisting of a forward LSTM and a backward LSTM. The forward LSTM unit generates the forward hidden state. The calculation process is as follows:
[0089]
[0090] Similarly, the backward LSTM unit generates the backward hidden state. .
[0091]
[0092] Then, the BiLSTM layer combines the forward and backward hidden states to generate a complete hidden state sequence. Then the hidden state vector is mapped to... A dimensional vector, where Indicates the number of relation tags.
[0093] 3. The Conditional Random Field (CRF) layer is input to the output of the BiLSTM layer and begins the sequence labeling process. It has the ability to utilize contextualized labeling information to produce better labeling accuracy. Therefore, relation labels PB are identified from the output labeled sequence to complete the relation extraction task.
[0094] Step 300: Based on the target triple, perform knowledge fusion and knowledge graph construction to obtain the knowledge graph of the target power scenario.
[0095] Specifically, step 300 includes:
[0096] Step 310: Based on the target triple, perform entity normalization, entity alignment, and relation fusion processing respectively to obtain the knowledge fusion result; wherein, the entity normalization is: removing duplicate similar entities and simplifying entities according to semantic similarity; the entity alignment is: fusing concepts in the entity set; the relation fusion is: removing redundant relations according to the result of entity normalization.
[0097] Step 320: Construct a knowledge graph from the knowledge fusion results to obtain the knowledge graph of the target power scenario.
[0098] In one embodiment, steps 310 to 330 are described.
[0099] 1. Knowledge Fusion. This includes entity normalization, entity alignment, and relation fusion. Entity normalization refers to removing duplicate similar entities and simplifying them based on semantic similarity; entity alignment refers to merging concepts in an entity set; and relation fusion refers to removing redundant relations based on the results of entity normalization.
[0100] 2. Construct a knowledge graph. Use open-source platforms such as Neo4j and CoWork to process the merged knowledge dataset and generate node, relationship, and query-based field models.
[0101] The knowledge graph construction method for power scenarios provided by this invention: 1) Employing deep learning technologies such as BiLSTM-CRF for entity recognition and relation extraction can improve the efficiency and accuracy of processing text data. BiLSTM can effectively capture contextual information in text, while CRF can utilize the dependencies between tags to improve the accuracy of annotation; 2) By using pre-trained BERT and GloVe algorithms to generate high-dimensional vectors of words, and combining steps such as entity normalization, entity alignment, and relation fusion, a universal and applicable knowledge graph construction method is formed, which can be applied to power scenario data of different types and scales; 3) Employing deep learning-based entity recognition and relation extraction methods can reduce the reliance on domain-specific expert knowledge, thereby reducing the cost and time consumption of manual intervention; 4) It can help the power industry better understand the development trends and electricity demand of power scenarios, provide more effective support for power supply and energy management, and promote the development and progress of the power industry.
[0102] The above describes the steps of the knowledge graph construction method for power scenarios provided by this invention. As can be seen from the above description, the knowledge graph construction method for power scenarios provided by this invention involves acquiring text and image data of a target power scenario, and performing entity recognition extraction on the text and image data of the target power scenario based on a Bidirectional Long Short-Term Memory (BiLSTM) network and a Conditional Random Field (CRF) to obtain multimodal data entity recognition results; performing relation extraction based on the multimodal data entity recognition results to determine corresponding target triples; wherein, the target triples include: entities and the relationships between entities; and performing knowledge fusion and knowledge graph construction based on the target triples to obtain the knowledge graph of the target power scenario. Therefore, this invention utilizes BiLSTM-CRF deep learning technology to achieve named entity recognition, while simultaneously forming a universal and applicable knowledge graph construction method to better manage and utilize relevant knowledge in the power field.
[0103] The knowledge graph construction device for power scenarios provided by the present invention is described below. The knowledge graph construction device for power scenarios described below and the knowledge graph construction method for power scenarios described above can be referred to in correspondence.
[0104] Figure 3 This is a schematic diagram of the knowledge graph construction device for power scenarios provided by the present invention, as shown below. Figure 3 As shown, the knowledge graph construction device for power scenarios provided by the present invention includes:
[0105] The entity recognition module 301 is used to acquire text data and image data of the target power scene, and to perform entity recognition extraction on the text data and image data of the target power scene based on the bidirectional long short time memory network BiLSTM and the conditional random field CRF to obtain multimodal data entity recognition results.
[0106] The relationship extraction module 302 is used to extract relationships based on the multimodal data entity recognition results and determine the corresponding target triples; wherein, the target triples include: entities and the association relationships between entities;
[0107] The knowledge graph construction module 303 is used to perform knowledge fusion and knowledge graph construction based on the target triples to obtain the knowledge graph of the target power scenario.
[0108] The knowledge graph construction device for power scenarios provided by this invention acquires text and image data of a target power scenario, and performs entity recognition extraction on the text and image data of the target power scenario based on a bidirectional long short-term memory network (BiLSTM) and a conditional random field (CRF) to obtain multimodal data entity recognition results. Based on the multimodal data entity recognition results, relation extraction is performed to determine corresponding target triples. The target triples include entities and the relationships between entities. Based on the target triples, knowledge fusion and knowledge graph construction are performed to obtain the knowledge graph of the target power scenario. Therefore, this invention utilizes BiLSTM-CRF deep learning technology to achieve named entity recognition and simultaneously forms a universal and applicable knowledge graph construction method for better management and utilization of relevant knowledge in the power field.
[0109] Based on the above embodiments, in this embodiment, the entity recognition module 301 is specifically used for:
[0110] The text data of the target power scenario is acquired, and the first entity recognition result corresponding to the text data is extracted based on the bidirectional long short-term memory network BiLSTM and the conditional random field CRF.
[0111] Image data of the target power scene is acquired, and the second entity recognition result corresponding to the image data is extracted based on the entity and image semantic matching method.
[0112] The multimodal data entity recognition result is constituted based on the first entity recognition result and the second entity recognition result.
[0113] Based on the above embodiments, in this embodiment, the device further includes a first identification module, specifically used for:
[0114] The text data of the target power scenario is acquired, and the text data is preprocessed to generate a first embedding vector corresponding to each word; wherein, the first embedding vector is obtained by concatenating the high-dimensional vector of the word generated based on the first pre-training algorithm and the unique encoding identifier of the word;
[0115] Based on the first embedding vector corresponding to each word, the contextual sentence-level features in the text data are extracted through the bidirectional long short-term memory network BiLSTM, and a complete labeled prediction sequence is obtained.
[0116] The neural unit dropout rate is set to obtain the predicted probability value of the label corresponding to each word; wherein, the neural unit dropout rate is used to prevent overfitting;
[0117] A Conditional Random Field (CRF) is constructed, and based on the complete labeled prediction sequence and the predicted probability value of the label corresponding to each word, the first entity recognition result corresponding to the text data is obtained.
[0118] Based on the above embodiments, in this embodiment, the device further includes a second identification module, specifically used for:
[0119] Image data of the target power scene is acquired, and image preprocessing is performed on the image data to obtain an overall image representation;
[0120] Based on the entity and image semantic matching method, target information is extracted from the overall image representation, and a second entity recognition result corresponding to the image data is obtained based on the target information.
[0121] Based on the above embodiments, in this embodiment, the relation extraction includes: an embedding layer, a bidirectional long short-term memory network (BiLSTM) layer, and a conditional random field (CRF) layer;
[0122] The relationship extraction module 302 is specifically used for:
[0123] For the embedding layer, a second embedding vector is generated for each word in the target power scene data, and a second embedding vector sequence is generated; wherein, the second embedding vector is obtained by concatenating the high-dimensional vector generated by the first pre-training algorithm and the high-dimensional vector generated by the second pre-training algorithm for each word;
[0124] The second embedded vector sequence is input into the Bidirectional Long Short-Term Memory (BiLSTM) layer to obtain a vector sequence of the target dimension; wherein, the target dimension is the number of relation labels;
[0125] The vector sequence of the target dimension is input into the Conditional Random Field (CRF) layer for sequence labeling to obtain the output label sequence, and the target triplet is identified based on the output label sequence.
[0126] Based on the above embodiments, in this embodiment, the knowledge fusion includes: entity normalization, entity alignment, and relationship fusion;
[0127] The knowledge graph construction module 303 is specifically used for:
[0128] Based on the target triples, entity normalization, entity alignment, and relation fusion are performed respectively to obtain the knowledge fusion result; wherein, entity normalization is: removing duplicate similar entities and simplifying entities according to semantic similarity; entity alignment is: fusing concepts in the entity set; relation fusion is: removing redundant relations based on the result of entity normalization;
[0129] A knowledge graph is constructed from the knowledge fusion results to obtain the knowledge graph of the target power scenario.
[0130] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device can be a robot or other electronic device, and may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. The processor 410, communication interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a knowledge graph construction method for the power scenario, including:
[0131] Text and image data of the target power scenario are acquired, and entity recognition extraction is performed on the text and image data of the target power scenario based on Bidirectional Long Short Time Memory Network (BiLSTM) and Conditional Random Field (CRF) to obtain multimodal data entity recognition results.
[0132] Relationships are extracted based on the multimodal data entity recognition results to determine the corresponding target triples; wherein, the target triples include: entities and the association relationships between entities;
[0133] Based on the target triples, knowledge fusion and knowledge graph construction are performed to obtain the knowledge graph of the target power scenario.
[0134] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0135] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program, which can be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer is able to execute the knowledge graph construction method for power scenarios provided by the above methods, including:
[0136] Text and image data of the target power scenario are acquired, and entity recognition extraction is performed on the text and image data of the target power scenario based on Bidirectional Long Short Time Memory Network (BiLSTM) and Conditional Random Field (CRF) to obtain multimodal data entity recognition results.
[0137] Relationships are extracted based on the multimodal data entity recognition results to determine the corresponding target triples; wherein, the target triples include: entities and the association relationships between entities;
[0138] Based on the target triples, knowledge fusion and knowledge graph construction are performed to obtain the knowledge graph of the target power scenario.
[0139] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the knowledge graph construction method for power scenarios provided by the above methods, including:
[0140] Text and image data of the target power scenario are acquired, and entity recognition extraction is performed on the text and image data of the target power scenario based on Bidirectional Long Short Time Memory Network (BiLSTM) and Conditional Random Field (CRF) to obtain multimodal data entity recognition results.
[0141] Relationships are extracted based on the multimodal data entity recognition results to determine the corresponding target triples; wherein, the target triples include: entities and the association relationships between entities;
[0142] Based on the target triples, knowledge fusion and knowledge graph construction are performed to obtain the knowledge graph of the target power scenario.
[0143] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0144] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a knowledge graph for a power scenario, characterized in that, include: Text and image data of the target power scenario are acquired, and entity recognition extraction is performed on the text and image data of the target power scenario based on Bidirectional Long Short Time Memory Network (BiLSTM) and Conditional Random Field (CRF) to obtain multimodal data entity recognition results. Relationships are extracted based on the multimodal data entity recognition results to determine the corresponding target triples; wherein, the target triples include: entities and the association relationships between entities; Based on the target triples, knowledge fusion and knowledge graph construction are performed to obtain the knowledge graph of the target power scenario; The process involves acquiring text and image data of the target power scenario, and then performing entity recognition extraction on the text and image data based on a Bidirectional Long Short-Term Memory (BiLSTM) network and a Conditional Random Field (CRF) to obtain multimodal data entity recognition results, including: The text data of the target power scenario is acquired, and the first entity recognition result corresponding to the text data is extracted based on the bidirectional long short-term memory network BiLSTM and the conditional random field CRF. Image data of the target power scene is acquired, and a second entity recognition result corresponding to the image data is extracted based on an entity and image semantic matching method. The entity and image semantic matching method includes: extracting entity bounding boxes from the image data using a pre-trained Mask R-CNN model, and outputting the entities present within the bounding boxes using a multi-label image classification model. The entities present within the bounding boxes are the entities in the first entity recognition result. The multi-label image classification model is a model based on a spatial transformation layer and an LSTM network. The spatial transformation layer locates attention regions from the convolutional feature map, and the LSTM network predicts semantic label scores on the located regions sequentially and captures the global dependencies of the located regions. The multimodal data entity recognition result is constituted based on the first entity recognition result and the second entity recognition result.
2. The knowledge graph construction method for power scenarios according to claim 1, characterized in that, The process of acquiring text data of the target power scenario, and extracting the first entity recognition result corresponding to the text data based on the Bidirectional Long Short-Term Memory (BiLSTM) network and the Conditional Random Field (CRF), includes: The text data of the target power scenario is acquired, and the text data is preprocessed to generate a first embedding vector corresponding to each word; wherein, the first embedding vector is obtained by concatenating the high-dimensional vector of the word generated based on the first pre-training algorithm and the unique encoding identifier of the word; Based on the first embedding vector corresponding to each word, the contextual sentence-level features in the text data are extracted through the bidirectional long short-term memory network BiLSTM, and a complete labeled prediction sequence is obtained. The neural unit dropout rate is set to obtain the predicted probability value of the label corresponding to each word; wherein, the neural unit dropout rate is used to prevent overfitting; A Conditional Random Field (CRF) is constructed, and based on the complete labeled prediction sequence and the predicted probability value of the label corresponding to each word, the first entity recognition result corresponding to the text data is obtained.
3. The knowledge graph construction method for power scenarios according to claim 1, characterized in that, The process of acquiring image data of the target power scene and extracting a second entity recognition result corresponding to the image data based on an entity and image semantic matching method includes: Image data of the target power scene is acquired, and image preprocessing is performed on the image data to obtain an overall image representation; Based on the entity and image semantic matching method, target information is extracted from the overall image representation, and a second entity recognition result corresponding to the image data is obtained based on the target information.
4. The knowledge graph construction method for power scenarios according to claim 1, characterized in that, The relation extraction includes: an embedding layer, a bidirectional long short-term memory network (BiLSTM) layer, and a conditional random field (CRF) layer; The step of extracting relationships based on the multimodal data entity recognition results and determining the corresponding target triples includes: For the embedding layer, a second embedding vector is generated for each word in the target power scene data, and a second embedding vector sequence is generated; wherein, the second embedding vector is obtained by concatenating the high-dimensional vector generated by the first pre-training algorithm and the high-dimensional vector generated by the second pre-training algorithm for each word; The second embedded vector sequence is input into the Bidirectional Long Short-Term Memory (BiLSTM) layer to obtain a vector sequence of the target dimension; wherein, the target dimension is the number of relation labels; The vector sequence of the target dimension is input into the Conditional Random Field (CRF) layer for sequence labeling to obtain the output label sequence, and the target triplet is identified based on the output label sequence.
5. The knowledge graph construction method for power scenarios according to claim 1, characterized in that, The knowledge fusion includes: entity normalization, entity alignment, and relation fusion; The process of knowledge fusion and knowledge graph construction based on the target triples to obtain the knowledge graph of the target power scenario includes: Based on the target triples, entity normalization, entity alignment, and relation fusion are performed respectively to obtain the knowledge fusion result; wherein, entity normalization is: removing duplicate similar entities and simplifying entities according to semantic similarity; entity alignment is: fusing concepts in the entity set; relation fusion is: removing redundant relations based on the result of entity normalization; A knowledge graph is constructed from the knowledge fusion results to obtain the knowledge graph of the target power scenario.
6. A knowledge graph construction device for a power scenario, characterized in that, include: The entity recognition module is used to acquire text data and image data of the target power scene, and to perform entity recognition extraction on the text data and image data of the target power scene based on the bidirectional long short time memory network BiLSTM and the conditional random field CRF to obtain multimodal data entity recognition results. The relationship extraction module is used to extract relationships based on the multimodal data entity recognition results and determine the corresponding target triples; wherein, the target triples include: entities and the association relationships between entities; The knowledge graph construction module is used to perform knowledge fusion and knowledge graph construction based on the target triples to obtain the knowledge graph of the target power scenario. The entity recognition module is specifically used for: The text data of the target power scenario is acquired, and the first entity recognition result corresponding to the text data is extracted based on the bidirectional long short-term memory network BiLSTM and the conditional random field CRF. Image data of the target power scene is acquired, and a second entity recognition result corresponding to the image data is extracted based on an entity and image semantic matching method. The entity and image semantic matching method includes: extracting entity bounding boxes from the image data using a pre-trained Mask R-CNN model, and outputting the entities present within the bounding boxes using a multi-label image classification model. The entities present within the bounding boxes are the entities in the first entity recognition result. The multi-label image classification model is a model based on a spatial transformation layer and an LSTM network. The spatial transformation layer locates attention regions from the convolutional feature map, and the LSTM network predicts semantic label scores on the located regions sequentially and captures the global dependencies of the located regions. The multimodal data entity recognition result is constituted based on the first entity recognition result and the second entity recognition result.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the knowledge graph construction method for the power scenario as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the knowledge graph construction method for the power scenario as described in any one of claims 1 to 5.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the knowledge graph construction method for the power scenario as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Electric power communication network knowledge graph construction method based on BERT model
CN112613314A
Operation safety risk identification method based on multi-modal knowledge graph
CN117408507A