Power equipment multi-modal relation extraction method and device based on deep neural network
By adopting a multimodal relationship extraction method based on deep neural network in power equipment news reports, combining image and text information, the problem of poor adaptability of the relationship extraction model in the prior art in the power field is solved, and higher relationship extraction accuracy and effect are achieved.
Patent Information
- Application Number
- CN202510107607.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-20
AI Technical Summary
In the news reports of power equipment, the existing relationship extraction model is difficult to effectively adapt to the characteristics of the power field due to the lack of semantics, lack of entities and many vocabulary in professional fields, resulting in the unsatisfactory accuracy and effect of relationship extraction.
The multimodal relationship extraction method of power equipment based on deep neural network is adopted, and the entity relationship recognition is performed using a pre-trained model. The model includes an image description generation layer, a dependency tree structure feature generation layer, a semantic feature alignment layer, a feature fusion layer, a BiLSTM model and a CRF model. Combining image information and text information, through feature fusion and alignment, the accuracy of relationship extraction is improved.
By integrating image information and dependency tree features, the missing semantic information in the text is supplemented, the accurate extraction of entity relationships is improved, the adaptability and effect of the relationship extraction model is enhanced, and the accuracy of identification of entity relationships in power equipment news reports is significantly improved.
Smart Images

Figure CN120181208A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal information extraction, and particularly relates to a method and device for extracting multimodal relationships of power equipment based on a deep neural network. Background Art
[0002] With the deep integration of Internet and artificial intelligence technologies with the power industry, using digital and information technologies to integrate power science and technology data resources, obtain industry knowledge, and support the construction of a new power system with a digital and intelligent power grid has become an important research direction for improving the safe and stable operation level of the power grid and improving grid services. In this process, the combination of smart grid and artificial intelligence deep learning is an important part of the digital transformation of the power industry. It promotes the intelligence and efficiency of the digital power grid through real-time monitoring and analysis of important power equipment and related public opinion information. At the same time, therefore, since the construction and operation of the smart grid involve collaborative operations in multiple departmental fields and a large amount of relevant data and information. The power department urgently needs to automatically process and analyze the correlation relationships of these massive data and information to assist in the construction and operation management of the smart grid. Relation Extraction (RE) is an important subtask of the information extraction task, which can help automatically mine the semantic relationships between entities from massive data, thereby improving the ability to understand and analyze information. Therefore, it is of great significance to improve the relation extraction model in combination with the characteristics of the power equipment news reporting field, and then assist in the construction of the smart grid, so as to realize the intelligence and efficient operation of the digital power grid.
[0003] Relation extraction is a fundamental task in information extraction. The current methods for relation extraction mainly fall into three categories: traditional rule-based methods, traditional machine learning methods, and finally deep learning-based methods. Rule-based methods require manually creating rules for each entity relationship. Although this approach can achieve high precision, the recall rate is relatively low, and it also consumes a large amount of human and time costs. With the widespread use of machine learning algorithms, methods based on artificial features have achieved better performance than rule-based methods. It regards extraction as a classification problem. First, entity relationships are predefined, and then a machine learning model is used to determine the relationship category between two entities. Currently, there are supervised learning and semi-supervised learning. However, regardless of the method, the criteria for artificial feature extraction overly rely on existing natural language processing tools and are prone to error propagation. In recent years, with the development of deep learning, neural networks can better mine the implicit semantic features of text without relying on artificial feature extraction, greatly improving the accuracy of relation extraction and achieving good results in relation extraction tasks. The BILSTM-CRF model is the mainstream deep learning framework for current sequence labeling. By combining the advantages of the BILSTM network and the CRF model, it can more accurately identify the boundaries of relation semantic chunks and extract more relation tuples than other models.
[0004] Although the research on relation extraction technology has become relatively mature in recent years, for information sources such as news reports on power equipment, when the news text lacks coherence, the accuracy of relation extraction will drop sharply, and the problem of missing entity extraction will also be very significant. Therefore, the performance of relation extraction models applied to actual data is not ideal. As a subtask of information extraction, relation extraction can lead to unsatisfactory results in converting data into knowledge in the power field due to error accumulation. The implementation of relation extraction in the power field enables power safety work to gain the ability to mine and analyze useful knowledge from large-scale text data. However, due to characteristics such as semantic missing, entity missing, and a large number of professional domain vocabulary in news reports on power equipment, the relation extraction model cannot promptly adapt to the characteristics of the power field and be transformed into social productivity, which has become a difficulty in relation extraction. Summary of the Invention
[0005] To overcome the above defects, the present invention proposes a method and device for multi-modal relation extraction of power equipment based on a deep neural network.
[0006] In a first aspect, there is provided a method for multi-modal relation extraction of power equipment based on a deep neural network. The method for multi-modal relation extraction of power equipment based on a deep neural network includes:
[0007] Obtain the text and image objects to be recognized;
[0008] Taking the image and text object to be identified as the input of the pre-trained multimodal relationship extraction model of electric power equipment, and obtaining the entity relationship recognition result of the image and text object to be identified output by the pre-trained multimodal relationship extraction model of electric power equipment;
[0009] The training process of the pre-trained multimodal relationship extraction model of electric power equipment includes:
[0010] Use a web crawler for news reports on power equipment to crawl power equipment news report data;
[0011] Using public news data in the public domain as basic data, and combining the power equipment news report data to build a power equipment news report forecast resource library;
[0012] Using the power equipment news report forecast resource library to construct training data;
[0013] The training data is used to train the multimodal relationship extraction model of electric power equipment to obtain the pre-trained multimodal relationship extraction model of electric power equipment.
[0014] Preferably, the multimodal relationship extraction model of power equipment includes: an image description generation layer, a dependency tree structure feature generation layer, a dependency tree structure feature alignment layer, a semantic feature alignment layer, a feature fusion layer, a first BiLSTM model, a sigmoid activation function layer, a BERT model, a second BiLSTM model, and a CRF model.
[0015] Furthermore, the image description generation layer is connected to the dependency tree structure feature generation layer and the semantic feature alignment layer respectively;
[0016] The dependency tree structure feature generation layer is connected to the dependency tree structure feature alignment layer;
[0017] The dependency tree structure feature alignment layer, the semantic feature alignment layer and the BERT model are all connected to the feature fusion layer;
[0018] The feature fusion layer, the first BiLSTM model, and the sigmoid activation function layer are connected in sequence;
[0019] The output features of the sigmoid activation function layer are concatenated with the output features of the BERT model and input into the second BiLSTM model;
[0020] The second BiLSTM model is connected to the CRF model.
[0021] Furthermore, the image description generation layer is used to generate description text corresponding to the visual content in the power equipment news report forecast resource library using an image description generation algorithm;
[0022] The dependency tree structure feature generation layer is used to generate the dependency tree structure features of the description text corresponding to the visual content and the dependency tree structure features of the actual text corresponding to the visual content by using a dependency tree generation algorithm;
[0023] The dependency tree structure feature alignment layer is used to perform structural alignment on the dependency tree structure features of the description text corresponding to the visual content and the dependency tree structure features of the actual text corresponding to the visual content to obtain a graph structure feature matrix;
[0024] The semantic feature alignment layer is used to perform semantic alignment on the description text corresponding to the visual content and the actual text corresponding to the visual content through an attention mechanism to obtain a semantic alignment weight matrix;
[0025] The BERT model is used to generate the semantic features of the description text corresponding to the visual content and the actual text corresponding to the visual content;
[0026] The feature fusion layer is used to fuse the graph structure feature matrix, the semantic alignment weight matrix, and the semantic features of the description text corresponding to the visual content and the actual text corresponding to the visual content to obtain a fused feature.
[0027] Further, the semantic alignment weight matrix is as follows:
[0028] β = softmax(QK T / d 1 / 2 )
[0029] In the above formula, β is the semantic alignment weight matrix, softmax is the activation function, Q is the query vector, K is the key vector, T is the transpose symbol, and d is the dimension of the key vector.
[0030] Further, the graph structure feature matrix is as follows:
[0031] α = (node ij ) |R1|×|R2|
[0032] In the above formula, α is the graph structure feature matrix, node ijis the element in the \(i\)-th row and \(j\)-th column of the graph structure feature matrix, \(R_1\) is the set of nodes in the dependency tree structure feature of the descriptive text corresponding to the visual content, and \(R_2\) is the set of nodes in the dependency tree structure feature of the actual text corresponding to the visual content. Among them, if the similarity between the \(i\)-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content and the \(j\)-th node in the dependency tree structure feature of the actual text corresponding to the visual content is the maximum value of the similarities between the \(i\)-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content and all nodes in the dependency tree structure feature of the actual text corresponding to the visual content, then the element in the \(i\)-th row and \(j\)-th column of the graph structure feature matrix is equal to the similarity between the \(i\)-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content and the \(j\)-th node in the dependency tree structure feature of the actual text corresponding to the visual content; otherwise, the element in the \(i\)-th row and \(j\)-th column of the graph structure feature matrix is equal to 0.
[0033] Further, the similarity between the \(i\)-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content and the \(j\)-th node in the dependency tree structure feature of the actual text corresponding to the visual content is as follows:
[0034]
[0035] In the above formula, \(sim(i, j)\) is the similarity between the \(i\)-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content and the \(j\)-th node in the dependency tree structure feature of the actual text corresponding to the visual content, \(e\) is the natural constant, \(y\) is a scalar parameter, \(d\) i is the structure vector of the \(i\)-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content, \(d\) j is the structure vector of the \(j\)-th node in the dependency tree structure feature of the actual text corresponding to the visual content.
[0036] Further, the structure vector of the \(i\)-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content is as follows:
[0037]
[0038] The structure vector of the \(j\)-th node in the dependency tree structure feature of the actual text corresponding to the visual content is as follows:
[0039]
[0040] In the above formula, is the vector of the node degrees in the set of \(k\)-hop neighbor nodes of the \(i\)-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content, is the vector of the node degrees in the set of \(k\)-hop neighbor nodes of the \(j\)-th node in the dependency tree structure feature of the actual text corresponding to the visual content, and \(K\) is a preset parameter.
[0041] Furthermore, the fusion features are as follows:
[0042] Q=(α T +β)V
[0043] In the above formula, Q is the fusion feature, α is the graph structure feature matrix, T is the transposition symbol, β is the semantic alignment weight matrix, and V is the semantic feature of the description text corresponding to the visual content and the actual text corresponding to the visual content.
[0044] In a second aspect, a device for extracting multimodal relations of electric power equipment based on a deep neural network is provided, wherein the device for extracting multimodal relations of electric power equipment based on a deep neural network comprises:
[0045] An acquisition module is used to acquire the graphic object to be identified;
[0046] A recognition module, used to take the graphic object to be recognized as the input of the pre-trained multimodal relationship extraction model of electric power equipment, and obtain the entity relationship recognition result of the graphic object to be recognized output by the pre-trained multimodal relationship extraction model of electric power equipment;
[0047] The training process of the pre-trained multimodal relationship extraction model of electric power equipment includes:
[0048] Use a web crawler for news reports on power equipment to crawl power equipment news report data;
[0049] Using public news data in the public domain as basic data, and combining the power equipment news report data to build a power equipment news report forecast resource library;
[0050] Using the power equipment news report forecast resource library to construct training data;
[0051] The training data is used to train the multimodal relationship extraction model of electric power equipment to obtain the pre-trained multimodal relationship extraction model of electric power equipment.
[0052] In a third aspect, a computer device is provided, comprising: one or more processors;
[0053] The processor is used to store one or more programs;
[0054] When the one or more programs are executed by the one or more processors, the method for extracting multimodal relationships of power equipment based on a deep neural network is implemented.
[0055] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed, the method for extracting multimodal relationships of power equipment based on a deep neural network is implemented.
[0056] The above one or more technical solutions of the present invention have at least one or more of the following beneficial effects:
[0057] The present invention provides a method and device for extracting multimodal relations of electric power equipment based on a deep neural network, comprising: obtaining a graphic object to be identified; using the graphic object to be identified as the input of a pre-trained multimodal relation extraction model for electric power equipment, and obtaining an entity relationship recognition result of the graphic object to be identified output by the pre-trained multimodal relation extraction model for electric power equipment; wherein the training process of the pre-trained multimodal relation extraction model for electric power equipment includes: crawling electric power equipment news report data using a web crawler for news reports on electric power equipment; using public domain news disclosure data as basic data, and building an expected resource library for electric power equipment news reports in combination with the electric power equipment news report data; building training data using the expected resource library for electric power equipment news reports; and training the multimodal relation extraction model for electric power equipment using the training data to obtain the pre-trained multimodal relation extraction model for electric power equipment. The technical solution provided by the present invention can automatically complete the extraction of specified entity relationships in different electric power equipment news reports, specifically:
[0058] The technical solution provided by the present invention supplements the missing semantic information in the text by integrating image information, so that entities can be extracted more accurately; by introducing a dependency tree and performing feature similarity calculation, the entity missing problem of the overall sense is compensated; the description-based method is used to integrate the entity relationship information of the image into the entire relationship task, thereby greatly improving the accuracy of relationship extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a flowchart of the main steps of a method for extracting multimodal relations of power equipment based on a deep neural network according to an embodiment of the present invention;
[0060] Figure 2 It is a main structural block diagram of the multimodal relationship extraction model of power equipment in an embodiment of the present invention. DETAILED DESCRIPTION
[0061] The specific implementation modes of the present invention will be further described in detail below in conjunction with the accompanying drawings.
[0062] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0063] As disclosed in the background art, with the deep integration of Internet and artificial intelligence technologies into the power industry, using digital and information technologies to integrate power science and technology data resources, obtain industry knowledge, and support the construction of a new power system with a digital and intelligent power grid has become an important research direction for improving the safe and stable operation level of the power grid and enhancing power grid services. In this process, the combination of smart grid and artificial intelligence deep learning is an important part of the digital transformation of the power industry. It promotes the intelligence and efficiency of the digital power grid through the real-time monitoring and analysis of important power equipment and related public opinion information. At the same time, therefore, since the construction and operation of the smart grid involve collaborative operations in multiple departmental fields and a large amount of relevant data and information. The power department urgently needs to automatically process and analyze the correlation relationships of these massive data and information to assist in the construction and operation management of the smart grid. Relation Extraction (RE) is an important subtask of information extraction tasks, which can help automatically mine the semantic relationships between entities from massive data, thereby improving the ability to understand and analyze information. Therefore, it is of great significance to improve the relation extraction model in combination with the characteristics of the power equipment news reporting field, and then assist in the construction of the smart grid, so as to realize the intelligence and efficient operation of the digital power grid.
[0064] Relation extraction is a basic task of information extraction. The current methods of relation extraction are mainly divided into traditional rule-based methods, traditional machine learning methods, and finally deep learning-based methods. Rule-based methods require manual creation of rules for each entity relationship. Although this mode can achieve a high precision rate, the recall rate is relatively low, and it also consumes a large amount of human and time costs. With the wide use of machine learning algorithms, the method based on artificial features has achieved better performance than rule-based methods. It regards extraction as a classification problem. First, entity relationships are predefined, and then a machine learning model is used to judge the relationship category between two entities. Currently, there are supervised learning and semi-supervised learning. However, no matter which method, the standard of artificial feature extraction overly relies on existing natural language processing tools and error propagation will also occur. In recent years, with the development of deep learning, neural networks can better mine the implicit semantic features of texts without relying on artificial feature extraction, which greatly improves the accuracy of relation extraction and achieves good results in relation extraction tasks. The BILSTM-CRF model is the mainstream deep learning framework for current sequence labeling. Combining the advantages of the BILSTM network and the CRF model, it can more accurately identify the boundaries of relation semantic blocks and extract more relation tuples than other models.
[0065] Although the research on relation extraction technology has matured in recent years, for information sources such as power equipment news reports, when the news text lacks coherence, the accuracy of relation extraction will drop sharply, and the problem of missing entity extraction will be very significant. Therefore, the performance of the relationship extraction model applied to actual data is not ideal. As a subtask of information extraction, relationship extraction will lead to unsatisfactory results in converting data into knowledge in the power field through error accumulation. The implementation of relationship extraction in the power field can enable power safety work to acquire the ability to mine and analyze useful knowledge in large-scale text data. However, due to the characteristics of semantic missing, entity missing, and many professional vocabulary in power equipment news reports, the relationship extraction model cannot adapt to the characteristics of the power field and transform into social productivity in a timely manner, which has become a difficulty in relationship extraction.
[0066] In order to improve the above-mentioned problems, the present invention provides a method and device for extracting multimodal relations of electric power equipment based on a deep neural network, including: obtaining a graphic object to be identified; using the graphic object to be identified as the input of a pre-trained multimodal relation extraction model for electric power equipment, and obtaining the entity relationship recognition result of the graphic object to be identified output by the pre-trained multimodal relation extraction model for electric power equipment; wherein the training process of the pre-trained multimodal relation extraction model for electric power equipment includes: crawling electric power equipment news report data using a web crawler for news reports on electric power equipment; using public domain news disclosure data as basic data, and combining the electric power equipment news report data to construct an expected resource library for electric power equipment news reports; using the expected resource library for electric power equipment news reports to construct training data; using the training data to train the multimodal relation extraction model for electric power equipment, and obtaining the pre-trained multimodal relation extraction model for electric power equipment. The technical solution provided by the present invention can automatically complete the extraction of specified entity relationships in different electric power equipment news reports, specifically:
[0067] The technical solution provided by the present invention supplements the missing semantic information in the text by integrating image information, so that entities can be extracted more accurately; by introducing dependency trees and performing feature similarity calculations, the entity missing problem of the overall perception is compensated; the description-based method is used to integrate the entity relationship information of the image into the entire relationship task, thereby greatly improving the accuracy of relationship extraction. The above solution is described in detail below.
[0068] Example 1
[0069] See attached Figure 1 , Figure 1 FIG. 1 is a flow chart of the main steps of a method for extracting multimodal relations of power equipment based on a deep neural network according to an embodiment of the present invention. Figure 1 As shown, the method for extracting multimodal relations of electric power equipment based on a deep neural network in an embodiment of the present invention mainly includes the following steps:
[0070] Step S101: obtaining a graphic object to be identified;
[0071] Step S102: using the image and text object to be identified as the input of the pre-trained multimodal relationship extraction model of electric power equipment, and obtaining the entity relationship recognition result of the image and text object to be identified output by the pre-trained multimodal relationship extraction model of electric power equipment;
[0072] The training process of the pre-trained multimodal relationship extraction model of electric power equipment includes:
[0073] Use a web crawler for power equipment news reports to crawl power equipment news report data; build a web crawler for power equipment news report text and images based on Python programming, and the crawling scope includes power equipment standard changes, technological frontier developments, fault reports, etc.;
[0074] Using public news data in the public domain as basic data, and combining the power equipment news report data to build a power equipment news report forecast resource library;
[0075] Using the power equipment news report forecast resource library to construct training data;
[0076] The training data is used to train the multimodal relationship extraction model of electric power equipment to obtain the pre-trained multimodal relationship extraction model of electric power equipment.
[0077] In this embodiment, the multimodal relationship extraction model of the power equipment is as follows: Figure 2 As shown, it includes: image description generation layer, dependency tree structure feature generation layer, dependency tree structure feature alignment layer, semantic feature alignment layer, feature fusion layer, first BiLSTM (Bidirectional Long Short-Term Memory) model, sigmoid activation function layer, BERT model (Bidirectional Encoder Representations from Transformers, a pre-trained language model based on Transformer architecture), second BiLSTM model, CRF (conditional random field) model. In the figure, Embeddings refers to the Embedding vector result of the word output by the BERT model, and Embedding refers to the process of mapping high-dimensional data (such as text, pictures, videos, etc.) to low-dimensional space.
[0078] In one embodiment, the image description generation layer is connected to the dependency tree structure feature generation layer and the semantic feature alignment layer respectively;
[0079] The dependency tree structure feature generation layer is connected to the dependency tree structure feature alignment layer;
[0080] The dependency tree structure feature alignment layer, the semantic feature alignment layer, and the BERT model are all connected to the feature fusion layer;
[0081] The feature fusion layer, the first BiLSTM model, and the sigmoid activation function layer are connected in sequence;
[0082] The output features of the sigmoid activation function layer and the output features of the BERT model are concatenated and then input into the second BiLSTM model;
[0083] The second BiLSTM model is connected to the CRF model.
[0084] In one embodiment, the image description generation layer is used to generate a description text corresponding to the visual content in the power equipment news report prediction resource library by using an image description generation algorithm;
[0085] Specifically, a basic convolutional neural network CNN - recurrent neural network RNN framework is adopted. First, the CNN performs object detection to obtain multiple picture frames and their corresponding features, and extracts image feature vectors. When predicting the first word, the image feature vector is used as an input to the RNN hidden layer, and at the same time, the RNN input layer is the feature corresponding to the start flag. When the predicted word is not the first one, the output of the previous hidden layer of the RNN is used as an input to the current hidden layer, and at the same time, the RNN input layer is the feature vector of the word predicted by the RNN at the previous moment. Finally, the previous step is repeated until the end flag is encountered. Finally, words are generated one by one and combined into a descriptive sentence with a certain grammatical structure.
[0086] The dependency tree structure feature generation layer is used to generate the dependency tree structure features of the description text corresponding to the visual content and the dependency tree structure features of the actual text corresponding to the visual content by using a dependency tree generation algorithm;
[0087] The dependency tree structure feature alignment layer is used to perform structural alignment on the dependency tree structure features of the description text corresponding to the visual content and the dependency tree structure features of the actual text corresponding to the visual content to obtain a graph structure feature matrix;
[0088] The semantic feature alignment layer is used to perform semantic alignment on the description text corresponding to the visual content and the actual text corresponding to the visual content through an attention mechanism to obtain a semantic alignment weight matrix;
[0089] The BERT model (Bidirectional Encoder Representations from Transformers, a pre-trained language model based on the Transformer architecture) is used to generate semantic features of the description text corresponding to the visual content and the actual text corresponding to the visual content; before the description text corresponding to the visual content and the actual text corresponding to the visual content are input, the tasks of annotating the text with "[CLS]" at the head of the sequence and "[SEP]" at the end of the sequence are performed. The CLS token is an abbreviation of "Classification" and is used to represent the main idea or central thought in a sentence. In the pre-training stage of BERT, the model learns the representation of language by predicting the next word in a sentence. However, for some tasks, we may be more concerned with the main idea of the whole sentence rather than the prediction of individual words. Therefore, the CLS token is used as the input token for the sentence main idea classification task. When we use BERT for the sentence main idea classification task in the fine-tuning stage, the representation of the CLS token is used as the output. The SEP token is an abbreviation of "Separator" and is used to separate different sentences or sequences. In the input of BERT, each sentence is represented as a sequence, and these sequences are separated by the SEP token. This allows the model to process multiple sentences or sequences as a single input without having to process each sentence or sequence independently. In some tasks, such as question answering or text generation, using the SEP token can help the model better understand the structure of the input text.
[0090] The feature fusion layer is used to fuse the graph structure feature matrix, the semantic alignment weight matrix, and the semantic features of the description text corresponding to the visual content and the actual text corresponding to the visual content to obtain fused features.
[0091] The input of the relationship extraction model based on BERT-BiLSTM-CRF composed of the BERT model, the second BiLSTM model, and the CRF model obtains all the relationships existing in the sentence through the entity relationship extraction model, and then entity pairs are extracted for each relationship in the sentence, that is, the main entity and the object entity corresponding to the relationship are extracted.
[0092] In the entity relationship extraction model, first, in the BERT network, characters obtain the token embedding tensor, the position encoding tensor, and the sentence chunk tensor through the token embedding layer, the segment embedding layer, and the position embedding layer.
[0093] When the sentence passes through the BiLSTM layer, only the correspondence between the text sequence and the labeled BIO tags is obtained. Finally, after passing through the CRF layer, the obtained predicted tags are constrained to reduce the number of invalid predicted tags, thereby obtaining the globally optimal BIO tag sequence.
[0094] There are two types of scores in the CRF layer. One is the emission matrix obtained through the BILSTM layer, which is the probability that a word token in a sentence is marked with each BIO tag, that is, the emission probability.
[0095] After obtaining the emission probability and transition probability in the CRF layer, the Viterbi algorithm is used to find the shortest path and obtain the predicted label corresponding to each character in the sentence. The main entity and object entity corresponding to the relationship are marked in this label, so that the corresponding entity relationship triple can be extracted.
[0096] In Figure 2 In the model diagram shown, first, image caption generation is realized through CNN-RNN. The text after image captioning and the input text are, on the one hand, represented by the pre-trained BERT model for semantic features, and on the other hand, parsed by the dependency tree tool. The generated dependency tree and text feature vectors are aligned using the graph alignment principle and text alignment principle respectively. After alignment, they are concatenated to generate a fused feature vector. After the full text vector is dot-multiplied, a feature representation vector is obtained. The fused feature vector and the feature representation vector are aligned and then spliced into the vector model after BERT training. Subsequently, the corresponding relationship between the text sequence and the marked BIO tags is obtained through BILSTM, and finally, the BIO tag sequence of the relationship is output through the CRF layer. B-COM represents the start of an entity, and I-COM represents the inside of an entity.
[0097] In one embodiment, the semantic alignment weight matrix is as follows:
[0098] β = softmax(QK T / d 1 / 2 )
[0099] In the above formula, β is the semantic alignment weight matrix, softmax is the activation function, Q is the query vector, K is the key vector, T is the transpose symbol, and d is the dimension of the key vector.
[0100] In one embodiment, the graph structure feature matrix is as follows:
[0101] α = (node ij ) |R1|×|R2|
[0102] In the above formula, α is the graph structure feature matrix, node ijis the element in the \(i\)-th row and \(j\)-th column of the graph structure feature matrix. \(R_1\) is the set of nodes in the dependency tree structure feature of the descriptive text corresponding to the visual content, and \(R_2\) is the set of nodes in the dependency tree structure feature of the actual text corresponding to the visual content. Among them, if the similarity between the \(i\)-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content and the \(j\)-th node in the dependency tree structure feature of the actual text corresponding to the visual content is the maximum value of the similarities between the \(i\)-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content and all nodes in the dependency tree structure feature of the actual text corresponding to the visual content, then the element in the \(i\)-th row and \(j\)-th column of the graph structure feature matrix is equal to the similarity between the \(i\)-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content and the \(j\)-th node in the dependency tree structure feature of the actual text corresponding to the visual content; otherwise, the element in the \(i\)-th row and \(j\)-th column of the graph structure feature matrix is equal to 0.
[0103] In one embodiment, the similarity between the \(i\)-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content and the \(j\)-th node in the dependency tree structure feature of the actual text corresponding to the visual content is as follows:
[0104]
[0105] In the above formula, \(sim(i, j)\) is the similarity between the \(i\)-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content and the \(j\)-th node in the dependency tree structure feature of the actual text corresponding to the visual content, \(e\) is the natural constant, \(y\) is a scalar parameter, \(d\) i is the structure vector of the \(i\)-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content, \(d\) j is the structure vector of the \(j\)-th node in the dependency tree structure feature of the actual text corresponding to the visual content.
[0106] In one embodiment, the structure vector of the \(i\)-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content is as follows:
[0107]
[0108] The structure vector of the \(j\)-th node in the dependency tree structure feature of the actual text corresponding to the visual content is as follows:
[0109]
[0110] In the above formula, is the vector of the node degrees in the set of \(k\)-hop neighbor nodes of the \(i\)-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content, is the vector of the node degrees in the set of \(k\)-hop neighbor nodes of the \(j\)-th node in the dependency tree structure feature of the actual text corresponding to the visual content, and \(K\) is a preset parameter.
[0111] In one embodiment, the fusion features are as follows:
[0112] Q=(α T +β)V
[0113] In the above formula, Q is the fusion feature, α is the graph structure feature matrix, T is the transposition symbol, β is the semantic alignment weight matrix, and V is the semantic feature of the description text corresponding to the visual content and the actual text corresponding to the visual content.
[0114] Example 2
[0115] Based on the same inventive concept, the present invention also provides a device for extracting multimodal relations of electric power equipment based on a deep neural network, and the device for extracting multimodal relations of electric power equipment based on a deep neural network comprises:
[0116] An acquisition module is used to acquire the graphic object to be identified;
[0117] A recognition module, used to take the graphic object to be recognized as the input of the pre-trained multimodal relationship extraction model of electric power equipment, and obtain the entity relationship recognition result of the graphic object to be recognized output by the pre-trained multimodal relationship extraction model of electric power equipment;
[0118] The training process of the pre-trained multimodal relationship extraction model of electric power equipment includes:
[0119] Use a web crawler for news reports on power equipment to crawl power equipment news report data;
[0120] Using public news data in the public domain as basic data, and combining the power equipment news report data to build a power equipment news report forecast resource library;
[0121] Using the power equipment news report forecast resource library to construct training data;
[0122] The training data is used to train the multimodal relationship extraction model of electric power equipment to obtain the pre-trained multimodal relationship extraction model of electric power equipment.
[0123] Preferably, the multimodal relationship extraction model of power equipment includes: an image description generation layer, a dependency tree structure feature generation layer, a dependency tree structure feature alignment layer, a semantic feature alignment layer, a feature fusion layer, a first BiLSTM model, a sigmoid activation function layer, a BERT model, a second BiLSTM model, and a CRF model.
[0124] Furthermore, the image description generation layer is connected to the dependency tree structure feature generation layer and the semantic feature alignment layer respectively;
[0125] The dependency tree structure feature generation layer is connected to the dependency tree structure feature alignment layer;
[0126] The dependency tree structure feature alignment layer, the semantic feature alignment layer, and the BERT model are all connected to the feature fusion layer;
[0127] The feature fusion layer, the first BiLSTM model, and the sigmoid activation function layer are connected in sequence;
[0128] The output features of the sigmoid activation function layer and the output features of the BERT model are concatenated and then input into the second BiLSTM model;
[0129] The second BiLSTM model is connected to the CRF model.
[0130] Furthermore, the image description generation layer is used to generate a description text corresponding to the visual content in the power equipment news report prediction resource library by using an image description generation algorithm;
[0131] The dependency tree structure feature generation layer is used to generate the dependency tree structure features of the description text corresponding to the visual content and the dependency tree structure features of the actual text corresponding to the visual content by using a dependency tree generation algorithm;
[0132] The dependency tree structure feature alignment layer is used to perform structural alignment on the dependency tree structure features of the description text corresponding to the visual content and the dependency tree structure features of the actual text corresponding to the visual content to obtain a graph structure feature matrix;
[0133] The semantic feature alignment layer is used to perform semantic alignment on the description text corresponding to the visual content and the actual text corresponding to the visual content through an attention mechanism to obtain a semantic alignment weight matrix;
[0134] The BERT model is used to generate the semantic features of the description text corresponding to the visual content and the actual text corresponding to the visual content;
[0135] The feature fusion layer is used to fuse the graph structure feature matrix, the semantic alignment weight matrix, and the semantic features of the description text corresponding to the visual content and the actual text corresponding to the visual content to obtain fused features.
[0136] Furthermore, the semantic alignment weight matrix is as follows:
[0137] β = softmax(QK T / d 1 / 2 )
[0138] In the above formula, β is the semantic alignment weight matrix, softmax is the activation function, Q is the query vector, K is the key vector, T is the transpose symbol, and d is the dimension of the key vector.
[0139] Furthermore, the graph structure feature matrix is as follows:
[0140] α = (node ij ) |R1|×|R2|
[0141] In the above formula, α is the graph structure feature matrix, node ij is the element in the i-th row and j-th column of the graph structure feature matrix, R1 is the set of nodes in the dependency tree structure feature of the descriptive text corresponding to the visual content, R2 is the set of nodes in the dependency tree structure feature of the actual text corresponding to the visual content. Among them, if the similarity between the i-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content and the j-th node in the dependency tree structure feature of the actual text corresponding to the visual content is the maximum value of the similarities between the i-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content and all nodes in the dependency tree structure feature of the actual text corresponding to the visual content, the element in the i-th row and j-th column of the graph structure feature matrix is equal to the similarity between the i-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content and the j-th node in the dependency tree structure feature of the actual text corresponding to the visual content; otherwise, the element in the i-th row and j-th column of the graph structure feature matrix is equal to 0.
[0142] Furthermore, the similarity between the i-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content and the j-th node in the dependency tree structure feature of the actual text corresponding to the visual content is as follows:
[0143]
[0144] In the above formula, sim(i, j) is the similarity between the i-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content and the j-th node in the dependency tree structure feature of the actual text corresponding to the visual content, e is the natural constant, y is a scalar parameter, d i is the structure vector of the i-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content, d j is the structure vector of the j-th node in the dependency tree structure feature of the actual text corresponding to the visual content.
[0145] Furthermore, the structure vector of the i-th node in the dependency tree structure feature of the descriptive text corresponding to the visual content is as follows:
[0146]
[0147] The structure vector of the j-th node in the dependency tree structure features of the actual text corresponding to the visual content is as follows:
[0148]
[0149] In the above formula, is the vector of node degrees in the set of k-hop neighbor nodes of the i-th node in the dependency tree structure features of the descriptive text corresponding to the visual content, is the vector of node degrees in the set of k-hop neighbor nodes of the j-th node in the dependency tree structure features of the actual text corresponding to the visual content, and K is a preset parameter.
[0150] Furthermore, the fusion feature is as follows:
[0151] Q = (α T + β)V
[0152] In the above formula, Q is the fusion feature, α is the graph structure feature matrix, T is the transpose symbol, β is the semantic alignment weight matrix, and V is the semantic feature of the descriptive text corresponding to the visual content and the actual text corresponding to the visual content.
[0153] Embodiment 3
[0154] Based on the same inventive concept, the present invention also provides a computer device, which includes a processor and a memory. The memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of a method for extracting multi-modal relationships of power equipment based on a deep neural network in the above embodiment.
[0155] Embodiment 4
[0156] Based on the same inventive concept, the present invention also provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device, used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and this storage space stores the operating system of the terminal. Moreover, in this storage space, there is also stored one or more instructions suitable for being loaded and executed by the processor. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the steps of a method for extracting multi-modal relationships of power equipment based on a deep neural network in the above embodiments.
[0157] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0158] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0159] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions in Figure 1 one flow or multiple flows and / or blocksFigure 1 The functions specified in one or more boxes.
[0160] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in one or more processes and / or boxes Figure 1 One process or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes.
[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent replacements. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A method for extracting multimodal relations of power equipment based on deep neural network, characterized in that: The method comprises: Get the image and text object to be recognized; Taking the image and text object to be identified as the input of the pre-trained multimodal relationship extraction model of electric power equipment, and obtaining the entity relationship recognition result of the image and text object to be identified output by the pre-trained multimodal relationship extraction model of electric power equipment; The training process of the pre-trained multimodal relationship extraction model of electric power equipment includes: Use a web crawler for news reports on power equipment to crawl power equipment news report data; Using public news data in the public domain as basic data, and combining the power equipment news report data to build a power equipment news report forecast resource library; Using the power equipment news report forecast resource library to construct training data; The training data is used to train the multimodal relationship extraction model of electric power equipment to obtain the pre-trained multimodal relationship extraction model of electric power equipment.
2. The method according to claim 1, characterized in that The multimodal relationship extraction model for electric power equipment includes: an image description generation layer, a dependency tree structure feature generation layer, a dependency tree structure feature alignment layer, a semantic feature alignment layer, a feature fusion layer, a first BiLSTM model, a sigmoid activation function layer, a BERT model, a second BiLSTM model, and a CRF model.
3. The method according to claim 2, characterized in that The image description generation layer is connected to the dependency tree structure feature generation layer and the semantic feature alignment layer respectively; The dependency tree structure feature generation layer is connected to the dependency tree structure feature alignment layer; The dependency tree structure feature alignment layer, the semantic feature alignment layer and the BERT model are all connected to the feature fusion layer; The feature fusion layer, the first BiLSTM model, and the sigmoid activation function layer are connected in sequence; The output features of the sigmoid activation function layer are concatenated with the output features of the BERT model and input into the second BiLSTM model; The second BiLSTM model is connected to the CRF model.
4. The method according to claim 3, characterized in that The image description generation layer is used to generate description text corresponding to the visual content in the power equipment news report forecast resource library using an image description generation algorithm; The dependency tree structure feature generation layer is used to generate dependency tree structure features of the description text corresponding to the visual content and dependency tree structure features of the actual text corresponding to the visual content using a dependency tree generation algorithm; The dependency tree structure feature alignment layer is used to perform structural alignment on the dependency tree structure features of the description text corresponding to the visual content and the dependency tree structure features of the actual text corresponding to the visual content to obtain a graph structure feature matrix; The semantic feature alignment layer is used to semantically align the description text corresponding to the visual content and the actual text corresponding to the visual content through an attention mechanism to obtain a semantic alignment weight matrix; The BERT model is used to generate semantic features of the description text corresponding to the visual content and the actual text corresponding to the visual content; The feature fusion layer is used to fuse the graph structure feature matrix, the semantic alignment weight matrix, and the semantic features of the description text corresponding to the visual content and the actual text corresponding to the visual content to obtain a fused feature.
5. The method according to claim 4, characterized in that The semantic alignment weight matrix is as follows: β=softmax(QK T / d 1 / 2 ) In the above formula, β is the semantic alignment weight matrix, softmax is the activation function, Q is the query vector, K is the key vector, T is the transposed symbol, and d is the dimension of the key vector.
6. The method according to claim 4, characterized in that The graph structure feature matrix is as follows: α=(node ij ) |R1|×|R2| In the above formula, α is the graph structure feature matrix, node ij is the element in the i-th row and j-th column in the graph structure feature matrix, R1 is the set of nodes in the dependency tree structure feature of the description text corresponding to the visual content, and R2 is the set of nodes in the dependency tree structure feature of the actual text corresponding to the visual content, wherein, if the similarity between the i-th node in the dependency tree structure feature of the description text corresponding to the visual content and the j-th node in the dependency tree structure feature of the actual text corresponding to the visual content is the maximum value of the similarities between the i-th node in the dependency tree structure feature of the description text corresponding to the visual content and all the nodes in the dependency tree structure feature of the actual text corresponding to the visual content, the element in the i-th row and j-th column in the graph structure feature matrix is equal to the similarity between the i-th node in the dependency tree structure feature of the description text corresponding to the visual content and the j-th node in the dependency tree structure feature of the actual text corresponding to the visual content, otherwise, the element in the i-th row and j-th column in the graph structure feature matrix is equal to 0.
7. The method according to claim 6, characterized in that The similarity between the i-th node in the dependency tree structure feature of the description text corresponding to the visual content and the j-th node in the dependency tree structure feature of the actual text corresponding to the visual content is as follows: In the above formula, sim(i,j) is the similarity between the i-th node in the dependency tree structure feature of the description text corresponding to the visual content and the j-th node in the dependency tree structure feature of the actual text corresponding to the visual content, e is a natural constant, y is a scalar parameter, and d i is the structural vector of the ith node in the dependency tree structure feature of the description text corresponding to the visual content, d j is the structural vector of the jth node in the dependency tree structure feature of the actual text corresponding to the visual content.
8. The method according to claim 7, characterized in that The structure vector of the i-th node in the dependency tree structure feature of the description text corresponding to the visual content is as follows: The structure vector of the jth node in the dependency tree structure feature of the actual text corresponding to the visual content is as follows: In the above formula, is the vector of node degrees in the set of k-hop neighbor nodes of the ith node in the dependency tree structure feature of the description text corresponding to the visual content, is the vector of node degrees in the set of k-hop neighbor nodes of the j-th node in the dependency tree structure feature of the actual text corresponding to the visual content, and K is a preset parameter.
9. The method according to claim 4, characterized in that The fusion features are as follows: Q=(a T +b)V In the above formula, Q is the fusion feature, α is the graph structure feature matrix, T is the transposition symbol, β is the semantic alignment weight matrix, and V is the semantic feature of the description text corresponding to the visual content and the actual text corresponding to the visual content.
10. A device for extracting multimodal relations of power equipment based on a deep neural network according to any one of claims 1 to 9, characterized in that: The device comprises: An acquisition module is used to acquire the graphic object to be identified; A recognition module, used to take the graphic object to be recognized as the input of the pre-trained multimodal relationship extraction model of electric power equipment, and obtain the entity relationship recognition result of the graphic object to be recognized output by the pre-trained multimodal relationship extraction model of electric power equipment; The training process of the pre-trained multimodal relationship extraction model of electric power equipment includes: Use a web crawler for news reports on power equipment to crawl power equipment news report data; Using public news data in the public domain as basic data, and combining the power equipment news report data to build a power equipment news report forecast resource library; Using the power equipment news report forecast resource library to construct training data; The training data is used to train the multimodal relationship extraction model of electric power equipment to obtain the pre-trained multimodal relationship extraction model of electric power equipment.
11. A computer device, characterized in that: include: one or more processors; The processor is configured to execute one or more programs; When the one or more programs are executed by the one or more processors, the method for extracting multimodal relationships of power equipment based on a deep neural network as described in any one of claims 1 to 9 is implemented.
12. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed, it implements the method for extracting multimodal relationships of power equipment based on deep neural networks as described in any one of claims 1 to 9.