Semantic analysis method and device based on power grid domain knowledge mining and electronic equipment
By building a pre-trained corpus and dynamic knowledge graph in the power grid field, combined with multimodal graph attention parsing network and knowledge distillation technology, the problem of fusing real-time status data of equipment with text semantic information is solved, and efficient semantic parsing and operation and maintenance decision support are achieved.
Patent Information
- Application Number
- CN202510614215.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies find it difficult to effectively integrate real-time device status data with textual semantic information, and are unable to build a multimodal correlation analysis system. This results in low analysis accuracy when faced with complex scenarios and is unable to provide accurate and reliable support for operation and maintenance decisions.
Based on knowledge mining in the power grid field, a pre-trained corpus is constructed and a dynamic knowledge graph and incremental learning fusion system are adopted. The multimodal graph attention parsing network is used to perform unified representation of multimodal corpus data and cross-modal semantic alignment. The parsing model is compressed through knowledge distillation technology and deployed to edge devices for semantic parsing.
Build a multimodal correlation analysis system to accurately capture complex semantic relationships, improve parsing accuracy, and provide accurate and reliable support for operation and maintenance decisions.
Smart Images

Figure CN120633664A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of semantic parsing technology, and in particular to a semantic parsing method, device and electronic equipment based on knowledge mining in the power grid field. Background Art
[0002] With the rapid development of smart grids, the scale of power grids continues to expand, and the types and number of equipment are becoming increasingly complex. Monitoring the operational status of power grids and ensuring precise operation and maintenance are crucial to ensuring a secure and stable power supply. Device status semantic parsing technology, a core component of intelligent device defect diagnosis and maintenance decision-making, can accurately extract key semantics from massive amounts of complex information, helping operators quickly locate equipment issues and develop scientifically sound operation and maintenance strategies. This technology is crucial for improving grid operation and maintenance efficiency, reducing costs, and ensuring safe grid operation.
[0003] In related technologies, natural language processing technology based on pre-trained models such as BERT is usually used to analyze text data such as equipment logs and maintenance reports. The word vector mapping method is used to convert each word in the text into a vector representation of a specific dimension, thereby mining the semantic associations between words in the vector space and realizing basic semantic analysis of the text.
[0004] However, the applicant recognizes that due to the high degree of professionalism and particularity of the power grid field, there are a large number of professional terms and key concepts such as "circuit breaker opening and closing delay" and "insulator surface discharge". These terms rarely appear in general corpus, resulting in deviations in the model's semantic understanding of these key concepts during the semantic parsing process, and the parsing accuracy is difficult to meet actual needs; moreover, a large amount of real-time status data will be generated during the operation of power grid equipment. Existing methods are difficult to effectively integrate the real-time status data of the equipment with text semantic information, and it is impossible to build a multimodal correlation analysis system. When faced with complex scenarios, due to the lack of support for the fusion of multimodal information, the model's ability to capture complex semantic relationships is insufficient, which further leads to low parsing accuracy and inability to provide accurate and reliable support for operation and maintenance decisions. Summary of the Invention
[0005] In view of this, the present application provides a semantic parsing method, device and electronic equipment based on knowledge mining in the power grid field. The main purpose is to solve the current problem that it is difficult to effectively integrate real-time status data of equipment with text semantic information, and it is impossible to build a multimodal correlation analysis system. When faced with complex scenarios, due to the lack of support for the fusion of multimodal information, the model's ability to capture complex semantic relationships is insufficient, which further leads to low parsing accuracy and inability to provide accurate and reliable support for operation and maintenance decisions.
[0006] According to the first aspect of the present application, a semantic parsing method based on power grid domain knowledge mining is provided, the method comprising:
[0007] Based on the original corpus related to the power grid field, a pre-trained corpus is constructed, and a dynamic knowledge graph and incremental learning fusion system are used to continuously update the corpus of the pre-trained corpus;
[0008] Using a multimodal graph attention parsing network, the corpus in the pre-trained corpus is subjected to unified representation processing of multimodal corpus data and cross-modal semantic alignment processing;
[0009] Building a parsing model based on the processed pre-training corpus, and compressing the parsing model through knowledge distillation technology;
[0010] The compressed parsing model is deployed to the edge device in the substation power grid, and the device status information of the edge device is continuously collected. The compressed parsing model is used to perform semantic parsing on the device status information, and the obtained parsing results are output.
[0011] According to a second aspect of the present application, a semantic parsing device based on power grid domain knowledge mining is provided, the device comprising:
[0012] A construction module is used to build a pre-trained corpus based on original corpus related to the power grid field, and to continuously update the pre-trained corpus using a dynamic knowledge graph and incremental learning fusion system;
[0013] An alignment processing module is used to perform unified representation processing of multimodal corpus data and cross-modal semantic alignment processing on the corpus in the pre-training corpus using a multimodal graph attention parsing network;
[0014] A compression processing module, configured to construct a parsing model based on the processed pre-trained corpus, and to compress the parsing model using a knowledge distillation technique;
[0015] The parsing module is used to deploy the compressed parsing model to the edge device in the substation power grid, and continuously collect the device status information of the edge device, use the compressed parsing model to perform semantic parsing on the device status information, and output the obtained parsing results.
[0016] According to a third aspect of the present application, an electronic device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the methods described in the first aspect when executing the computer program.
[0017] According to a fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the methods in the first aspect are implemented.
[0018] Through the above technical solution, the present application provides a semantic parsing method, device and electronic device based on knowledge mining in the power grid field. The present application constructs a pre-trained corpus based on original corpus related to the power grid field, and adopts a dynamic knowledge graph and incremental learning fusion system to continuously update the corpus of the pre-trained corpus. The multimodal graph attention parsing network is used to perform unified representation processing of multimodal corpus data and cross-modal semantic alignment processing on the corpus in the pre-trained corpus. Based on the processed pre-trained corpus, a parsing model is constructed, and the parsing model is compressed through knowledge distillation technology. The compressed parsing model is deployed to the edge device in the substation power grid, and the device status information of the edge device is continuously collected. The compressed parsing model is used to perform semantic parsing on the device status information, and the obtained parsing results are output. The dynamic knowledge graph and incremental learning are used to realize continuous injection of domain knowledge, and the multimodal graph attention parsing network is combined to accurately align semantics to build a multimodal association analysis system, so that complex semantic relationships can be accurately captured even in complex scenarios, the parsing accuracy is improved, and accurate and reliable support is provided for operation and maintenance decision-making.
[0019] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present application. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0021] Figure 1A A flowchart of a semantic parsing method based on power grid domain knowledge mining provided by an embodiment of the present application is shown;
[0022] Figure 1B A schematic diagram of a semantic parsing method based on power grid domain knowledge mining provided by an embodiment of the present application is shown;
[0023] Figure 2 A schematic diagram of the structure of a semantic parsing device based on power grid domain knowledge mining provided by an embodiment of the present application is shown;
[0024] Figure 3 A schematic diagram of the device structure of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0025] The following describes exemplary embodiments of the present application in more detail with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0026] The embodiment of the present application provides a semantic parsing method based on power grid domain knowledge mining, such as Figure 1A As shown, the method includes:
[0027] 101. Based on the original corpus related to the power grid field, a pre-training corpus is constructed, and a dynamic knowledge graph and incremental learning fusion system are used to continuously update the pre-training corpus.
[0028] In an embodiment of the present application, it is first necessary to collect information related to the power grid field as raw corpus, and extract structured data streams from the raw corpus, so as to achieve the purpose of obtaining multi-source structured data streams and building a multi-source data fusion pipeline. Among them, the raw corpus can include monitoring system telemetering data, work order system transaction table data and equipment parameter document data. In terms of monitoring system telemetering data collection, by deploying an industrial protocol conversion gateway, the IEC protocol is used to collect SCADA (Supervisory and Control System) telemetering data in real time, and the sampling frequency can be flexibly set according to actual needs to ensure the real-time and accuracy of the data. For the work order system transaction table data, a JDBC connector is used to batch extract the MySQL transaction table of the OMS (Operation and Maintenance Management System) work order system on an hourly basis, so as to obtain comprehensive and timely work order processing information. In addition, the equipment parameter document data obtains the XML format parameter document of the equipment ledger by calling the REST API (Representational State Transfer Application Programming Interface), thereby realizing efficient management of equipment parameters.
[0029] After successfully acquiring the original corpus, in order to ensure the accurate alignment and fusion of multi-source data in the time dimension, the embodiment of the present application further carries out data preprocessing. One of the core links of preprocessing is time axis alignment. By creating a global timestamp and using the Lamport logical clock algorithm to effectively compensate for network delays, the time synchronization error of multi-source data is strictly controlled within 5ms, thereby greatly improving the consistency and availability of data, providing a reliable guarantee for subsequent multi-source data fusion and analysis. The specific process of timestamp alignment and outputting structured data streams is as follows:
[0030] First, calculate the global timestamp using the following formula 1:
[0031] Formula 1: t global=max(t local , t received , Δ netwdrk )
[0032] Among them, t global Represents the global timestamp, t local Indicates the local clock time, Δ netwdrk Indicates the dynamic network delay measured by two-way ping;
[0033] Subsequently, a sliding window of a preset duration is initialized. The corpus whose local clock time falls within the sliding window is captured from the original corpus as the corpus to be aligned. The preset duration can be 60 seconds. Next, the corpus to be aligned is time-aligned according to the global timestamp, and the aligned corpus is output as a structured data stream. The sliding window is then slid again at the preset sliding interval, and the corpus whose local clock time falls within the sliding window after the sliding is recaptured from the original corpus as the new corpus to be aligned. The new corpus to be aligned is time-aligned according to the global timestamp and output as a structured data stream. This process continues until the entire original corpus has been traversed, completing the structured data stream extraction from the original corpus. In practical applications, Apache Flink can be used as the core real-time processing engine. By setting the sliding window parameters, such as a window size of 60 seconds and a sliding step size of 10 seconds, the engine can efficiently process and analyze the input data in real time, ultimately outputting a structured data stream, providing accurate and orderly data support for subsequent data processing steps.
[0034] After obtaining the structured data stream through the above process, the embodiment of the present application will carry out multimodal feature fusion through a three-level neural network architecture, taking into account general semantics and power grid field characteristics. Specifically, the input layer of the three-level neural network architecture carries out word segmentation processing on structured data streams such as work order texts and ledger descriptions, and converts them into 768-dimensional word vector representations. In the self-attention calculation link, a 12-head attention mechanism is introduced, that is, a multi-dimensional word vector is generated based on the structured data stream after word segmentation processing, and the following formula 2 is used to calculate the attention head output results of the multi-dimensional word vector, thereby obtaining a preset number (that is, 12) of attention head outputs, and by splicing these attention head output results, a vector with general semantic representation capabilities (that is, general semantic representation) is constructed, laying a solid foundation for subsequent feature fusion and in-depth analysis:
[0035] Formula 2:
[0036] Among them, Q i =EW i Q , K i =EW i K, V i =EW i V , head i represents the output of the i-th attention head, softmax represents the activation function used to convert the input value into a probability distribution, W i Q 、W i K 、W i V is a trainable parameter, E represents a multidimensional word vector, d k is the key vector K i Dimensions, is the key vector K i The transposed matrix of .
[0037] Next, based on the connection relationship of the devices in the substation power grid, an adjacency matrix is generated, the general semantic representation is integrated into the initial node feature matrix in the graph convolutional network, and the adjacency matrix is processed using the integrated graph convolutional network to obtain the topological enhanced feature vector H between the devices in the substation network. gcn , where the formula for each layer of graph convolution operation in the graph convolution network is as follows:
[0038] Formula 3:
[0039] Among them, H (l+1) represents the node feature matrix obtained after the l+1th layer graph convolution operation, σ represents the activation function used to perform nonlinear transformation on the calculation results, represents the adjacency matrix with self-loops added, represents the degree matrix and The i-th diagonal element of is equal to the adjacency matrix with the self-loop added The sum of the elements in the i-th row of Degree matrix The square root of the inverse of (l) Represents the node feature matrix before the l-th layer graph convolution operation; W (l) Represents the trainable weight matrix of the l-th layer graph convolution operation; after the last layer of graph convolution operation in the graph convolution network is calculated, the topology enhancement feature vector H is obtained gcn It should be noted that, in the embodiment of the present application, the adjacency matrix generated according to the connection relationship of the devices in the substation power grid can be A, which is an N×N matrix, where N represents the number of nodes in the adjacency matrix, and the elements in the matrix are 0 or 1, which are used to represent the connection relationship between the devices. If A ij =1, it means there is a connection between node i and node j; if A ij = 0, it means that there is no connection between node i and node j. It is an adjacency matrix with self-loops added. Setting an N×N identity matrix I, adding A to the identity matrix I is equivalent to adding a connection pointing to itself (self-loop) to each node in the original adjacency relationship, thereby considering the characteristics of the node itself in the graph convolution operation. Furthermore, W (l) The dimension is represents the set of real numbers, d in is the dimension of the input feature, d out is the dimension of the output feature; the topologically enhanced feature vector H gcn The dimension is N×512, where N is the number of nodes, 512 represents the feature dimension of each node’s final output, and the topology-enhanced feature vector H gcn It contains node feature information learned by the graph convolutional network and can be used for subsequent tasks such as node classification and graph classification.
[0040] Subsequently, the present embodiment will combine the topological enhancement feature vector H gcn , the following loss function, i.e., the following formula 4, is used to perform adversarial training on the preset domain classifier D and feature generator F, adjust the term space, and obtain the adjusted feature generator F:
[0041] Formula 4:
[0042] in, represents the loss function of adversarial training, Represents the expectation operator.
[0043] Next, the embodiment of the present application determines the pluggable parameter matrix M, and uses the following formula 5 to enhance the topology feature vector H through the pluggable parameter matrix M. gcn Perform linear transformation to obtain the transformed topological enhancement feature vector H gcn , the topologically enhanced eigenvector H after the transformation gcn That is, a vector that integrates domain characteristics:
[0044] Formula 5: H adapt =H gcn M+b
[0045] Among them, H adapt Represents the topological enhancement feature vector H after linear transformation gcn , b represents the preset bias vector, and the dimension of the pluggable parameter matrix M is 512×512.
[0046] Finally, the adjusted feature generator F is used to adapt Processing is performed to obtain the vector H of the fusion domain characteristics output by the adjusted feature generator F final , and using the vector Hfinal Construct a pre-training corpus. Among them, the vector H final The dimension is N×512, where N represents the number of nodes, and the vector H final It integrates the feature information obtained through various previous operations, has domain adaptability, and can be used for subsequent node classification, graph classification and other related tasks.
[0047] In another optional implementation scheme, the embodiment of the present application will also extract triples through a hybrid model, and combine it with dynamic strategies to achieve continuous evolution of the graph. Specifically, the hybrid model is used to accurately and efficiently extract triple information, and the dynamic strategy is combined to flexibly adjust the graph construction and update rules according to actual needs and data changes, thereby building a system that integrates dynamic knowledge graphs and incremental learning, and continuously updating the pre-trained corpus to ensure that the corpus always keeps up with the dynamic changes of the data and incorporates new knowledge and new information in a timely manner. In this way, on the one hand, the dynamic knowledge graph can reflect the latest status and associations of knowledge in real time, and enhance the adaptability to complex and changeable knowledge scenarios; on the other hand, the incremental learning mechanism eliminates the need for the model to be trained from scratch, and can quickly absorb new corpus based on existing knowledge, effectively improving learning efficiency and model performance, and better meeting the timeliness and accuracy requirements of knowledge updating and processing in practical applications. The specific process is as follows:
[0048] First, obtain the text sequence to be updated and call the hybrid model built in advance with reference to the dynamic knowledge graph strategy. The hybrid model can be a hybrid model of BiLSTM-CRF and PCNN built by the joint learning architect, so that the BiLSTM-CRF layer in the hybrid model uses the following formula 6 to calculate the entity boundary probability of the text sequence to be updated:
[0049] Formula 6: P(y t |h t )=softmax(W crf h t +b)
[0050] Among them, P(y t |h t ) represents the entity boundary probability, y t represents the predicted label of the text sequence to be updated at time step t, h t represents the hidden state of BiLSTM at time step t in the hybrid model, W crf represents the weight matrix of the CRF layer in the hybrid model, and b represents the preset bias vector.
[0051] Then, based on the entity boundary probability, the main entity vector e is extracted from the text sequence to be updated s , customer entity vector e o and the context vector e c, the PCNN layer in the hybrid model uses the following formula 7 to calculate the inter-entity relationship feature vector of the text sequence to be updated:
[0052] Formula 7: p=maxpool(CNN([e s ;e o ;e c ]))
[0053] Among them, p represents the feature vector of the relationship between entities, maxpool represents the maximum pooling function, CNN represents the convolutional neural network function, and e s Represents the main entity vector of the text sequence to be updated, e o Represents the object entity vector of the text sequence to be updated, e c It should be noted that a dynamic threshold mechanism is set in the above process, and the adaptive relationship confidence can be set.
[0054] Next, take the principal entity vector e s , customer entity vector e o And the feature vector p of the relationship between entities, construct the triples of the text sequence to be updated. Among them, if the above process fails to generate triples for the text sequence to be updated, the incremental learning fusion system Few-shot is used to generate pseudo labels for the text sequence to be updated, and the pseudo labels are used as triples of the text sequence to be updated. For example, templated triples (new entity, superordinate concept, device type) are inserted. In the early stage of incremental learning, when new entities and other data appear but lack sufficient annotation information, Few-shot learning is used to generate pseudo labels using a small number of samples, providing initial annotation data for subsequent learning and model updates, helping the model to quickly adapt to new data and start the incremental learning process.
[0055] In addition, the embodiment of the present application also provides a conflict resolution mechanism, so that in the incremental learning process, when multiple possible relationships or knowledge representations appear, the most reasonable and semantically consistent relationship is selected by calculating the semantic consistency score, and possible conflicts are resolved to ensure the consistency and accuracy of the knowledge graph. The specific process is as follows: First, multiple non-conflicting relationships are extracted from the triples, and a path set is constructed based on the triples. The following formula 8 is used to calculate the score for each non-conflicting relationship to determine the target non-conflicting relationship with the highest score:
[0056] Formula 8:
[0057] Where score(r) represents the score of the currently calculated non-conflicting relation r, p represents a path in the path set, Paths represents the path set, and sim(p, r) represents the semantic similarity between path p and the non-conflicting relation r. In this way, the non-conflicting relations other than the target non-conflicting relation are considered as the non-conflicting relations to be removed from the triples. The removed triples are then used to update the pre-training corpus.
[0058] In another optional implementation scheme, the embodiment of the present application will also dynamically adjust the edge weight by combining the edge weight at the previous moment and the ratio of the number of relationship triggering times in the current time period, thereby dynamically adjusting the weight of the connection between them according to the activity level of the entity in different time periods, that is, the relationship triggering frequency, so that the knowledge graph can better reflect the dynamic changes of the relationship between entities and improve the timeliness and accuracy of the knowledge graph. The specific process is as follows: the following formula 9 is used to calculate the edge weights between the entities of the continuously updated pre-training corpus, so that the dynamic changes of the relationship between the entities of the pre-training corpus are reflected based on the calculated dynamic edge weights:
[0059] Formula 9:
[0060] in, represents the edge weight between entity i and entity j at time step t, represents the edge weight between entity i and entity j at time step t-1, α represents the preset forgetting factor, freq ij Indicates the number of times the relationship between entity i and entity j is triggered within a specified time period, ∑ k freq ik Represents the sum of the number of relationship triggers between entity i and all other entities.
[0061] It should be noted that in actual application, the above process can successfully capture 83% of new fault modes, the relationship extraction F1 value reaches 92.7%, and the graph update delay is controlled within 300ms. This further shows that the use of the above incremental learning method can effectively discover new fault modes and achieve high performance in relationship extraction tasks. At the same time, it can quickly update the knowledge graph, meet the needs of timely knowledge updates in actual applications, and improve the practicality and reliability of the system.
[0062] 102. Using the multimodal graph attention parsing network, the corpus in the pre-training corpus is processed into a unified representation of multimodal corpus data and cross-modal semantic alignment.
[0063] In an embodiment of the present application, IEC61850 protocol data (structured device parameters) are encoded into low-dimensional vectors to achieve unified representation of multimodal data, thereby converting complex structured device parameters into a low-dimensional vector form that is convenient for subsequent processing, facilitating unified processing and analysis in different models and algorithms, and improving the efficiency and compatibility of data processing. The specific process is as follows: First, for each corpus in the pre-training corpus, a data object DO node metadata descriptor is generated. The DO node metadata descriptor includes a data type type, a dimension unit, a valid value range lower limit min, a valid value range upper limit max, and a data timestamp timestamp. The data type (type) specifies the type of data represented by the DO node, such as integer or floating-point. The dimension (unit) represents the unit of data, such as meters for length or seconds for time, and helps understand the actual physical meaning of the data. The lower limit (min) of the valid range (min) is the minimum value in the valid range, defining the lower limit of the possible values for the DO node data. The upper limit (max) of the valid range (max) is the maximum value in the valid range, defining the upper limit of the possible values for the DO node data. The timestamp (timestamp) is the data timestamp, recording the time when the data was generated or collected, which is crucial for analyzing the data's temporal characteristics. By generating such metadata descriptors, we can comprehensively and concisely describe the key information of the DO node, providing a foundation for subsequent data processing and analysis.
[0064] Next, the following formula 10 is used to generate the DO feature vector at the current moment for each corpus. This effectively captures the dynamic changes of device parameters over time, mines the temporal dependencies in the data, and better understands the operating status and parameter change trends of the equipment, providing strong support for applications such as equipment monitoring and fault diagnosis:
[0065] Formula 10:
[0066] in, represents the DO feature vector of the corpus at the current moment, GRU represents the GRU network formula, x t represents the DO metadata vector of the corpus at time t, Indicates the hidden state of the corpus at the previous moment.
[0067] Then, the following formula 11 is used to calculate the attention weight between each corpus and its neighboring corpora. Based on the connection relationship and feature similarity between the corpora, different attention weights are adaptively assigned to the neighboring corpora of each corpus, so as to more effectively aggregate the features of the neighboring corpora and obtain more representative and discriminative corpus feature representation:
[0068] Formula 11:
[0069] Among them, α ij Represents the attention weight between corpus i and neighbor corpus j, LeakyReLU represents the linear rectification function with leakage, a T represents the transpose of the preset attention coefficient vector α, W represents the trainable weight matrix, represents the DO feature vector of corpus i, represents the DO feature vector of neighbor corpus j, Represents the set of neighboring corpora of corpus i.
[0070] Next, we use the following formula 12 of the multimodal graph attention parsing network to aggregate the DO feature vectors of all subordinate corpora of each corpus to obtain the subordinate DO feature vector of each corpus:
[0071] Formula 12:
[0072] Among them, h ln represents the subordinate DO feature vector, α ij represents the attention weight between corpus i and neighbor corpus j, and W represents the trainable weight matrix represents the DO feature vector of neighbor corpus j, N(i) Represents the set of neighboring corpora of corpus i.
[0073] Furthermore, in the embodiment of the present application, the following formula 13 is used to calculate the subordinate DO feature vector of each corpus to obtain the gating vector g:
[0074] Formula 13:
[0075] Among them, W g represents the trainable weight matrix, and Represents the feature vector of each subordinate DO, and σ represents the activation function.
[0076] Then, the following formula 14 is used to calculate the subordinate DO feature vector of each corpus to obtain the integrated feature vector h ld :
[0077] Formula 14:
[0078] Where g represents the gate vector, and Represents the feature vectors of each subordinate DO.
[0079] Finally, the DO node metadata descriptor of each corpus is associated with the subordinate DO feature vector of each corpus to complete the unified representation processing of the multimodal corpus data in the pre-training corpus.
[0080] In another optional implementation scheme, in order to enable the model to learn the semantic associations between different modal data (natural language work orders, equipment parameters, and defect maps), by constructing positive and negative sample pairs, the model can distinguish between matching and mismatching data combinations, thereby better aligning their semantic spaces. In this embodiment of the application, the semantic spaces of natural language work orders, equipment parameters, and defect maps are also aligned through comparative learning. The specific process is as follows:
[0081] First, the historical data of the substation power grid is obtained, and matching triples (work order text T, equipment E, defective node D) are extracted from the historical data. The matching triples (work order text T, equipment E, defective node D) are used to construct positive sample pairs (T, ), where T represents the work order text, represents the device feature vector of device E, Defective node feature vector representing defective node D;
[0082] The following formula 15 is used to calculate the matching triple (work order text T, equipment E, defect node D) to obtain the similarity score of the matching triple (work order text T, equipment E, defect node D):
[0083] Formula 15:
[0084] Among them, s represents the similarity score, h T The feature vector representing the work order text T, h E Represents the operating condition vector of equipment E, h D The feature vector representing the defective node D.
[0085] Then, referring to the positive sample pair (T, ), construct negative sample pairs (T k , E k , D k ), and the following contrastive learning loss function, that is, Formula 16, is used to calculate the matching triples (work order text T, equipment E, defect node D) and the negative sample pairs (T k , E k , D k ) between sample pairs:
[0086] Formula 16:
[0087] in, represents the contrastive learning loss function, τ is used to control the distribution of similarity scores, s represents the similarity score of the matching triple (work order text T, equipment E, defective node D), and k represents the number of negative sample pairs. In addition, when sampling negative samples, the device-level negative samples can be replaced with vectors of other equipment in the same substation, while the text-level negative samples remain unchanged and T is replaced with irrelevant work order text to obtain negative sample pairs.
[0088] Finally, referring to the sample pair similarity, the corpus in the pre-training corpus is subjected to cross-modal semantic alignment, and then the feature information from different sources is integrated, and the degree of association between different data is measured by calculating the similarity score, thereby optimizing the loss function of contrastive learning, so that the model can better learn the semantic relationship between data, make full use of multi-source feature information, improve the model's ability to understand the relationship between different modal data, and enhance the model's generalization ability and accuracy in complex scenarios.
[0089] In another optional implementation scheme, the embodiment of the present application will also construct a heterogeneous graph based on the substation power grid, and use the following formula 17 to calculate the feature vector of each node in the heterogeneous graph in the multimodal graph attention parsing network, so that the feature vector of each node contains its local context information, wherein the heterogeneous graph includes multiple nodes, the multiple nodes include device nodes, defect nodes, and treatment measure nodes, and the edges in the heterogeneous graph include edges between device nodes and defect nodes, and edges between defect nodes and measure nodes, wherein the weight of the edge between the device node and the defect node is equal to the co-occurrence probability, and the weight of the edge between the defect node and the measure node is equal to the conditional probability:
[0090] Formula 17:
[0091] in, represents the feature vector of node i in the l+1th layer of the multimodal graph attention parsing network, σ represents the activation function, R represents the edge set of the edges in the heterogeneous graph, N r (i) represents the set of neighbor nodes of node i under edge r, represents the attention weight of node i to neighbor node j under edge r, represents the weight matrix of edge r in layer l, Represents the feature vector of node j at layer l. In this way, through the graph attention mechanism, different attention weights are adaptively assigned to each node's neighboring nodes based on the connection relationship and feature similarity between nodes, thereby more effectively aggregating the features of neighboring nodes and obtaining more representative and discriminative node feature representations. Moreover, it captures the complex relationships and importance differences between nodes, enhances the model's understanding and processing capabilities of graph-structured data, and helps improve the performance of subsequent tasks (such as node classification and graph classification).
[0092] In another optional embodiment, the embodiment of the present application will also determine the threshold control index N according to the substation power grid, and read the basic similarity threshold τ preset in the multimodal graph attention parsing network. base , the following formula 18 is used to calculate the basic similarity threshold τ base Calculate and obtain the similarity threshold τ after dynamic adjustment dynamic , and the similarity threshold τ after dynamic adjustment dynamic , the preset basic similarity threshold τ in the multimodal graph attention parsing network base To update:
[0093] Formula 18:
[0094] Among them, Δ τ Represents the threshold adjustment amount. In this way, by flexibly adjusting the threshold according to changes in the actual scenario, while ensuring a certain accuracy, the operating efficiency of the model is improved, unnecessary computing resource consumption is reduced, and the model is made more practical and adaptable.
[0095] In another optional implementation scheme, the embodiment of the present application further uses the following formula 19 to calculate the risk value of the substation power grid and output the calculated risk value:
[0096] Formula 19:
[0097] Among them, Risk(t) represents the risk value of the substation power grid at time t, v wind (t) represents the wind speed of the substation grid at time t, load r ate(t) represents the load rate of the substation grid at time t. By considering environmental risks (such as wind speed) and work order load, and dynamically adjusting the semantic similarity threshold to adapt to the needs of different scenarios, the model's accuracy and computational efficiency are balanced.
[0098] In the actual application process, when v wind (t) is greater than or equal to 14m / s and load r When ate(t) is greater than 85%, Risk(t) calculation will begin so that when the wind speed is too high and the work order load is too heavy, specific measures can be taken by calculating the risk value, such as adjusting the equipment operating status and issuing early warnings.
[0099] 103. Based on the processed pre-trained corpus, a parsing model is constructed, and the parsing model is compressed using knowledge distillation technology.
[0100] In the embodiment of the present application, a parsing model is constructed based on the processed pre-trained corpus. Furthermore, in order to achieve model compression while retaining effective model performance, the parsing model will continue to be compressed using knowledge distillation technology, that is, the knowledge of the 100 billion parameter teacher model will be first transferred to the lightweight student model. The specific process is as follows:
[0101] In the embodiment of the present application, the teacher model with hundreds of billions of parameters refers to the parsing model. In actual application, the parsing model adopts a multimodal Transformer network with hundreds of billions of parameters, integrating three modules of power grid equipment parameter parsing, protocol decoding, and fault mode recognition. The input is a 1024-dimensional IEC61850 protocol message feature vector. Its powerful performance and rich knowledge reserve provide a good foundation for subsequent knowledge migration; and the lightweight student model refers to the compressed parsing model obtained by compressing the parsing model through knowledge distillation technology. The compressed parsing model constructs a hybrid architecture of lightweight graph convolutional network (GCN) and bidirectional LSTM, wherein the GCN layer is used to parse the device topology relationship, and the bidirectional LSTM processes the timing signal, and the model parameter amount is controlled at 28.7MB. This design aims to reduce the complexity of the model and the computing resource requirements, making it more suitable for actual application scenarios, especially in resource-constrained environments.
[0102] Specifically, it is necessary to first build a hybrid architecture of a lightweight graph convolutional network and a bidirectional LSTM, and then migrate the knowledge of the parsing model to the hybrid architecture. The hybrid architecture is then used to train the migrated knowledge to obtain the initial compression model.
[0103] Then, the KL divergence loss between the initial compression model and the analytical model is calculated using the following formula 20:
[0104] Formula 20:
[0105] in, represents the KL divergence loss value, T represents the temperature coefficient used to control the smoothness of the probability distribution, N represents the number of samples, represents the raw output of the parsing model, represents the original output of the initial compression model, and softmax represents the function used to convert the output into a probability distribution, thereby measuring the difference in output probability distribution between the parsing model and the initial compression model. By minimizing the KL divergence, the initial compression model can better fit the output of the parsing model, thereby retaining the knowledge of the parsing model.
[0106] At the same time, the embodiment of the present application performs grid-specific attention migration and introduces attention matrix alignment loss at the protocol decoding layer. That is, the following formula 21 is used to calculate the attention matrix alignment loss value between the initial compression model and the parsing model:
[0107] Formula 21:
[0108] in, represents the attention matrix alignment loss value, L represents the number of attention layers, l represents the index of the attention layer, represents the attention query matrix of the parsing model at layer l, The attention query matrix of the initial compression model at layer l is expressed, thereby fully leveraging the advantages of the parsing model in the attention mechanism. By forcing the initial compression model to inherit the attention distribution pattern of the parsing model on the voltage phase features in the SV message (Sampled Values), the performance of the initial compression model on specific tasks is improved, especially the accuracy and reliability when processing power grid-related data.
[0109] Finally, the initial compression model is optimized based on the KL divergence loss and attention matrix alignment loss. The KL divergence loss and attention matrix alignment loss between the optimized initial compression model and the analytical model are recalculated and optimized again until the initial compression model reaches the preset convergence condition. The optimized initial compression model is then used as the compressed analytical model. This series of operations enables the lightweighting of the model while preserving and inheriting the knowledge and performance of the analytical model to the greatest extent possible, meeting the power grid industry's demand for efficient and accurate models.
[0110] 104. Deploy the compressed parsing model to the edge device in the substation power grid, continuously collect device status information of the edge device, use the compressed parsing model to perform semantic parsing on the device status information, and output the obtained parsing results.
[0111] In an embodiment of the present application, the compressed parsing model is deployed to the edge device in the substation power grid, and the device status information of the edge device is continuously collected. The compressed parsing model is used to perform semantic parsing on the device status information, and the obtained parsing results are output. Specifically, device status monitoring and strategy generation are achieved through time series analysis. The process is as follows: first, mutation detection is performed, and the device status information sampled at 10Hz (covering temperature, vibration, current, voltage, power factor, insulation resistance, etc.) is used as input and input into the compressed parsing model, so that the first layer LSTM in the compressed parsing model calculates the device status information using the following formula 22 to obtain the first hidden state vector of the device status information:
[0112] Formula 22:
[0113] in, Represents the first hidden state vector of the first layer LSTM at time t, x t Indicates device status information. The hidden state vector of the first LSTM layer represents the device status information at time t-1. It aims to capture the sudden change characteristics of the device status from the time series data and detect device anomalies in a timely manner.
[0114] Next, the second LSTM layer in the compressed parsing model processes the first hidden state vector using the following formula 23 to obtain the second hidden state vector:
[0115] Formula 23:
[0116] in, represents the second hidden state vector of the first hidden state vector at time t in the second layer of LSTM, g t Represents the gate vector of the second layer LSTM at time t, It represents the hidden state vector of the first hidden state vector in the second layer of LSTM at time t-1, thereby focusing on the feature information that is more important for equipment status monitoring and improving the efficiency and accuracy of subsequent processing.
[0117] Then, the second hidden state vector is processed by a multi-layer perceptron to generate a defect response strategy knowledge subgraph, and the confidence of the defect response strategy knowledge subgraph is calculated using the following formula 24:
[0118] Formula 24:
[0119] Among them, c ij represents the confidence between node i and node j in the defect response strategy knowledge subgraph, W c represents the weight matrix, represents the second hidden state vector corresponding to node i, Represents the second hidden state vector corresponding to node j. It should be noted that when generating the strategy knowledge subgraph, a five-tuple can be defined, where nodes are equipment components, defect modes, and disposal measures, and edges are fault propagation paths. The priority of each node in the defect response strategy knowledge subgraph is calculated using the following formula 25 through a multi-layer perceptron, and the defect response strategy knowledge subgraph is annotated using the calculated priority of each node:
[0120] Formula 25:
[0121] Among them, p i represents the priority of node i in the defect response strategy knowledge subgraph, MLP represents multi-layer perceptron, The second hidden state vector corresponding to node i is represented by the labeled defect response strategy knowledge subgraph and confidence level as the parsing result, and the parsing result is output. This systematically organizes and analyzes information related to the device status, providing a comprehensive and targeted strategy for addressing device defects. Through the above steps, knowledge distillation and quantization are used to reduce the model size to 1 / 35, increasing inference speed by 4 times. At the same time, a two-layer LSTM is used to achieve high sensitivity in detecting sudden changes in device status. This effectively improves the accuracy and reliability of device status monitoring and strategy generation while maintaining processing efficiency, providing strong support for the stable operation of the equipment.
[0122] The actual application process, such as Figure 1B As shown in the figure, in edge computing scenarios, when an edge server receives a task request, it first loads the model into the model loader. The model loader, as an intermediary, is responsible for adapting the model and preparing it for transmission to the subsequent processing engine. The model is then passed to the FP16 quantization engine, which performs FP16 quantization on the model. This operation significantly reduces the model's storage space and computing resource requirements without significantly compromising model accuracy, thereby improving the model's efficiency on edge devices. The quantized model then enters the TensorRT inference engine, a high-performance deep learning inference engine provided by NVIDIA. It optimizes inference on the quantized model, further improving inference speed and performance. After inference is complete, the results are stored in the result buffer pool, which temporarily stores the inference results for subsequent calls and processing. Finally, the work order response interface retrieves the inference results from the result buffer pool and generates the corresponding work order response based on the results. This completes the entire process from model loading, quantization, inference, to result response, achieving efficient and fast model inference and task response on edge devices.
[0123] The method provided in the embodiment of the present application is to construct a pre-trained corpus based on original corpus related to the power grid field, and adopt a dynamic knowledge graph and incremental learning fusion system to continuously update the corpus of the pre-trained corpus, and use a multimodal graph attention parsing network to perform unified representation processing of multimodal corpus data and cross-modal semantic alignment processing on the corpus in the pre-trained corpus, build a parsing model based on the processed pre-trained corpus, and compress the parsing model through knowledge distillation technology, deploy the compressed parsing model to the edge device in the substation power grid, and continuously collect device status information of the edge device, use the compressed parsing model to perform semantic parsing on the device status information, and output the obtained parsing results, realize continuous injection of domain knowledge through dynamic knowledge graph and incremental learning, combine the multimodal graph attention parsing network to accurately align semantics, and build a multimodal association analysis system, so that complex semantic relationships can be accurately captured even in complex scenarios, improve the parsing accuracy, and provide accurate and reliable support for operation and maintenance decisions.
[0124] Furthermore, as a specific implementation of the method described in FIG1 , the embodiment of the present application provides a semantic parsing device based on power grid domain knowledge mining, such as Figure 2 As shown, the device includes: a construction module 201, an alignment processing module 202, a compression processing module 203 and a parsing module 204.
[0125] The construction module 201 is used to construct a pre-training corpus based on original corpus related to the power grid field, and continuously update the corpus of the pre-training corpus using a dynamic knowledge graph and incremental learning fusion system;
[0126] The alignment processing module 202 is used to use a multimodal graph attention parsing network to perform unified representation processing of multimodal corpus data and cross-modal semantic alignment processing on the corpus in the pre-training corpus;
[0127] The compression processing module 203 is used to construct a parsing model based on the processed pre-training corpus, and compress the parsing model using knowledge distillation technology;
[0128] The parsing module 204 is used to deploy the compressed parsing model to the edge device in the substation power grid, and continuously collect the device status information of the edge device, use the compressed parsing model to perform semantic parsing on the device status information, and output the obtained parsing results.
[0129] In a specific application scenario, the construction module 201 is used to collect information related to the power grid field as the original corpus, and extract a structured data stream from the original corpus; perform word segmentation on the structured data stream, generate a multidimensional word vector using the processed structured data stream, and calculate the attention head output results of the multidimensional word vector using the following formula to obtain a preset number of attention head output results, and splice the preset number of attention head output results to obtain a common semantic representation.
[0130]
[0131] Among them, Q i =EW i Q , K i =EW i K , V i =EW i V , head i represents the output of the i-th attention head, softmax represents the activation function used to convert the input value into a probability distribution, W i Q 、W i K 、W i V is a trainable parameter, E represents the multidimensional word vector, d k is the key vector K i Dimensions, is the key vector K i The transposed matrix of
[0132] According to the connection relationship of the equipment in the substation power grid, an adjacency matrix is generated, the general semantic representation is integrated into the initial node feature matrix in the graph convolution network, and the adjacency matrix is processed using the integrated graph convolution network to obtain the topology enhanced feature vector H between the equipment in the substation network. gcn , where the formula for each layer of graph convolution operation in the graph convolution network is as follows:
[0133]
[0134] Among them, H (l+1) represents the node feature matrix obtained after the l+1th layer graph convolution operation, σ represents the activation function used to perform nonlinear transformation on the calculation results, represents the adjacency matrix with self-loops added, represents the degree matrix and The i-th diagonal element is equal to the adjacency matrix with the self-loop added The sum of the elements in the i-th row of Degree matrix The square root of the inverse of (l) Represents the node feature matrix before the l-th layer graph convolution operation; W (l) Represents the trainable weight matrix of the l-th layer graph convolution operation; after the last layer of graph convolution operation in the graph convolution network is calculated, the topology enhancement feature vector H is obtained gcn ; Combined with the topological enhancement feature vector H gcn , the following loss function is used to perform adversarial training on the preset domain classifier D and feature generator F to obtain the adjusted feature generator F:
[0135]
[0136] in, represents the loss function of adversarial training, Represents the expectation operator;
[0137] Determine the pluggable parameter matrix M, and use the following formula to enhance the topology feature vector H through the pluggable parameter matrix M gcn Perform linear transformation to obtain the transformed topology enhancement feature vector H gcn :
[0138] H adapt =H gcn M+b
[0139] Among them, H adapt Represents the topological enhancement feature vector H after linear transformation gcn , b represents the preset bias vector;
[0140] Use the adjusted feature generator F to H adapt Processing is performed to obtain the vector H of the fusion domain characteristics output by the adjusted feature generator F final , and using the vector H final Construct the pre-training corpus.
[0141] In a specific application scenario, the construction module 201 is used to collect information related to the power grid field, use all the collected information as the original corpus, and calculate the global timestamp using the following formula. The original corpus includes monitoring system telesignaling data, work order system transaction table data, and equipment parameter document data:
[0142] t global =max(t local , t received , Δ netwdrk )
[0143] Among them, t global represents the global timestamp, the t local Indicates the local clock time, Δ netwdrk Indicates the dynamic network delay measured by two-way ping;
[0144] Initialize a sliding window of a preset time length, capture the corpus whose local clock time is within the sliding window in the original corpus as the corpus to be aligned; perform time alignment on the corpus to be aligned according to the global timestamp, and output the aligned corpus to be aligned as a structured data stream; slide the sliding window according to a preset sliding time interval, and re-capture the corpus whose local clock is within the sliding window after sliding in the original corpus as a new corpus to be aligned, and perform time alignment on the new corpus to be aligned according to the global timestamp and output it as a structured data stream, until all the original corpus are traversed and the structured data stream extraction of the original corpus is completed.
[0145] In a specific application scenario, the construction module 201 is used to obtain a text sequence to be updated, and call a hybrid model constructed in advance with reference to a dynamic knowledge graph strategy, so that the hybrid model calculates the entity boundary probability of the text sequence to be updated using the following formula:
[0146] P(y t |h t )=softmax(W crf h t +b)
[0147] Among them, P(y t |h t ) represents the entity boundary probability, y t represents the predicted label of the text sequence to be updated at time step t, h t represents the hidden state of BiLSTM at time step t in the hybrid model, W crf represents the weight matrix of the CRF layer in the hybrid model, and b represents the preset bias vector;
[0148] Based on the entity boundary probability, the main entity vector e is extracted from the text sequence to be updated. s , customer entity vector e o and the context vector e c , the hybrid model uses the following formula to calculate the inter-entity relationship feature vector of the text sequence to be updated:
[0149] p=maxpool(CNN([e s ;e o ;e c ]))
[0150] Wherein, p represents the feature vector of the relationship between entities, maxpool represents the maximum pooling function, CNN represents the convolutional neural network function, and e s Represents the main entity vector of the text sequence to be updated, e o Represents the object entity vector of the text sequence to be updated, e c Represents the context vector of the text sequence to be updated; adopts the main entity vector e s , the object entity vector e o and the inter-entity relationship feature vector p, constructing a triple of the text sequence to be updated, wherein, if the above process fails to generate the triple for the text sequence to be updated, the incremental learning fusion system Few-shot is used to generate a pseudo label for the text sequence to be updated, and the pseudo label is used as the triple of the text sequence to be updated; extracting multiple non-conflicting relations from the triple, and constructing a path set based on the triple, using the following formula to calculate the score for each non-conflicting relationship, and determining the target non-conflicting relationship with the highest score:
[0151]
[0152] Wherein, score(r) represents the score of the currently calculated non-conflicting relation r, p represents a path in the path set, Paths represents the path set, and sim(p, r) represents the semantic similarity between path p and the non-conflicting relation r;
[0153] The non-conflicting relationships other than the target non-conflicting relationship among the multiple non-conflicting relationships are taken as non-conflicting relationships to be eliminated, the non-conflicting relationships to be eliminated are eliminated from the triples to obtain the eliminated triples, and the pre-training corpus is updated using the eliminated triples.
[0154] In a specific application scenario, the construction module 201 is further configured to calculate edge weights between entities in the continuously updated pre-training corpus using the following formula, so as to reflect the dynamic changes in the relationships between entities in the pre-training corpus based on the calculated dynamic edge weights:
[0155]
[0156] in, represents the edge weight between entity i and entity j at time step t, represents the edge weight between entity i and entity j at time step t-1, α represents the preset forgetting factor, freq ij Indicates the number of times the relationship between entity i and entity j is triggered within a specified time period, Σ k freqik Represents the sum of the number of relationship triggers between entity i and all other entities.
[0157] In a specific application scenario, the alignment processing module 202 is used to generate a data object DO node metadata descriptor for each corpus in the pre-training corpus. The DO node metadata descriptor includes a data type type, a dimension unit, a valid value range lower limit min, a valid value range upper limit max, and a data timestamp timestamp; and the following formula is used to generate a DO feature vector at the current moment for each corpus:
[0158]
[0159] in, represents the DO feature vector of the corpus at the current moment, GRU represents the GRU network formula, x t represents the DO metadata vector of the corpus at time t, Indicates the hidden state of the corpus at the previous moment;
[0160] The following formula is used to calculate the attention weight between each corpus and its neighboring corpora:
[0161]
[0162] Among them, α ij Represents the attention weight between corpus i and neighbor corpus j, LeakyReLU represents the linear rectification function with leakage, α T represents the transpose of the preset attention coefficient vector α, W represents the trainable weight matrix, represents the DO feature vector of the corpus i, represents the DO feature vector of the neighbor corpus j, represents the neighbor corpus set of the corpus i;
[0163] The following formula of the multimodal graph attention parsing network is used to aggregate the DO feature vectors of all subordinate corpora of each corpus to obtain the subordinate DO feature vector of each corpus:
[0164]
[0165] Among them, h ln represents the subordinate DO feature vector, α ij represents the attention weight between corpus i and neighbor corpus j, and W represents the trainable weight matrix represents the DO feature vector of the neighbor corpus j, N(i) represents the neighbor corpus set of the corpus i;
[0166] The following formula is used to calculate the subordinate DO feature vector of each of the corpus to obtain the gate vector g:
[0167]
[0168] Among them, W g represents the trainable weight matrix, and represents the feature vector of each subordinate DO, and σ represents the activation function;
[0169] The following formula is used to calculate the subordinate DO feature vector of each corpus to obtain the integrated feature vector h ld :
[0170]
[0171] Where g represents the gate vector, and Represents each subordinate DO feature vector;
[0172] Associating and storing the DO node metadata descriptor of each of the corpora with the subordinate DO feature vectors of each of the corpora, thereby completing the unified representation processing of the multimodal corpus data in the pre-training corpus;
[0173] Obtain the historical data of the substation power grid, extract matching triples (work order text T, equipment E, defective node D) from the historical data, and use the matching triples (work order text T, equipment E, defective node D) to construct positive sample pairs (T, ), where T represents the work order text, represents the device feature vector of device E, Defective node feature vector representing defective node D;
[0174] The following formula is used to calculate the matching triple (work order text T, equipment E, defective node D) to obtain the similarity score of the matching triple (work order text T, equipment E, defective node D):
[0175]
[0176] Among them, s represents the similarity score, h T The feature vector representing the work order text T, h E Represents the operating condition vector of equipment E, h D The feature vector representing the defective node D;
[0177] Referring to the positive sample pair (T, ), construct negative sample pairs (T k , E k , D k), and the following contrastive learning loss function is used to calculate the matching triple (work order text T, equipment E, defective node D) and the negative sample pair (T k , E k , D k ) between sample pairs:
[0178]
[0179] in, represents the contrastive learning loss function, τ is used to control the distribution of similarity scores, s represents the similarity score of the matching triple (work order text T, equipment E, defect node D), and k represents the number of negative sample pairs;
[0180] Referring to the sample pair similarity, cross-modal semantic alignment processing is performed on the corpus in the pre-training corpus.
[0181] In a specific application scenario, the alignment processing module 202 is further used to construct a heterogeneous graph based on the substation power grid, and use the following formula to calculate the feature vector of each node in the heterogeneous graph in the multimodal graph attention parsing network, so that the feature vector of each node contains its local context information, the heterogeneous graph includes multiple nodes, the multiple nodes include device nodes, defect nodes and treatment measure nodes, the edges in the heterogeneous graph include edges between device nodes and defect nodes, and edges between defect nodes and measure nodes, wherein the weight of the edge between the device node and the defect node is equal to the co-occurrence probability, and the weight of the edge between the defect node and the measure node is equal to the conditional probability:
[0182]
[0183] in, represents the feature vector of node i in the l+1th layer of the multimodal graph attention parsing network, σ represents the activation function, R represents the edge set of the edges in the heterogeneous graph, N r (i) represents the set of neighbor nodes of node i under edge r, represents the attention weight of node i to neighbor node j under edge r, represents the weight matrix of edge r in layer l, represents the feature vector of node j at layer l; and / or,
[0184] According to the substation power grid, determine the threshold control index N, and read the basic similarity threshold τ preset in the multimodal graph attention parsing network base , the basic similarity threshold τ is calculated using the following formula base Calculate and obtain the similarity threshold τ after dynamic adjustment dynamic, and the similarity threshold τ after dynamic adjustment dynamic , the basic similarity threshold τ preset in the multimodal graph attention parsing network base To update:
[0185]
[0186] Among them, Δ τ represents a threshold adjustment amount; and / or,
[0187] The following formula is used to calculate the risk value of the substation power grid and the calculated risk value is output:
[0188]
[0189] Among them, Risk(t) represents the risk value of the substation power grid at time t, v wind (t) represents the wind speed of the substation grid at time t, load r ate(t) represents the load rate of the substation power grid at time t.
[0190] In a specific application scenario, the compression processing module 203 is used to construct a hybrid architecture of a lightweight graph convolutional network and a bidirectional LSTM, and to migrate the knowledge of the analytical model to the hybrid architecture, and to perform model training on the migrated knowledge using the hybrid architecture to obtain an initial compression model;
[0191] The KL divergence loss between the initial compression model and the analytical model is calculated using the following formula:
[0192]
[0193] in, represents the KL divergence loss value, T represents the temperature coefficient used to control the smoothness of the probability distribution, N represents the number of samples, represents the raw output of the parsing model, represents the original output of the initial compression model, and softmax represents the function used to convert the output into a probability distribution;
[0194] The following formula is used to calculate the attention matrix alignment loss between the initial compression model and the parsing model:
[0195]
[0196] in, represents the attention matrix alignment loss value, L represents the number of attention layers, l represents the index of the attention layer, represents the attention query matrix of the parsing model at layer l, represents the attention query matrix of the initial compression model at layer l;
[0197] Referring to the KL divergence loss value and the attention matrix alignment loss value, the initial compression model is optimized, and the KL divergence loss value and the attention matrix alignment loss value between the optimized initial compression model and the analytical model are recalculated and the optimization is continued until the initial compression model reaches the preset convergence condition, and the currently optimized initial compression model is used as the compressed analytical model.
[0198] In a specific application scenario, the parsing module 204 is configured to, when the device status information of the edge device is collected, input the device status information into the compressed parsing model, so that the first LSTM layer in the compressed parsing model calculates the device status information using the following formula to obtain a first hidden state vector of the device status information:
[0199]
[0200] in, represents the first hidden state vector of the first layer LSTM at time t, x t Indicates the device status information, The hidden state vector representing the device state information at time t-1 in the first LSTM layer;
[0201] The second hidden state vector is obtained by processing the first hidden state vector using the following formula by the second LSTM layer in the compressed parsing model:
[0202]
[0203] in, represents the second hidden state vector of the first hidden state vector in the second layer LSTM at time t, g t represents the gate vector of the second layer LSTM at time t, represents the hidden state vector of the first hidden state vector in the second layer LSTM at time t-1;
[0204] The second hidden state vector is processed by a multi-layer perceptron to generate a defect response strategy knowledge subgraph, and the confidence of the defect response strategy knowledge subgraph is calculated using the following formula:
[0205]
[0206] Among them, c ijrepresents the confidence between node i and node j in the defect response strategy knowledge subgraph, W c represents the weight matrix, represents the second hidden state vector corresponding to node i, represents the second hidden state vector corresponding to node j;
[0207] The multi-layer perceptron is used to calculate the priority of each node in the defect response strategy knowledge subgraph using the following formula, and the defect response strategy knowledge subgraph is labeled using the calculated priority of each node:
[0208]
[0209] Among them, p i represents the priority of node i in the defect response strategy knowledge subgraph, MLP represents a multi-layer perceptron, represents the second hidden state vector corresponding to node i;
[0210] The labeled defect response strategy knowledge subgraph and the confidence level are used as the parsing result, and the parsing result is output.
[0211] The device provided in the embodiment of the present application constructs a pre-trained corpus based on original corpus related to the power grid field, and adopts a dynamic knowledge graph and incremental learning fusion system to continuously update the corpus of the pre-trained corpus, and uses a multimodal graph attention parsing network to perform unified representation processing of multimodal corpus data and cross-modal semantic alignment processing on the corpus in the pre-trained corpus. Based on the processed pre-trained corpus, a parsing model is constructed, and the parsing model is compressed through knowledge distillation technology, and the compressed parsing model is deployed to the edge device in the substation power grid, and the device status information of the edge device is continuously collected, and the compressed parsing model is used to perform semantic parsing on the device status information, and the obtained parsing results are output. Continuous injection of domain knowledge is achieved through dynamic knowledge graph and incremental learning, and semantics are accurately aligned with the multimodal graph attention parsing network to build a multimodal association analysis system, so that complex semantic relationships can be accurately captured even in complex scenarios, the parsing accuracy is improved, and accurate and reliable support is provided for operation and maintenance decisions.
[0212] It should be noted that for other corresponding descriptions of the functional units involved in the semantic parsing device based on power grid domain knowledge mining provided in the embodiment of the present application, please refer to Figure 1A to Figure 1B The corresponding description in will not be repeated here.
[0213] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0214] The above embodiments and the technical features in the embodiments can be arbitrarily combined with each other. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0215] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
[0216] In an exemplary embodiment, see Figure 3 Also provided is an electronic device comprising a bus, a processor, a memory, and a communication interface. The electronic device may also include an input / output interface and a display device, wherein the various functional units can communicate with each other via the bus. The memory stores a computer program, and the processor is configured to execute the program stored in the memory and implement the semantic parsing method based on power grid domain knowledge mining in the above-described embodiment.
[0217] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the semantic parsing method based on power grid domain knowledge mining.
[0218] Through the description of the above implementation methods, those skilled in the art can clearly understand that the present application can be implemented through hardware or by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including a number of instructions for enabling an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each implementation scenario of the present application.
[0219] Those skilled in the art will understand that the accompanying drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the accompanying drawings are not necessarily required to implement the present application.
[0220] Those skilled in the art will appreciate that the modules in the devices in the implementation scenario can be distributed in the devices of the implementation scenario according to the implementation scenario description, or can be modified accordingly and located in one or more devices different from the implementation scenario. The modules in the above implementation scenario can be combined into one module or further split into multiple submodules.
[0221] The above application serial numbers are for description only and do not represent the advantages or disadvantages of the implementation scenarios.
[0222] The above disclosure only describes several specific implementation scenarios of the present application. However, the present application is not limited thereto, and any changes that can be conceived by those skilled in the art should fall within the scope of protection of the present application.
Claims
1. A semantic parsing method based on power grid domain knowledge mining, characterized in that: include: Based on the original corpus related to the power grid field, a pre-trained corpus is constructed, and a dynamic knowledge graph and incremental learning fusion system are used to continuously update the corpus of the pre-trained corpus; Using a multimodal graph attention parsing network, the corpus in the pre-trained corpus is subjected to unified representation processing of multimodal corpus data and cross-modal semantic alignment processing; Building a parsing model based on the processed pre-training corpus, and compressing the parsing model through knowledge distillation technology; The compressed parsing model is deployed to the edge device in the substation power grid, and the device status information of the edge device is continuously collected. The compressed parsing model is used to perform semantic parsing on the device status information, and the obtained parsing results are output.
2. The method according to claim 1, characterized in that The pre-training corpus is constructed based on the original corpus related to the power grid field, including: Collecting information related to the power grid field as the original corpus, and extracting a structured data stream from the original corpus; Perform word segmentation on the structured data stream, generate a multidimensional word vector using the processed structured data stream, calculate the attention head output result of the multidimensional word vector using the following formula, obtain a preset number of attention head output results, and splice the preset number of attention head output results to obtain a universal semantic representation, Among them, Q i =EW i Q , K i =EW i K , V i =EW i V , head i represents the output of the i-th attention head, softmax represents the activation function used to convert the input value into a probability distribution, W i Q 、W i K 、W i V is a trainable parameter, E represents the multidimensional word vector, d k is the key vector K i Dimensions, is the key vector K i The transposed matrix of According to the connection relationship of the equipment in the substation power grid, an adjacency matrix is generated, the general semantic representation is integrated into the initial node feature matrix in the graph convolution network, and the adjacency matrix is processed using the integrated graph convolution network to obtain the topology enhanced feature vector H between the equipment in the substation network. gcn , where the formula for each layer of graph convolution operation in the graph convolution network is as follows: Among them, H (l+1) represents the node feature matrix obtained after the l+1th layer graph convolution operation, σ represents the activation function used to perform nonlinear transformation on the calculation results, represents the adjacency matrix with self-loops added, represents the degree matrix and The i-th diagonal element is equal to the adjacency matrix with the self-loop added The sum of the elements in the i-th row of Degree matrix The square root of the inverse of (l) Represents the node feature matrix before the l-th layer graph convolution operation; W (l) Represents the trainable weight matrix of the l-th layer graph convolution operation; after the last layer of graph convolution operation in the graph convolution network is calculated, the topology enhancement feature vector H is obtained gcn ; Combined with the topological enhancement feature vector H gcn , the following loss function is used to perform adversarial training on the preset domain classifier D and feature generator F to obtain the adjusted feature generator F: in, represents the loss function of adversarial training, Represents the expectation operator; Determine the pluggable parameter matrix M, and use the following formula to enhance the topology feature vector H through the pluggable parameter matrix M gcn Perform linear transformation to obtain the transformed topology enhancement feature vector H gcn : H adapt =H gcn ·M+b Among them, H adapt Represents the topological enhancement feature vector H after linear transformation gcn , b represents the preset bias vector; Use the adjusted feature generator F to H adapt Processing is performed to obtain the vector H of the fusion domain characteristics output by the adjusted feature generator F final , and using the vector H final Construct the pre-training corpus.
3. The method according to claim 2, characterized in that The collecting of information related to the power grid field as the original corpus and extracting a structured data stream from the original corpus includes: Collect information related to the power grid field, use all the collected information as the original corpus, and calculate the global timestamp using the following formula. The original corpus includes monitoring system telesignaling data, work order system transaction table data, and equipment parameter document data: t global =max(t local ,t received ,Δ netwdrk ) Among them, t global represents the global timestamp, the t local Indicates the local clock time, Δ netwdrk Indicates the dynamic network delay measured by two-way ping; Initializing a sliding window of a preset length, and capturing corpora whose local clock time is within the sliding window in the original corpus as corpora to be aligned; Performing time alignment on the corpus to be aligned according to the global timestamp, and outputting the aligned corpus to be aligned as a structured data stream; The sliding window is slid according to a preset sliding time interval, and the corpus in the sliding window whose local clock is in the sliding window after the sliding is recaptured in the original corpus as a new corpus to be aligned, and the new corpus to be aligned is time-aligned according to the global timestamp and output as a structured data stream, until all the original corpus is traversed, and the structured data stream extraction of the original corpus is completed.
4. The method according to claim 1, wherein The dynamic knowledge graph and incremental learning fusion system are used to continuously update the pre-training corpus, including: Obtain the text sequence to be updated, and call the hybrid model constructed in advance with reference to the dynamic knowledge graph strategy, so that the hybrid model calculates the entity boundary probability of the text sequence to be updated using the following formula: P(y t |h t )=softmax(W crf h t +b) Among them, P(y t |h t ) represents the entity boundary probability, y t represents the predicted label of the text sequence to be updated at time step t, h t represents the hidden state of BiLSTM at time step t in the hybrid model, W crf represents the weight matrix of the CRF layer in the hybrid model, and b represents the preset bias vector; Based on the entity boundary probability, the main entity vector e is extracted from the text sequence to be updated. s , customer entity vector e o and the context vector e c , the hybrid model uses the following formula to calculate the inter-entity relationship feature vector of the text sequence to be updated: p=maxpool(CNN([e s ;e o ;e c ])) Wherein, p represents the feature vector of the relationship between entities, maxpool represents the maximum pooling function, CNN represents the convolutional neural network function, and e s Represents the main entity vector of the text sequence to be updated, e o Represents the object entity vector of the text sequence to be updated, e c A context vector representing the text sequence to be updated; Using the main entity vector e s , the object entity vector e o and the inter-entity relationship feature vector p, constructing a triple of the text sequence to be updated, wherein, if the above process fails to generate the triple for the text sequence to be updated, using the incremental learning fusion system Few-shot to generate a pseudo label for the text sequence to be updated, and using the pseudo label as the triple of the text sequence to be updated; Extract multiple non-conflicting relationships from the triples, construct a path set based on the triples, calculate a score for each non-conflicting relationship using the following formula, and determine the target non-conflicting relationship with the highest score: Wherein, score(r) represents the score of the currently calculated non-conflicting relation r, p represents a path in the path set, Paths represents the path set, and sim(p, r) represents the semantic similarity between path p and the non-conflicting relation r; The non-conflicting relationships other than the target non-conflicting relationship among the multiple non-conflicting relationships are taken as non-conflicting relationships to be eliminated, the non-conflicting relationships to be eliminated are eliminated from the triples to obtain the eliminated triples, and the pre-training corpus is updated using the eliminated triples.
5. The method according to claim 4, characterized in that The method further comprises: The following formula is used to calculate the edge weights between entities in the continuously updated pre-training corpus, so that the dynamic changes in the relationships between entities in the pre-training corpus are reflected based on the calculated dynamic edge weights: in, represents the edge weight between entity i and entity j at time step t, represents the edge weight between entity i and entity j at time step t-1, α represents the preset forgetting factor, freq ij Indicates the number of times the relationship between entity i and entity j is triggered within a specified time period, Σ k freq ik Represents the sum of the number of relationship triggers between entity i and all other entities.
6. The method according to claim 1, characterized in that The method of using a multimodal graph attention parsing network to perform cross-modal semantic alignment processing on the corpus in the pre-training corpus includes: For each corpus in the pre-training corpus, generate a data object DO node metadata descriptor, wherein the DO node metadata descriptor includes a data type type, a dimension unit, a valid value range lower limit min, a valid value range upper limit max, and a data timestamp timestamp; The following formula is used to generate the DO feature vector at the current moment for each of the corpus: in, represents the DO feature vector of the corpus at the current moment, GRU represents the GRU network formula, x t represents the DO metadata vector of the corpus at time t, Indicates the hidden state of the corpus at the previous moment; The following formula is used to calculate the attention weight between each corpus and its neighboring corpus: Among them, α ij Represents the attention weight between corpus i and neighbor corpus j, LeakyReLU represents the linear rectification function with leakage, a T represents the transpose of the preset attention coefficient vector a, W represents the trainable weight matrix, represents the DO feature vector of the corpus i, represents the DO feature vector of the neighbor corpus j, represents the neighbor corpus set of the corpus i; The following formula of the multimodal graph attention parsing network is used to aggregate the DO feature vectors of all subordinate corpora of each corpus to obtain the subordinate DO feature vector of each corpus: Among them, h ln represents the subordinate DO feature vector, α ij represents the attention weight between corpus i and neighbor corpus j, and W represents the trainable weight matrix represents the DO feature vector of the neighbor corpus j, N(i) represents the neighbor corpus set of the corpus i; The following formula is used to calculate the subordinate DO feature vector of each of the corpus to obtain the gate vector g: Among them, W g represents the trainable weight matrix, and represents the feature vector of each subordinate DO, and σ represents the activation function; The following formula is used to calculate the subordinate DO feature vector of each corpus to obtain the integrated feature vector h ld : Where g represents the gate vector, and Represents each subordinate DO feature vector; Associating and storing the DO node metadata descriptor of each of the corpora with the subordinate DO feature vectors of each of the corpora, thereby completing the unified representation processing of the multimodal corpus data in the pre-training corpus; Obtain the historical data of the substation power grid, extract matching triples (work order text T, equipment E, defective node D) from the historical data, and use the matching triples (work order text T, equipment E, defective node D) to construct a positive sample pair. Among them, T represents the work order text, represents the device feature vector of device E, Defective node feature vector representing defective node D; The following formula is used to calculate the matching triple (work order text T, equipment E, defective node D) to obtain the similarity score of the matching triple (work order text T, equipment E, defective node D): Among them, s represents the similarity score, h T The feature vector representing the work order text T, h E Represents the operating condition vector of equipment E, h D The feature vector representing the defective node D; Refer to the positive sample pair Construct negative sample pairs (T k , E k , D k ), and the following contrastive learning loss function is used to calculate the matching triple (work order text T, equipment E, defective node D) and the negative sample pair (T k , E k , D k ) between sample pairs: in, represents the contrastive learning loss function, τ is used to control the distribution of similarity scores, s represents the similarity score of the matching triple (work order text T, equipment E, defective node D), and k represents the number of negative sample pairs; Referring to the sample pair similarity, cross-modal semantic alignment processing is performed on the corpus in the pre-training corpus.
7. The method according to claim 6, characterized in that The method further comprises: According to the substation power grid, a heterogeneous graph is constructed, and the following formula is used to calculate the feature vector of each node in the heterogeneous graph in the multimodal graph attention parsing network, so that the feature vector of each node contains its local context information, the heterogeneous graph includes multiple nodes, the multiple nodes include device nodes, defect nodes, and treatment measure nodes, the edges in the heterogeneous graph include edges between device nodes and defect nodes, and edges between defect nodes and measure nodes, wherein the weight of the edge between the device node and the defect node is equal to the co-occurrence probability, and the weight of the edge between the defect node and the measure node is equal to the conditional probability: in, represents the feature vector of node i in the l+1th layer of the multimodal graph attention parsing network, σ represents the activation function, R represents the edge set of the edges in the heterogeneous graph, N r (i) represents the set of neighbor nodes of node i under edge r, represents the attention weight of node i to neighbor node j under edge r, represents the weight matrix of edge r in layer l, represents the feature vector of node j at layer l; and / or, According to the substation power grid, determine the threshold control index N, and read the basic similarity threshold τ preset in the multimodal graph attention parsing network base , the basic similarity threshold τ is calculated using the following formula base Calculate and obtain the similarity threshold τ after dynamic adjustment dynamic , and the similarity threshold τ after dynamic adjustment dynamic , the basic similarity threshold τ preset in the multimodal graph attention parsing network base To update: Among them, Δ τ represents a threshold adjustment amount; and / or, The following formula is used to calculate the risk value of the substation power grid and the calculated risk value is output: Among them, Risk(t) represents the risk value of the substation power grid at time t, v wind (t) represents the wind speed of the substation grid at time t, load r ate(t) represents the load rate of the substation power grid at time t.
8. The method according to claim 1, characterized in that The compressing process of the analytical model by using the knowledge distillation technology includes: Constructing a hybrid architecture of a lightweight graph convolutional network and a bidirectional LSTM, migrating the knowledge of the analytical model to the hybrid architecture, and using the hybrid architecture to perform model training on the migrated knowledge to obtain an initial compression model; The KL divergence loss between the initial compression model and the analytical model is calculated using the following formula: in, represents the KL divergence loss value, T represents the temperature coefficient used to control the smoothness of the probability distribution, N represents the number of samples, represents the raw output of the parsing model, represents the original output of the initial compression model, and softmax represents the function used to convert the output into a probability distribution; The following formula is used to calculate the attention matrix alignment loss between the initial compression model and the parsing model: in, represents the attention matrix alignment loss value, L represents the number of attention layers, l represents the index of the attention layer, represents the attention query matrix of the parsing model at layer l, represents the attention query matrix of the initial compression model at layer l; Referring to the KL divergence loss value and the attention matrix alignment loss value, the initial compression model is optimized, and the KL divergence loss value and the attention matrix alignment loss value between the optimized initial compression model and the analytical model are recalculated and the optimization is continued until the initial compression model reaches the preset convergence condition, and the currently optimized initial compression model is used as the compressed analytical model.
9. The method according to claim 1, characterized in that The continuously collecting the device status information of the edge device, performing semantic parsing on the device status information using the compressed parsing model, and outputting the obtained parsing results include: When the device status information of the edge device is collected, the device status information is input into the compressed parsing model, so that the first layer LSTM in the compressed parsing model calculates the device status information using the following formula to obtain the first hidden state vector of the device status information: in, represents the first hidden state vector of the first layer LSTM at time t, x t Indicates the device status information, The hidden state vector representing the device state information at time t-1 in the first LSTM layer; The second hidden state vector is obtained by processing the first hidden state vector using the following formula by the second LSTM layer in the compressed parsing model: in, represents the second hidden state vector of the first hidden state vector in the second layer LSTM at time t, g t represents the gate vector of the second layer LSTM at time t, represents the hidden state vector of the first hidden state vector in the second layer LSTM at time t-1; The second hidden state vector is processed by a multi-layer perceptron to generate a defect response strategy knowledge subgraph, and the confidence of the defect response strategy knowledge subgraph is calculated using the following formula: Among them, c ij represents the confidence between node i and node j in the defect response strategy knowledge subgraph, W c represents the weight matrix, represents the second hidden state vector corresponding to node i, represents the second hidden state vector corresponding to node j; The multi-layer perceptron is used to calculate the priority of each node in the defect response strategy knowledge subgraph using the following formula, and the defect response strategy knowledge subgraph is labeled using the calculated priority of each node: Among them, p i represents the priority of node i in the defect response strategy knowledge subgraph, MLP represents a multi-layer perceptron, represents the second hidden state vector corresponding to node i; The labeled defect response strategy knowledge subgraph and the confidence level are used as the parsing result, and the parsing result is output.
10. A semantic parsing device based on power grid domain knowledge mining, characterized in that: include: A construction module is used to build a pre-trained corpus based on original corpus related to the power grid field, and to continuously update the pre-trained corpus using a dynamic knowledge graph and incremental learning fusion system; An alignment processing module is used to perform unified representation processing of multimodal corpus data and cross-modal semantic alignment processing on the corpus in the pre-training corpus using a multimodal graph attention parsing network; A compression processing module, configured to construct a parsing model based on the processed pre-trained corpus, and to compress the parsing model using a knowledge distillation technique; The parsing module is used to deploy the compressed parsing model to the edge device in the substation power grid, and continuously collect the device status information of the edge device, use the compressed parsing model to perform semantic parsing on the device status information, and output the obtained parsing results.
Citation Information
Cited By
Power distribution network multi-mode fault reasoning method and device and medium
CN121302297A
A method, device and medium for multimodal fault reasoning in power distribution networks
CN121302297B
Power green supply chain scheduling optimization method based on multi-modal data fusion
CN121684547A
Voice interaction method, related device, electronic equipment and storage medium
CN121687050A