Knowledge conversion and fusion processing method and system for massive power grid operation data

By building an event semantic link network and regional ontology knowledge base, the multi-source heterogeneity and semantic inconsistency of massive power grid operation data are solved, the deep integration and intelligent understanding of data are achieved, fault diagnosis and processing efficiency is improved, and the grid intelligence process is promoted.

CN120012884APending Publication Date: 2025-05-16CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +2

Patent Information

Application Number
CN202411611223.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Due to the multi-source heterogeneity, structural complexity and semantic inconsistency, the massive data generated in modern power grid operations are difficult to deeply integrate data and efficiently resolve, hindering the process of intelligent power grids.

Method used

A knowledge conversion and fusion processing method of massive power grid operation data is adopted. By obtaining the independent alarm event collection after deduplication and structured operation event data, timing, spatial association and causal association are performed, event semantic link network is constructed, and regional ontology knowledge base is constructed based on the geographical topology of the distribution network to form a distribution network semantic model to achieve in-depth integration and intelligent understanding of data.

Benefits of technology

It effectively solves the problems of data heterogeneity and processing complexity, improves the accuracy of fault diagnosis and the efficiency of fault processing, provides a solid data foundation and decision-making support, and promotes the power grid to a higher level of safety, operational efficiency and intelligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012884A_ABST
    Figure CN120012884A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge conversion and fusion processing method and system for massive power grid operation data, and the method specifically comprises the steps: obtaining the attributes, including the alarm type, timestamp and severity, of equipment granularity based on alarm event data collected by a power distribution automation terminal, and carrying out the data cleaning and standardization processing, converting the multi-source heterogeneous alarm data into uniform event granularity representation; fusing the alarm event data and the operation event data, carrying out space-time alignment and causal association on the two types of events based on timestamps and equipment ID attributes, constructing a cross-data-source event semantic link network, and carrying out association analysis on a semantic relationship between the events; and according to a fault diagnosis result, historical records related to fault processing are extracted from the operation event knowledge base, similar fault processing modes are matched, and a specific fault processing scheme is generated in combination with factors of severity and influence range of the current fault.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of electric power information technology, and in particular relates to a knowledge conversion and fusion processing method and system for massive power grid operation data. Background Art

[0002] In modern power grid operation scenarios, with the deep integration of information technology and power systems, unprecedented massive data has been generated. These data are not only diverse in source and complex in structure, but also widely distributed in time and space dimensions. In addition, the inconsistency of semantic expression seriously hinders the deep integration and efficient analysis of data, and constitutes a major obstacle in the current process of power grid intelligence. Specifically, the multi-source heterogeneous nature of data leads to the frequent occurrence of information islands, and the lack of a unified semantic framework further limits the full release of data value. In addition, due to the large scale and high dimensionality of data, traditional methods are very likely to fall into a low efficiency trough when processing, and it is difficult to meet the needs of real-time analysis and decision support. Therefore, how to design and implement a technical solution that can cross the heterogeneity of data sources, unify semantic descriptions, efficiently organize and mine the knowledge contained in power grid operation data, and build a semantic association network to achieve deep data integration and intelligent understanding has become a key technical problem that needs to be solved urgently. The breakthrough of this technology will provide a solid data foundation and decision support for the fields of intelligent operation and maintenance, fault prediction and resource optimization allocation of power grids, and is an important step in promoting the power grid to a higher level of safety, operational efficiency and intelligence.

[0003] Prior art document 1 (CN112507035B) discloses a unified standardized processing system and method for multi-source heterogeneous data of transmission lines. However, its shortcoming is that the relevant processing methods for transmission line data are mostly targeted at structured data, ignoring unstructured data such as text data, images, and videos. Unstructured data is not integrated into the transmission line data, and data integration and unified management cannot be achieved. Summary of the invention

[0004] In order to address the deficiencies in the prior art, the present invention provides a knowledge conversion and fusion processing method and system for massive power grid operation data, which can overcome the heterogeneity of data sources, unify semantic descriptions, efficiently organize and mine the knowledge contained in power grid operation data, and build a semantic association network to achieve a technical solution for deep data integration and intelligent understanding. It provides a solid data foundation and decision-making support for the fields of intelligent operation and maintenance, fault prediction and optimal resource allocation of power grids, and promotes the power grid to move towards a higher level of safety, operational efficiency and intelligence.

[0005] The present invention adopts the following technical solution. The present invention provides a method for knowledge conversion and fusion processing of massive power grid operation data, which specifically includes:

[0006] Obtain the deduplicated independent alarm event set and structured operation event data;

[0007] Mapping the operation event data to a unified data model, performing temporal association, spatial association, and causal association, and then vectorizing the association results to obtain an event semantic link network;

[0008] According to the geographical topological structure of the distribution network, a regional ontology knowledge base is constructed, and combined with the event semantic link network, a distribution network semantic model is formed;

[0009] For distribution network fault scenarios, semantic knowledge of faulty equipment is obtained based on the regional ontology knowledge base and event semantic link network in the distribution network semantic model, and a fault handling solution is generated;

[0010] In the real-time fault handling process, the newly discovered equipment association and event causal relationship knowledge in the processing process is recorded in real time and integrated into the regional ontology knowledge base and event semantic link network of the distribution network semantic model. The distribution network semantic model is optimized through incremental learning and active learning to obtain the distribution network semantic model, which is used for knowledge conversion and fusion processing of massive power grid operation data.

[0011] Preferably, obtaining a deduplicated independent alarm event set and structured operation event data includes:

[0012] Collect alarm event data, clean and standardize the alarm event data, and then convert it into a unified event granularity. Match and merge the alarm event data of the unified event granularity according to the attributes of the alarm event data to obtain a deduplicated independent alarm event set. The attributes of the alarm event data include the device ID, alarm type, timestamp, and severity at the device granularity.

[0013] Natural language processing is used to perform text mining on the unstructured operation events uploaded by the distribution automation terminal, and the unstructured operation events are converted into structured representations of event granularity to obtain structured operation event data.

[0014] Preferably, the unstructured operation event includes operation logs and voice recording data, and the specific implementation of structuring the voice recording data includes:

[0015] Perform speech recognition processing on the collected speech recording data to obtain a text sequence corresponding to the speech;

[0016] Perform natural language processing on the text sequence corresponding to the speech, identify the entity types of time, place, equipment and personnel in the text sequence according to the preset feature template and annotation rules, and obtain the text sequence of the speech based on the entity recognition model;

[0017] According to the pre-built rule library and dictionary library, structured event attribute fields are extracted from the text sequence of speech based on the entity recognition model. The event attribute fields include time, place, equipment and personnel.

[0018] The attribute values ​​of the extracted structured event attribute fields are standardized, and the standardized attribute values ​​are converted into structured voice recording data.

[0019] Preferably, the operation event data is mapped to a unified data model, and temporal association, spatial association, and causal association are performed, and then the association results are vectorized to obtain an event semantic link network, including:

[0020] Perform data cleaning and data conversion on the deduplicated independent alarm event set and operation event data to obtain heterogeneous data, map the heterogeneous data to a unified data model and format, and obtain alarm event data and operation events in a unified data model and format;

[0021] Construct the time series, spatial and causal associations of alarm event data and operation events in a unified data model and format;

[0022] Alarm events and operation events with unified data models and formats are used as nodes of the graph, and the temporal associations, spatial associations, and causal associations between events are used as edges of the graph. The Cypher query language is used to map event data to the graph model, and a semantic link network across data sources is constructed.

[0023] By vectorizing the event nodes in the semantic link network across data sources, an event semantic link network is obtained.

[0024] Preferably, the time series association, spatial association and causal association of alarm event data and operation events in a unified data model and format are constructed, including:

[0025] Set a time window of fixed length, sort the alarm events and operation events of the unified data model and format by timestamp, and for each alarm event of the unified data model and format, search for the operation event of the unified data model and format within its time window, and establish a temporal association between the two;

[0026] Use device topology information to build a device connection diagram, find the adjacency relationship between devices through depth-first search, and spatially associate alarm events and operation events with unified data models and formats on the same device and adjacent devices;

[0027] The event association sequence obtained by time-space alignment is matched with the rule base, the causal relationship between alarm events and operation events in a unified data model and format is identified, the event causal chain is constructed, and the causal association is obtained.

[0028] Preferably, building a regional ontology knowledge base includes:

[0029] Obtain the geographic topology data of the distribution network and extract the attribute information of the equipment, including the voltage level of the substation, the length of the feeder, the current capacity, the model and status of the switch;

[0030] Parse the collected distribution network geographic topology equipment attribute information data and extract the equipment coordinates, convert the equipment coordinates into geometric objects, extract the spatial geometric information of the equipment based on the geometric objects, find the adjacency relationship between the equipment, and build the spatial topology model of the distribution network equipment;

[0031] Converting the spatial topological model of the distribution network equipment into an RDF graph, mapping the spatial topological relationship into an RDF triple, and obtaining the RDF triple;

[0032] Convert the RDF triples into an OWL ontology to obtain an OWL ontology;

[0033] OWL ontology is used to define the logical relationship between entity concepts, formally define the ontology relationship, obtain the OWL ontology description logic, and build a regional ontology knowledge base.

[0034] Preferably, for the distribution network fault scenario, the semantic knowledge of the faulty device is obtained according to the regional ontology knowledge base and the event semantic link network in the distribution network semantic model, the fault cause and impact scope are determined, and a fault diagnosis report is generated, including:

[0035] Receive a fault alarm message carrying a unique identification ID of the fault event, the fault alarm message is sent by a real-time data interface of the distribution automation system, and obtain the fault alarm message with the unique identification ID of the fault event;

[0036] According to the fault alarm message with the unique identification ID of the fault event, the semantic knowledge associated with the fault alarm message with the unique identification ID of the fault event is searched in the event semantic link network to obtain the semantic knowledge associated with the event;

[0037] According to the semantic knowledge associated with the event, the spatial position coordinates of the faulty device are obtained, and the power distribution equipment within the set range is searched to obtain the power distribution equipment within the set range;

[0038] Determine whether current and voltage are abnormal, and if so, determine the equipment within the set range affected by the fault to obtain the impact range of the fault;

[0039] Determine the impact scope of the outage of the faulty equipment on downstream users according to the impact scope of the fault, the impact scope on downstream users includes the power outage scope and important user types, and obtain the impact scope of the outage of the faulty equipment on downstream users;

[0040] According to the impact scope of the outage of the faulty equipment on downstream users, historical alarm events directly related to the current faulty equipment are extracted from the event semantic link network to obtain the diagnosis results and processing solutions of the historical alarm events;

[0041] According to the diagnosis results and processing schemes of the historical alarm events, the fault diagnosis results are obtained by integrating the alarm information of the faulty equipment, the real-time operation status, the analysis results of historical similar events and the multi-source heterogeneous knowledge of the status of the equipment within the set range;

[0042] A fault diagnosis report is generated according to the result of the fault diagnosis.

[0043] Preferably, according to the fault diagnosis result of generating the fault diagnosis report, the historical records of fault handling are extracted from the operation event knowledge base, and a fault handling solution is generated in combination with the current fault, including:

[0044] Extract information from the fault diagnosis report, including the name of the faulty device, fault type, and cause analysis. Based on the extracted information, semantic modeling is performed in the operation event knowledge base using an ontology to define the concepts and attributes of fault type, device type, and processing effect, forming a domain ontology.

[0045] Instantiate and match the statement template with the domain ontology to obtain historical operation events related to the current fault;

[0046] For the historical operation events related to the current fault, by converting the fault information and the text description of the operation event into a vector representation, calculating the similarity with the current fault, and obtaining the historical case with the highest similarity with the current fault;

[0047] Extracting elements including operation method, tool usage, personnel arrangement and time consumption from the historical case with the highest degree of similarity to the current fault, generating a decision tree from fault to operation, and obtaining a preliminary generated fault handling solution;

[0048] With respect to the initially generated fault handling solution, combined with resource restrictions of fault repair, the initial fault handling solution is optimized with fault recovery time and resource consumption as optimization targets to obtain a fault handling solution.

[0049] Preferably, in the real-time fault handling process, the newly discovered equipment association and event causal relationship knowledge in the processing process is recorded in real time and integrated into the regional ontology knowledge base and event semantic link network of the distribution network semantic model, and the distribution network semantic model is optimized through incremental learning and active learning to obtain the distribution network semantic model, including:

[0050] Based on the on-site operation data collected during the execution of the fault handling work order, multimodal perception is used to extract the operation steps to form a structured record and obtain the fault handling text;

[0051] From the fault handling text, according to predefined entity and relationship templates, entity attributes and entity associations are extracted from subject-verb-object, attributive and adverbial structures, and a dictionary and relationship template library are constructed to obtain new fault handling data;

[0052] For the new fault processing data, an ontology fragment that can represent the new fault processing process is constructed, and the ontology fusion technology is used to construct the mapping relationship between the new and old ontologies. The newly discovered equipment association information and event causal relationship knowledge are compared and integrated with the original regional ontology knowledge base to form a new regional fault ontology knowledge representation, and a new regional ontology knowledge base is obtained;

[0053] For the event semantic link network, the newly discovered events, device nodes and their relationship edges are added to the original network topology, and a new event semantic link network is obtained by dynamically updating the event semantic link network;

[0054] By dynamically optimizing the new regional ontology knowledge base and the new event semantic link network, the distribution network semantic model is updated through incremental learning and active learning, and the distribution network semantic model after knowledge conversion and fusion is obtained.

[0055] A second aspect of the present invention provides a system for knowledge conversion and fusion processing of massive power grid operation data, which is used to execute the method for knowledge conversion and fusion processing of massive power grid operation data, including:

[0056] Data acquisition module: used to obtain the deduplicated independent alarm event set and structured operation event data;

[0057] Event semantic link network construction module: used to map the deduplicated independent alarm event set and structured operation event data to a unified data model, and then perform temporal association, spatial association, and causal association, and then vectorize the association results to obtain the event semantic link network;

[0058] Distribution network semantic model construction module: used to build a regional ontology knowledge base based on the geographical topology of the distribution network, and combine the event semantic link network to form a distribution network semantic model;

[0059] Fault diagnosis report generation module: for distribution network fault scenarios, the module obtains the semantic knowledge of faulty equipment from the regional ontology knowledge base and event semantic link network in the distribution network semantic model, determines the fault cause and impact range, and generates a fault diagnosis report;

[0060] Fault handling solution generation module: for distribution network fault scenarios, the module obtains semantic knowledge of faulty equipment based on the regional ontology knowledge base and event semantic link network in the distribution network semantic model, and generates a fault handling solution;

[0061] Distribution network semantic model optimization module: It is used to record the newly discovered equipment association and event causal relationship knowledge in real time during the implementation of the fault handling plan and integrate it into the regional ontology knowledge base and event semantic link network of the distribution network semantic model. It also optimizes the distribution network semantic model through incremental learning and active learning to obtain the distribution network semantic model, which is used for knowledge conversion and fusion processing of massive power grid operation data.

[0062] Preferably, obtaining a deduplicated independent alarm event set and structured operation event data includes:

[0063] Collect alarm event data, clean and standardize the alarm event data, and then convert it into a unified event granularity. Match and merge the alarm event data of the unified event granularity according to the attributes of the alarm event data to obtain a deduplicated independent alarm event set. The attributes of the alarm event data include the device ID, alarm type, timestamp, and severity at the device granularity.

[0064] Natural language processing is used to perform text mining on the unstructured operation events uploaded by the distribution automation terminal, and the unstructured operation events are converted into structured representations of event granularity to obtain structured operation event data.

[0065] Preferably, the unstructured operation event includes operation logs and voice recording data, and the specific implementation of structuring the voice recording data includes:

[0066] Perform speech recognition processing on the collected speech recording data to obtain a text sequence corresponding to the speech;

[0067] Perform natural language processing on the text sequence corresponding to the speech, identify the entity types of time, place, equipment and personnel in the text sequence according to the preset feature template and annotation rules, and obtain the text sequence of the speech based on the entity recognition model;

[0068] According to the pre-built rule library and dictionary library, structured event attribute fields are extracted from the text sequence of speech based on the entity recognition model. The event attribute fields include time, place, equipment and personnel.

[0069] The attribute values ​​of the extracted structured event attribute fields are standardized, and the standardized attribute values ​​are converted into structured voice recording data.

[0070] Preferably, the operation event data is mapped to a unified data model, and temporal association, spatial association, and causal association are performed, and then the association results are vectorized to obtain an event semantic link network, including:

[0071] Perform data cleaning and data conversion on the deduplicated independent alarm event set and operation event data to obtain heterogeneous data, map the heterogeneous data to a unified data model and format, and obtain alarm event data and operation events in a unified data model and format;

[0072] Construct the time series, spatial and causal associations of alarm event data and operation events in a unified data model and format;

[0073] Alarm events and operation events with unified data models and formats are used as nodes of the graph, and the temporal associations, spatial associations, and causal associations between events are used as edges of the graph. The Cypher query language is used to map event data to the graph model, and a semantic link network across data sources is constructed.

[0074] By vectorizing the event nodes in the semantic link network across data sources, an event semantic link network is obtained.

[0075] Preferably, the time series association, spatial association and causal association of alarm event data and operation events in a unified data model and format are constructed, including:

[0076] Set a time window of fixed length, sort the alarm events and operation events of the unified data model and format by timestamp, and for each alarm event of the unified data model and format, search for the operation event of the unified data model and format within its time window, and establish a temporal association between the two;

[0077] Use device topology information to build a device connection diagram, find the adjacency relationship between devices through depth-first search, and spatially associate alarm events and operation events with unified data models and formats on the same device and adjacent devices;

[0078] The event association sequence obtained by time-space alignment is matched with the rule base, the causal relationship between alarm events and operation events in a unified data model and format is identified, the event causal chain is constructed, and the causal association is obtained.

[0079] Preferably, building a regional ontology knowledge base includes:

[0080] Obtain the geographic topology data of the distribution network and extract the attribute information of the equipment, including the voltage level of the substation, the length of the feeder, the current capacity, the model and status of the switch;

[0081] Parse the collected distribution network geographic topology equipment attribute information data and extract the equipment coordinates, convert the equipment coordinates into geometric objects, extract the spatial geometric information of the equipment based on the geometric objects, find the adjacency relationship between the equipment, and build the spatial topology model of the distribution network equipment;

[0082] Converting the spatial topological model of the distribution network equipment into an RDF graph, mapping the spatial topological relationship into an RDF triple, and obtaining the RDF triple;

[0083] Convert the RDF triples into an OWL ontology to obtain an OWL ontology;

[0084] OWL ontology is used to define the logical relationship between entity concepts, formally define the ontology relationship, obtain the OWL ontology description logic, and build a regional ontology knowledge base.

[0085] Preferably, for the distribution network fault scenario, the semantic knowledge of the faulty device is obtained according to the regional ontology knowledge base and the event semantic link network in the distribution network semantic model, the fault cause and impact scope are determined, and a fault diagnosis report is generated, including:

[0086] Receive a fault alarm message carrying a unique identification ID of the fault event, the fault alarm message is sent by a real-time data interface of the distribution automation system, and obtain the fault alarm message with the unique identification ID of the fault event;

[0087] According to the fault alarm message with the unique identification ID of the fault event, the semantic knowledge associated with the fault alarm message with the unique identification ID of the fault event is searched in the event semantic link network to obtain the semantic knowledge associated with the event;

[0088] According to the semantic knowledge associated with the event, the spatial position coordinates of the faulty device are obtained, and the power distribution equipment within the set range is searched to obtain the power distribution equipment within the set range;

[0089] Determine whether current and voltage are abnormal, and if so, determine the equipment within the set range affected by the fault to obtain the impact range of the fault;

[0090] Determine the impact scope of the outage of the faulty equipment on downstream users according to the impact scope of the fault, the impact scope on downstream users includes the power outage scope and important user types, and obtain the impact scope of the outage of the faulty equipment on downstream users;

[0091] According to the impact scope of the outage of the faulty equipment on downstream users, historical alarm events directly related to the current faulty equipment are extracted from the event semantic link network to obtain the diagnosis results and processing solutions of the historical alarm events;

[0092] According to the diagnosis results and processing schemes of the historical alarm events, the fault diagnosis results are obtained by integrating the alarm information of the faulty equipment, the real-time operation status, the analysis results of historical similar events and the multi-source heterogeneous knowledge of the status of the equipment within the set range;

[0093] A fault diagnosis report is generated according to the result of the fault diagnosis.

[0094] Preferably, according to the fault diagnosis result of generating the fault diagnosis report, the historical records of fault handling are extracted from the operation event knowledge base, and a fault handling solution is generated in combination with the current fault, including:

[0095] Extract information from the fault diagnosis report, including the name of the faulty device, fault type, and cause analysis. Based on the extracted information, semantic modeling is performed in the operation event knowledge base using an ontology to define the concepts and attributes of fault type, device type, and processing effect, forming a domain ontology.

[0096] Instantiate and match the statement template with the domain ontology to obtain historical operation events related to the current fault;

[0097] For the historical operation events related to the current fault, by converting the fault information and the text description of the operation event into a vector representation, calculating the similarity with the current fault, and obtaining the historical case with the highest similarity with the current fault;

[0098] Extracting elements including operation method, tool usage, personnel arrangement and time consumption from the historical case with the highest degree of similarity to the current fault, generating a decision tree from fault to operation, and obtaining a preliminary generated fault handling solution;

[0099] With respect to the initially generated fault handling solution, combined with resource restrictions of fault repair, the initial fault handling solution is optimized with fault recovery time and resource consumption as optimization targets to obtain a fault handling solution.

[0100] Preferably, in the real-time fault handling process, the newly discovered equipment association and event causal relationship knowledge in the processing process is recorded in real time and integrated into the regional ontology knowledge base and event semantic link network of the distribution network semantic model, and the distribution network semantic model is optimized through incremental learning and active learning to obtain the distribution network semantic model, including:

[0101] Based on the on-site operation data collected during the execution of the fault handling work order, multimodal perception is used to extract the operation steps to form a structured record and obtain the fault handling text;

[0102] From the fault handling text, according to predefined entity and relationship templates, entity attributes and entity associations are extracted from subject-verb-object, attributive and adverbial structures, and a dictionary and relationship template library are constructed to obtain new fault handling data;

[0103] For the new fault processing data, an ontology fragment that can represent the new fault processing process is constructed, and the ontology fusion technology is used to construct the mapping relationship between the new and old ontologies. The newly discovered equipment association information and event causal relationship knowledge are compared and integrated with the original regional ontology knowledge base to form a new regional fault ontology knowledge representation, and a new regional ontology knowledge base is obtained;

[0104] For the event semantic link network, the newly discovered events, device nodes and their relationship edges are added to the original network topology, and a new event semantic link network is obtained by dynamically updating the event semantic link network;

[0105] By dynamically optimizing the new regional ontology knowledge base and the new event semantic link network, the distribution network semantic model is updated through incremental learning and active learning, and the distribution network semantic model after knowledge conversion and fusion is obtained.

[0106] The beneficial effect of the present invention is that, compared with the prior art, the present invention discloses a knowledge conversion and fusion processing method for massive power grid operation data. For multi-source heterogeneous alarm data collected from distribution automation terminals, the alarm data of different formats and sources are converted into a unified event granularity representation through data cleaning and standardization processing, thereby ensuring the accuracy of the alarm information; text mining is performed on the operation logs and voice records uploaded by the distribution automation terminal to extract key event elements such as time, location, equipment and personnel, and these unstructured data are converted into structured event granularity representations, thereby improving the availability of data; by using timestamps and device ID attributes The integration of alarm event data and operation event data is carried out in time and space, and causally related, to build a semantic link network of events across data sources, and use the graph database to store the semantic relationship between these events, realizing efficient association analysis; according to the geographical topology of the distribution network, a regional ontology knowledge base is built, and the entity concepts such as substations, feeders, switches, and their spatial adjacency and logical connection relationships are defined using the ontology language OWL, forming a semantic model of the distribution network and deepening the understanding of the relationship between events; in short, it effectively solves the problems of data heterogeneity and processing complexity in the distribution automation system, and improves the accuracy of fault diagnosis and the efficiency of fault handling. At the same time, through dynamic update and learning mechanisms, it can continuously optimize system performance, adapt to the dynamic changes of the distribution network, and improve the intelligence and automation level of the entire distribution system. BRIEF DESCRIPTION OF THE DRAWINGS

[0107] Figure 1 A flow chart of a method for knowledge conversion and fusion processing of massive power grid operation data according to the present invention;

[0108] Figure 2A schematic diagram of a method for knowledge conversion and fusion processing of massive power grid operation data according to the present invention;

[0109] Figure 3 It is another schematic diagram of a knowledge conversion and fusion processing method of massive power grid operation data of the present invention. DETAILED DESCRIPTION

[0110] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0111] like Figure 1 As shown, Embodiment 1 of the present invention proposes a method for knowledge conversion and fusion processing of massive power grid operation data, which specifically includes:

[0112] Step 1: Obtain the deduplicated independent alarm event set and structured operation event data;

[0113] Step 1.1: Figure 2 As shown, the distribution automation terminal collects multi-source heterogeneous alarm event data, cleans and standardizes the multi-source heterogeneous alarm event data, and then converts it into a unified event granularity. At the same time, the alarm event data of the unified event granularity are matched and merged to obtain a deduplicated independent alarm event set.

[0114] Obtain alarm event data collected by distribution automation terminals;

[0115] Clean and standardize the acquired alarm event data, convert the multi-source heterogeneous alarm data into a unified event granularity representation, and obtain multi-source heterogeneous alarm data of unified granularity;

[0116] Through the attributes of the alarm event data, multi-source heterogeneous alarm data of uniform granularity are matched and merged to form a unified event representation, and a deduplicated independent alarm event set is obtained.

[0117] The attributes of the alarm event data include device ID, alarm type, timestamp and severity at device granularity.

[0118] Specifically, in the data preprocessing stage, regular expressions can be used to perform pattern matching and extraction on the alarm description text to filter out noise data such as non-ASCII characters and special symbols. For timestamp fields, they can be uniformly converted to standard date and time formats, such as ISO8601. The severity can be mapped to a numerical representation, such as 4 for severe, 3 for high, 2 for medium, and 1 for low. String similarity algorithms, such as edit distance or Jaccard similarity, can be used to perform fuzzy matching on key fields such as device ID and alarm type, and the similarity threshold is set to 0.8. Data above this threshold is considered to be the same event.

[0119] Step 1.2: Apply natural language processing methods to perform text mining on the operation logs and voice records uploaded by the distribution automation terminal, extract structured, standardized and consistent event attribute fields, convert unstructured operation events into structured representations of event granularity, and obtain structured operation event data.

[0120] Step 1.2.1: Obtain unstructured operation logs and convert the unstructured operation logs into structured operation log data;

[0121] Unstructured operation logs and voice recording data, including unstructured text and audio;

[0122] Operation logs and voice recording data are collected through distribution automation terminals, and distributed storage methods are used to efficiently store and manage massive amounts of unstructured data, converting unstructured operation logs into structured representations at event granularity; the distributed computing framework is used to perform denoising and format conversion preprocessing on raw data, and extract key features of audio and text;

[0123] The distribution automation terminal collects operation logs and voice recording data, including unstructured text and audio. The HDFS distributed file system is used to store unstructured data, with a storage capacity of up to PB level and horizontal expansion support. The MapReduce distributed computing framework is used to pre-process the raw data, and the audio data is denoised and key features are extracted through digital signal processing algorithms such as FFT. The MFCC feature vector of the audio is extracted. This feature vector represents the spectral characteristics of the sound and is used to identify different sound patterns; the text data is segmented and stop words are removed, and the TF-IDF feature vector of the text is extracted. The TF-IDF feature vector quantifies the importance of vocabulary in the document by calculating TF (Term Frequency) and IDF (Inverse Document Frequency), which helps to identify keywords related to faults, maintenance or special events, thereby quickly locating and classifying key information in the log.

[0124] The key features of the text (such as TF-IDF feature vectors) are mainly used to help quickly locate and classify key information in the log. For example, these features can help identify key words in the text, especially those related to failures, maintenance, or special events.

[0125] Step 1.2.2: Obtain unstructured voice recording data, convert the unstructured voice recording data into structured voice recording data, and combine it with the structured operation log data obtained in step 1.2.1 into structured operation event data.

[0126] Step 1.2.2.1 performs speech recognition processing on the collected unstructured speech recording data to obtain a text sequence corresponding to the speech;

[0127] A domain-specific language model is built based on the professional terminology and language characteristics in the power distribution field, and speech recognition processing is performed on the uploaded voice recording data to convert the voice signal into a text sequence.

[0128] For professional terms in the field of power distribution, such as "trip" and "maintenance", a Transformer neural network language model is constructed. By pre-training on a large amount of field-related corpus, the model learns how to map text to vector representations in a high-dimensional feature space. These vectors contain the semantic information and syntactic structure of the vocabulary. When applied to speech recognition, the model can convert audio streams into text, and through the optimization of the CTC loss function, it ensures a high recognition accuracy rate, and can maintain good performance even in mixed background noise. Pre-training was performed on 5 million corpora in the field of power distribution, and the vocabulary coverage of the language model reached more than 95%. The language model is used to perform speech recognition on speech data, and the CTC loss function is used for end-to-end training. The word error rate of speech recognition is less than 5%.

[0129] Step 1.2.2.2: Perform natural language processing on the text sequence corresponding to the speech obtained in step 1.2.2.1, and identify different types of entities such as time, place, equipment, and personnel in the text sequence according to the preset feature template and annotation rules, and obtain the text sequence corresponding to the speech based on the entity recognition model;

[0130] Natural language processing methods are used to analyze the text sequence corresponding to the speech, and different feature templates and annotation rules are designed for entities of different types such as time, place, equipment, and personnel.

[0131] Natural language processing is performed on the text sequence obtained by speech recognition, and the BiLSTM-CRF named entity recognition model is adopted. For the four entity types of time, place, equipment, and personnel, feature templates such as regular expressions and dictionary matching are designed. The model is trained with manually annotated corpus, and the F1 value of entity recognition reaches more than 90%.

[0132] Step 1.2.2.3: According to the pre-built rule library and dictionary library, extract structured event attribute fields from the text sequence corresponding to the speech based on the entity recognition model obtained in step 1.2.2.2, including time, place, equipment, and personnel;

[0133] Step 1.2.2.4: Standardize the attribute values ​​of the event attribute fields of step 1.2.2.3, and convert the attribute values ​​into a standardized and consistent representation.

[0134] According to the business rules in the power distribution field, a rule base and dictionary base for extracting operation event elements are constructed, and structured time, location, equipment, and personnel attribute fields are extracted from the operation event description text; based on the ontology Schema, the concepts and relationships of events, time, location, equipment, and personnel are defined, the extracted event element attribute values ​​are standardized, and the extracted event element data are verified and repaired.

[0135] Based on the 200 summarized operating procedures and 5,000 professional vocabulary, a rule base and a dictionary base were constructed, and 20 sentence templates were designed using the dependency syntax analysis method to extract structured event attribute fields from the text. Based on the ontology schema, the extracted attribute values ​​are mapped to standardized concepts, such as the time format is unified as "YYYY-MM-DDHH:MM:SS", and the location is unified as a triplet of substation-voltage level-equipment number. Data quality detection algorithms such as value range detection and consistency detection are used to verify event attribute data, repair missing values ​​or abnormal values, and monitor data quality. The data quality pass rate reaches more than 98%. The structured event data is stored in the MongoDB database, and the query-oriented data model and index are used to control the average query response time within 100ms, providing efficient data support for distribution operation behavior analysis and risk warning applications.

[0136] Step 2: Integrate the deduplicated independent alarm event set obtained in step 1 with the obtained operation event data. Based on the timestamp and device ID attributes, align the two types of events in time and space and causally associate them to build an event semantic link network across data sources. At the same time, perform association analysis on the semantic relationships between events to obtain an event semantic link network.

[0137] Step 2.1: Preprocess the deduplicated independent alarm event set and operation event data, clean and convert the preprocessed data to obtain heterogeneous data, map the heterogeneous data to a unified data model and format, obtain alarm event data and operation events in a unified data model and format, and form a standardized event attribute representation;

[0138] Specifically, in the preprocessing stage, regular expressions can be used to extract the timestamps in the alarm events and operation events, such as "2023-04-2014:30:25", and key attributes such as device ID, from the deduplicated independent alarm event set and operation event data. Then, data cleaning methods are used on the extracted key attribute data, such as removing missing values ​​and outliers, to ensure data quality. Then, data conversion methods are used on the key attribute data after data cleaning, such as one-hot encoding to convert discrete attributes into numerical attributes to obtain heterogeneous data, and then the heterogeneous data is mapped to a unified data model to form a standardized event attribute representation.

[0139] Step 2.2: Construct the time series association, spatial association and causal association of the alarm event data and operation events in the unified data model and format obtained in step 2.1;

[0140] Based on the timestamp attribute, a fixed-length time window is set, and the alarm events and operation events of the unified data model and format are sorted by timestamp. For each alarm event, the operation event within its time window is searched to establish a temporal association between the two.

[0141] A fixed time window of 30 minutes can be set, and a double-pointer sliding window algorithm can be used to sort alarm events and operation events by timestamp. For each alarm event, operation events are searched within 30 minutes before and after it to establish a temporal association between the two.

[0142] Based on the device ID attributes and using the device topology information, a device connection diagram is constructed. The adjacency relationship between devices is found through the depth-first search algorithm, and the alarm events and operation events on the same device or adjacent devices are spatially associated. When spatially associating, the depth-first search algorithm can be used, starting from the alarm event device, traversing its one-hop and two-hop neighbor devices in the topology diagram, and associating the events on the same device or adjacent devices.

[0143] The event association sequence obtained by time-space alignment is matched with the rule base to identify the causal relationship between events and build an event causal chain. Causal association can construct IF-THEN rules based on expert knowledge, such as "IF the transformer temperature is too high THEN the transformer fails". The forward reasoning algorithm is used to match the event sequence aligned in time and space with the rules to identify the causal chain.

[0144] Step 2.3: Use the alarm events and operation events with unified data model and format as the nodes of the graph, and the temporal association, spatial association, and causal association between events obtained in step 2.2 as the edges of the graph. Use the Cypher query language to map event data to the graph model, and build an initial semantic link network across data sources.

[0145] With the help of Neo4j, a powerful and flexible graph database platform, the Cypher query language is used to map massive event information into nodes in the graph, and the complex relationships between events are carefully depicted as edges connecting these nodes, thus building a large and sophisticated semantic link network. This network not only intuitively shows the distribution and connection of events, but also implies the deep-level relationships and patterns between events. In order to further explore the intrinsic connection between events, the shortest path algorithm is used to find the most direct connection path between two points in the graph. By analyzing the shortest path between alarm events and operation events, the context of a series of event developments can be revealed, the root causes that may lead to failures can be identified, and accurate clues can be provided for fault troubleshooting and prevention. This process is like laying a clear navigation route in a complex event network, making the root cause analysis of failures both intuitive and efficient. In addition, the Louvain algorithm can be used to identify closely connected node groups, namely communities, in the graph. By dividing event nodes into communities, a series of typical alarm scenarios can be identified, which often represent specific types of failure modes or operation scenarios. The event nodes within each community reveal common characteristics and behavior patterns due to their high connectivity, providing an important perspective for understanding the nature of event clusters.

[0146] Step 2.4: By vectorizing the event nodes in the initial event semantic link network in step 2.3, a semantic link network is obtained, which can predict the implicit relationship between events.

[0147] Using the TransE model in the OpenKE library, each event node and relationship edge is transformed into a compact 50-dimensional vector representation. These vectors not only contain the intrinsic characteristics of the event, but also imply the complex interaction patterns between events. The TransE model assumes that the vector representation of two related events should satisfy a translation invariance principle in the vector space, that is, the vector of one event plus the vector representing the relationship between the two should be close to the vector of the other event. This principle is achieved by minimizing the Euclidean distance between vectors, thereby capturing the correlation and similarity of events in a multi-dimensional space. With the help of the TransE model, a high-dimensional vector space can be constructed, in which each vector becomes the semantic embodiment of the event, and the distance between vectors reflects the closeness of the event association, which greatly enriches our understanding of the intrinsic structure of events and opens up new paths for predicting potential associations between events. By calculating the distance between vectors, we can intuitively identify which events have frequently co-occurred in history, which events are tightly coupled in time or space, and even predict possible event combinations in the future.

[0148] Step 3: According to the geographic topology of the distribution network, a regional ontology knowledge base is constructed to define the entity concepts of substations, feeders, and switches in the regional ontology knowledge base, as well as the spatial adjacency and logical connection relationships between them. The ontology language OWL is used to represent the semantic knowledge of regional granularity, and the event semantic link network obtained in step 2 is combined to form a distribution network semantic model, and finally the distribution network semantic model is obtained.

[0149] Step 3.1: Obtain the geographical topology data of the distribution network and extract the attribute information of the equipment, including the voltage level of the substation, the length and current capacity of the feeder, and the model and status of the switch;

[0150] The geographic topology refers to the spatial location and connection relationship of physical equipment (such as substations, feeders, switches, etc.) in the distribution network. It describes the physical layout and connection method of the power grid.

[0151] When constructing the geographic topology structure, the device ID, spatial coordinates, topological connection and attribute information are obtained.

[0152] Events are linked to specific devices in the geographic topology through device IDs and spatial coordinates. For example, if an alarm event is associated with a specific device ID, the location and connection relationship of the device in the geographic topology can be queried to analyze the possible impact of the event.

[0153] Distribution network geographic information data collection can be obtained through the data interface of the power GIS system, such as WebService, API, etc., in XML, JSON and other formats, including the spatial coordinates, topological connections, and attribute information of substations, feeders, switches, etc. By calling the GetFeature interface, the layer name and attribute field of the substation are passed in, and the geographic location coordinates, name, voltage level, capacity and other attributes of the substation are returned.

[0154] Step 3.2: Parse the collected distribution network geographic topology device attribute information data and extract the device coordinates, convert the device coordinates into geometric objects, extract the spatial geometric information of the devices based on the geometric objects, find the adjacency relationship between the devices, and build the spatial topology model of the distribution network equipment;

[0155] Parse and extract the acquired JSON format data, use Python's Shapely library to convert device coordinates into geometric objects such as points, lines, and polygons, extract the spatial geometric information of the device, find the adjacency relationship between devices, and build a spatial topology model.

[0156] Step 3.3: Convert the spatial topology model in step 3.2 into an RDF graph, map the spatial topology relationship into an RDF triple, form the basic data of the distribution network semantic model, and obtain the RDF triple;

[0157] R-tree index and STRtree algorithm are used to perform spatial partitioning and hierarchical nesting according to the MBR (Minimum Bounding Rectangle) of the equipment geometric objects in the spatial topology model, so as to achieve fast spatial adjacency query with a time complexity of O(logn). In the construction of the ontology knowledge base, the D2RQ tool is used to map the distribution network equipment data table in MySQL into RDF triples, such as <substation instance, name, "110kV substation A">, <substation instance, voltage level, "110kV">, etc.

[0159] Step 3.4: Based on the physical characteristics and logical functions of the distribution network equipment, define the entity concepts and attributes in the ontology knowledge base, convert the RDF triples into OWL ontology, form an instance of the distribution network knowledge base, and obtain the OWL ontology;

[0160] Then, the ontology visualization interface of protege software is used to import RDF data into OWL ontology, and the classes of substation, feeder, switch, etc. and their DP (Data Property) and OP (Object Property) are defined.

[0161] Step 3.5: Use OWL ontology to define the logical relationship between entity concepts, formally define the ontology relationship, form a distribution network semantic model, obtain the OWL ontology description logic, construct a regional ontology knowledge base, and obtain a regional ontology knowledge base;

[0162] The regional ontology knowledge base is constructed according to the geographical topology of the distribution network, and the concepts of substation, feeder, switch entity, as well as the spatial adjacency and logical connection relationship between them are defined. The ontology language OWL is used to represent the semantic knowledge of regional granularity.

[0163] Step 3.6: Associate the regional ontology knowledge base with the event semantic link network to obtain the topological structure and equipment relationship of the distribution network and obtain the distribution network semantic model.

[0164] As a key object attribute, "supplyPowerTo" clearly indicates the energy transmission relationship from the substation to the feeder. Its domain is set to "substation" and its range corresponds to "feeder". Using Pellet, an OWLDL reasoning engine, the entire ontology is comprehensively checked and verified to timely discover possible logical conflicts or inconsistencies, ensuring the high quality and availability of the ontology. At the same time, Pellet also provides DLQuery function, which can execute SPARQL queries, such as

[0165] The statement "SELECT?station?lineWHERE{?station:supplyPowerTo?line.?stationrdf:type:Subst ation.}" retrieves all substations and their corresponding power feeders to form a preliminary distribution network semantic model. Based on the preliminary distribution network semantic model, the regional ontology knowledge base is combined with the event semantic link network. Specifically, the events in the preliminary distribution network semantic model are connected to the specific devices in the geographic topology through the device ID and spatial coordinates to analyze the possible impact of the events. The events in the event semantic link network are associated with the entities in the regional ontology knowledge base, and the specific devices involved in the event are determined through the device ID and spatial coordinates. The impact range of the event is further analyzed based on the location and connection relationship of these devices in the ontology. Combining the information of the two, an in-depth exploration and understanding of the power grid structure is achieved, forming a complete distribution network semantic model.

[0166] You can also use GeoServer, an open source geospatial data server, to publish the semantic data of the distribution network through its WFS (WebFeatureService) service. With the support of the open source JavaScript library of OpenLayers, you can load the WFS layer on the map to show the geographical distribution of the power grid. You can also use the ol.format.GeoJSON module to parse feature attributes to ensure the data integrity of each geographic feature. Furthermore, by integrating the ol.interaction.Select interactive module, you can click on any feature on the map and display the information box. This function enables users to understand and analyze the semantic knowledge of the distribution network in an intuitive and interactive way, thus providing a powerful auxiliary tool for power grid operation and maintenance decision-making.

[0167] Step 4: For distribution network fault scenarios, obtain semantic knowledge about spatially adjacent devices related to the faulty device, historical alarm events of the same type, and potential impact range from the regional ontology knowledge base of the distribution network semantic model formed in step 3 and the event semantic link network in step 3, determine the cause of the fault and the impact range, and generate a fault diagnosis report.

[0168] Step 4.1: receiving a fault alarm message carrying a unique ID of a fault event, wherein the fault alarm message is sent by a real-time data interface of a distribution automation system;

[0169] Specifically, when a fault occurs in the distribution network, the real-time data interface of the distribution automation system is used to subscribe to the fault alarm event message through WebSocket and other methods. The latest fault record can also be obtained from the historical alarm event table through regular database queries, including the time, location, equipment name, alarm type, etc. of the fault.

[0170] Step 4.2: According to the unique identification ID of the fault event in step 4.1, query the semantic knowledge associated with the event in the event semantic link network;

[0171] Through the unique identification ID of the fault event, the semantic knowledge of other events, devices, regions, etc. associated with the event is queried in the event semantic link network.

[0172] Step 4.3: Obtain the semantic information of the spatial location coordinates of the faulty device in step 4.2, and search for the power distribution equipment within the set range;

[0173] According to the spatial location coordinates of the faulty equipment, common spatial data index structures such as R-tree and quadtree are used to quickly find the adjacent distribution equipment within the set range through the calculation of spatial distance and intersection relationship. The specific implementation can use the spatial index function provided by the PostGIS plug-in of the PostgreSQL database.

[0174] Step 4.4: Obtain semantic information of the power distribution equipment within the setting range of step 4.3, determine whether current and voltage abnormalities occur, and if abnormalities occur, determine the equipment within the setting range affected by the fault, and further clarify the impact range of the fault;

[0175] Obtain semantic information such as the type, status, and electrical connection relationship of these devices to determine the set range of devices that the fault may affect.

[0176] Step 4.5: Obtain semantic knowledge of the devices within the set range affected by the fault in step 4.4, and determine the potential impact of the outage of the faulty equipment on downstream users;

[0177] Based on the semantic knowledge such as the distribution network topology and power supply range in the regional ontology knowledge base, the reasoning API provided by the Apache Jena framework, such as the OWLReasoner class, is used to load the distribution network ontology model and rule files. Then, an ontology model with reasoning function is created through the InfModel interface. SPARQL query statements are run to determine the potential impact of the outage of faulty equipment on downstream users, such as the power outage range and important user types, and to estimate the severity and impact of the fault.

[0178] Step 4.6: Obtain the semantic information of historical alarm events of the potential impact scope of the faulty equipment outage on downstream users in step 4.5, and analyze the diagnosis results and treatment solutions of these historical alarm events;

[0179] In the event semantic link network, with the fault event as the central node, graph traversal algorithms such as breadth-first search and shortest path search are used to extract historical alarm events that are directly or indirectly related to the fault event, focusing on events that are the same or similar to the current faulty equipment type and alarm type, analyzing the diagnostic results and processing solutions of these historical events, and drawing on their lessons.

[0180] Step 4.7: Infer the cause of the fault by integrating the multi-source heterogeneous knowledge such as the diagnosis results of historical similar events obtained in step 4.6 of the faulty device, alarm information, the real-time operating status of the faulty device, and the status of devices within the set range of the faulty device;

[0181] According to the empirical rules and decision trees of distribution network fault diagnosis, a fault diagnosis rule base is constructed. A combination of forward reasoning and reverse reasoning is used to integrate the alarm information of the faulty equipment, real-time operating status, analysis results of historical similar events, status of adjacent equipment and other multi-source heterogeneous knowledge to infer the possible causes of the fault, such as hardware damage of the equipment itself, external environmental factors, and adjacent equipment failure. The fuzzy comprehensive evaluation method is used to sort the fault causes obtained by diagnosis. For each cause, evaluation indicators are set from the aspects of faulty equipment status, historical similar events, and expert experience. A membership score between 0 and 1 is given, and then weighted summation is performed to obtain the comprehensive membership of each cause, which is sorted from large to small. At the same time, a risk matrix is ​​constructed with reference to the IEC60812 standard. A 5x5 matrix is ​​constructed from the two dimensions of the severity of the fault impact and the probability of occurrence. The severity is divided into 5 levels from negligible to catastrophic, and the probability of occurrence is divided into 5 levels from almost impossible to very frequent. The risk matrix is ​​formed by combining two values. The larger the value, the higher the risk level. The scope and severity of the fault impact corresponding to each cause are evaluated. According to the results of fault diagnosis, a fault diagnosis report is automatically generated, including the time and location of the fault, the name and type of the equipment involved, the description of the alarm information, the judged fault cause and its possibility ranking, the estimated fault impact range and the number of power outage users, the processing priority according to the risk matrix, the recommended fault isolation plan and emergency repair measures, etc., for reference and decision-making by distribution dispatchers. The WebSocket message stream is used to subscribe to fault alarm events to ensure real-time capture of network fluctuations, such as switch tripping. Whenever such an event is triggered, the system immediately receives a JSON format message, which includes key data such as event ID, precise timestamp, specific switch number, and specific type of tripping. In the construction of the ontology knowledge base, the D2RQ tool is used to map the distribution network equipment data table into RDF triples, such as <substation instance, name, "110kV substation A">, <substation instance, voltage level, "110kV">, etc., to form an instance of the ontology knowledge base. Then, using the ST_DWithin function in the PostGIS tool, with the faulty switch as the center point and a radius of 500 meters, the surrounding distribution equipment nodes are accurately located. By examining the attributes such as equipment type, rated voltage and operating status, the impact of the fault on the adjacent facilities is deeply analyzed. Next, with the help of the Jena reasoning engine, relying on the OWL document and SWRL rule set of the distribution network entity, logical expressions such as "faulty device (?x)^connection (?x,?y)->affected device (?y)" are used to infer the power outage area that may be caused by the switch fault. Execute SPARQL query statements

[0182] "SELECT?feeder?stationWHERE{?switchX:locate?feeder.?feeder:connect?station}" can obtain the names of the affected feeders and substations. Using the Cypher query language of the Neo4j graph database, starting from the fault event ID, traverse the "MATCH(n)-

[0183] ->(m)WHEREn.event_id={id}RETURNm", trace the related events of similar faults in history, and use machine learning models such as k-NN and random forest to deeply analyze the feature vectors of past cases, providing strong support for the classification and cause identification of the current fault.

[0184] Step 4.8: Generate a fault diagnosis report based on the fault diagnosis results of step 4.7.

[0185] It can also conduct detailed comparative scoring for eight indicators including fault frequency, maintenance cycle, and economic profit and loss, generate weight vectors, and clarify the priority sequence of various causes. With the help of probabilistic risk assessment (PRA), the risk level of each cause is quantified from the perspectives of accident severity and potential probability, and then a detailed risk matrix chart is drawn to form a comprehensive and accurate fault diagnosis report.

[0186] Step 5: Based on the results of the fault diagnosis in step 4, extract the historical records related to fault handling from the operation event knowledge base and match the handling methods of similar faults; combine the severity and impact scope of the current fault to generate a fault handling plan.

[0187] Step 5.1: Extract information from the fault diagnosis report. The extracted information includes the name of the faulty device, the fault type and the cause analysis. For the extracted information, semantic modeling is performed in the operation event knowledge base using an ontology to define the concepts and attributes of the fault type, device type and processing effect, thus forming a domain ontology.

[0188] Specifically, according to the fault equipment name, fault type, cause analysis and other information in the fault diagnosis report, the historical operation records related to the current fault are retrieved in the operation event knowledge base. The operation event knowledge base adopts the ontology method for semantic modeling, defines concepts and attributes such as fault type, equipment type, operation steps, and processing effect, and forms a domain ontology.

[0189] Step 5.2: Based on the extracted keywords of fault equipment, phenomenon, and cause, instantiate and match the sentence template with the domain ontology of step 5.1 to obtain historical operation events related to the current fault;

[0190] Using SPARQL statement templates, we instantiate and match the concepts in the ontology based on keywords such as the equipment, phenomenon, and cause of the current fault, focusing on operation events with the same fault equipment type, similar fault phenomenon, and the same cause judgment.

[0191] Step 5.3: For the historical fault handling events retrieved in step 5.2, convert the text description of the fault information and the operation event into a vector representation, calculate the similarity, and obtain historical cases similar to the current fault;

[0192] The retrieved historical fault handling events are similarity calculated and sorted. A content-based recommendation algorithm is used to convert the text descriptions of fault information and operation events into vector representations. The similarity between them is calculated using metrics such as cosine similarity to obtain the Top-N historical cases most similar to the current fault.

[0193] Step 5.4: Extract key operating methods, tool usage, personnel arrangements, and time consumption factors from similar historical fault cases in step 5.3, use a decision tree algorithm to generate a decision tree from fault to operation, and obtain a recommended fault handling solution;

[0194] For the sorted similar historical fault cases, analyze their processing and effects one by one according to the similarity from high to low, extract key operating steps, tool use, personnel arrangement, time consumption and other factors, and form fault handling experience knowledge in the case library. Associate the operating elements extracted from the case library with the attributes of the current fault, build a decision table, each row represents a case, and the columns are fault attributes and operating elements, and then use decision tree algorithms such as C4.5, CART, etc. to select the optimal partitioning attributes according to information gain or Gini index, generate a decision tree from fault to operation, and then substitute the attribute value of the current fault into the tree to obtain the recommended operating steps and form a preliminary fault handling plan.

[0195] Step 5.5: Based on the preliminary fault handling plan in step 5.4, combined with the resource constraints of fault repair, with fault recovery time and resource consumption as optimization goals, the preliminary fault handling plan is optimized to obtain a fault handling plan.

[0196] The generated fault handling scheme is feasibility analyzed and optimized. Considering the resource constraints of fault repair, such as spare parts, manpower arrangement, weather conditions, etc., a multi-objective genetic algorithm is used. The fault recovery time and resource consumption are optimized, and the manpower, materials, transportation, etc. are used as constraints. The initial population is randomly generated, and the evolution is carried out through genetic operators such as selection, crossover, and mutation. Then, the non-inferior solution on the Pareto front is selected by non-dominated sorting, congestion, etc., for dispatchers to use. The optimized fault handling scheme is issued to the on-site repair personnel in the form of a work order. The workflow engine such as Activiti and Flowable is used to convert the fault handling scheme into a process definition file, which contains various task nodes, execution order, assigned roles, etc., and then the process file is parsed by the engine and instantiated into an executable work order instance, which is then pushed to the mobile APP through Web services, message queues, etc. The key information in the scheme, such as fault judgment, operation steps, precautions, completion time, etc., is displayed to the repair personnel in an interactive and friendly way to guide them to carry out on-site processing. When extracting key information such as the faulty equipment name, fault type, and cause analysis from the fault diagnosis report, regular expressions such as "\w+The equipment has a (\w+) fault, and the reason is due to (\w+)" can be used to extract key fields. At the same time, named entity recognition models such as BiLSTM-CRF are used to identify key entities such as faulty equipment, fault type, cause analysis, and their categories by training on the power equipment fault corpus, and the extraction accuracy can reach more than 90%. For the operation event knowledge base, the ontology construction tool Protege is used to define class concepts such as fault type and equipment type, and attributes such as hasOperation and hasTool, forming a domain ontology containing 120 concepts and 55 attributes. According to the extracted keywords of faulty equipment, phenomenon, and cause, SPARQL statements are used to instantiate and match with the concepts in the ontology, and an average of more than 85% of relevant historical operation events are retrieved. For the retrieved historical fault handling events, the text descriptions of fault information and operation events are converted into 200-dimensional vector representations by using the Word2Vec word embedding model. The cosine similarity measurement method is used to calculate the similarity, and the Top-10 historical cases with the highest similarity to the current fault are obtained, with an average similarity of more than 8. Key elements are extracted from similar historical fault cases, and the C5 decision tree algorithm is used to select split attributes through the information gain ratio to generate a fault-to-operation decision tree with an average depth of 8. 15% of the test faults are recommended with an accuracy of 85%.For the generated fault handling scheme, the NSGA-II multi-objective genetic algorithm is used, with the population size set to 50, the crossover probability to 8, the mutation probability to 1, and 200 generations of iterations. The fault recovery time and resource consumption are optimized, and the genetic operators such as tournament selection, single-point crossover, and uniform mutation are evolved. The non-inferior solution on the Pareto front is selected by fast non-dominated sorting and congestion calculation, which shortens the fault recovery time by 20% and saves resource consumption by 15% on average. The optimized scheme is converted into a BPMN0 standard process definition file, deployed to the Flowable workflow engine, instantiated into an executable work order containing 5 to 10 task nodes, and pushed to the mobile APP of the repair personnel to guide on-site processing in an interactive and friendly way.

[0197] Step 6: During the implementation of the fault handling solution in step 5, the newly discovered equipment associations during the processing are recorded in real time, the regional ontology knowledge base and the event semantic link network are dynamically updated, and the newly discovered equipment associations and event causal relationship knowledge are fed back to the regional ontology knowledge base. The distribution network semantic model is optimized through incremental learning and / or active learning methods to achieve knowledge conversion and fusion, and finally the optimized distribution network semantic model is obtained. Figure 3 shown.

[0198] Step 6.1: Extract key information from the field operation data collected during the execution of the fault handling work order to form a structured record and obtain the fault handling text;

[0199] Specifically, during the execution of fault handling work orders, on-site operation data is collected through IoT devices, mobile terminals, etc., and information such as personnel, time, location, operation content, tools used, etc. of each operation step is recorded in real time. Multimodal perception methods such as voice recognition and image recognition are used to extract key information to form structured records.

[0200] Step 6.2: Parse the fault handling text in step 6.1, extract entity attributes and entity associations from the subject-predicate-object, attributive, and adverbial structures according to the predefined entity and relationship templates, construct a dictionary and relationship template library, and obtain new fault handling data;

[0201] Parse the fault handling text and extract entity attributes and entity associations from structures such as subject, predicate, object, attributive, and adverbial. Through the knowledge graph representation method, the core entities in the fault handling job record, such as equipment, parts, materials, personnel, tools, etc. and relationships, such as located, connected, used, and replaced, are formed into RDF triples to construct a job-specific ontology knowledge fragment. The ontology knowledge fragment constructed in this step analyzes the newly discovered facts and phenomena in the fault handling process, and extracts entities such as equipment, parts, materials, alarms, and operations from unstructured text through knowledge extraction methods such as named entity recognition and relationship extraction. Use a transfer-based syntactic analyzer, such as the Stanford parser, to perform dependency syntactic tree parsing on the fault handling text, and extract entity attributes and entity associations from structures such as subject, predicate, object, attributive, and adverbial according to predefined entity and relationship templates to build a dictionary and relationship template library specific to the fault handling field.

[0202] The dictionary and relation template library constructed in step 6.2 are built based on newly discovered facts and phenomena, while step 6.3 is the process of integrating this newly discovered knowledge with the existing regional ontology knowledge base.

[0203] Construct an ontology fragment that can represent the new fault handling process, and use ontology fusion technology to compare and fuse the newly discovered equipment association information and event causal relationship knowledge with the original regional ontology knowledge base to form an expanded and consistent regional fault ontology knowledge representation. The ontology fragment of the new fault handling process here is closely related to the dictionary and relationship template library, but it involves fusing these newly discovered knowledge with the existing regional ontology knowledge base to form an expanded and consistent regional fault ontology knowledge representation.

[0204] Step 6.3: For the new fault processing data obtained in step 6.2, construct an ontology fragment that can represent the new fault processing process, use the ontology fusion technology to build the mapping relationship between the new and old ontologies, compare and fuse the newly discovered equipment association information and event causal relationship knowledge with the original regional ontology knowledge base, form an expanded and consistent regional fault ontology knowledge representation, and obtain a consistent regional ontology knowledge base;

[0205] The ontology fragments of the new fault handling process refer to the newly discovered facts and phenomena in the fault handling process, which are extracted from unstructured text through knowledge extraction methods (such as named entity recognition, relationship extraction, etc.) and constructed through syntactic analysis (such as dependency syntactic analysis). These ontology fragments are usually represented in the form of RDF triples, including the core entities in the fault handling operation record (such as equipment, parts, materials, personnel, tools, etc.) and the relationships between them (such as located, connected, used, replaced, etc.).

[0206] Newly discovered device association information and event causal relationship knowledge refer to the newly discovered association information between devices and the causal relationship between events during the troubleshooting process. Device association information may include the connection relationship and location relationship between devices. Event causal relationship knowledge may include the cause of the event, the scope of impact, the temporal relationship between events, etc.

[0207] The relationship between the ontology fragments of the new fault handling process and the newly discovered equipment association information and event causal relationship knowledge is that the ontology fragments are formed by structuring and formalizing these newly discovered knowledge. The newly discovered equipment association information and event causal relationship knowledge are the basis for constructing ontology fragments, and ontology fragments are the structured representation of these knowledge. Through ontology fusion technology, these newly discovered knowledge will be integrated into the original regional ontology knowledge base to form an expanded and consistent regional fault ontology knowledge representation.

[0208] By adopting ontology fusion technology and drawing on ontology matching and fusion tools such as FCA-Merge and PROMPT, a mapping relationship between the new and old ontologies is established through ontology concept semantic similarity calculation and ontology concept logical relationship reasoning. The newly discovered equipment association information, event causal relationship and other knowledge are compared and integrated with the original regional ontology knowledge base to form an expanded and consistent regional fault ontology knowledge representation.

[0209] Step 6.4: For the event semantic link network, the newly discovered events, device nodes and their relationship edges obtained in step 6.2 are added to the original network topology, and a new event semantic link network is obtained by dynamically updating the event semantic link network;

[0210] For the event semantic link network, newly discovered events, device nodes and their relationship edges are added to the original network topology. When adding new nodes and edges to the event semantic link network, the existing concept nodes and relationship edges are reused in combination with the semantic similarity with the existing nodes. The batch import function of the graph database is used to convert the extracted entity and relationship tuples into vertices and edges of the graph model, set the corresponding attributes, and add them to the graph database at one time to ensure the connectivity and completeness of the network topology.

[0211] Step 6.5: By dynamically optimizing the consistent regional ontology knowledge base obtained in step 6.3 and the new event semantic link network obtained in step 6.4, the distribution network semantic model is updated through incremental learning and active learning methods.

[0212] Through incremental learning and / or active learning methods, the regional ontology knowledge base is dynamically optimized based on new fault processing samples, and the distribution network semantic model is updated. The distribution network semantic model is updated on the basis of the original model by learning new fault processing samples and knowledge feedback.

[0213] Dynamically updating the regional ontology knowledge base and event semantic link network and updating the distribution network semantic model through incremental learning and / or active learning methods are complementary processes and are usually carried out simultaneously. First, based on the newly discovered fault processing data, the regional ontology knowledge base and event semantic link network are updated to reflect the latest equipment association information and event causal relationship.

[0214] Then, using this updated knowledge, the distribution network semantic model is optimized through incremental learning and / or active learning methods, so that the model can better adapt to new fault scenarios.

[0215] Using incremental learning technologies such as incremental decision trees and incremental SVM, on the basis of the original fault diagnosis and processing model, new fault processing samples and knowledge feedback are used for distribution network semantic model training and optimization. Through sample enhancement and parameter fine-tuning, the distribution network semantic model is adapted to new fault scenarios, improving the accuracy and efficiency of fault diagnosis and processing. At the same time, active learning methods are introduced. Through uncertainty sampling strategies, the operation records that have the greatest impact on the prediction results of the existing distribution network semantic model are screened out from massive historical data such as operation and maintenance work orders and operation records. Then, high-quality annotations are performed and added to the sample set for incremental model training, which improves the model performance in a targeted manner, so that incremental learning and active learning can cooperate with each other to optimize the distribution network semantic model training process.

[0216] Comprehensively use technologies such as ontology reasoning, graph computing, and cognitive computing to dynamically optimize the regional ontology knowledge base and event semantic link network. Formulate quality metrics such as consistency, completeness, and simplicity based on ontology. Regularly evaluate and improve the regional fault ontology through technologies such as ontology verification and ontology pruning. Dynamically update the fault diagnosis rule base and case base to adapt the knowledge representation and reasoning mechanism to new fault handling processes and experiences. On the basis of offline evaluation and online feedback, continuously track key performance indicators such as classification accuracy and response time of fault diagnosis and processing models, dynamically update the model through methods such as feature selection and hyperparameter tuning, perform model version management, record the input and output of each model iteration and effect evaluation, and facilitate the backtracking and auditing of fault diagnosis and processing processes. During the fault handling process, wearable devices such as smart bracelets can automatically collect the voice commands and hand operations of operators, and combine the deep learning model DBNet to perform text detection and recognition on the work order image of 1 frame per minute, extract equipment nameplates, indications, alarm information, etc., and form a structured information flow with an average length of 20 words. Then, the spaCy toolkit was used to perform dependency syntax analysis on the information flow, and a job knowledge graph containing 500 faulty equipment entities, 800 processing operation entities, and 1,200 entity relationships was constructed. The existing regional ontology was matched, and the Doc2Vec unsupervised learning algorithm was used to calculate the semantic similarity between the newly added entities and the ontology concepts. Those greater than 0.8 were included in the same concept, and those less than 0.5 were added as new concepts. The Cypher statement of the graph database Neo4j was used to import the newly discovered device nodes, relationship edges, and attributes at a speed of 200,000 per second and integrate them into the original event link network.

[0217] Dynamically updating the regional ontology knowledge base and the event semantic link network can make the distribution network semantic model more adaptable to new fault scenarios and processing procedures, and improve the accuracy and efficiency of fault diagnosis and processing. By continuously learning new knowledge, the distribution network semantic model can better understand the complex relationship between fault events, thereby making more accurate fault diagnosis and proposing more effective processing solutions.

[0218] The beneficial effect of the present invention is that, compared with the prior art, the present invention discloses a knowledge conversion and fusion processing method for massive power grid operation data. For multi-source heterogeneous alarm data collected from distribution automation terminals, the alarm data of different formats and sources are converted into a unified event granularity representation through data cleaning and standardization processing, thereby ensuring the accuracy of the alarm information; text mining is performed on the operation logs and voice records uploaded by the distribution automation terminal to extract key event elements such as time, location, equipment and personnel, and these unstructured data are converted into structured event granularity representations, thereby improving the availability of data; by using timestamps and device ID attributes The integration of alarm event data and operation event data is carried out in time and space, and causally related, to build a semantic link network of events across data sources, and use the graph database to store the semantic relationship between these events, realizing efficient association analysis; according to the geographical topology of the distribution network, a regional ontology knowledge base is built, and the entity concepts such as substations, feeders, switches, and their spatial adjacency and logical connection relationships are defined using the ontology language OWL, forming a semantic model of the distribution network and deepening the understanding of the relationship between events; in short, it effectively solves the problems of data heterogeneity and processing complexity in the distribution automation system, and improves the accuracy of fault diagnosis and the efficiency of fault handling. At the same time, through dynamic update and learning mechanisms, it can continuously optimize system performance, adapt to the dynamic changes of the distribution network, and improve the intelligence and automation level of the entire distribution system.

[0219] Embodiment 2 of the present invention provides a system for knowledge conversion and fusion processing of massive power grid operation data, which is used to execute the method for knowledge conversion and fusion processing of massive power grid operation data described in embodiment 1, including:

[0220] Data acquisition module: used to obtain the deduplicated independent alarm event set and structured operation event data;

[0221] Event semantic link network construction module: used to map the deduplicated independent alarm event set and structured operation event data to a unified data model, and then perform temporal association, spatial association, and causal association, and then vectorize the association results to obtain the event semantic link network;

[0222] Distribution network semantic model construction module: used to build a regional ontology knowledge base based on the geographical topology of the distribution network, and combine the event semantic link network to form a distribution network semantic model;

[0223] Fault diagnosis report generation module: for distribution network fault scenarios, the module obtains the semantic knowledge of faulty equipment from the regional ontology knowledge base and event semantic link network in the distribution network semantic model, determines the fault cause and impact range, and generates a fault diagnosis report;

[0224] Fault handling solution generation module: for distribution network fault scenarios, the module obtains semantic knowledge of faulty equipment based on the regional ontology knowledge base and event semantic link network in the distribution network semantic model, and generates a fault handling solution;

[0225] Distribution network semantic model optimization module: It is used to record the newly discovered equipment association and event causal relationship knowledge in real time during the implementation of the fault handling plan and integrate it into the regional ontology knowledge base and event semantic link network of the distribution network semantic model. It also optimizes the distribution network semantic model through incremental learning and active learning to obtain the distribution network semantic model, which is used for knowledge conversion and fusion processing of massive power grid operation data.

[0226] Obtain the deduplicated independent alarm event set and structured operation event data, including:

[0227] Collect alarm event data, clean and standardize the alarm event data, and then convert it into a unified event granularity. Match and merge the alarm event data of the unified event granularity according to the attributes of the alarm event data to obtain a deduplicated independent alarm event set. The attributes of the alarm event data include the device ID, alarm type, timestamp, and severity at the device granularity.

[0228] Natural language processing is used to perform text mining on the unstructured operation events uploaded by the distribution automation terminal, and the unstructured operation events are converted into structured representations of event granularity to obtain structured operation event data.

[0229] Unstructured operation events include operation logs and voice recording data. The specific implementation of structuring voice recording data includes:

[0230] Perform speech recognition processing on the collected speech recording data to obtain a text sequence corresponding to the speech;

[0231] Perform natural language processing on the text sequence corresponding to the speech, identify the entity types of time, place, equipment and personnel in the text sequence according to the preset feature template and annotation rules, and obtain the text sequence of the speech based on the entity recognition model;

[0232] According to the pre-built rule library and dictionary library, structured event attribute fields are extracted from the text sequence of speech based on the entity recognition model. The event attribute fields include time, place, equipment and personnel.

[0233] The attribute values ​​of the extracted structured event attribute fields are standardized, and the standardized attribute values ​​are converted into structured voice recording data.

[0234] Map the operation event data to a unified data model, perform temporal association, spatial association, and causal association, and then vectorize the association results to obtain an event semantic link network, including:

[0235] Perform data cleaning and data conversion on the deduplicated independent alarm event set and operation event data to obtain heterogeneous data, map the heterogeneous data to a unified data model and format, and obtain alarm event data and operation events in a unified data model and format;

[0236] Construct the time series, spatial and causal associations of alarm event data and operation events in a unified data model and format;

[0237] Alarm events and operation events with unified data models and formats are used as nodes of the graph, and the temporal associations, spatial associations, and causal associations between events are used as edges of the graph. The Cypher query language is used to map event data to the graph model, and a semantic link network across data sources is constructed.

[0238] By vectorizing the event nodes in the semantic link network across data sources, an event semantic link network is obtained.

[0239] Construct the time series, spatial and causal associations of alarm event data and operation events in a unified data model and format, including:

[0240] Set a time window of fixed length, sort the alarm events and operation events of the unified data model and format by timestamp, and for each alarm event of the unified data model and format, search for the operation event of the unified data model and format within its time window, and establish a temporal association between the two;

[0241] Use device topology information to build a device connection diagram, find the adjacency relationship between devices through depth-first search, and spatially associate alarm events and operation events with unified data models and formats on the same device and adjacent devices;

[0242] The event association sequence obtained by time-space alignment is matched with the rule base, the causal relationship between alarm events and operation events in a unified data model and format is identified, the event causal chain is constructed, and the causal association is obtained.

[0243] Build a regional ontology knowledge base, including:

[0244] Obtain the geographic topology data of the distribution network and extract the attribute information of the equipment, including the voltage level of the substation, the length of the feeder, the current capacity, the model and status of the switch;

[0245] Parse the collected distribution network geographic topology equipment attribute information data and extract the equipment coordinates, convert the equipment coordinates into geometric objects, extract the spatial geometric information of the equipment based on the geometric objects, find the adjacency relationship between the equipment, and build the spatial topology model of the distribution network equipment;

[0246] Converting the spatial topological model of the distribution network equipment into an RDF graph, mapping the spatial topological relationship into an RDF triple, and obtaining the RDF triple;

[0247] Convert the RDF triples into an OWL ontology to obtain an OWL ontology;

[0248] OWL ontology is used to define the logical relationship between entity concepts, formally define the ontology relationship, obtain the OWL ontology description logic, and build a regional ontology knowledge base.

[0249] For distribution network fault scenarios, the semantic knowledge of the faulty equipment is obtained based on the regional ontology knowledge base and event semantic link network in the distribution network semantic model, the fault cause and impact range are determined, and a fault diagnosis report is generated, including:

[0250] Receive a fault alarm message carrying a unique identification ID of the fault event, the fault alarm message is sent by a real-time data interface of the distribution automation system, and obtain the fault alarm message with the unique identification ID of the fault event;

[0251] According to the fault alarm message with the unique identification ID of the fault event, the semantic knowledge associated with the fault alarm message with the unique identification ID of the fault event is searched in the event semantic link network to obtain the semantic knowledge associated with the event;

[0252] According to the semantic knowledge associated with the event, the spatial position coordinates of the faulty device are obtained, and the power distribution equipment within the set range is searched to obtain the power distribution equipment within the set range;

[0253] Determine whether current and voltage are abnormal, and if so, determine the equipment within the set range affected by the fault to obtain the impact range of the fault;

[0254] Determine the impact scope of the outage of the faulty equipment on downstream users according to the impact scope of the fault, the impact scope on downstream users includes the power outage scope and important user types, and obtain the impact scope of the outage of the faulty equipment on downstream users;

[0255] According to the impact scope of the outage of the faulty equipment on downstream users, historical alarm events directly related to the current faulty equipment are extracted from the event semantic link network to obtain the diagnosis results and processing solutions of the historical alarm events;

[0256] According to the diagnosis results and processing schemes of the historical alarm events, the fault diagnosis results are obtained by integrating the alarm information of the faulty equipment, the real-time operation status, the analysis results of historical similar events and the multi-source heterogeneous knowledge of the status of the equipment within the set range;

[0257] A fault diagnosis report is generated according to the result of the fault diagnosis.

[0258] According to the fault diagnosis result of generating the fault diagnosis report, the fault handling history is extracted from the operation event knowledge base, and combined with the current fault, a fault handling solution is generated, including:

[0259] Extract information from the fault diagnosis report, including the name of the faulty device, fault type, and cause analysis. Based on the extracted information, semantic modeling is performed in the operation event knowledge base using an ontology to define the concepts and attributes of fault type, device type, and processing effect, forming a domain ontology.

[0260] Instantiate and match the statement template with the domain ontology to obtain historical operation events related to the current fault;

[0261] For the historical operation events related to the current fault, by converting the fault information and the text description of the operation event into a vector representation, calculating the similarity with the current fault, and obtaining the historical case with the highest similarity with the current fault;

[0262] Extracting elements including operation method, tool usage, personnel arrangement and time consumption from the historical case with the highest degree of similarity to the current fault, generating a decision tree from fault to operation, and obtaining a preliminary generated fault handling solution;

[0263] With respect to the initially generated fault handling solution, combined with resource restrictions of fault repair, the initial fault handling solution is optimized with fault recovery time and resource consumption as optimization targets to obtain a fault handling solution.

[0264] In the real-time fault handling process, the newly discovered equipment association and event causal relationship knowledge in the processing process is recorded in real time and integrated into the regional ontology knowledge base and event semantic link network of the distribution network semantic model. The distribution network semantic model is optimized through incremental learning and active learning to obtain the distribution network semantic model, including:

[0265] Based on the on-site operation data collected during the execution of the fault handling work order, multimodal perception is used to extract the operation steps to form a structured record and obtain the fault handling text;

[0266] From the fault handling text, according to predefined entity and relationship templates, entity attributes and entity associations are extracted from subject-verb-object, attributive and adverbial structures, and a dictionary and relationship template library are constructed to obtain new fault handling data;

[0267] For the new fault processing data, an ontology fragment that can represent the new fault processing process is constructed, and the ontology fusion technology is used to construct the mapping relationship between the new and old ontologies. The newly discovered equipment association information and event causal relationship knowledge are compared and integrated with the original regional ontology knowledge base to form a new regional fault ontology knowledge representation, and a new regional ontology knowledge base is obtained;

[0268] For the event semantic link network, the newly discovered events, device nodes and their relationship edges are added to the original network topology, and a new event semantic link network is obtained by dynamically updating the event semantic link network;

[0269] By dynamically optimizing the new regional ontology knowledge base and the new event semantic link network, the distribution network semantic model is updated through incremental learning and active learning, and the distribution network semantic model after knowledge conversion and fusion is obtained.

[0270] The above only lists some preferred embodiments of the present invention, but the present invention is not limited thereto, and many improvements and changes can be made. As long as the improvements and changes are made on the basis of the basic principles of the present invention, they should be regarded as falling within the protection scope of the present invention.

Claims

1. A method for knowledge conversion and fusion processing of massive power grid operation data, characterized in that: include: Obtain the deduplicated independent alarm event set and structured operation event data; Mapping the operation event data to a unified data model, performing temporal association, spatial association, and causal association, and then vectorizing the association results to obtain an event semantic link network; According to the geographical topological structure of the distribution network, a regional ontology knowledge base is constructed, and combined with the event semantic link network, a distribution network semantic model is formed; For distribution network fault scenarios, semantic knowledge of faulty equipment is obtained based on the regional ontology knowledge base and event semantic link network in the distribution network semantic model, and a fault handling solution is generated; In the real-time fault handling process, the newly discovered equipment association and event causal relationship knowledge in the processing process is recorded in real time and integrated into the regional ontology knowledge base and event semantic link network of the distribution network semantic model. The distribution network semantic model is optimized through incremental learning and active learning to obtain the distribution network semantic model, which is used for knowledge conversion and fusion processing of massive power grid operation data.

2. The method for knowledge conversion and fusion processing of massive power grid operation data according to claim 1, characterized in that: Obtain the deduplicated independent alarm event set and structured operation event data, including: Collect alarm event data, clean and standardize the alarm event data, and then convert it into a unified event granularity. Match and merge the alarm event data of the unified event granularity according to the attributes of the alarm event data to obtain a deduplicated independent alarm event set. The attributes of the alarm event data include the device ID, alarm type, timestamp, and severity at the device granularity. Natural language processing is used to perform text mining on the unstructured operation events uploaded by the distribution automation terminal, and the unstructured operation events are converted into structured representations of event granularity to obtain structured operation event data.

3. The method for knowledge conversion and fusion processing of massive power grid operation data according to claim 2 is characterized by: Unstructured operation events include operation logs and voice recording data. The specific implementation of structuring voice recording data includes: Perform speech recognition processing on the collected speech recording data to obtain a text sequence corresponding to the speech; Perform natural language processing on the text sequence corresponding to the speech, identify the entity types of time, place, equipment and personnel in the text sequence according to the preset feature template and annotation rules, and obtain the text sequence of the speech based on the entity recognition model; According to the pre-built rule library and dictionary library, structured event attribute fields are extracted from the text sequence of speech based on the entity recognition model. The event attribute fields include time, place, equipment and personnel. The attribute values ​​of the extracted structured event attribute fields are standardized, and the standardized attribute values ​​are converted into structured voice recording data.

4. The method for knowledge conversion and fusion processing of massive power grid operation data according to claim 1 is characterized by: Map the operation event data to a unified data model, perform temporal association, spatial association, and causal association, and then vectorize the association results to obtain an event semantic link network, including: Perform data cleaning and data conversion on the deduplicated independent alarm event set and operation event data to obtain heterogeneous data, map the heterogeneous data to a unified data model and format, and obtain alarm event data and operation events in a unified data model and format; Construct the time series, spatial and causal associations of alarm event data and operation events in a unified data model and format; Alarm events and operation events with unified data models and formats are used as nodes of the graph, and the temporal associations, spatial associations, and causal associations between events are used as edges of the graph. The Cypher query language is used to map event data to the graph model, and a semantic link network across data sources is constructed. By vectorizing the event nodes in the semantic link network across data sources, an event semantic link network is obtained.

5. The method for knowledge conversion and fusion processing of massive power grid operation data according to claim 4 is characterized by: Construct the time series, spatial and causal associations of alarm event data and operation events in a unified data model and format, including: Set a time window of fixed length, sort the alarm events and operation events of the unified data model and format by timestamp, and for each alarm event of the unified data model and format, search for the operation event of the unified data model and format within its time window, and establish a temporal association between the two; Use device topology information to build a device connection diagram, find the adjacency relationship between devices through depth-first search, and spatially associate alarm events and operation events with unified data models and formats on the same device and adjacent devices; The event association sequence obtained by time-space alignment is matched with the rule base, the causal relationship between alarm events and operation events in a unified data model and format is identified, the event causal chain is constructed, and the causal association is obtained.

6. The method for knowledge conversion and fusion processing of massive power grid operation data according to claim 1, characterized in that: Build a regional ontology knowledge base, including: Obtain distribution network geographic topology data and extract equipment attribute information, including substation voltage level, feeder length, current capacity, switch model and status; Parse the collected distribution network geographic topology equipment attribute information data and extract the equipment coordinates, convert the equipment coordinates into geometric objects, extract the spatial geometric information of the equipment based on the geometric objects, find the adjacency relationship between the equipment, and build the spatial topology model of the distribution network equipment; Converting the spatial topological model of the distribution network equipment into an RDF graph, mapping the spatial topological relationship into an RDF triple, and obtaining the RDF triple; Convert the RDF triples into an OWL ontology to obtain an OWL ontology; OWL ontology is used to define the logical relationship between entity concepts, formally define the ontology relationship, obtain the OWL ontology description logic, and build a regional ontology knowledge base.

7. The method for knowledge conversion and fusion processing of massive power grid operation data according to claim 1 is characterized by: For distribution network fault scenarios, the semantic knowledge of the faulty equipment is obtained based on the regional ontology knowledge base and event semantic link network in the distribution network semantic model, the fault cause and impact range are determined, and a fault diagnosis report is generated, including: Receive a fault alarm message carrying a unique identification ID of the fault event, the fault alarm message is sent by a real-time data interface of the distribution automation system, and obtain the fault alarm message with the unique identification ID of the fault event; According to the fault alarm message with the unique identification ID of the fault event, the semantic knowledge associated with the fault alarm message with the unique identification ID of the fault event is searched in the event semantic link network to obtain the semantic knowledge associated with the event; According to the semantic knowledge associated with the event, the spatial position coordinates of the faulty device are obtained, and the power distribution equipment within the set range is searched to obtain the power distribution equipment within the set range; Determine whether current and voltage are abnormal, and if so, determine the equipment within the set range affected by the fault to obtain the impact range of the fault; Determine the impact scope of the outage of the faulty equipment on downstream users based on the impact scope of the fault, where the impact scope on downstream users includes the power outage scope and important user types, and obtain the impact scope of the outage of the faulty equipment on downstream users; According to the impact scope of the outage of the faulty equipment on downstream users, historical alarm events directly related to the current faulty equipment are extracted from the event semantic link network to obtain the diagnosis results and processing solutions of the historical alarm events; According to the diagnosis results and processing schemes of the historical alarm events, the fault diagnosis results are obtained by integrating the alarm information of the faulty equipment, the real-time operation status, the analysis results of historical similar events and the multi-source heterogeneous knowledge of the status of the equipment within the set range; A fault diagnosis report is generated according to the result of the fault diagnosis.

8. The method for knowledge conversion and fusion processing of massive power grid operation data according to claim 7 is characterized by: According to the fault diagnosis result of generating the fault diagnosis report, the fault handling history is extracted from the operation event knowledge base, and combined with the current fault, a fault handling solution is generated, including: Extract information from the fault diagnosis report, including the name of the faulty device, fault type, and cause analysis. Based on the extracted information, semantic modeling is performed in the operation event knowledge base using an ontology to define the concepts and attributes of fault type, device type, and processing effect, forming a domain ontology. Instantiate and match the statement template with the domain ontology to obtain historical operation events related to the current fault; For the historical operation events related to the current fault, by converting the fault information and the text description of the operation event into a vector representation, calculating the similarity with the current fault, and obtaining the historical case with the highest similarity with the current fault; Extracting elements including operation method, tool usage, personnel arrangement and time consumption from the historical case with the highest degree of similarity to the current fault, generating a decision tree from fault to operation, and obtaining a preliminary generated fault handling solution; With respect to the initially generated fault handling solution, combined with resource restrictions of fault repair, the initial fault handling solution is optimized with fault recovery time and resource consumption as optimization targets to obtain a fault handling solution.

9. The method for knowledge conversion and fusion processing of massive power grid operation data according to claim 1, characterized in that: In the real-time fault handling process, the newly discovered equipment association and event causal relationship knowledge in the processing process is recorded in real time and integrated into the regional ontology knowledge base and event semantic link network of the distribution network semantic model. The distribution network semantic model is optimized through incremental learning and active learning to obtain the distribution network semantic model, including: Based on the on-site operation data collected during the execution of the fault handling work order, multimodal perception is used to extract the operation steps to form a structured record and obtain the fault handling text; From the fault handling text, according to predefined entity and relationship templates, entity attributes and entity associations are extracted from subject-verb-object, attributive and adverbial structures, and a dictionary and relationship template library are constructed to obtain new fault handling data; For the new fault processing data, an ontology fragment that can represent the new fault processing process is constructed, and the ontology fusion technology is used to construct the mapping relationship between the new and old ontologies. The newly discovered equipment association information and event causal relationship knowledge are compared and integrated with the original regional ontology knowledge base to form a new regional fault ontology knowledge representation, and a new regional ontology knowledge base is obtained; For the event semantic link network, the newly discovered events, device nodes and their relationship edges are added to the original network topology, and a new event semantic link network is obtained by dynamically updating the event semantic link network; By dynamically optimizing the new regional ontology knowledge base and the new event semantic link network, the distribution network semantic model is updated through incremental learning and active learning, and the distribution network semantic model after knowledge conversion and fusion is obtained.

10. A knowledge conversion and fusion processing system for massive power grid operation data, characterized in that: include: Data acquisition module: used to obtain the deduplicated independent alarm event set and structured operation event data; Event semantic link network construction module: used to map the deduplicated independent alarm event set and structured operation event data to a unified data model, and then perform temporal association, spatial association, and causal association, and then vectorize the association results to obtain the event semantic link network; Distribution network semantic model construction module: used to build a regional ontology knowledge base based on the geographical topology of the distribution network, and combine the event semantic link network to form a distribution network semantic model; Fault handling solution generation module: for distribution network fault scenarios, the module obtains semantic knowledge of faulty equipment based on the regional ontology knowledge base and event semantic link network in the distribution network semantic model, and generates a fault handling solution; Distribution network semantic model optimization module: It is used to record the newly discovered equipment association and event causal relationship knowledge in real time during the implementation of the fault handling plan and integrate it into the regional ontology knowledge base and event semantic link network of the distribution network semantic model. It also optimizes the distribution network semantic model through incremental learning and active learning to obtain the distribution network semantic model, which is used for knowledge conversion and fusion processing of massive power grid operation data.

11. A knowledge conversion and fusion processing system for massive power grid operation data according to claim 10, characterized in that: Obtain the deduplicated independent alarm event set and structured operation event data, including: Collect alarm event data, clean and standardize the alarm event data, and then convert it into a unified event granularity. Match and merge the alarm event data of the unified event granularity according to the attributes of the alarm event data to obtain a deduplicated independent alarm event set. The attributes of the alarm event data include the device ID, alarm type, timestamp, and severity at the device granularity. Natural language processing is used to perform text mining on the unstructured operation events uploaded by the distribution automation terminal, and the unstructured operation events are converted into structured representations of event granularity to obtain structured operation event data.

12. The knowledge conversion and fusion processing system for massive power grid operation data according to claim 10, characterized in that: Unstructured operation events include operation logs and voice recording data. The specific implementation of structuring voice recording data includes: Perform speech recognition processing on the collected speech recording data to obtain a text sequence corresponding to the speech; Perform natural language processing on the text sequence corresponding to the speech, identify the entity types of time, place, equipment and personnel in the text sequence according to the preset feature template and annotation rules, and obtain the text sequence of the speech based on the entity recognition model; According to the pre-built rule library and dictionary library, structured event attribute fields are extracted from the text sequence of speech based on the entity recognition model. The event attribute fields include time, place, equipment and personnel. The attribute values ​​of the extracted structured event attribute fields are standardized, and the standardized attribute values ​​are converted into structured voice recording data.

13. The knowledge conversion and fusion processing system for massive power grid operation data according to claim 10, characterized in that: Map the operation event data to a unified data model, perform temporal association, spatial association, and causal association, and then vectorize the association results to obtain an event semantic link network, including: Perform data cleaning and data conversion on the deduplicated independent alarm event set and operation event data to obtain heterogeneous data, map the heterogeneous data to a unified data model and format, and obtain alarm event data and operation events in a unified data model and format; Construct the time series, spatial and causal associations of alarm event data and operation events in a unified data model and format; Alarm events and operation events with unified data models and formats are used as nodes of the graph, and the temporal associations, spatial associations, and causal associations between events are used as edges of the graph. The Cypher query language is used to map event data to the graph model, and a semantic link network across data sources is constructed. By vectorizing the event nodes in the semantic link network across data sources, an event semantic link network is obtained.

14. A knowledge conversion and fusion processing system for massive power grid operation data according to claim 13, characterized in that: Construct the time series, spatial and causal associations of alarm event data and operation events in a unified data model and format, including: Set a time window of fixed length, sort the alarm events and operation events of the unified data model and format by timestamp, and for each alarm event of the unified data model and format, search for the operation event of the unified data model and format within its time window, and establish a temporal association between the two; Use device topology information to build a device connection diagram, find the adjacency relationship between devices through depth-first search, and spatially associate alarm events and operation events with unified data models and formats on the same device and adjacent devices; The event association sequence obtained by time-space alignment is matched with the rule base, the causal relationship between alarm events and operation events in a unified data model and format is identified, the event causal chain is constructed, and the causal association is obtained.

15. The knowledge conversion and fusion processing system for massive power grid operation data according to claim 10, characterized in that: Build a regional ontology knowledge base, including: Obtain distribution network geographic topology data and extract equipment attribute information, including substation voltage level, feeder length, current capacity, switch model and status; Parse the collected distribution network geographic topology equipment attribute information data and extract the equipment coordinates, convert the equipment coordinates into geometric objects, extract the spatial geometric information of the equipment based on the geometric objects, find the adjacency relationship between the equipment, and build the spatial topology model of the distribution network equipment; Converting the spatial topological model of the distribution network equipment into an RDF graph, mapping the spatial topological relationship into an RDF triple, and obtaining the RDF triple; Convert the RDF triples into an OWL ontology to obtain an OWL ontology; OWL ontology is used to define the logical relationship between entity concepts, formally define the ontology relationship, obtain the OWL ontology description logic, and build a regional ontology knowledge base.

16. The knowledge conversion and fusion processing system for massive power grid operation data according to claim 10, characterized in that: For distribution network fault scenarios, the semantic knowledge of the faulty equipment is obtained based on the regional ontology knowledge base and event semantic link network in the distribution network semantic model, the fault cause and impact range are determined, and a fault diagnosis report is generated, including: Receive a fault alarm message carrying a unique identification ID of the fault event, the fault alarm message is sent by a real-time data interface of the distribution automation system, and obtain the fault alarm message with the unique identification ID of the fault event; According to the fault alarm message with the unique identification ID of the fault event, the semantic knowledge associated with the fault alarm message with the unique identification ID of the fault event is searched in the event semantic link network to obtain the semantic knowledge associated with the event; According to the semantic knowledge associated with the event, the spatial position coordinates of the faulty device are obtained, and the power distribution equipment within the set range is searched to obtain the power distribution equipment within the set range; Determine whether current and voltage are abnormal, and if so, determine the equipment within the set range affected by the fault to obtain the impact range of the fault; Determine the impact scope of the outage of the faulty equipment on downstream users based on the impact scope of the fault, where the impact scope on downstream users includes the power outage scope and important user types, and obtain the impact scope of the outage of the faulty equipment on downstream users; According to the impact scope of the outage of the faulty equipment on downstream users, historical alarm events directly related to the current faulty equipment are extracted from the event semantic link network to obtain the diagnosis results and processing solutions of the historical alarm events; According to the diagnosis results and processing schemes of the historical alarm events, the fault diagnosis results are obtained by integrating the alarm information of the faulty equipment, the real-time operation status, the analysis results of historical similar events and the multi-source heterogeneous knowledge of the status of the equipment within the set range; A fault diagnosis report is generated according to the result of the fault diagnosis.

17. A knowledge conversion and fusion processing system for massive power grid operation data according to claim 16, characterized in that: According to the fault diagnosis result of generating the fault diagnosis report, the fault handling history is extracted from the operation event knowledge base, and combined with the current fault, a fault handling solution is generated, including: Extract information from the fault diagnosis report, including the name of the faulty device, fault type, and cause analysis. Based on the extracted information, semantic modeling is performed in the operation event knowledge base using an ontology to define the concepts and attributes of fault type, device type, and processing effect, forming a domain ontology. Instantiate and match the statement template with the domain ontology to obtain historical operation events related to the current fault; For the historical operation events related to the current fault, by converting the fault information and the text description of the operation event into a vector representation, calculating the similarity with the current fault, and obtaining the historical case with the highest similarity with the current fault; Extracting elements including operation method, tool usage, personnel arrangement and time consumption from the historical case with the highest degree of similarity to the current fault, generating a decision tree from fault to operation, and obtaining a preliminary generated fault handling solution; With respect to the initially generated fault handling solution, combined with resource restrictions of fault repair, the initial fault handling solution is optimized with fault recovery time and resource consumption as optimization targets to obtain a fault handling solution.

18. The knowledge conversion and fusion processing system for massive power grid operation data according to claim 10, characterized in that: In the real-time fault handling process, the newly discovered equipment association and event causal relationship knowledge in the processing process is recorded in real time and integrated into the regional ontology knowledge base and event semantic link network of the distribution network semantic model. The distribution network semantic model is optimized through incremental learning and active learning to obtain the distribution network semantic model, including: Based on the on-site operation data collected during the execution of the fault handling work order, multimodal perception is used to extract the operation steps to form a structured record and obtain the fault handling text; From the fault handling text, according to predefined entity and relationship templates, entity attributes and entity associations are extracted from subject-verb-object, attributive and adverbial structures, and a dictionary and relationship template library are constructed to obtain new fault handling data; For the new fault processing data, an ontology fragment that can represent the new fault processing process is constructed, and the ontology fusion technology is used to construct the mapping relationship between the new and old ontologies. The newly discovered equipment association information and event causal relationship knowledge are compared and integrated with the original regional ontology knowledge base to form a new regional fault ontology knowledge representation, and a new regional ontology knowledge base is obtained; For the event semantic link network, the newly discovered events, device nodes and their relationship edges are added to the original network topology, and a new event semantic link network is obtained by dynamically updating the event semantic link network; By dynamically optimizing the new regional ontology knowledge base and the new event semantic link network, the distribution network semantic model is updated through incremental learning and active learning, and the distribution network semantic model after knowledge conversion and fusion is obtained.

Citation Information

Patent Citations

  • Unified Standardized Processing System and Method for Multi-Source Heterogeneous Data of Transmission Lines

    CN112507035B

Cited By

  • Power distribution network system fault monitoring method and system based on multi-source information

    CN120405328A

  • Method for associating operation and maintenance alarm with knowledge base based on large model

    CN120596679A

  • Power plant equipment fault intelligent diagnosis method

    CN120632175A

  • Converter transformer fault case structured processing method, system, device and medium

    CN120687730A

  • Knowledge graph generation method and system applied to safety monitoring of thermal power plant

    CN120822589A