Power grid fault event fusion method and system and computer storage medium
Through the structured semantic analysis and hierarchical clustering mechanism driven by a large language model, the subject-predicate logical relationship of grid fault events is decoupled, and the efficient fusion of multi-source heterogeneous data is achieved, which solves the problems of semantic information mining and cross-system fusion in grid fault event processing, and improves the accuracy of fault root cause positioning and the efficiency of event correlation analysis.
Patent Information
- Application Number
- CN202510779331.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The existing power grid fault event processing technology is difficult to effectively mine deep semantic information of multi-source heterogeneous power failure data, resulting in low location efficiency of fault root cause, poor cross-system event fusion capability, and insufficient professional knowledge adaptability and insufficient semantic representation capability for the application of large language models in the field of power failure.
The structured semantic analysis and hierarchical clustering mechanism driven by a large language model are adopted to decouple the subject-predicate logical relationship of fault events through dual-channel coding technology, build a composite similarity function, merge similar event clusters from bottom to top, generate multi-grained clustering results, and realize the precise fusion of power fault events.
It significantly improves the semantic correlation and domain adaptability of power grid fault events, improves the accuracy of fault root cause traceability and the efficiency of constructing the matter map, and provides a highly explanatory decision-making basis.
Smart Images

Figure CN120296682A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power grid fault data processing, and in particular to a power grid fault event fusion method, system and computer storage medium. Background Art
[0002] With the in-depth development of the intelligent power system, the fault event data generated during the operation of the power grid is growing exponentially. These data include equipment alarm logs, sensor monitoring data, manual maintenance reports, etc., with the characteristics of multi-source heterogeneity, complex semantics, and strong spatial and temporal correlation. Traditional fault event analysis methods rely on manual experience or rule-based systems to classify and associate events, which makes it difficult to effectively mine the deep semantic information of massive unstructured text data, resulting in low efficiency in fault root cause location and poor cross-system event fusion capabilities. In terms of event graph construction, existing technologies mainly focus on static entity relationship modeling, and lack causal logic analysis of dynamic event links. As an advanced form of event graph, the event graph requires accurate description of the evolution of fault events and the law of causal transmission. However, traditional methods have many problems: event chain extraction based on rule templates is difficult to cover the complex and changeable fault scenarios in the power field; the event fusion process does not consider the semantic hierarchical characteristics of the subject-predicate structure, resulting in high redundancy of event nodes in the graph, which seriously affects the accuracy of fault deduction analysis. At present, power grid fault event processing technology faces the following bottlenecks: First, existing semantic clustering methods are mostly based on shallow text features, such as keyword matching, TF-IDF weight calculation, etc., which cannot accurately grasp the contextual semantics of power professional terms and the logical connection between events. The difference in description of the same equipment in different fault scenarios can easily lead to clustering deviations. Second, the heterogeneity of multi-source event data makes it difficult for traditional methods to achieve efficient fusion of cross-modal information, resulting in information fragmentation problems in event correlation analysis. In addition, there is a lack of effective modeling methods for the dynamic association between fault events and power grid topology and equipment knowledge base, which limits the real-time assessment of the impact range of events and collaborative disposal decisions. In recent years, large language models have shown obvious advantages in text semantic understanding, but their application in the field of power failure is still limited: on the one hand, the general domain pre-training model is not adaptable enough to the professional knowledge of the power grid, resulting in low accuracy of event entity recognition and relationship extraction; on the other hand, existing methods fail to fully explore the semantic space characteristics of large language models and fail to deeply cooperate with clustering algorithms, which restricts the ability to discover complex event patterns. Therefore, how to combine the semantic representation ability of large language models with the knowledge of power grid faults and build an efficient event fusion framework has become a key problem in improving the level of intelligent diagnosis of power system faults. Summary of the invention
[0003] To solve the problems of the accuracy and logical consistency of event fusion in the construction of power grid fault event graphs, as well as the misjudgment of semantically equivalent events caused by differences in subject-predicate structures in power equipment fault texts, the present invention provides a power grid fault event fusion method, system, and computer storage medium. Through a structured semantic parsing and hierarchical clustering mechanism driven by a large language model, it realizes the decoupling and precise fusion of the subject and predicate semantics of power fault events, improves the event relevance and domain adaptability of the event graph, and provides technical support for power grid fault root cause tracing. Specifically, the method includes the following steps: S1: Obtain multi-source heterogeneous data of the power system, standardize the definition of power fault events, construct a standardized event text structure, extract original events from the multi-source heterogeneous data, and integrate them according to the standardized event text structure to form an original event text set; S2: Based on the original event set, perform semantic parsing on each original event text, extract the subject and predicate of the event, where the subject includes the equipment, components, or state variables involved in the fault event, and the predicate includes abnormal actions or state changes of the fault event; S3: Respectively perform vectorized representation on the subject and predicate of the event to obtain dual-channel encoding vectors of the subject and predicate; S4: Calculate the similarity of the dual-channel encoding vectors of the subject and predicate, construct a composite similarity function based on the similarity, and calculate the semantic similarity matrix between events based on the composite similarity function; S5: Based on the semantic similarity matrix, merge similar event clusters from bottom to top, and generate multi-granularity clustering results according to the preset clustering granularity level; S6: Based on the multi-granularity clustering results, generate multiple candidate standardized events for each clustering cluster, and assign a unique standardized event to each original event in the clustering cluster to complete the fusion of power grid fault events.
[0004] In an embodiment of the present invention, in S5, the method for generating multi-granularity clustering results according to the preset clustering granularity level is as follows: S51: Use the dual-channel encoding vectors of the subject and predicate extracted in S3 to update the original event text set to obtain an updated fault event set , where each event is a vector representation of the subject and predicate after dual-channel semantic encoding; and, calculate the semantic similarity matrix between events in the updated fault event set according to the composite similarity function , where the matrix element represents the semantic similarity between event and event ; S52: Cluster each of the updated fault events separately to construct an initial cluster set , where the number of clusters is equal to the number of fault events; S53: Start iterative calculation from the initial cluster set to calculate the inter-cluster similarity matrix at the current time t , for cluster and cluster , calculate their similarity using the following formula : , where and represent the number of events in cluster and cluster respectively, represents the similarity between event and event calculated according to the composite similarity function; S54: Based on the inter-cluster similarity matrix , obtain the cluster pair with the maximum current inter-cluster similarity , i.e.: ; S55: Determine whether the condition that the number of clusters after the current iterative update reaches the target value K or the maximum inter-cluster similarity is less than or equal to the similarity threshold is satisfied: If not, return to execute step S53; If so, compare the inter-cluster similarity with the preset similarity threshold . If the inter-cluster similarity is greater than the similarity threshold , then merge cluster and cluster into a new cluster , and update the cluster set , represents the removal operation; S56: According to the preset clustering granularity level, gradually construct a dendrogram for each cluster merged in each iteration, and cut the dendrogram to obtain a multi-granularity clustering result .
[0005] In an embodiment of the present invention, in S56, according to the preset clustering granularity level, gradually construct a dendrogram for each cluster merged in each iteration, and cut the dendrogram to obtain a multi-granularity clustering result The method is as follows: Taking the initial single event cluster as the root node, each newly generated cluster during merging as a child node, and the similarity during merging as the node height, a dendrogram is gradually constructed. Each node of the dendrogram represents a cluster, and the height of the node represents the similarity during merging; According to the preset clustering granularity level, determine the corresponding cutting height in the dendrogram; perform a horizontal cut at the specified height in the dendrogram. According to the intersection points of the cutting line and the branches of the dendrogram, divide the dendrogram into multiple subtrees. Each subtree corresponds to a clustering result. By performing cuts at different heights, multi-granularity clustering results are obtained 。
[0006] In an embodiment of the present invention, in S1, form the original event text set The method is as follows: Define the equipment, components, or state variables involved in the fault event as the subject, and the abnormal actions or state changes of the fault event as the predicate. Construct a binary array according to the subject and the predicate to represent the fault event text structure, as follows: , Extract the original events from the multi-source heterogeneous data and integrate them according to the standardized event text structure to form the original event text set : ,wherein, represents the subject array, represents the predicate array, represents the set of original events of the k-th data source, represents the k-th data source, k = 1,..., K, and K represents the total number of the multi-source heterogeneous data.
[0007] In an embodiment of the present invention, in S2, the method for extracting the subject and predicate of an event is as follows: Construct an event parsing function: ,where e represents the fault event, s is the subject, and p is the predicate; Guide the event parsing function through text instructions to perform semantic parsing on each original event text, output the fault event text in JSON format, and extract the subject and predicate of the event from the JSON format fault event text.
[0008] In an embodiment of the present invention, in S2, the method for extracting the subject and predicate of an event further includes: making the parsing result of the event parsing function fully conform to the JSON format through regular expressions: , wherein, represents the output data, A regular expression is represented, and JSON represents data consisting of property-value pairs. It represents a constraint item.
[0009] In one embodiment of the present invention, in S3, the method for obtaining the dual-channel encoding vectors of the subject and the predicate is as follows: Perform structured event dual-channel encoding, perform semantic embedding on the parsed subject s and predicate p respectively, use the pre-trained SBERT model to encode the structured input, and the subject s and the predicate p share the same encoder to ensure vector space alignment, and generate the dual-channel encoding vectors of the subject and the predicate: , , Among them, represents the subject channel encoding vector, represents the predicate channel encoding vector, represents the dimension, and d is the output dimension size of the SBERT model.
[0010] In one embodiment of the present invention, in S4, the method for constructing a composite similarity function according to the similarity is as follows: Calculate the weighted cosine similarity of the dual-channel encoding vectors of the subject and the predicate of the fault event and the fault event , and construct a composite similarity function according to the weighted cosine similarity : , Among them, α∈[0,1] is an adjustable weight coefficient, represents the subject channel encoding vector of the fault event , represents the subject channel encoding vector of the fault event , represents the predicate channel encoding vector of the fault event , represents the predicate channel encoding vector of the fault event .
[0011] Based on the same inventive concept, the present invention also provides a power grid fault event fusion system for implementing the steps of the power grid fault event fusion method described above. The power grid fault event fusion system includes the following modules: A data standardization module, which is used to obtain multi-source heterogeneous data of the power system, define the power grid fault event in a standardized manner, construct a standardized event text structure, extract the original event from the multi-source heterogeneous data, and integrate it according to the standardized event text structure to form a set of original event texts; An event semantic analysis module, for performing semantic analysis on each original event text based on the original event set, and extracting the subject and predicate of the event, wherein the subject includes the equipment, components or state quantity involved in the fault event, and the predicate includes the abnormal action or state change of the fault event; An event vectorization representation module is used to vectorize the subject and predicate of an event respectively, and obtain a dual-channel encoding vector of the subject and predicate; A semantic similarity calculation module, used to calculate the similarity of the dual-channel encoding vectors of the subject and the predicate, construct a composite similarity function according to the similarity, and calculate the semantic similarity matrix between events based on the composite similarity function; A multi-granularity clustering result acquisition module, used to merge similar event clusters from bottom to top based on the semantic similarity matrix, and generate multi-granularity clustering results according to a preset clustering granularity level; The event fusion module is used to generate multiple candidate standardized events for each cluster based on the multi-granularity clustering results, and assign a unique standardized event to each original event in the cluster from the multiple candidate standardized events to complete the fusion of power grid fault events.
[0012] The present invention also provides a computer storage medium, wherein the computer storage medium stores a computer software product, wherein the computer software product includes several instructions for enabling a computer device to execute the power grid fault event fusion method.
[0013] The above technical solution of the present invention has the following advantages compared with the prior art: 1. Accurate semantic parsing capability: Based on the structured semantic parsing technology of the large language model, combined with the characteristics of domain data, it effectively decouples the subject-predicate logical relationship of the fault event, significantly improves the semantic recognition accuracy of the equipment subject and state action, and overcomes the limitation of traditional methods in capturing the contextual association of professional terms.
[0014] 2. Multi-dimensional semantic representation optimization: Dual-channel semantic encoding technology is used to construct independent vector spaces for device entities and state actions respectively. Through the fusion of multi-dimensional semantic features, the differentiated representation capability of "synonymous and heterogeneous" texts is enhanced, providing more accurate semantic support for event association analysis.
[0015] 3. Intelligent event fusion architecture: It integrates hierarchical clustering and generative language model technology, automatically divides event clusters through multi-granularity semantic analysis, and generates standardized event descriptions that comply with industry norms, significantly improving the efficiency of event graph construction and the systematic nature of knowledge accumulation, providing highly interpretable decision-making basis for fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] To make the content of the present invention easier to understand clearly, the following further elaborates on the present invention according to specific embodiments of the present invention in conjunction with the accompanying drawings. Among them, Figure 1 is a schematic flowchart of a power grid fault event fusion method provided in an embodiment of the present invention; Figure 2 is a schematic flowchart of a process for generating multi-granularity clustering results provided in an embodiment of the present invention; Figure 3 is a schematic structural diagram of a power grid fault event fusion system provided in an embodiment of the present invention; Explanation of reference numerals in the accompanying drawings: 100, data standardization module; 200, event semantic parsing module; 300, event vectorized representation module; 400, semantic similarity calculation module; 500, multi-granularity clustering result acquisition module; 600, event fusion module. Specific embodiments
[0017] The following further elaborates on the present invention in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited do not limit the present invention.
[0018] Embodiment 1:
[0019] Referring to Figure 1 as shown, the present invention provides a power grid fault event fusion method, and this method includes the following steps: S1: Standardize the definition of power failure events and construct a standardized event text structure. Collect original events from various data sources of the power system (such as equipment alarm logs, sensor monitoring records, manual maintenance reports, etc.) and integrate them to form an original event text set, providing a data basis with a unified format for subsequent processing; S2: Based on the original event text set, use a large language model to perform semantic parsing on each original event text, accurately extract the subject (i.e., the equipment, component, or state quantity involved in the fault, etc.) and predicate (such as state change, abnormal action, etc.) in the event, and construct standardized semantic features to prepare for subsequent accurate clustering analysis; S3: Apply the dual-channel semantic encoding technology to vectorize and represent the subject and predicate of the event respectively, obtain the dual-channel encoding vectors of the subject and predicate, and then construct a multi-dimensional semantic feature space of the event to more comprehensively and accurately reflect the event semantic information; S4: Calculate the similarity between the dual-channel encoding vectors of the subject and predicate, construct a composite similarity function in combination with the dynamic weight allocation mechanism, and calculate the semantic similarity between each event according to this composite similarity function to generate a semantic similarity matrix, quantifying the semantic association degree between events; S5: Use a hierarchical clustering algorithm. Based on the semantic similarity matrix, iteratively merge similar event clusters in a bottom-up manner, and control and adjust according to the preset clustering granularity levels to generate a clustering result with multi-granularity characteristics, providing a reasonable classification basis for event fusion; S6: According to the multi-granularity clustering result, use a large language model to generate multiple candidate standardized events for each clustering cluster. Through a semantic matching algorithm, accurately assign a unique standardized event to each original event within the clustering cluster to complete the fusion of power grid fault events and meet the actual application requirements.
[0020] Further, in S1, the method of forming the original event text set is as follows: Define the equipment, components, or state variables involved in the fault event as the subject, and define the abnormal actions or state changes of the fault event as the predicate. Construct a binary array according to the subject and the predicate to represent the fault event text structure, as follows: , Extract the original events from the multi-source heterogeneous data and integrate them according to the standardized event text structure to form the original event text set : , where represents the subject array, represents the predicate array, represents the set of original events of the k-th data source, represents the k-th data source, k = 1,..., K, and K represents the total number of the multi-source heterogeneous data.
[0021] Further, in S2, the method of extracting the subject and predicate of the event is as follows: Construct an event parsing function: , where e represents the fault event, s is the subject, and p is the predicate; Through the text instruction Prompt, guide the event parsing function to perform semantic parsing on each original event text, output the fault event text in JSON format, and extract the subject and predicate of the event from the JSON format fault event text.
[0022] Assume the dedicated Prompt is designed as: "You are an expert in extracting the subject and predicate of power grid fault events. Given an event, identify the subject and predicate in it and return the set of these elements. The subject of the event is the equipment main body or state variable representing the event, and the predicate is the abnormal action or state change representing the event. The following are some examples: Input: The guide vane is aging Output: {"Event": "Deflector aging", "Subject": "Deflector", "Predicate": "aging"}
[0023] Specifically, in S2, based on the original event set, the method for semantic parsing of each original event text to extract the subject and predicate of the event further includes: making the parsing result of the event parsing function fully conform to the JSON format through regular expressions to ensure robustness: , where, represents the output data, represents the regular expression, and JSON represents the data composed of property-value pairs. represents the constraint term. The above formula means that the input JSON data (e) is verified or processed according to the regular expression rules defined in the constraint term and the result is returned.
[0024] Furthermore, in S3 of this embodiment, the method for vectorizing and characterizing the subject and predicate of the event respectively to obtain the dual-channel encoding vectors of the subject and predicate is as follows: Perform structured event dual-channel encoding, perform semantic embedding on the parsed subject s and predicate p respectively, use the pre-trained SBERT model to encode the structured input, and the subject s and the predicate p share the same encoder to ensure the alignment of the vector space and generate the dual-channel encoding vectors of the subject and predicate: , , where, represents the subject channel encoding vector, represents the predicate channel encoding vector, represents the dimension, and d is the output dimension size of the SBERT model, usually set to 768.
[0025] Furthermore, in S4, calculate the similarity of the dual-channel encoding vectors and of the subject and predicate, and the method for constructing the composite similarity function according to the similarity is as follows: Calculate the weighted cosine similarity of the subject and predicate dual-channel encoding vectors of the fault event and the fault event , and construct the composite similarity function : , where, α∈[0,1] is an adjustable weight coefficient, and the contribution ratio of the subject and predicate is dynamically adjusted through domain knowledge (such as device criticality). Indicates a fault event The subject channel encoding vector of Indicates a fault event The subject channel encoding vector of Indicates a fault event The predicate channel encoding vector of Indicates a fault event The predicate channel encoding vector of the fault event
[0026] Furthermore, as shown in Figure 2 In S5, the method for generating a multi-granularity clustering result according to the preset clustering granularity level is as follows:
[0027] S51: Use the dual-channel encoding vectors of the subject and predicate extracted in S3 to update the original event text set to obtain an updated fault event set , where each event Is the vector representation of the subject and predicate after dual-channel semantic encoding; and, calculate the semantic similarity matrix between events in the updated fault event set according to the composite similarity function , where the matrix element Indicates event And event The semantic similarity between; S52: Cluster each fault event in the updated fault events separately to construct an initial cluster set , the number of clusters is equal to the number of fault events; S53: Start iterative calculation from the initial cluster set, and calculate the inter-cluster similarity matrix at the current time t , for cluster And cluster , use the following formula to calculate their similarity : , where, And Respectively represent cluster And cluster The number of events in, Indicates the event calculated according to the composite similarity function And event The similarity between; S54: Based on the inter-cluster similarity matrix , obtain the cluster pair with the maximum current inter-cluster similarity , that is: ; ; S55: Determine whether the number of clusters after the current iterative update reaches the target value K or the maximum inter-cluster similarity is less than or equal to the similarity threshold Condition: If not, return to execute step S53; If so, compare the inter-cluster similarity with a preset similarity threshold in terms of size. If the inter-cluster similarity is greater than the similarity threshold , then combine cluster and cluster into a new cluster , and update the cluster set . indicates a removal operation; S56: According to the preset clustering granularity level, gradually construct a dendrogram for each iteratively merged cluster, and cut the dendrogram to obtain multi-granularity clustering results .
[0028] Furthermore, in S56, according to the preset clustering granularity level, gradually construct a dendrogram for each iteratively merged cluster, and cut the dendrogram to obtain multi-granularity clustering results The method is as follows: Take the initial single-event cluster as the root node, the newly generated cluster in each merge as the child node, and the similarity at the time of merge as the node height, and gradually construct a dendrogram. Each node of the dendrogram represents a cluster, and the height of the node represents the similarity at the time of merge; According to the preset clustering granularity level, determine the corresponding cutting height in the dendrogram. The higher the cutting height, the fewer the number of generated clusters and the coarser the granularity; the lower the cutting height, the more the number of generated clusters and the finer the granularity.
[0029] Perform a horizontal cut at the specified height in the dendrogram. According to the intersection points of the cutting line and the branches of the dendrogram, divide the dendrogram into multiple subtrees. Each subtree corresponds to a clustering result. By cutting at different heights, multi-granularity clustering results are obtained .
[0030] Taking the event set as an example, through the iterative merging process, merge and to form cluster , merge and to form cluster , merge and to form the final cluster .
[0031] If cutting at the position with a height of 2, the obtained cluster division result is: ; if cutting at the position with a height of 1, the obtained cluster division result is: ; If cutting at the position with a height of 3, the obtained cluster division result is: .
[0032] Furthermore, in S6, based on the multi-granularity clustering result, multiple candidate standardized events are generated for each clustering cluster, and a unique standardized event is assigned to each original event in the clustering cluster to complete the fusion of power grid fault events. The specific method is as follows: The final cluster division result obtained through the hierarchical clustering in S5 , each clustering cluster contains events with similar semantics; For each clustering cluster , a dedicated text instruction Prompt is designed to guide the large language model to generate a standardized event description. An example of Prompt is as follows: "You are an expert in fusing power equipment fault events. I am building a fault event logic graph of power equipment, and you will help me summarize multiple events in an event cluster into a standardized event. These events may be similar, or they may share certain features and semantics. Your task is to summarize them into one or more standardized events. Each standardized event description should contain as much key information as possible to ensure clear meaning. The following is the list of events in the event cluster: Event 1: [Specific description of Event 1] Event 2: [Specific description of Event 2] ... Please generate n candidate standardized event descriptions (n≥1), and each description should be concise and cover key information." Input the above Prompt into the large language model to generate n candidate standardized events (n≥1): .
[0033] For each original event and each candidate standardized event , m = 1,..., n, use the pre-trained SBERT model for semantic embedding to obtain semantic vectors and , calculate the semantic similarity between each original event e and each candidate standardized event , and adopt the cosine similarity formula: ; For each original event , select the candidate standardized event with the largest semantic similarity to it as the unique standardized event for this event, that is: ; Assign its corresponding unique standardized event , forming a semantic mapping relationship M: M→ .
[0034] For each cluster , output a standardized event set after semantic mapping The standardized event descriptions in each cluster are consistent and concise, which is convenient for subsequent fault diagnosis and event graph construction. At this point, the fusion of power grid fault events is completed.
[0035] Embodiment 2: Based on the same inventive concept as that of the first embodiment, the present invention further provides a power grid fault event fusion system, which is used to implement the steps of the power grid fault event fusion method described in the first embodiment. Figure 3 As shown, the power grid fault event fusion system includes the following modules: The data standardization module 100 is used to obtain multi-source heterogeneous data of the power system, standardize the definition of power fault events, construct a standardized event text structure, extract original events from the multi-source heterogeneous data, and integrate them according to the standardized event text structure to form an original event text set; An event semantic analysis module 200 is used to perform semantic analysis on each original event text based on the original event set, and extract the subject and predicate of the event, wherein the subject includes the equipment, components or state quantity involved in the fault event, and the predicate includes the abnormal action or state change of the fault event; An event vectorization representation module 300 is used to vectorize the subject and predicate of the event respectively to obtain dual-channel encoding vectors of the subject and predicate; A semantic similarity calculation module 400 is used to calculate the similarity of the dual-channel encoding vectors of the subject and the predicate, construct a composite similarity function according to the similarity, and calculate the semantic similarity matrix between events based on the composite similarity function; A multi-granularity clustering result acquisition module 500 is used to merge similar event clusters from bottom to top based on the semantic similarity matrix and generate a multi-granularity clustering result according to a preset clustering granularity level; The event fusion module 600 is used to generate multiple candidate standardized events for each cluster based on the multi-granularity clustering results, and assign a unique standardized event to each original event in the cluster from the multiple candidate standardized events to complete power grid fault event fusion.
[0036] A power grid fault event fusion system proposed in this embodiment is used to implement the aforementioned power grid fault event fusion method. Therefore, the specific implementation manners in the power grid fault event fusion system can be seen in the embodiment part of the aforementioned power grid fault event fusion method. For example, the data standardization module 100, the event semantic parsing module 200, the event vectorization characterization module 300, the semantic similarity calculation module 400, the multi-granularity clustering result acquisition module 500, and the event fusion module 600 are respectively used to correspondingly implement steps S1, S2, S3, S4, S5, and S6 in the power grid fault event fusion method described in Embodiment 1. Therefore, its specific implementation manners can be referred to the descriptions of the corresponding individual embodiments. To avoid redundancy, they will not be elaborated here.
[0037] Embodiment 3:
[0038] The present invention also provides a computer storage medium. The computer storage medium stores a computer software product. The computer software product includes a number of instructions for causing a computer device to execute the power grid fault event fusion method described in Embodiment 1.
[0039] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0040] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0041] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the specified functions in the process Figure 1One process or multiple processes and / or boxes Figure 1 The functions specified in one box or multiple boxes.
[0042] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 One process or multiple processes and / or boxes Figure 1 The steps of the functions specified in one box or multiple boxes.
[0043] Obviously, the above embodiments are only examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. A method for fusing power grid fault events, characterized in that, It includes the following steps: S1: Obtain multi-source heterogeneous data of the power system, standardize the definition of power fault events, construct a standardized event text structure, extract the original events from the multi-source heterogeneous data, and integrate them according to the standardized event text structure to form a set of original event texts; S2: Based on the set of original events, perform semantic parsing on each original event text, extract the subject and predicate of the event. The subject includes the equipment, components or state variables involved in the fault event, and the predicate includes the abnormal actions or state changes of the fault event; S3: Respectively perform vectorized characterization on the subject and predicate of the event to obtain the dual-channel encoding vectors of the subject and predicate; S4: Calculate the similarity of the dual-channel encoding vectors of the subject and predicate, construct a composite similarity function according to the similarity, and calculate the semantic similarity matrix between events based on the composite similarity function; S5: Merge similar event clusters bottom-up based on the semantic similarity matrix, and generate multi-granularity clustering results according to the preset clustering granularity level; S6: Based on the multi-granularity clustering results, generate multiple candidate standardized events for each clustering cluster, and assign a unique standardized event to each original event in the clustering cluster to complete the fusion of power grid fault events.
2. The power grid fault event fusion method according to claim 1, wherein In S5, the method for generating multi-granularity clustering results according to the preset clustering granularity level is as follows: S51: Update the original event text set using the dual-channel encoding vectors of the subject and predicate extracted in S3 to obtain an updated fault event set , where each event is the vector representation of the subject and predicate after dual-channel semantic encoding; and, calculate the semantic similarity matrix between events in the updated fault event set according to the composite similarity function , where the matrix element represents the semantic similarity between event and event . S52: Cluster each of the updated fault events individually to construct an initial cluster set , where the number of clusters is equal to the number of fault events; S53: Starting from the initial cluster set, perform iterative calculations to calculate the inter-cluster similarity matrix at the current time t , for cluster and cluster , calculate the similarity between them using the following formula : , where and respectively represent the number of events in cluster and cluster ; represents the similarity between event and event calculated according to the composite similarity function; S54: Based on the inter-cluster similarity matrix , obtain the current inter-cluster similarity of the cluster pair with the maximum value , that is: ; S55: Determine whether the condition that the number of clusters after the current iterative update reaches the target value K or the maximum inter-cluster similarity is less than or equal to the similarity threshold is satisfied is met: If not, return to execute step S53; If so, compare the similarity between clusters and a preset similarity threshold in terms of magnitude. If the similarity between clusters is greater than the similarity threshold , then combine cluster and cluster into a new cluster , and update the cluster set . Denote the removal operation; S56: According to the preset clustering granularity levels, gradually construct a dendrogram for each cluster merged in each iteration, and cut the dendrogram to obtain multi-granularity clustering results .
3. The grid fault event fusion method according to claim 2, wherein In S56, according to a preset clustering granularity level, a dendrogram is gradually constructed for each cluster merged in each iteration, and the dendrogram is cut to obtain a multi-granularity clustering result The method is as follows: Use the initial single event cluster as the root node, each newly generated cluster during each merge as a sub-node, and the similarity during the merge as the node height, and gradually construct a dendrogram. Each node of the dendrogram represents a cluster, and the height of the node represents the similarity during the merge; According to the preset clustering granularity hierarchy, determine the corresponding cutting height in the dendrogram; perform a horizontal cut at the specified height in the dendrogram, and divide the dendrogram into multiple subtrees according to the intersection points of the cutting line and the branches of the dendrogram. Each subtree corresponds to a clustering result. By performing cuts at different heights, multi-granularity clustering results are obtained .
4. The power grid fault event fusion method according to claim 1, characterized in that In S1, a set of original event texts is formed The method is as follows: Define the equipment, components or state variables involved in the fault event as the subject, and the abnormal actions or state changes of the fault event as the predicate. Construct a binary array according to the subject and the predicate to represent the fault event text structure, as follows: , Extract the original events from the multi-source heterogeneous data and integrate them according to the standardized event text structure to form a set of original event texts : , where represents the subject array, represents the predicate array, represents the set of original events of the k-th data source, represents the k-th data source, k = 1,..., K, and K represents the total number of the multi-source heterogeneous data.
5. The power grid fault event fusion method according to claim 1, wherein In S2, the method for extracting the subject and predicate of the event is as follows: Build an event parsing function: , where e represents a fault event, s is the subject, and p is the predicate; Guide the event parsing function through text instructions to perform semantic parsing on each original event text, output the fault event text in JSON format, and extract the subject and predicate of the event from the JSON format fault event text.
6. The power grid fault event fusion method according to claim 5, wherein, In S2, the method for extracting the subject and predicate of the event further includes: making the parsing result of the event parsing function fully conform to the JSON format through regular expressions; , Among them, represents the output data, represents the regular expression, and JSON represents the data composed of property-value pairs, represents the constraint item.
7. The grid fault event fusion method according to claim 1, characterized in that In S3, the method for obtaining the dual-channel encoding vectors of the subject and predicate is as follows: Perform structured event dual-channel encoding, perform semantic embedding on the parsed subject s and predicate p respectively, use the pre-trained SBERT model to encode the structured input. The subject s and the predicate p share the same encoder to ensure vector space alignment, and generate the dual-channel encoding vectors of the subject and predicate; , , Among them, represents the subject channel encoding vector, represents the predicate channel encoding vector, represents the dimension, and d is the output dimension size of the SBERT model.
8. The power grid fault event fusion method according to claim 1, characterized in that In S4, the method for constructing a composite similarity function according to the similarity is as follows: Calculation of fault events and fault events weighted cosine similarity of the subject and predicate dual-channel coding vectors, and construct a composite similarity function according to the weighted cosine similarity : , where α ∈ [0, 1] is an adjustable weight coefficient, represents the subject channel encoding vector of the fault event ; represents the subject channel encoding vector of the fault event ; represents the predicate channel encoding vector of the fault event ; represents the predicate channel encoding vector of the fault event .
9. A power grid fault event fusion system, characterized in that, For implementing the steps of the power grid fault event fusion method described in any one of claims 1 to 8, the power grid fault event fusion system includes the following modules: A data standardization module, which is used to obtain multi-source heterogeneous data of the power system, define power failure events in a standardized manner, construct a standardized event text structure, extract original events from the multi-source heterogeneous data, and integrate them according to the standardized event text structure to form a set of original event texts; An event semantic parsing module, which is used to perform semantic parsing on each original event text based on the original event set, extract the subject and predicate of the event, where the subject includes the equipment, components or state quantities involved in the failure event, and the predicate includes abnormal actions or state changes of the failure event; An event vectorization representation module, which is used to perform vectorization representation on the subject and predicate of the event respectively to obtain dual-channel coding vectors of the subject and predicate; A semantic similarity calculation module, which is used to calculate the similarity of the dual-channel coding vectors of the subject and predicate, construct a composite similarity function according to the similarity, and calculate an inter-event semantic similarity matrix based on the composite similarity function; A multi-granularity clustering result acquisition module, which is used to merge similar event clusters from bottom to top based on the semantic similarity matrix, and generate multi-granularity clustering results according to a preset clustering granularity level; An event fusion module, which is used to generate multiple candidate standardized events for each clustering cluster based on the multi-granularity clustering result, and assign a unique standardized event to each original event in the clustering cluster from the multiple candidate standardized events to complete the grid fault event fusion.
10. A computer storage medium, characterized in that, The computer storage medium stores a computer software product, and the computer software product includes several instructions for causing a computer device to execute the grid fault event fusion method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Power grid environment model supporting man-machine two-way understanding and modeling method
CN112434532A
Power equipment fault analysis method and device, equipment and storage medium
CN117235254A
Fault detection method and device
CN119004065A
Structured Knowledge Modeling and Extraction from Images
US20170132526A1
News event search method and system based on multi-level image-text semantic alignment model
WO2023093574A1
Cited By
Marine diesel engine anomaly detection method and system, storage medium and equipment
CN121524646A
Open domain Chinese event pattern induction method and system based on multi-dimensional feature fusion
CN121835691A
Automated cross platform event handling systems, methods, and computer program products
TWI931311B