A Power System Anomaly Detection Method Based on Dual Control Masks
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]本申请实施例提供了一种基于双重控制掩码的电力系统异常检测方法,可以解决电力系统的异常检测延迟高的问题
在本申请的实施例中,通过基于电力设备的电气信号序列和环境工况序列生成多个样本,并基于生成的样本、故障类别、输电线路、调度区和电力系统构建由原始节点、类别原型节点、线路节点、区域节点全局节点构成的多关系图,然后基于生成的样本构建增广特征矩阵和增广邻接矩阵,并基于这些样本构建双重控制掩码矩阵,接着将增广特征矩阵、增广邻接矩阵和双重控制掩码矩阵输入特征提取网络进行特征提取,最终基于提取到的数据特征得到电力系统中各电力设备的异常检测结果。其中,由于双重控制掩码矩阵用于对增广邻接矩阵中的信息传播路径进行结构约束与语义约束的双重控制,使得特征提取网络在特征提取的过程中,能够基于双重控制掩码矩阵进行注意力计算,限制在线推理时控制支持集与查询集之间的信息流动方向,控制信息仅在同设备、同类别等合理结构之间传播,减少不必要的节点交互,降低注意力计算复杂度,从而大大提升检测速度,降低电力系统的异常检测延迟。
Smart Images

Figure CN122365309B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of power system anomaly detection technology, and in particular relates to a power system anomaly detection method based on dual control masks. Background Technology
[0002] Equipment in power systems (such as transformers, circuit breakers, and instrument transformers) is susceptible to overload, insulation aging, and partial discharge during long-term operation, leading to the accumulation of internal defects and potential faults. To ensure reliable power grid operation, real-time monitoring and early warning of equipment status are necessary. Traditional monitoring methods are mostly based on physical models or threshold rules, such as frequency domain analysis of partial discharge and empirical curve prediction of temperature rise. These methods rely on large amounts of historical data and expert experience and are difficult to adapt to complex scenarios involving multiple operating conditions and equipment types.
[0003] Current deep learning-based methods for detecting power system equipment anomalies include time-series feature extraction methods based on the Transformer model. This method directly models signals such as voltage and current and uses multi-head self-attention to capture long-short-term dependencies to identify faults. However, the computational complexity of the global attention mechanism is high, resulting in high latency in power system anomaly detection and making it difficult to meet the low-latency requirements of online monitoring. Summary of the Invention
[0004] This application provides a power system anomaly detection method based on dual control masks, which can solve the problem of high delay in power system anomaly detection.
[0005] This application provides a power system anomaly detection method based on dual control masks, including: Acquire the electrical signal sequence and environmental condition sequence of each power device in the power system during the detection cycle; Multiple samples are generated based on the acquired electrical signal sequence and environmental condition sequence; All generated samples are divided into a support set and a query set; Each sample is treated as a raw node, each fault category as a category prototype node, each transmission line in the power system as a line node, each dispatch area in the power system as a region node, and the power system as a global node. An augmented feature matrix and an augmented adjacency matrix are constructed based on the generated samples. The augmented feature matrix includes the features of the original nodes, the features of the category prototype nodes, the features of the line nodes, the features of the region nodes, and the features of the global nodes. The augmented adjacency matrix is used to characterize the structural relationships and semantic associations between multiple types of nodes, including the original nodes, category prototype nodes, line nodes, region nodes, and global nodes. Based on the support set and query set, a dual control mask matrix is constructed; the dual control mask matrix is used to perform dual control of structural and semantic constraints on the information propagation path in the augmented adjacency matrix. The augmented feature matrix, augmented adjacency matrix, and dual control mask matrix are input into the feature extraction network for feature extraction to obtain the data features of the power system during the detection period. Anomaly detection results of various power equipment in the power system are obtained based on data characteristics.
[0006] Optionally, multiple samples are generated based on the acquired electrical signal sequence and environmental condition sequence, including: For power equipment in power systems Perform the following steps: Power equipment will be placed in a window of a preset length. The electrical signal sequence is divided into multiple electrical signal segments, and the power equipment is... The environmental operating condition sequence is divided into multiple environmental operating condition segments; , This indicates the total number of electrical devices in the power system; The first The electrical signal segments and environmental condition segments corresponding to each window are input into a long short-term memory network for processing to obtain the first window. Preliminary feature representation corresponding to each window; , This indicates the number of time windows corresponding to each power device; A linear transformation is performed on the preliminary feature representation using a linear layer to obtain the linear transformation result; The linear transformation result is input into an activation function for processing to obtain the power equipment. In the Samples under each window .
[0007] Optionally, all generated samples can be divided into a support set and a query set, including: For each generated sample, if the fault category label corresponding to the sample is stored in the detection database, the sample is assigned to the support set; if the fault category label corresponding to the sample is not stored in the detection database, the sample is assigned to the query set. Fault category labels are used to indicate the fault category.
[0008] Optional, augmented feature matrix for: ; in, Represents the original node feature matrix. Represents the feature matrix of the category prototype nodes. Represents the feature matrix of line nodes. Represents the feature matrix of the region nodes. Represents the global node feature vector; , Represents the original node eigenvectors, = , , , Indicates the number of samples generated; , Represents the category prototype node eigenvectors, , , This indicates the number of fault categories corresponding to all samples in the support set. Represents the category prototype node Typical sample representation in feature space, , Represents the category prototype node The set of all corresponding original node indices, express The number of original node indexes in the middle. This indicates the preset adjustable coefficient. Indicates risk weight. , Represents the category prototype node The corresponding number of repairs, Represents the category prototype node The corresponding average recovery time, This represents a pre-defined mechanism vector; , Represents line nodes eigenvectors, , , This indicates the number of transmission lines in the power system. Represents line nodes The set of all corresponding original node indices, , Represents the original node The score, Indicates to The result after normalization Represents the original node The score, , Indicates learnable parameters, Indicates learnable parameters, Indicates learnable parameters; , Represents a region node eigenvectors, , This indicates the number of dispatch zones in the power system.
[0009] Optional, augmented adjacency matrix for: ; in, This represents the connection matrix between the original nodes. This represents the normalized similarity weight matrix from the original node to the category prototype node. This represents the pooling matrix of line nodes. Represents the pooling matrix of region nodes. express The vector, express The matrix, express The matrix, express The vector, express The matrix, Represents a binary relation matrix. express The matrix, express The vector, express The vector, express The vector, express The vector, , , , , , , , , and All elements in the middle are 1; Connection matrix The Middle Line number Column elements for: ; Represents the original node eigenvectors With the original node eigenvectors Similarity score, , , , This indicates the preset bandwidth parameter; Line node pooling matrix The Middle Line number Column elements for: ; Represents the original node Corresponding transmission line index, express The number of original node indexes in the middle; Region Node Pooling Matrix The Middle Line number Column elements for: ; Represents a region node The set of all corresponding original node indices, express The number of original node indexes in the middle. Represents the original node The corresponding region node index; Binary relation matrix The Middle Line number Column elements for: ; This indicates that there is at least one original node. Belongs to line node And belongs to the regional node , Represents line nodes Not a regional node ; Normalized similarity weight matrix The Middle Line number Column elements for: ; ; Represents the original node With category prototype nodes Unnormalized similarity between them Represents the original node With category prototype nodes The unnormalized similarity between them.
[0010] Optionally, a dual control mask matrix can be constructed based on the support set and the query set, including: The structure mask matrix and pattern mask matrix are determined based on the support set and query set; The dual control mask matrix is calculated using the following formula. : ; in, Represents the structure mask matrix. Represents the pattern mask matrix.
[0011] Optional, pattern mask matrix The Middle Line number Column elements for: ; in, Indicates the number of samples supporting the set. , , .
[0012] Optional, structure mask matrix The Middle Line number Column elements for: when Corresponding original node , Corresponding original node At that time, if the original node and the original node Corresponding to the same type of power equipment and the original node and the original node When all corresponding samples belong to the support set or query set, then =1, otherwise, =0; when Corresponding original node , Corresponding category prototype node At that time, if the original node The corresponding category prototype node index is ,but =1, otherwise, =0; when Corresponding original node , When corresponding to line nodes, regional nodes, or global nodes. =1; when , When all correspond to a category prototype node, line node, region node, or global node =1.
[0013] Optionally, if the pattern mask matrix The Middle Line number Column elements With structural mask matrix The Middle Line number Column elements If both are 1, then the feature extraction network performs well in the feature extraction process. and Attention is calculated on the feature vectors of the corresponding nodes; otherwise, the feature extraction network will not perform attention calculations during the feature extraction process. and Attention is calculated using the feature vectors of the corresponding nodes.
[0014] Optionally, the anomaly detection results of each power device in the power system can be obtained based on data characteristics, including: Calculate the features of the data in the sample Output embedding vectors and category prototype nodes Euclidean distance between eigenvectors ; The Euclidean distance is expressed by the following formula. Convert to probability distribution : ; The predicted fault category is calculated using the following formula. : ; The basic anomaly is calculated using the following formula. : ; The risk-weighted outlier score is calculated using the following formula. : ; Based on the risk-weighted anomaly scores corresponding to all samples, the anomaly detection results for each power device in the power system are obtained; the anomaly detection result for each power device is used to indicate whether the power device has an anomaly; where, if the risk-weighted anomaly score... If the value is greater than or equal to the preset threshold, then the sample is determined. The corresponding power equipment is malfunctioning; in, This indicates the preset temperature parameter. Representing the sample in the data features Output embedding vectors and category prototype nodes The Euclidean distance between the eigenvectors, express The probability distribution at time, This indicates that the fault category is no fault. This represents the risk weight hyperparameter. This represents the risk weight corresponding to the predicted fault category. , Indicates the predicted fault category The corresponding number of repairs, Indicates the predicted fault category The corresponding average recovery time.
[0015] The above-mentioned solution in this application has the following beneficial effects: In the embodiments of this application, multiple samples are generated based on the electrical signal sequences and environmental condition sequences of power equipment. A multi-relationship graph, consisting of original nodes, category prototype nodes, line nodes, regional nodes, and global nodes, is constructed based on the generated samples, fault categories, transmission lines, dispatch areas, and the power system. Then, an augmented feature matrix and an augmented adjacency matrix are constructed based on the generated samples, and a dual control mask matrix is constructed based on these samples. The augmented feature matrix, augmented adjacency matrix, and dual control mask matrix are then input into a feature extraction network for feature extraction. Finally, the anomaly detection results for each power equipment in the power system are obtained based on the extracted data features. The dual control mask matrix is used to impose both structural and semantic constraints on the information propagation path in the augmented adjacency matrix. This allows the feature extraction network to perform attention calculations based on the dual control mask matrix during feature extraction, limiting the direction of information flow between the support set and the query set during online inference. Information is controlled to propagate only between reasonable structures such as devices and categories, reducing unnecessary node interactions and lowering the computational complexity of attention calculations. This significantly improves detection speed and reduces the anomaly detection latency of the power system.
[0016] Other beneficial effects of this application will be described in detail in the following detailed description section. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart of a power system anomaly detection method based on dual control masks provided in an embodiment of this application. Detailed Implementation
[0019] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0020] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0021] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0022] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0023] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0024] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0025] To address the high latency issue in current power system anomaly detection, this application provides a power system anomaly detection method based on a dual control mask. This method generates multiple samples based on electrical signal sequences and environmental condition sequences of power equipment. Based on the generated samples, fault categories, transmission lines, dispatch areas, and the power system, a multi-relationship graph is constructed, consisting of original nodes, category prototype nodes, line nodes, regional nodes, and global nodes. Then, an augmented feature matrix and an augmented adjacency matrix are constructed based on the generated samples, and a dual control mask matrix is built based on these samples. Next, the augmented feature matrix, augmented adjacency matrix, and dual control mask matrix are input into a feature extraction network for feature extraction. Finally, the anomaly detection results for each power device in the power system are obtained based on the extracted data features. In particular, since the dual control mask matrix is used to perform dual structural and semantic constraints on the information propagation path in the augmented adjacency matrix, the feature extraction network can perform attention calculation based on the dual control mask matrix during the feature extraction process. This restricts the direction of information flow between the support set and the query set during online inference, and controls information to propagate only between reasonable structures such as the same device and the same category. This reduces unnecessary node interactions, lowers the computational complexity of attention, and thus greatly improves the detection speed and reduces the anomaly detection latency of the power system.
[0026] The following describes the power system anomaly detection method based on dual control masks provided in this application by way of specific embodiments.
[0027] like Figure 1 As shown in the embodiments of this application, the power system anomaly detection method based on dual control masks includes the following steps: Step 11: Obtain the electrical signal sequence and environmental condition sequence of each power device in the power system during the detection cycle.
[0028] The aforementioned power equipment can be transformers, circuit breakers, instrument transformers, etc. The aforementioned testing cycle can be set according to actual conditions, such as setting it to the nearest day or the nearest week.
[0029] The aforementioned electrical signal sequence includes basic electrical quantities of the power equipment during the detection period, such as three-phase voltage, three-phase current, active and reactive power. In some optional embodiments, frequency and power quality indicators may also be included.
[0030] The aforementioned environmental condition sequence includes external meteorological factors such as temperature and humidity of the environment in which the power equipment is located during the detection period. In some optional embodiments, it may also include internal operating condition information such as load rate, switch status, and operating mode.
[0031] Step 12: Generate multiple samples based on the acquired electrical signal sequence and environmental condition sequence.
[0032] In some embodiments of this application, step 12 is specifically implemented as follows: For power equipment in power systems ( , (This represents the total number of electrical devices in the power system), so proceed with steps 12.1 to 12.4: Step 12.1: Place the electrical equipment in a window of the preset length. The electrical signal sequence is divided into multiple electrical signal segments, and the power equipment is... The environmental operating condition sequence is divided into multiple environmental operating condition segments.
[0033] The preset length mentioned above can be set according to actual conditions. It is understood that the preset length window is used to monitor electrical equipment. electrical signal sequence and environmental operating condition sequence By dividing the electrical signal sequence, Divide into multiple electrical signal segments and environmental operating condition sequence It is divided into multiple environmental condition segments. Among them, the first... The signals within each window include electrical signal segments. and environmental operating condition segments .
[0034] Step 12.2, the first The electrical signal segments and environmental condition segments corresponding to each window are input into a long short-term memory network for processing to obtain the first window. Preliminary feature representation corresponding to each window; , This indicates the number of time windows corresponding to each power device.
[0035] Specifically, the first The electrical signal segment corresponding to each window and environmental operating condition segments The input is processed by a single-layer long short-term memory network to obtain the first... Preliminary feature representation corresponding to each window.
[0036] It is understandable that by traversing all windows, a preliminary feature representation for each window can be obtained.
[0037] Step 12.3: Perform a linear transformation on the preliminary feature representation using a linear layer to obtain the linear transformation result.
[0038] It should be noted that, for each window, a linear layer is used to perform a linear transformation on the preliminary feature representation corresponding to that window, thus obtaining the linear transformation result for that window. The role of the linear layer here is to perform a linear mapping on the preliminary feature representation and adjust the feature dimensions.
[0039] Step 12.4: Input the linear transformation result into the activation function for processing to obtain the power equipment. In the Samples under each window .
[0040] For example, the linear transformation result corresponding to each window can be input into the activation function separately. After processing, the sample under this window can be obtained, which is the power equipment. In the Samples under each window .
[0041] In some embodiments of this application, after obtaining samples for each power device, a globally unique index is assigned to each sample. ( , This refers to the number of samples generated, which is the total number of windows across all power devices and all time periods. Simultaneously, indexes will be obtained from external systems (such as the power system's management system) or tags. The system stores information about the associated equipment (power equipment), fault category, transmission line, and dispatch area in a detection database. Each power device, fault category, transmission line, and dispatch area has a unique and verifiable identifier. It should be noted that in the embodiments of this application, for the purpose of accurate anomaly detection, normal operation (i.e., no fault) is also considered a fault category. Fault categories include transformer faults, circuit breaker faults, and instrument transformer faults. It should be noted that the aforementioned dispatch area refers to a dispatch management unit formed within the power system based on operation management, geographical region, or grid topology. Each dispatch area is typically monitored and dispatched uniformly by the power grid dispatch center and includes several substations, transmission lines, and their associated power equipment. Different dispatch areas are relatively independent in operation control but can be interconnected through the main power grid. In practical applications, the dispatch area to which a power equipment belongs is determined by the power grid topology. Based on the power grid partition to which the substation or busbar to which the equipment is connected belongs, the dispatch area number to which the power equipment belongs can be directly obtained through the equipment-dispatch area mapping table in the power dispatch automation system (SCADA), and the power equipment is assigned to the corresponding dispatch area.
[0042] To facilitate data processing and identification, each fault category in the detection data has a corresponding fault category label. For example, a fault category label of 0 indicates a normal fault category, while a label of 1 indicates a circuit breaker fault. In practical applications, there are time periods within the aforementioned detection cycle where the abnormal conditions of various power equipment are known. In other words, for samples within these time periods, corresponding fault category labels are assigned. For instance, if a circuit breaker operates normally during this period, the fault category for the circuit breaker samples within that period is "normal"; if a current transformer malfunctions during this period, the fault category for the current transformer samples within that period is "current transformer fault."
[0043] Step 13: Divide all generated samples into a support set and a query set.
[0044] In some embodiments of this application, step 13 can be implemented as follows: for each generated sample, if the detection database stores a fault category label corresponding to the sample, the sample is assigned to the support set; if the detection database does not store a fault category label corresponding to the sample, the sample is assigned to the query set. The fault category label indicates the fault category. The support set samples are from historically labeled time periods, and the query set samples are from the current or future time periods to be detected. The two sets are strictly isolated in time to avoid data leakage and ensure the reliability of the anomaly detection results.
[0045] Step 14: Treat each sample as a raw node, each fault category as a category prototype node, each transmission line in the power system as a line node, each dispatch area in the power system as a region node, and the power system as a global node.
[0046] It should be noted that after defining the above-mentioned original nodes, category prototype nodes, line nodes, region nodes, and global nodes, these nodes constitute a multi-relationship graph. This multi-relationship graph is an augmented task graph that includes risk correction prototype nodes and hierarchical convergence nodes, which can achieve high-precision matching with small samples.
[0047] Step 15: Construct an augmented feature matrix and an augmented adjacency matrix based on the generated samples.
[0048] The augmented feature matrix mentioned above includes features of the original nodes, features of the category prototype nodes, features of the line nodes, features of the region nodes, and features of the global nodes; the augmented adjacency matrix mentioned above is used to characterize the structural relationships and semantic associations between multiple types of nodes, including original nodes, category prototype nodes, line nodes, region nodes, and global nodes.
[0049] In some embodiments of this application, the above-mentioned augmented feature matrix for: ; in, Represents the original node feature matrix. Represents the feature matrix of the category prototype nodes. Represents the feature matrix of line nodes. Represents the feature matrix of the region nodes. This represents the feature vector of the global node.
[0050] , Represents the original node eigenvectors, = , , , This indicates the number of samples generated. It can be understood as a sample index, and also as a raw node index.
[0051] , Represents the category prototype node eigenvectors, , , This indicates the number of fault categories corresponding to all samples in the support set. Represents the category prototype node Typical sample representation in feature space, , It contains only samples from the support set, representing category prototype nodes. The set of all corresponding original node indices (category prototype nodes) Essentially, it refers to the fault category. Category prototype node The corresponding raw node indexes refer to those that support centralized fault categories. The index of all original nodes (i.e., samples), where the fault category of each sample can be queried from the detection database. express The number of original node indexes in the middle. This indicates the preset adjustable coefficient. Determine the overall weight of risk knowledge in prototype adjustments. Indicates risk weight. , Represents the category prototype node The corresponding number of repairs (i.e., the number of maintenance visits). Represents the category prototype node The corresponding mean recovery time (i.e., mean repair time). This represents a pre-defined mechanism vector. Here... This can be understood as a fault category index, and also as a category prototype node index. Among them, the mechanism vector... These are parameters pre-set by domain experts based on the fault mechanism to reflect the prior distribution direction of different fault types in the feature space.
[0052] The number of repairs and the average recovery time mentioned above can be obtained by querying the maintenance log. This maintenance log records all fault repair information for the power system, specifically including the time of each repair, the fault category of each repair, the number of repairs for each fault category, and the average repair time.
[0053] The above This is used to reflect the frequency and severity of the fault. The larger the value, the more common and slow-recovering the fault is, posing a higher risk to the power system and thus deserving greater attention in the model. A predefined vector is generated by domain experts to characterize the fault category. The desired directional bias in the feature space compensates for possible bias or scarcity of data samples.
[0054] , Represents line nodes eigenvectors, , , This indicates the number of transmission lines in the power system. Represents line nodes The set of all corresponding original node indices (line nodes) Essentially a power transmission line (here This can be understood as a transmission line index, and also a line node index. (Line node) All corresponding original node indices refer to transmission lines (Index of all original nodes (i.e., samples) of all power devices). , Represents the original node The score is obtained by assigning weights to each node in the set using a parameterized attention scorer, and then performing a weighted sum. Indicates to The result after normalization Represents the original node The score, , Indicates learnable parameters, , Indicates learnable parameters, , Indicates learnable parameters, , Hiding dimensions for attention For feature vectors The length (obtained by projection in step 12.4). and The calculation methods are the same, the only difference being that one uses the original node. The feature vectors, one of which is the original node The eigenvectors are not described in detail here.
[0055] , Represents a region node eigenvectors, , This indicates the number of dispatch areas in the power system. It can be understood as a scheduling area index, and also a region node index.
[0056] It should be noted that regional nodes eigenvectors and The calculation method is the same, the difference lies in... In the calculation formula Replace with , Replace with Specifically, , Represents a region node The set of all corresponding original node indices (region nodes) Essentially a scheduling area Dispatch area All corresponding raw node indices refer to the scheduling area. (Index of all original nodes (i.e., samples) of all power devices within the system). , Represents the original node The score, Indicates to The result after normalization Represents the original node The score, , Indicates learnable parameters, Indicates learnable parameters, This represents the learnable parameters.
[0057] The above for ,Should The calculation method and The calculation method is the same, so the calculation formula will not be repeated here. The difference lies in the calculation... Attention pooling covers all nodes in the entire network (i.e., all original nodes). Replace with , Replace with , Replace with , , Replace with , , Represents the set of all original node indices. Indicates to The result after normalization Represents the original node The score.
[0058] In some embodiments of this application, the above-described augmented adjacency matrix for: ; in, This represents the connection matrix between the original nodes. This represents the normalized similarity weight matrix from the original node to the category prototype node. This represents the pooling matrix of line nodes. , Represents the pooling matrix of region nodes. , express The vector, express The matrix, express The matrix, express The vector, express The matrix, Represents a binary relation matrix. , express The matrix, express The vector, express The vector, express The vector, express The vector, , , , , , , , , and All elements in the matrix are 1. The matrix or vector with all elements being 1 is a column vector with all elements being 1, from the original node to the global node. This indicates that the global node receives equal weighted information from each original node. Category nodes, line nodes, and region nodes are connected to each other with all elements being 1, indicating that their interactions are fully connected.
[0059] , : Represents the connection relationship between all original nodes and the global node, that is, the global node receives equal weight information from all original nodes. , : Represents a bidirectional fully connected relationship between the category prototype node and the global node. , : Indicates a bidirectional fully connected relationship between a regional node and a global node. , : Indicates a bidirectional full connectivity relationship between category nodes and line nodes. , : Represents a bidirectional fully connected relationship between category prototype nodes and region nodes.
[0060] Connection matrix It is based on the Gaussian kernel similarity definition and is used to represent the similarity between the original nodes. This connection matrix The Middle Line number Column elements Represents the original node With the original node This matrix indicates whether there are connections between elements and the strength of those connections. A value of 0 indicates no connection, while a non-zero value indicates a connection. Larger values indicate stronger connections. The Middle Line number Column elements for: ; Represents the original node eigenvectors With the original node eigenvectors Similarity score, , , Original node With the original node It can be used for the same or different power equipment. This represents the preset bandwidth parameter used to measure sample similarity. The smaller the value, the stronger the sensitivity. Indicates other. When A similarity value greater than 0.5 indicates high similarity, meaning the electrical behavior or fault modes match, and a connection can be made; otherwise, the connection is set to 0. In practical applications, a connection matrix can be constructed according to the node order. .
[0061] The above line node pooling matrix This line node pooling matrix is used to represent the attribution relationship from the original node to the line node and its normalized aggregation weights. The Middle Line number Column elements Represents the original node To the line node The attribution relationship and its normalized aggregation weight, the pooling matrix of the line nodes. The Middle Line number Column elements for: ; Represents the original node Corresponding transmission line index, express The number of original node indexes in the middle.
[0062] The above region node pooling matrix This is used to represent the attribution relationship from the original node to the region node and its normalized aggregation weight; the region node pooling matrix is used for this purpose. The Middle Line number Column elements Represents the original node In the relevant dispatch area The weights borne by nodes in the aggregation process, and the node pooling matrix of this region. The Middle Line number Column elements for: ; Represents a region node The set of all corresponding original node indices, express The number of original node indexes in the middle. Represents the original node The corresponding region node index.
[0063] The above binary relation matrix row index Corresponding line nodes, column index The corresponding regional nodes represent the correspondence between transmission lines and dispatch areas. This binary relation matrix... The Middle Line number Column elements Indicates transmission line Does it belong to the dispatch area? The attribution relationship, the binary relation matrix The Middle Line number Column elements for: ; This indicates that there is at least one original node. Belongs to line node And belongs to the regional node (i.e., the original node) The corresponding power equipment belongs to the transmission line. And the original node The corresponding power equipment belongs to the dispatch area. ), Represents line nodes Not a regional node (i.e., line nodes) Corresponding transmission lines Not part of the scheduling area ).
[0064] The above normalized similarity weight matrix It is the normalized similarity weight matrix from the original node to the category prototype node. The Middle Line number Column elements Represents the original node With category prototype nodes The unnormalized similarity between them (a Gaussian kernel-based similarity measure), and the normalized similarity weight matrix. The Middle Line number Column elements for: ; ; Represents the original node With category prototype nodes Unnormalized similarity between them (calculated in the same way as) (Same). It should be noted that, and The calculation method is the same; you only need to add the category prototype node. Replace the relevant data with category prototype nodes The relevant data.
[0065] In some embodiments of this application, the augmented feature matrix , This represents the total number of nodes in the graph. .
[0066] Step 16: Construct a dual control mask matrix based on the support set and query set.
[0067] The aforementioned dual control mask matrix is used to perform dual control of structural and semantic constraints on the information propagation path in the augmented adjacency matrix.
[0068] In some embodiments of this application, the structure mask matrix and the pattern mask matrix can be determined based on the support set and the query set; then, the dual control mask matrix can be calculated using the following formula. : ; in, Represents the structure mask matrix. It is used to control whether attention computing or information interaction is allowed between nodes. It constrains the connection relationship between nodes based on the type of power equipment, fault category and task hierarchy structure. The task hierarchy consists of line nodes, regional nodes and global nodes. Represents the pattern mask matrix. Used to control the communication mode between nodes in real-time or batch inference scenarios, ensuring computational efficiency and minimal latency, dual control mask matrix. .
[0069] In some embodiments of this application, in low-latency real-time monitoring scenarios, newly acquired samples in the query set need to be immediately assessed for anomalies. To address the real-time monitoring requirements of this scenario and minimize node communication while ensuring rapid inference, this pattern mask matrix is defined. The Middle Line number Column elements for: ; in, Indicates the number of samples supporting the set. , Supports centralized sample indexing. The index of the sample in the query set is Ultimately, whether the actual information is transmitted still needs to be determined jointly with the structure mask matrix: only when... =1、 At time 1, node Only then will the information be received by the node Received and used for attention calculation or feature aggregation. The constraint rules regarding the support set and query set in the pattern mask matrix only apply to the connection relationships between the original nodes. or When the corresponding node is a category prototype node, line node, region node, or global node, it is not affected. Restrictions, their corresponding elements The default value is 1, that is The corresponding node can be one of the following: original node, category prototype node, line node, region node, or global node. The corresponding node can be one of the following: original node, category prototype node, line node, region node, and global node. Corresponding nodes and The corresponding nodes cannot all be the original nodes corresponding to the samples in the query set.
[0070] Combining the support set and the sample index in the query set, the above pattern mask matrix This indicates that information exchange is only allowed between support set nodes and between support set and query set nodes, while communication between query set nodes is prohibited to achieve the lowest possible latency. Specifically, all category prototype nodes and task-level aggregation nodes (line nodes, region nodes, global nodes), as well as they and any original node (support or query), are fully connected and set to 1.
[0071] In other embodiments of this application, during the offline batch inference and training phases, historical information between query node sequences can be used for feature enhancement. Query nodes can access information from previous windows, which is beneficial for capturing abnormal trends in long sequences and improving overall detection accuracy. Batch backtesting analysis is performed daily on electrical signals collected over the past 24 hours, processing all query windows in chronological order and utilizing prior information to improve detection accuracy. In this scenario, the aforementioned pattern mask matrix... The Middle Line number Column elements for: ; in, , , , This indicates that subsequent query nodes do not pass information to previous query nodes. Represents a node Feature information can flow to nodes In graph convolutional or graph attention networks, It will receive from the node The features are used to update their own representations.
[0072] The above structure mask matrix The Middle Line number Column elements Depend on Corresponding nodes and The type of the corresponding node determines this. , .
[0073] Specifically, when Corresponding original node , Corresponding original node At that time, if the original node and the original node Corresponding to the same type of power equipment and the original node and the original node When all corresponding samples belong to the support set or query set, then =1, otherwise, =0.
[0074] Among them, the aforementioned original nodes With the original node Each corresponds to one of the previously generated samples. That is, when and When both correspond to the original node, a connection is allowed only if they belong to the same type of power equipment and both come from the support set or both come from the query set.
[0075] when Corresponding original node , Corresponding category prototype node At that time, if the original node The corresponding category prototype node index is ,but =1, otherwise, =0. That is, for any original node and category prototype node, as long as the fault category is the same, it is connected to the corresponding category prototype node; that is, if... The corresponding original node category is It connects to the category prototype node. .
[0076] when Corresponding original node , When corresponding to line nodes, regional nodes, or global nodes. =1. That is, any task node (i.e., any line node, area node, or global node) is set to 1 when communicating with all original nodes.
[0077] when , When all correspond to a category prototype node, line node, region node, or global node =1. In other words, the connections between category prototype nodes and task nodes are not limited by the sample-level structure, maintaining full connectivity to achieve cross-category and cross-level information interaction, structure mask matrix. Attention is only turned on between original node pairs that are "on the same device and have the same type of fault", completely blocking direct interaction between different devices or heterogeneous samples, while preserving the global connectivity between class nodes, task nodes and original nodes.
[0078] It should be noted that, regardless of , Which node it corresponds to, and whether the actual information is ultimately transmitted, still needs to be determined jointly with the pattern mask matrix: only when... =1、 When =1, node Only then will the information be received by the node It receives and uses the information for attention calculation or feature aggregation. The structure mask only controls which information is allowed to flow between nodes structurally; it does not directly represent the actual information transfer, but rather restricts the connections between specific node pairs. The corresponding node can be one of the following: original node, category prototype node, line node, region node, or global node. The corresponding node can be one of the following: original node, category prototype node, line node, region node, and global node.
[0079] Furthermore, in some embodiments of this application, only when the pattern mask matrix The Middle Line number Column elements With structural mask matrix The Middle Line number Column elements When both are 1, Corresponding nodes and Attention calculations are only allowed between corresponding nodes. It should be noted that the pattern mask matrix... The Middle Line number Column elements With structural mask matrix The Middle Line number Column elements When both are 1, it indicates that both are 1, representing a node. With nodes This allows for connectivity both structurally and in the current mode, at which point the node... Feature information can be obtained from nodes Receiving and processing. It's important to note that graph neural networks or attention aggregation need to determine the contribution weights of information between nodes; if information transmission between nodes is not allowed, attention is not needed and the nodes are directly blocked; when both are 1, it indicates that a legal and allowed communication path exists between nodes, and attention needs to be calculated to determine the node's contribution weight. Information for nodes The actual contribution of features is used to perform feature aggregation or information transmission.
[0080] Step 17: Input the augmented feature matrix, augmented adjacency matrix and dual control mask matrix into the feature extraction network for feature extraction to obtain the data features of the power system during the detection period.
[0081] In some embodiments of this application, the feature extraction network described above may be specifically a Transformer model.
[0082] When the augmented feature matrix, augmented adjacency matrix, and dual control mask matrix are input into the feature extraction network for feature extraction, the dual control mask matrix is used to exert both structural and semantic constraints on the information propagation path in the augmented adjacency matrix. Specifically, if the pattern mask matrix... The Middle Line number Column elements With structural mask matrix The Middle Line number Column elements If both are 1, then the feature extraction network performs well in the feature extraction process. and Attention is calculated on the feature vectors of the corresponding nodes; otherwise, the feature extraction network will not perform attention calculations during the feature extraction process. and Attention is calculated using the feature vectors of the corresponding nodes. The corresponding node can be one of the following: original node, category prototype node, line node, region node, or global node. The corresponding node can be one of the following: original node, category prototype node, line node, region node, and global node.
[0083] It should be noted that the Transformer model is a commonly used feature extraction model, and its principles will not be elaborated on here.
[0084] Step 18: Obtain the anomaly detection results of each power device in the power system based on data characteristics.
[0085] The above data features are composed of the output embedding vectors of all samples, and the output embedding vector of each sample expresses the feature representation after the synthesis of the corresponding time window.
[0086] In some embodiments of this application, step 18 is specifically implemented as follows: steps 18.1 to 18.6: Step 18.1, calculate the sample features in the data. Output embedding vectors and category prototype nodes Euclidean distance between eigenvectors .
[0087] Specifically, it can be done through formulas Calculate the features of the data in the sample The output embedding vector With category prototype nodes eigenvectors ( Euclidean distance between ) . , The data features output by the feature extraction network are matrices. .
[0088] Step 18.2, probabilistic matching, using temperature parameters The controlled Softmax function transforms similarity into a probability distribution; specifically, it converts Euclidean distance using the following formula. Convert to probability distribution : ; Step 18.3: Calculate the predicted fault category using the following formula. : ; Step 18.4, calculate the basic anomalies using the following formula. : ; Step 18.5: Calculate the risk-weighted anomaly score using the following formula. : ; Step 18.6: Based on the risk-weighted anomaly scores corresponding to all samples, obtain the anomaly detection results for each power device in the power system; the anomaly detection result for each power device is used to indicate whether the power device has an anomaly; wherein, if the risk-weighted anomaly score... If the value is greater than or equal to the preset threshold, then the sample is determined. The corresponding power equipment has an anomaly, if the risk-weighted anomaly score is... If the sample size is less than the preset threshold, then the sample is determined to be a valid sample. The corresponding power equipment is normal (i.e., there is no abnormality).
[0089] in, This indicates the preset temperature parameter (which can be set according to the actual situation). Representing the sample in the data features Output embedding vectors and category prototype nodes The Euclidean distance between the eigenvectors, express The probability distribution at time, This indicates that the fault category is no fault. This represents the risk weight hyperparameter. , Used to balance the impact of data and engineering risks, when When relying solely on model probabilities as the data driver, when At times, relying solely on risk history as an expert driver, This represents the risk weight corresponding to the predicted fault category (to ensure consistency of measurement, the weight will be adjusted accordingly). (Normalization process) , Indicates the predicted fault category (Predicted Fault Category) For the aforementioned The number of repairs corresponding to one of the fault categories. Indicates the predicted fault category The corresponding average recovery time.
[0090] in, The value combines current observational evidence of anomalies with the historical risk of this type of failure. The higher the value, the more suspicious the sample is and the higher the priority it should be given.
[0091] It should be noted that the learnable parameters in the embodiments of this application are all obtained through a training process. For example, training methods commonly used in deep learning can be used to train the relevant model parameters (such as the learnable parameters of the attention scorer and the relevant parameters of the feature extraction network) in the embodiments of this application.
[0092] The following example illustrates the power system anomaly detection method based on dual control masks proposed in this application.
[0093] In this example, a comparative experiment is conducted using a real monitoring dataset of 222 devices and 132,000 time window samples from a power grid. Temporal features are extracted using LSTM. , , The dimensions representing the hidden states and cell states of the LSTM. (Representing the dimension of the projected features), constructing a multi-level graph structure of device-line-region, and introducing risk-weighted prototype nodes ( , (representing risk weighting coefficients) and the three-level task aggregation node ( , (Representing the hidden dimension of the attention mechanism for line / region / global level nodes), using a 4-layer 8-head graph neural network (GNN) - Transformer ( , , This indicates the hidden dimension of the feedforward layer. (Representing the mask weighting coefficients) for feature inference, and setting the temperature during the fault matching stage. Distance threshold Confidence threshold Risk weight hyperparameter .
[0094] Experiments conducted under the aforementioned conditions revealed that the model employing the proposed method achieved an unknown anomaly alarm rate of 0.710, compared to only 0.320 for the baseline Transformer model without graph structure and memory matching, representing an improvement of approximately 121.9% in the unknown fault recall rate. Simultaneously, the known fault recall rate reached 0.835, with a false alarm rate as low as 0.125, demonstrating significant superiority over the comparative methods in all core metrics. Therefore, the proposed power equipment anomaly detection method is both reasonable and feasible.
[0095] In summary, the method provided in this application achieves high-precision matching with small samples by constructing a multi-relationship graph; and designs a dual control mask mechanism to restrict the direction of information flow between the support set and the query set during online inference, ensuring that control information only spreads between reasonable structures such as the same device and the same category, reducing unnecessary node interactions and lowering the computational complexity of attention, thereby significantly reducing online inference latency while ensuring detection recall, and thus meeting the engineering requirements for real-time monitoring and risk warning of power system equipment.
[0096] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles described in this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A power system anomaly detection method based on dual control masks, characterized in that, include: Acquire the electrical signal sequence and environmental condition sequence of each power device in the power system during the detection cycle; Multiple samples are generated based on the acquired electrical signal sequence and environmental condition sequence; All generated samples are divided into a support set and a query set; Each sample is treated as a raw node, each fault category as a category prototype node, each transmission line in the power system as a line node, each dispatch area in the power system as a region node, and the power system as a global node. An augmented feature matrix and an augmented adjacency matrix are constructed based on the generated samples. The augmented feature matrix includes features of the original nodes, features of the category prototype nodes, features of the line nodes, features of the region nodes, and features of the global nodes. The augmented adjacency matrix is used to characterize the structural relationships and semantic associations between multiple types of nodes, including the original nodes, the category prototype nodes, the line nodes, the region nodes, and the global nodes. Based on the support set and query set, a dual control mask matrix is constructed; the dual control mask matrix is used to perform dual control of structural and semantic constraints on the information propagation path in the augmented adjacency matrix; The augmented feature matrix, the augmented adjacency matrix, and the dual control mask matrix are input into a feature extraction network for feature extraction to obtain the data features of the power system within the detection period. Based on the data characteristics, obtain the anomaly detection results of each power device in the power system; The construction of a dual control mask matrix based on the support set and query set includes: The structure mask matrix and pattern mask matrix are determined based on the support set and query set; The dual control mask matrix is calculated using the following formula. : ; in, Represents the structure mask matrix. Represents the pattern mask matrix; Pattern mask matrix The Middle Line 1 Column elements for: ; in, This indicates the number of samples in the support set. , , ; Structure mask matrix The Middle Line 1 Column elements for: when Corresponding original node , Corresponding original node At that time, if the original node and the original node Corresponding to the same type of power equipment and the original node and the original node When all corresponding samples belong to the support set or query set, then =1, otherwise, =0; when Corresponding original node , Corresponding category prototype node At that time, if the original node The corresponding category prototype node index is ,but =1, otherwise, =0; when Corresponding original node , When corresponding to line nodes, regional nodes, or global nodes. =1; when , When all correspond to a category prototype node, line node, region node, or global node =1.
2. The power system anomaly detection method according to claim 1, characterized in that, The process generates multiple samples based on the acquired electrical signal sequence and environmental condition sequence, including: For the power equipment in the power system Perform the following steps: The power equipment is arranged according to a window of preset length. The electrical signal sequence is divided into multiple electrical signal segments, and the power equipment The environmental operating condition sequence is divided into multiple environmental operating condition segments; , This indicates the total number of electrical devices in the power system; The first The electrical signal segments and environmental condition segments corresponding to each window are input into a long short-term memory network for processing to obtain the first window. Preliminary feature representation corresponding to each window; , This indicates the number of time windows corresponding to each power device; The preliminary feature representation is linearly transformed using a linear layer to obtain the linear transformation result; The linear transformation result is input into an activation function for processing to obtain the power equipment. In the Samples under each window .
3. The power system anomaly detection method according to claim 2, characterized in that, The process of dividing all generated samples into a support set and a query set includes: For each generated sample, if the fault category label corresponding to the sample is stored in the detection database, the sample is assigned to the support set; if the fault category label corresponding to the sample is not stored in the detection database, the sample is assigned to the query set. The fault category label is used to indicate the fault category.
4. The power system anomaly detection method according to claim 3, characterized in that, Augmented feature matrix for: ; in, Represents the original node feature matrix. Represents the feature matrix of the category prototype nodes. Represents the feature matrix of line nodes. Represents the feature matrix of the region nodes. Represents the global node feature vector; , Represents the original node eigenvectors, = , , , Indicates the number of samples generated; , Represents the category prototype node eigenvectors, , , This indicates the number of fault categories corresponding to all samples in the support set. Represents the category prototype node Typical sample representation in feature space, , Represents the category prototype node The set of all corresponding original node indices, express The number of original node indexes in the middle. This indicates the preset adjustable coefficient. Indicates risk weight. , Represents the category prototype node The corresponding number of repairs, Represents the category prototype node The corresponding average recovery time, This represents a pre-defined mechanism vector; , Represents line nodes eigenvectors, , , This indicates the number of transmission lines in the power system. Represents line nodes The set of all corresponding original node indices, , Represents the original node The score, Indicates to The result after normalization Represents the original node The score, , Indicates learnable parameters, Indicates learnable parameters, Indicates learnable parameters; , Represents a region node eigenvectors, , This indicates the number of dispatch zones in the power system.
5. The power system anomaly detection method according to claim 4, characterized in that, Augmented adjacency matrix for: ; in, This represents the connection matrix between the original nodes. This represents the normalized similarity weight matrix from the original node to the category prototype node. This represents the pooling matrix of line nodes. Represents the pooling matrix of region nodes. express The vector, express The matrix, express The matrix, express The vector, express The matrix, Represents a binary relation matrix. express The matrix, express The vector, express The vector, express The vector, express The vector, , , , , , , , , and All elements in the middle are 1; Connection matrix The Middle Line 1 Column elements for: ; Represents the original node eigenvectors With the original node eigenvectors Similarity score, , , , This indicates the preset bandwidth parameter; Line node pooling matrix The Middle Line 1 Column elements for: ; Represents the original node Corresponding transmission line index, express The number of original node indexes in the middle; Region Node Pooling Matrix The Middle Line 1 Column elements for: ; Represents a region node The set of all corresponding original node indices, express The number of original node indexes in the middle. Represents the original node The corresponding region node index; Binary relation matrix The Middle Line 1 Column elements for: ; This indicates that there is at least one original node. Belongs to line node And belongs to the regional node , Represents line nodes Not a regional node ; Normalized similarity weight matrix The Middle Line 1 Column elements for: ; ; Represents the original node With category prototype nodes Unnormalized similarity between them Represents the original node With category prototype nodes The unnormalized similarity between them.
6. The power system anomaly detection method according to claim 1, characterized in that, If the mode mask matrix The Middle Line 1 Column elements With structural mask matrix The Middle Line 1 Column elements If both are 1, then the feature extraction network performs the feature extraction process on... and Attention is calculated on the feature vectors of the corresponding nodes; otherwise, the feature extraction network does not perform attention calculations during the feature extraction process. and Attention is calculated using the feature vectors of the corresponding nodes.
7. The power system anomaly detection method according to claim 1, characterized in that, The step of obtaining the anomaly detection results of each power device in the power system based on the data features includes: Calculate the sample in the data features Output embedding vectors and category prototype nodes Euclidean distance between eigenvectors ; The Euclidean distance is expressed by the following formula. Convert to probability distribution : ; The predicted fault category is calculated using the following formula. : ; The basic anomaly is calculated using the following formula. : ; The risk-weighted outlier score is calculated using the following formula. : ; Based on the risk-weighted anomaly scores corresponding to all samples, the anomaly detection results for each power device in the power system are obtained; the anomaly detection result for each power device is used to indicate whether the power device has an anomaly; wherein, if the risk-weighted anomaly score... If the value is greater than or equal to a preset threshold, then the sample is determined. The corresponding power equipment is malfunctioning; in, This indicates the preset temperature parameter. Indicating the sample in the data features Output embedding vectors and category prototype nodes The Euclidean distance between the eigenvectors, express The probability distribution at time, This indicates that the fault category is no fault. This represents the risk weight hyperparameter. This represents the risk weight corresponding to the predicted fault category. , Indicates the predicted fault category The corresponding number of repairs, Indicates the predicted fault category The corresponding average recovery time.
Citation Information
Patent Citations
Power equipment fault detection method and system based on deep learning network
CN120632701A
Processing method and device based on double mask sparse attention, equipment and medium
CN120874910A