Extra-high voltage equipment abnormal data aggregation method
By performing entropy calculations and random forest classification on current sensors, voltage transformers, and partial discharge detection units of UHV equipment, and combining graph neural network analysis, a continuous chain structure was constructed. This solved the problem of unified management of abnormal data and fault early warning for UHV equipment, thereby improving the safety and stability of the power system.
Patent Information
- Application Number
- CN202511321620.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2026-01-09
Smart Images

Figure CN121302157A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of anomaly aggregation technology, and in particular to a method for aggregating anomaly data from ultra-high voltage equipment. Background Technology
[0002] The field of anomaly aggregation technology aims to solve the problems of scattered, heterogeneous, and redundant anomaly information in power systems or other large and complex systems by centrally collecting, integrating, classifying, and modeling multiple anomaly data generated during operation. It forms a unified anomaly data expression framework, enabling the monitoring platform to quickly identify anomaly characteristics, locate potential fault sources, and reveal the correlation between anomalies. On this basis, it improves fault early warning capabilities, operational risk assessment capabilities, and maintenance decision support capabilities.
[0003] The purpose of an ultra-high voltage (UHV) equipment anomaly data aggregation method is to establish a unified aggregation processing mechanism for multiple anomaly data generated during the operation of UHV transmission equipment. This mechanism aims to eliminate the dispersion of anomaly data in terms of source, format, and indicators, and to achieve centralized management and correlation analysis of anomaly information. It can form a complete anomaly data view in a large-scale power grid environment, thereby improving the accuracy of equipment operation status monitoring, enabling early warning of potential faults, and enhancing the safety and stability of power system operation.
[0004] While existing technologies can achieve centralized collection and classification of abnormal data, in practical applications, the lack of uniformity in numerical scale and indicator weights can easily lead to comparison barriers between abnormal information from different sources, resulting in biased analysis results. Furthermore, existing technologies rely on static modeling, which cannot characterize the evolution of risks over time. This makes it easy for early risk signals to be submerged in the overall data, reducing the timeliness of early warnings. At the level of abnormal information expression, existing technologies can only output isolated indicator anomalies, failing to form a continuous chain structure. This limits the ability to trace complex abnormal events, resulting in insufficient accuracy of identification results and delayed risk prediction when monitoring platforms face large-scale and complex power grid environments, thus affecting the effectiveness of maintenance decisions. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing a method for aggregating abnormal data from ultra-high voltage equipment.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: a method for aggregating abnormal data from ultra-high voltage equipment, comprising the following steps: S1: Based on the current sensor, voltage transformer, and partial discharge detection unit of UHV equipment, calculate the entropy value of the measurement sequence, compare the difference values and assign weights, judge and correct the change amplitude of the weight set item by item, and generate an abnormal index weight matrix. S2: Based on the anomaly index weight matrix, input the weighted values of current sensor, voltage transformer and partial discharge detection unit, perform node splitting and classification, use random forest, set the depth and sampling ratio, cross-divide and compare the training and verification sets, perform current sensor numerical constraint judgment, and obtain equipment anomaly risk classification output; S3: Based on the equipment anomaly risk classification output, the probability result range is mapped to 0 to 100, the extreme values of each item in the historical risk sequence are extracted and threshold judgment and comparison are performed to generate the equipment risk index sequence; S4: Based on the equipment risk index sequence, a graph neural network is used to calculate the weights of the current sensor, voltage transformer and partial discharge detection unit, compare the results and input them into the graph nodes and perform neighbor sampling and updating, transfer and aggregate the feature values and judge the association probability threshold to obtain the index contribution and association distribution map. S5: Based on the contribution and correlation distribution map of the aforementioned indicators, multi-hop path extraction is performed on the nodes of the current sensor, voltage transformer, and partial discharge detection unit. Sequence combination and path sorting are then performed to construct a continuous chain and output abnormal chain structure information.
[0007] As a further embodiment of the present invention, the abnormal indicator weight matrix includes current offset weight, voltage fluctuation amplitude weight, and partial discharge pulse count weight; the equipment abnormal risk classification output includes risk level label, risk probability value, and classification boundary conditions; the equipment risk index sequence includes risk value range, risk threshold point, and risk trend curve; the indicator contribution and correlation distribution map includes indicator proportion relationship, indicator correlation probability, and indicator neighbor link; and the abnormal chain structure information includes multi-hop propagation path, event chain sorting, and event chain combination.
[0008] As a further aspect of the present invention, the specific steps for generating the anomaly index weight matrix are as follows: Based on the current sensor, voltage transformer, and partial discharge detection unit of UHV equipment, multiple measurement sequence values are extracted and the entropy value is calculated for each value. By comparing the differences between the values and determining their magnitude, values with larger differences are assigned higher weights to generate a set of difference weights. Based on the set of differential weights, the change range of each weight value in the set is detected. By comparing with a set threshold and correcting the part exceeding the threshold, the corrected values are recombined to generate an abnormal indicator weight matrix.
[0009] As a further aspect of the present invention, the specific steps for generating the equipment anomaly risk classification output are as follows: Based on the aforementioned anomaly index weight matrix, the weighted values of the input current sensor are combined with the weighted values of the voltage transformer and the partial discharge detection unit. The values are then used to divide the nodes into different branches at the node level, and the differences between the branches are recorded to generate the node splitting results. Based on the node splitting results, a random forest is used, with a limited tree depth and a set sampling ratio, to split the dataset into a training part and a validation part. The values of the two parts are cross-compared at the classification boundary to generate cross-comparison results. Based on the cross-comparison results, the risk output of the detection current sensor value is increased and a non-decreasing constraint is set. The classification boundary conditions are compared and the risk level range is divided to obtain the equipment abnormality risk classification output.
[0010] As a further aspect of the present invention, the random forest takes sensor-weighted features and labels as input. Each tree obtains a subset by self-sampling from the sample set, randomly extracts k candidate features from all features, enumerates the split threshold of each candidate feature and calculates the split index, selects the split with the best index to establish a child node, and repeats the extraction of candidate features and threshold search until a set depth, a set number of leaf node samples and a purity threshold are reached and then stopped. T trees are independently trained according to the same process. During prediction, the sample falls to the leaf node of each tree and outputs the category and category probability. The classification result is aggregated by majority voting and probability mean. At the same time, the generalization error is estimated by using out-of-bag samples and the hyperparameters are adjusted accordingly. Class weights are set to handle class imbalance, and non-reduction constraints such as current deviation and risk non-reduction can be applied to specified features. Splits that violate the constraints are rejected and monotonicity correction is performed during the aggregation stage.
[0011] As a further aspect of the present invention, the specific steps for generating the equipment risk index sequence are as follows: Based on the equipment anomaly risk classification output, the probability values in the classification results are extracted and the proportions are converted item by item. The values are restricted to the range of 0 to 100 through a mapping function and linear distribution adjustment is made to generate a mapping probability sequence. Based on the mapped probability sequence, the extreme values of the historical risk sequence are calculated. By comparing the magnitudes of each value and combining them with a threshold, the judgment results are reorganized into a sequence to generate the equipment risk index sequence.
[0012] As a further aspect of the present invention, the method of comparing the magnitudes of each value and making a judgment in combination with a threshold is as follows: when processing the mapped probability sequence, each value in the sequence is first compared with other values in the same sequence, the size relationship is calculated and the difference magnitude is marked, and then a threshold is set as the judgment boundary. The threshold is derived from the statistical distribution range of the historical risk sequence and is determined by the upper and lower quantiles of the risk value. Values exceeding the upper quantile or falling below the lower quantile are marked as abnormal, while values within the threshold range are marked as normal.
[0013] As a further aspect of the present invention, the specific steps for generating the above-mentioned product are as follows: Based on the equipment risk index sequence, a graph neural network is used to extract the weights of the current sensor and voltage transformer and compare them with the weights of the partial discharge detection unit. The results are sorted by the ratio difference and the values are recorded to generate the ratio comparison results. Based on the ratio comparison results, input values into the graph node set and sample adjacent nodes, update the feature values of each node and merge them into the target node features, accumulate feature values, and generate node aggregation results; Based on the node aggregation results, the association probability between nodes is calculated and compared with a set threshold. Node relationships exceeding the threshold are identified and recorded in the link table to obtain the indicator contribution and association distribution map.
[0014] As a further aspect of the present invention, the graph neural network first aggregates the features of adjacent nodes during the message passing phase and assigns different weighting coefficients according to the weight and type of the edges. The aggregated information is then concatenated and summed with the node's own features to obtain an updated representation. In the multi-layer iterative propagation, the range of neighbors is expanded layer by layer, gradually capturing the dependencies on multi-hop paths. Under the action of the attention mechanism, the contributions of different neighbors are distinguished by learnable weights and assigned weights during aggregation. In the subsequent output phase, risk scores can be predicted for each node, the association probability between node pairs can be calculated, and strong and weak connections can be judged using thresholds. At the same time, the propagation links are extracted to form an index contribution and association distribution map.
[0015] As a further aspect of the present invention, the specific steps for generating the abnormal chain structure information are as follows: Based on the contribution and correlation distribution map of the aforementioned indicators, multi-hop paths of the node sets of current sensors, voltage transformers, and partial discharge detection units are extracted. Path sequences are formed by sequentially connecting the nodes, and the path sequences are combined to generate node path combination results. Based on the node path combination results, the path set order is compared and priority is set item by item. High-frequency node chains are combined with adjacent paths and merged into a continuous structure, and abnormal chain structure information is output.
[0016] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, by calculating the entropy value of the measurement sequence and constructing a differential weight system, multiple abnormal indicators are quantitatively mapped in a unified matrix, ensuring the consistency of numerical comparison of data from different sources, and eliminating bias interference by correcting each item to form a stable weight expression. In this invention, node splitting and classification are performed using random forests. By combining depth and sampling ratio settings with cross-division of training and validation sets, risk classification achieves hierarchical accuracy and enhances the reliability of results under constraints. Risk results are mapped to intervals and compared with extreme value thresholds to form a dynamic risk index sequence, enabling potential fault risks to have the ability to track time series and characterize stage evolution. In this invention, a graph neural network is introduced to perform neighbor sampling and feature aggregation on the correlation strength between indicators, realizing the transformation from single-point anomalies to coupling patterns between multiple indicators. This can reveal the causal relationships and propagation paths between different types of anomalies. Through multi-hop path extraction and sequence combination, a continuous anomaly chain structure is formed, transforming scattered anomaly information into a chain expression with evolutionary logic, thereby improving the ability to identify potential faults in complex systems in advance and ensure overall operational safety. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the main steps of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0019] Example 1 Please see Figure 1 This invention provides a technical solution: a method for aggregating abnormal data from ultra-high voltage (UHV) equipment, comprising the following steps: S1: Based on the current sensor, voltage transformer, and partial discharge detection unit of UHV equipment, calculate the entropy value of the measurement sequence, compare the difference values and assign weights, judge and correct the change amplitude of the weight set item by item, and generate an abnormal index weight matrix. S2: Based on the anomaly index weight matrix, the weighted values of current sensor, voltage transformer and partial discharge detection unit are input, node splitting and classification are performed, random forest is used, depth and sampling ratio are set, training and verification sets are cross-partitioned and compared, current sensor numerical constraint judgment is performed, and equipment anomaly risk classification output is obtained. S3: Based on the output of equipment anomaly risk classification, the probability result range is mapped to 0 to 100, the extreme values of each item in the historical risk sequence are extracted and threshold judgment and comparison are performed to generate the equipment risk index sequence. S4: Based on the equipment risk index sequence, a graph neural network is used to calculate the weights of current sensors, voltage transformers, and partial discharge detection units. The comparison results are input into graph nodes and neighbor sampling and updating are performed. Feature values are passed and aggregated, and association probability threshold judgment is performed to obtain the index contribution and association distribution map. S5: Based on the contribution and correlation distribution map of the indicators, multi-hop path extraction is performed on the nodes of current sensor, voltage transformer and partial discharge detection unit, sequence combination and path sorting are performed to construct a continuous chain and output abnormal chain structure information.
[0020] The abnormal indicator weight matrix includes current offset weight, voltage fluctuation amplitude weight, and partial discharge pulse count weight. The equipment abnormal risk classification output includes risk level label, risk probability value, and classification boundary conditions. The equipment risk index sequence includes risk value range, risk threshold point, and risk trend curve. The indicator contribution and correlation distribution map includes indicator proportion relationship, indicator correlation probability, and indicator neighbor link. The abnormal chain structure information includes multi-hop propagation path, event chain sorting, and event chain combination.
[0021] The specific steps for generating the anomaly indicator weight matrix are as follows: Based on the current sensor, voltage transformer, and partial discharge detection unit of UHV equipment, multiple measurement sequence values are extracted and the entropy value is calculated for each value. By comparing the differences between the values and determining their magnitude, values with larger differences are assigned higher weights to generate a set of difference weights. Based on the differential weight set, the change range of each weight value in the set is detected. By comparing with a set threshold and correcting the part that exceeds the threshold, the corrected values are recombined to generate an abnormal indicator weight matrix. Based on current sensors, voltage transformers, and partial discharge detection units of UHV equipment, multiple measurement sequence values are extracted and entropy values are calculated for each item. The Shannon entropy calculation method is used to perform probability distribution statistics on each sequence, with the calculation precision set to six decimal places and the input sequence length set to 1024 points. A sliding window method is used for processing, with a window size of 256 points and a step size of 64 points. The entropy value is calculated for each item in each window and the results are recorded. All window results are then weighted and averaged to obtain the overall entropy value. Difference comparison between sequences is performed, and the Euclidean distance calculation method is used to compare the entropy value differences between any two sequences item by item. The difference threshold is set to 0.05. If the difference exceeds the threshold, it is judged as a large difference value, and a higher weight is assigned to the large difference value. The linear normalization method is used to unify the weights to the range of 0.1 to 1.0. All processed results are recombined to generate a difference weight set. Based on the differential weight set, the variation range of each weight value in the set is detected. The sequence is smoothed using the moving average method with a window size of 10. The difference between each weight and the moving average is calculated, and the magnitude is compared with a threshold of 0.15. The part exceeding the threshold is corrected by using the Z-score standardization method to adjust the weights that are out of range to within the range of the mean plus or minus two standard deviations. After the correction is completed, all the corrected weights are recombined, and the combination result is arranged into a two-dimensional matrix using the matrix reconstruction method. The number of rows in the matrix is set to the number of sensor types, and the number of columns is set to the total number of measurement sequences corresponding to each type of sensor. Finally, an abnormal index weight matrix is generated.
[0022] The specific steps for generating equipment anomaly risk classification output are as follows: Based on the anomaly index weight matrix, the weighted values of the input current sensor are merged with the weighted values of the voltage transformer and the partial discharge detection unit. The values are then used to divide the nodes into different branches at the node level, and the differences between the branches are recorded to generate the node splitting results. Based on the node splitting results, a random forest is used, with a limited tree depth and a set sampling ratio, to split the dataset into a training part and a validation part. The values of the two parts are cross-compared at the classification boundary to generate cross-comparison results. Based on the cross-comparison results, the risk output of the current sensor value is detected when it increases and a non-decreasing constraint is set. The classification boundary conditions are compared and the risk level range is divided to obtain the equipment abnormality risk classification output. Based on the anomaly index weight matrix, the weighted values of the input current sensor are merged with the weighted values of the voltage transformer and the partial discharge detection unit. A hierarchical segmentation method is used to divide the nodes into branches. Specifically, the input weighted values are divided into equal intervals within the range of 0.0 to 1.0, with an interval step size of 0.05. Each weighted value is categorized according to the interval and a corresponding branch is generated in the node. The values in each branch are rearranged and stored in ascending order. The difference judgment method is used to compare the numerical differences between any two branches item by item. The threshold is set to 0.1. Values exceeding the threshold are recorded as difference marks. After merging all branches and difference marks, the node splitting result is output. Based on the node splitting results, a random forest method was used for processing. The number of trees in the model was set to 200, the tree depth was limited to 20 layers, and the sampling ratio was set to 0.7. The input dataset was split into training and validation parts according to the ratio, with the training part accounting for 80% and the validation part accounting for 20%. Cross-comparison was performed on the training and validation parts at the classification boundary. Specifically, the boundary values generated by the training part were compared with the boundary values generated by the validation part item by item. When the difference exceeded 0.05, the cross-bias was recorded, and the matching and bias status of all comparison items were output, and the final cross-comparison results were generated. Based on the cross-comparison results, the risk output of the detection current sensor value is increased. A monotonic constraint method is used to set a non-decreasing constraint condition. Specifically, when the value of the next value in the input sequence is greater than the value of the previous value, the risk output remains no less than the previous output. All classification boundary conditions are compared, and the threshold intervals 0 to 30, 31 to 60, and 61 to 100 are used as the level division intervals. Each output is recorded in the corresponding interval, and finally the equipment abnormality risk classification output is obtained.
[0023] Random forest takes sensor-weighted features and labels as input. Each tree obtains a subset by self-sampling from the sample set, randomly selects k candidate features from all features, enumerates the split threshold of each candidate feature and calculates the split index, selects the split with the best index to establish a child node, and repeats the extraction of candidate features and threshold search until the set depth, set number of leaf node samples and purity threshold are reached and stopped. T trees are trained independently according to the same process. During prediction, the sample falls to the leaf node of each tree and outputs the class and class probability. The classification result is aggregated by majority voting and probability mean. At the same time, the generalization error is estimated by using out-of-bag samples and the hyperparameters are adjusted accordingly. Class weights are set to handle class imbalance, and non-reduction constraints such as current deviation and risk non-reduction can be applied to specified features. This is achieved by rejecting splits that violate the constraints and performing monotonicity correction during the aggregation stage. Random forest, according to the formula: ; in: This represents the cross-comparison result value of the aggregated outlier data. This represents the total number of decision trees constructed. Indicates the first The weight coefficient of each tree, Indicates the first The classification function of the tree, This represents a sample of operating data for ultra-high voltage (UHV) equipment. Indicates the defined tree depth. This indicates the set sampling ratio. This represents the correction coefficient for node splitting imbalance. This represents the sensitivity coefficient between the training and validation parts at the classification boundary; Execution process: First, the equipment risk index sequence is used as the input sample. Introduce a random forest classifier, and then classify the samples according to the sampling ratio. The input samples are divided into training and validation parts, and multiple decision trees are generated sequentially on the training part, with a total number of trees. The depth of each tree layer is limited during the construction process. To prevent overfitting, an imbalance correction coefficient is then introduced during node splitting. To adjust for differences in the distribution of class samples, a boundary sensitivity coefficient is introduced during the cross-comparison between the training and validation parts to amplify the influence of critical point samples. Subsequently, the classification accuracy of each tree on the validation part is calculated and normalized to obtain the corresponding weight coefficients. Then output the classification function for each tree. Multiply by weighting factor The results are summed and averaged across all trees to obtain the overall prediction output, which is the result of cross-comparison of the equipment risk index sequence in the training and validation parts. This is used to establish the proportional comparison value and support the aggregation and judgment of abnormal data.
[0024] The specific steps for generating the equipment risk index sequence are as follows: Based on the output of equipment anomaly risk classification, the probability values in the classification results are extracted and the proportions are converted item by item. The values are restricted to the range of 0 to 100 through a mapping function and linear distribution adjustment is made to generate a mapping probability sequence. Based on the mapped probability sequence, the extreme values of the historical risk sequence are calculated. By comparing the magnitudes of each value and combining them with the threshold, the judgment results are reorganized into a sequence to generate the equipment risk index sequence. Based on the output of equipment anomaly risk classification, the probability values in the classification results are extracted and the proportions are converted item by item. The minimum-maximum normalization method is used to standardize the probability values. All probability values are scaled according to the minimum and maximum values in the formula, and the value range is limited to 0 to 1. After scaling, the value range is expanded to 0 to 100 using a linear mapping function. A linear transformation is performed on each value and two decimal places are retained. The distribution of the critical value range is adjusted using a linear interpolation method, and the number of interpolation points is set to 5. All adjusted values are arranged and recombine in sequence to finally generate a mapped probability sequence. Based on the mapping probability sequence, the extreme value detection method is used to process the historical risk sequence. Each value in the sequence is compared one by one, with a comparison window size of 50 records. The maximum and minimum values in the window are extracted and recorded. Then, each value is compared with the window extreme value. When the value is greater than the upper threshold, it is marked as a high-risk point, and when the value is less than the lower threshold, it is marked as a low-risk point. The threshold range is set to an upper threshold of 90 and a lower threshold of 10. All marking results are recombined in chronological order. The sequence reconstruction method is used to convert the marked value sequence into a standardized index form. The index length is equal to the length of the original input sequence, and finally, the equipment risk index sequence is generated.
[0025] By comparing the magnitudes of each value and combining them with a threshold, when processing the mapped probability sequence, each value in the sequence is first compared with other values in the same sequence, the size relationship is calculated and the difference is marked, and then a threshold is set as the judgment boundary. The threshold is derived from the statistical distribution range of the historical risk sequence and is determined by the upper and lower quantiles of the risk value. Values exceeding the upper quantile or falling below the lower quantile are marked as abnormal, while values within the threshold range are marked as normal.
[0026] The specific steps for generating the product are as follows: Based on the equipment risk index sequence, a graph neural network is used to extract the weights of current sensors and voltage transformers and compare them with the weights of partial discharge detection units. The results are sorted by the ratio difference and the values are recorded to generate the ratio comparison results. Based on the proportional comparison results, input values into the graph node set and sample adjacent nodes, update the feature values of each node and merge them into the target node features, accumulate feature values, and generate node aggregation results; Based on the node aggregation results, the probability of association between nodes is calculated and compared with a set threshold. Node relationships exceeding the threshold are identified and recorded in the link table to obtain the indicator contribution and association distribution map. Based on the equipment risk index sequence, a graph neural network method is used to extract the weights of the current sensor and voltage transformer and compare them with the weights of the partial discharge detection unit. The weight values of the three types are sorted by proportional difference. Specifically, all input weights are first normalized to the range of 0 to 1, and the sorting method is set to ascending order. Each weight difference value is recorded item by item, and the sorting result is retained to four decimal places. A proportional difference label is generated for each comparison item and stored in an array. All results are reorganized into an ordered set, and finally, the proportional comparison result is generated. Based on the proportional comparison results, the obtained values are input into the graph node set. The neighbor sampling method is used to process each node. The number of sampled neighbors is set to 10. For each node, 10 neighbor nodes are selected and the corresponding feature values are extracted. Then, the node update method is used to perform a weighted operation with the weight parameter set to 0.25. The features of the neighbor nodes and the features of the target node are weighted and added together. The updated features are stored in the target node structure. The operation is repeated for all target nodes and the sum is accumulated in turn to finally generate the node aggregation result. Based on the node aggregation results, a probability calculation method is used to calculate the association strength between any two nodes. The accuracy of the probability value is set to three decimal places. Each probability value is compared with a threshold of 0.6. All node relationships exceeding the threshold are marked and recorded in the link table. The link table is stored in the form of an adjacency matrix, with the row and column dimensions corresponding to the number of nodes and the matrix elements storing the marked status. Finally, the indicator contribution and association distribution map are obtained.
[0027] The graph neural network first aggregates the features of neighboring nodes in the message passing stage and assigns different weighting coefficients according to the weight and type of the edges. The aggregated information is then concatenated and summed with the node's own features to obtain an updated representation. In the multi-layer iterative propagation, the range of neighbors is expanded layer by layer, gradually capturing the dependencies on multi-hop paths. Under the action of the attention mechanism, the contributions of different neighbors are distinguished by learnable weights and are assigned weights during aggregation. In the output stage, risk scores can be predicted for each node, the association probability between node pairs can be calculated, and the strong and weak connections can be judged using thresholds. At the same time, the propagation links are extracted to form an index contribution and association distribution map. Graph neural networks, according to the formula: ; in: This represents the feature representation of node v at layer k. Represents a non-linear activation function. Let v represent the set of neighbors of node v. This represents the attention weights of node v when it aggregates its neighbor u at layer k. This represents the sensor reliability coefficient of neighbor node u. This represents the time decay coefficient of the neighbor node u. This represents the metering calibration gain of neighbor node u. This represents the normalization constant between neighboring node u and node v. This represents the weight matrix of the k-th layer. This represents the feature representation of neighbor node u at layer (k-1). Execution process: First, the equipment risk index sequence is input into the graph structure construction module, including nodes such as current sensor nodes, voltage transformer nodes, and partial discharge detection unit nodes. Then, the neighbor set is determined based on the adjacency relationship. And calculate the normalization constant. Next, a sensor reliability coefficient is assigned to each neighboring node. The error was calculated by verifying historical cycles to reduce the impact of unstable sensors, and the time decay coefficient was also calculated. Based on the time interval between the sample and the current time, the contribution of outdated data is controlled using an exponential decay method, and then the measurement calibration gain is determined. The sensor output is scaled to match the reference signal using least-squares fitting, and then the attention weights are calculated using an attention mechanism. This is used to characterize the importance of different types of nodes in structural dependency and localized states, and then the features of neighboring nodes are used. via weight matrix After mapping, multiply by the aggregation coefficient and sum within the neighborhood. The result is then activated by the activation function. Transformation to form updated node features After multi-level iteration, the final representation of each node can be obtained and the weights of the current sensor and voltage transformer can be read. Then, the weights of the partial discharge detection unit are compared with the ratio difference. The results are sorted according to the degree of difference and recorded. Finally, the ratio comparison results are generated for abnormal data aggregation.
[0028] The specific steps for generating the anomaly chain structure information are as follows: Based on the index contribution and correlation distribution map, the multi-hop path of the node set of current sensor, voltage transformer and partial discharge detection unit is extracted, the path sequence is formed by sequential connection between nodes, and the path sequence is combined to generate the node path combination result. Based on the node path combination results, the path set order is compared and priority is set item by item. High-frequency node chains are combined with adjacent paths and merged into a continuous structure, and abnormal chain structure information is output. Based on the contribution and correlation distribution map of the indicators, multi-hop paths of the current sensor node set, voltage transformer node set, and partial discharge detection unit node set are extracted. The depth-first search method is used to expand the path of the node set. The search depth is set to 5. During the traversal, the next node directly connected to the current node is selected each time. The visited nodes are recorded in the path stack. Each path is numbered according to the access order. The numbering format is P plus the sequence number, which is incremented from P1. Duplicate nodes are skipped and marked in the exclusion table. After all paths are traversed, they are recombined into a path sequence according to the node order. The path sequence is stored in the form of an array and the node path combination result is generated. Based on the node path combination results, a priority sorting method is used to compare the path set in sequence. The priority is set according to the frequency of the node in the path set. Nodes with a frequency higher than 10 times are marked as high-frequency nodes. The chain containing the high-frequency node is extracted and compared with the adjacent path. The adjacent path determination rule is that the end node of the path is the same as the first node of another path. For the paths that meet the determination conditions, a merging operation is performed. The merging method is to sequentially concatenate the nodes of the two paths and remove duplicate nodes. The merged result is stored in a continuous structure array. The operation is repeated for all paths until no further merging is possible. Finally, the abnormal chain structure information is output.
[0029] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. An ultra-high voltage equipment abnormal data aggregation method, characterized in that, The method comprises the following steps: S1: based on the ultra-high voltage equipment current sensor, voltage transformer, partial discharge detection unit, the measurement sequence entropy value is calculated, the difference value is compared and weighted, the weight set change amplitude is judged and corrected item by item, and the abnormal index weight matrix is generated; S2: based on the abnormal index weight matrix, the current sensor, voltage transformer, partial discharge detection unit weighted value is input, node splitting classification is carried out, random forest is adopted, depth and sampling ratio are set, training and verification set cross division comparison is carried out, current sensor numerical constraint judgment is carried out, and equipment abnormal risk classification output is obtained; S3: based on the equipment abnormal risk classification output, the probability result interval is mapped to 0 to 100, the extreme value of the historical risk sequence is extracted item by item and the threshold value is judged and compared, and the equipment risk index sequence is generated; S4: based on the equipment risk index sequence, the graph neural network is adopted, the current sensor, voltage transformer, partial discharge detection unit weight is calculated in proportion, the comparison result is input into the graph node, and the neighbor sampling and updating are carried out, the characteristic value is transmitted, aggregated and associated probability threshold value is judged, and the index contribution and associated distribution graph is obtained; S5: based on the index contribution and associated distribution graph, the current sensor, voltage transformer, partial discharge detection unit node is extracted by multi-hop path, sequence combination and path sorting are carried out, continuous chain is constructed, and abnormal chain structure information is output.
2. The method of claim 1, wherein, The abnormal index weight matrix includes current offset weight, voltage fluctuation amplitude weight and partial discharge pulse number weight, the equipment abnormal risk classification output includes risk level label, risk probability value and classification boundary condition, the equipment risk index sequence includes risk numerical interval, risk threshold point and risk trend curve, the index contribution and associated distribution graph includes index proportion relationship, index associated probability and index neighbor link, and the abnormal chain structure information includes multi-hop propagation path, event chain sorting and event chain combination.
3. The method of claim 1, wherein the method further comprises: The specific steps for generating the abnormal index weight matrix are: Based on the ultra-high voltage equipment current sensor, voltage transformer and partial discharge detection unit, the multiple measurement sequence values are extracted and the entropy values are calculated item by item, the difference of each value is compared and judged, the values with larger difference are assigned higher weight, and the difference weight set is generated; Based on the difference weight set, the change amplitude of each weight value in the set is detected, the values exceeding the threshold value are compared and modified, the modified values are recombined, and the abnormal index weight matrix is generated.
4. The method of claim 1, wherein, The specific steps for generating the equipment abnormal risk classification output are: Based on the abnormal index weight matrix, the current sensor weighted value is input and combined with the voltage transformer weighted value and the partial discharge detection unit weighted value, the values are divided into different branches at the node level, and the differences between the branches are recorded, and the node splitting result is generated; Based on the node splitting result, random forest is adopted, the tree layer depth is limited and the sampling ratio is set, the data set is divided into training part and verification part, the values of the two parts are compared and contrasted at the classification boundary, and the cross comparison result is generated; Based on the cross-comparison result, the risk output of the current sensor value in the increasing case is detected, the non-decreasing constraint is set, the comparison classification boundary condition is divided, and the risk level interval is divided to obtain the equipment abnormal risk classification output.
5. The method of claim 4, wherein, The random forest takes the sensor weighted features and labels as input, each tree is obtained from the subset of the sample set by bootstrap sampling, k candidate features are randomly selected from all features, the split threshold of each candidate feature is enumerated and the split index is calculated, the split with the optimal index is selected to establish the child node, the candidate feature and threshold search are repeated until the set depth, the set leaf node sample number and the purity threshold are stopped, T trees are independently trained according to the same process, and the class category and class probability are output when the sample falls to the leaf node of each tree during prediction, the classification result is aggregated by majority voting and probability mean, the generalization error is estimated by using out-of-bag samples, the class weight is set to process the class imbalance, and the non-decreasing constraint can be applied to the specified feature, such as current deviation and risk non-decreasing, the split that violates the constraint is rejected, and the monotonicity correction is performed in the aggregation stage.
6. The method of claim 1, wherein, The specific steps for generating the device risk index sequence are: Based on the device abnormal risk classification output, the probability values in the classification result are extracted and converted proportionally, the values are limited in the interval of 0 to 100 by a mapping function, and linear distribution adjustment is made to generate a mapping probability sequence; Based on the mapping probability sequence, the extreme value of the historical risk sequence is calculated, the judgment is made by comparing the values and combining the threshold, the judgment result is reorganized into a sequence to generate a device risk index sequence.
7. The method of claim 6, wherein the method further comprises: When processing the mapping probability sequence, each value in the sequence is compared with other values in the sequence, the size relationship is calculated and the difference amplitude is marked, and then a threshold is set as a judgment boundary. The threshold is derived from the statistical distribution interval of the historical risk sequence, which is determined by the upper and lower quantile points of the risk value. Values exceeding the upper quantile point and below the lower quantile point are marked as abnormal, and values within the threshold interval are marked as normal.
8. The method of claim 1, wherein, The specific steps for generating the specific steps are: Based on the device risk index sequence, the current sensor weight and the voltage transformer weight are extracted and compared with the partial discharge detection unit weight by using the graph neural network, the proportion difference is sorted and the value is recorded to generate a proportion comparison result; Based on the proportion comparison result, the input value is input to the graph node set and the adjacent nodes are sampled, the feature values of each node are updated and merged into the target node feature, the feature values are accumulated, and the node aggregation result is generated; Based on the node aggregation result, the correlation probability between nodes is calculated and compared with the set threshold, the node relationship of the nodes exceeding the threshold is identified and recorded in the link table to obtain the index contribution and correlation distribution map.
9. The method of claim 8, wherein, The graph neural network firstly aggregates features of adjacent nodes in a message passing stage and assigns different weighting coefficients according to weights and types of edges, concatenates and sums the aggregated information with the node's own features to obtain an updated representation, expands the neighbor range layer by layer in multi-layer iterative propagation, gradually captures the dependency relationship on the multi-hop path, and under the action of the attention mechanism, the contributions of different neighbors are distinguished by learnable weights and are assigned weights when aggregated. In the output stage, the risk score of each node can be predicted, the correlation probability between node pairs is calculated, the strong and weak connections are judged by using a threshold, and the index contribution and correlation distribution atlas are extracted to form a link.
10. The method of claim 1, wherein, The specific steps for generating the abnormal chain structure information are as follows: Based on the index contribution and correlation distribution atlas, the multi-hop path of the node set of the current sensor, the voltage transformer and the partial discharge detection unit is extracted, the path sequence is formed through the sequential connection between nodes, and the path sequence is combined to generate a node path combination result; Based on the node path combination result, the path set order is compared and the priority is set item by item, the high-frequency node chain is combined with the adjacent path to form a continuous structure, and the abnormal chain structure information is output.