Substation equipment abnormal operation diagnosis method based on data analysis
By utilizing the feature extraction and LSTM model of voltage, current, and temperature series in the diagnosis of abnormalities in substation equipment, combined with random forest classifiers and cluster analysis, the problems of low efficiency and misjudgment in traditional methods are solved, and higher diagnostic accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510600245.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional methods for diagnosing abnormalities in substation equipment are inefficient and inaccurate, and deep learning models are prone to misjudgment, making it difficult to meet the requirements of modern power systems for equipment operation reliability.
By collecting voltage, current, and temperature series with abnormal duration, extracting manual features and time series features, combining LSTM and random forest classifiers, using HDBSCAN clustering and Mahalanobis distance to judge abnormalities, and optimizing the hyperparameters of the random forest classifier, the diagnostic accuracy is improved.
It improves the accuracy and reliability of abnormal diagnosis of substation equipment, reduces misjudgments and missed judgments, and ensures safe and stable operation of equipment.
Smart Images

Figure CN120671001A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power systems, and in particular to a method for diagnosing abnormal operation of substation equipment based on data analysis. Background Art
[0002] In the power system, the stable operation of substation equipment is crucial to the reliability and safety of the entire power grid. Once abnormal operation occurs, it may cause large-scale power outages, bringing serious impacts on social production and daily life.
[0003] Traditional methods for diagnosing substation equipment anomalies rely primarily on manual inspections and empirical judgment. This approach is inefficient and inaccurate, making it difficult to meet the reliability requirements of modern power systems. With the advancement of data analysis technology, using data analysis to diagnose substation equipment anomalies has become a research hotspot. Many existing methods for detecting substation equipment anomalies rely on historical data fed into deep learning models. Leveraging the powerful feature learning capabilities of deep learning, these methods not only significantly improve diagnostic efficiency but also enable rapid processing of large amounts of data. However, misjudgments still occur in practical applications because deep learning models have extremely high data quality requirements. If there are issues with the historical data, the model can easily learn incorrect features. Summary of the Invention
[0004] In response to the shortcomings of the existing technology, the present invention provides a method for diagnosing abnormal operation of substation equipment based on data analysis, which solves the problems of low efficiency and poor accuracy of traditional abnormal diagnosis and the easy misjudgment of existing deep learning models.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: A method for diagnosing abnormal operation of substation equipment based on data analysis, comprising:
[0006] S1. For each anomaly, collect the duration data of each anomaly to obtain a duration sequence. From this sequence, find the minimum and maximum values to form the duration interval corresponding to the anomaly. After several similar processes, obtain the duration intervals of different anomalies. These duration intervals are divided from small to large according to the lower limit of the interval. Determine the anomalies corresponding to different duration intervals. Perform HDBSCAN clustering on the historical data of each anomaly and calculate the centroid of these clusters.
[0007] S2. When a real-time abnormality occurs in substation equipment, extract the feature vectors of the voltage sequence, current sequence, and temperature sequence, as well as the duration of the abnormality. Substitute the feature vectors into the random forest classifier to find abnormalities with a probability greater than a threshold. Place the abnormality in the first judgment abnormality set. Locate the abnormality corresponding to the duration interval based on the abnormality duration. Calculate the Mahalanobis distance between the feature vector of the real-time abnormality and the centroid of each located abnormal cluster. Mark clusters with a Mahalanobis distance less than a specified threshold distance. Calculate the ratio of the number of marked clusters to the total number of clusters. If the ratio is greater than the specified ratio threshold, place the abnormality in the second judgment abnormality set. Intersect the first judgment abnormality set with the second judgment abnormality set to obtain the final diagnosed abnormality.
[0008] As a further solution of the present invention, S0 is also included before S1. S0 obtains historical data of abnormalities in substation equipment, calculates manually extracted features of voltage series, current series, and temperature series within the duration of the abnormality, and uses LSTM to extract time series features of the three. The manually extracted features and time series features are spliced together to obtain input features, and the input features are substituted into the random forest model for training to obtain a trained random forest classifier.
[0009] As a further solution of the present invention, the manually extracted features specifically include mean, variance, energy distribution entropy, and fractal box dimension.
[0010] As a further solution of the present invention, the specific steps of using LSTM to extract the three sequence time series features are as follows:
[0011] Collect the voltage series V, current series I, and temperature series T within each abnormal duration, check for missing values and abnormal values in each series, and perform Z-score standardization to obtain V', I', and T';
[0012] Determine the size of the sliding window w, and divide the normalized sequence into sliding windows with a step size of 1. For each window, combine the data of the voltage, current, and temperature series at the same time step to form the kth sample Xk: k=1,2,...,n-w+1, where n is the sequence length of voltage, current, and temperature;
[0013] The samples are divided into training set and validation set in a ratio of 7:3, and substituted into the constructed LSTM to obtain the trained LSTM model;
[0014] For real-time or historical time series data, the sample is input into the trained LSTM model, and the hidden state hw of the last time step output by the LSTM model is extracted. The dimension of the hidden state is d, and hw is the extracted time series feature.
[0015] As a further solution of the present invention, the value range of the size w of the sliding window is [0.25n, 0.3n].
[0016] As a further solution of the present invention, the number of layers in the constructed LSTM is set to 2, the number of units in each layer is 64, the optimizer is selected as Adam, the learning rate is 0.001, the batch size is 32, and the training cycle is 50.
[0017] As a further solution of the present invention, a root system optimization algorithm is used to optimize the number of trees, the maximum depth, and the minimum number of sample splits in the random forest.
[0018] As a further solution of the present invention, the duration interval can be divided into a shared time interval and a non-shared time interval. The shared time interval is the intersection interval obtained by calculating the duration intervals corresponding to two or more different anomalies, and the non-shared time interval is the corresponding duration interval containing only one anomaly.
[0019] As a further solution of the present invention, the duration interval is located according to the duration of the real-time anomaly. If the duration interval is a shared time interval, the prior probability of each anomaly is calculated according to the formula P(anomaly i|t∈[a,b])=the number of times anomaly i occurs in the interval / the total number of times anomaly i occurs in each interval, and the anomalies are sorted from large to small according to the prior probability. The feature vectors of the anomalies and the real-time anomalies are selected one by one from the front of the sort and compared until the first matching anomaly is found and stopped; if the duration is not in the shared time interval but in the duration interval, and there is only one anomaly corresponding to the duration interval but it does not meet the requirements, then it is transferred to the shared time interval containing the anomaly for judgment. If it still does not meet the requirements, the distance between the duration and other duration intervals is calculated, and the duration interval is located according to the distance from small to large until a matching anomaly is found, where a and b are the upper and lower limits of the shared time interval.
[0020] As a further solution of the present invention, the intersection of the first abnormality judgment set and the second abnormality judgment set is calculated. If the intersection is an empty set, the Mahalanobis distance between the real-time abnormality feature vector and the abnormal clusters after the sorting needs to be calculated until a satisfactory one is found.
[0021] The present invention provides a method for diagnosing abnormal operation of substation equipment based on data analysis, which has the following beneficial effects compared with the prior art:
[0022] (1) The present invention extracts the time series features of voltage, current, and temperature during the abnormal time period of the equipment through LSTM, combines the time series features with the manually extracted features, comprehensively reflects the operating status of the equipment, and optimizes the random forest classifier through the swarm intelligence optimization algorithm, thereby improving the accuracy of the diagnostic model;
[0023] (2) The present invention uses cluster analysis of abnormal duration intervals and historical data to judge abnormalities in shared time intervals and special circumstances, and combines the abnormalities obtained by the random forest classifier to obtain the overall abnormality, effectively improving the accuracy and reliability of abnormal operation diagnosis of substation equipment and reducing misjudgments and missed judgments. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a flow chart of the steps of the present invention. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0026] like Figure 1 The present invention provides a method for diagnosing abnormal operation of substation equipment based on data analysis, comprising:
[0027] Obtain historical data on abnormalities in substation equipment, mainly including voltage, current, and temperature series during the duration of the abnormality, and calculate the manually extracted features and time series features of the three series respectively;
[0028] The reasons for selecting voltage, current, and temperature as key variables for determining abnormalities in transformer equipment are as follows: Voltage, which is either too high or too low, can damage the equipment. For example, long-term overvoltage conditions can accelerate the aging of the equipment's insulation materials and shorten its service life, while undervoltage can cause the equipment to fail to start normally and operate unstably. Furthermore, voltage fluctuations can affect the stability of the power system, potentially triggering cascading failures and causing large-scale power outages. Current reflects the equipment's load and circuit operating conditions. When an abnormality such as a short circuit or overload occurs, the current increases dramatically. For example, when a transformer short-circuits, the short-circuit current can far exceed the rated current. The abnormal increase in current can be detected by monitoring temperature. Equipment generates heat during operation. If heat dissipation is poor or abnormal, the temperature rises. Excessive temperature can accelerate equipment aging, reduce performance, and even cause safety accidents such as fires. For example, transformer winding overheating may be caused by a short circuit, overload, or abnormal cooling system. By monitoring temperature changes, the type of equipment abnormality can be detected in a timely manner, allowing maintenance and repairs to be performed in advance to ensure safe operation of the equipment.
[0029] Most existing data analysis methods focus on the moment when voltage, current and temperature suddenly change, while the present invention extracts the voltage, current and temperature sequence during the duration of abnormality of transformer equipment;
[0030] When an anomaly occurs, the operating status of the equipment changes dynamically. Data at a single point in time can only reflect the instantaneous situation and cannot reflect the evolution of the status. However, voltage, current, and temperature sequences over a period of time can show the changing trends of these physical quantities over time. For example, in the early stages of a transformer winding short circuit, the current may only rise slowly. By observing the current sequence, this gradual change can be detected in a timely manner. Relying only on data at certain points in time may miss the subtle changes in the early stages of the anomaly, resulting in delayed detection.
[0031] Substation equipment anomalies may be periodic or intermittent. For example, some equipment may experience periodic temperature increases after a period of operation due to poor heat dissipation. Sequential data over a period of time can clearly capture this periodic change and accurately identify anomalies. However, data at certain points in time may be normal, making intermittent anomalies undetectable.
[0032] Sequential data within the duration of an abnormality contains more information and can provide richer features for the diagnostic model. The correlation of data at multiple time points helps the model learn more complex patterns from a temporal perspective, thereby improving diagnostic accuracy.
[0033] The manually extracted features specifically include mean, variance, energy distribution entropy, and fractal box dimension;
[0034] From the perspective of mean values, under normal operating conditions, the mean values of the voltage, current, and temperature series of substation equipment usually remain within a relatively stable range. If the mean values show a significant deviation, it may indicate that there is an abnormality in the equipment. Different abnormalities also correspond to different mean values. For example, a continuous increase in the mean value of the voltage series may be due to grid voltage fluctuations, abnormal transformer voltage regulation, etc.; an increase in the mean value of the current series may indicate that the equipment is overloaded; and an increase in the mean value of the temperature series may indicate a fault such as a short circuit or poor contact inside the equipment, resulting in increased heat generation.
[0035] From the perspective of variance, under normal circumstances, the variance of voltage, current, and temperature series should fluctuate within a certain range. If the variance is large, it means that the fluctuation of the series is aggravated, and there may be abnormal conditions. For example, a large variance of the voltage series may be due to the presence of a large number of nonlinear loads in the power grid, resulting in increased voltage fluctuations; an increase in the variance of the current series may be due to unstable loads on the equipment or intermittent short circuit anomalies; an increase in the variance of the temperature series may be due to problems with the equipment's cooling system, resulting in severe temperature fluctuations.
[0036] From the perspective of energy distribution entropy, during normal operation, the energy distribution of voltage, current, and temperature series should be relatively uniform. If the energy distribution entropy value decreases, it means that the energy distribution is more concentrated, and there may be abnormal conditions. For example, in the voltage series, if the energy is concentrated in certain time periods, it may be due to harmonic interference in the power grid, resulting in uneven voltage energy distribution. In the temperature series, if the energy is concentrated in a local area of the equipment, it may indicate overheating in that area.
[0037] From the perspective of fractal box dimension, during normal operation, voltage, current, and temperature sequences usually have a certain regularity, and their fractal box dimension is relatively stable. If the fractal box dimension increases, it means that the sequence waveform has become more complex and the fluctuations have become more irregular, which may indicate an abnormality. For example, in the voltage sequence, if the fractal box dimension increases, it may be due to a large amount of random interference in the power grid, resulting in abnormal voltage fluctuations. In the current sequence, an increase in the fractal box dimension may indicate a change in the load characteristics of the equipment or the presence of electromagnetic interference.
[0038] The methods for calculating the mean and variance of voltage, current, and temperature series are not described here in detail. Instead, the specific methods for calculating the energy distribution entropy and fractal box dimension are introduced:
[0039] Energy distribution entropy, for the sequence {xi,i∈[1,n]}, calculate the energy of each point according to the formula Ei=xi^2, obtain the normalized energy distribution according to the formula pi=Ei / sum(Ei), and calculate the energy distribution entropy according to the formula EDE=-sum(pi*ln(pi));
[0040] The fractal box dimension is to treat the sequence as a point (i, xi) on a two-dimensional plane, and cover it with a grid with a side length of ε. The minimum number of grids N(ε) that covers all points is counted. For multiple ε values, a straight line is fitted between ln(N(ε)) and ln(1 / ε). The slope is the box dimension. The specific formula is:
[0041] The specific approach of using LSTM to extract the three sequence time series features is as follows:
[0042] (1) Multi-dimensional time series data preprocessing: Collect the voltage sequence V = [v1, v2, ..., vn], the current sequence I = [i 1, i2, ..., in], and the temperature sequence T = [tt1, tt2, ..., ttn] within each abnormal duration, and check whether there are missing values or abnormal values in the data. For missing values, linear interpolation, mean filling and other methods can be used to process them. For abnormal values, they can be identified and corrected by setting thresholds or using statistical methods. In order to eliminate the influence of the dimensions of different physical quantities, each sequence is Z-score standardized to obtain the standardized sequences V', I', and T';
[0043] (2) Sliding window division and sample construction: According to the actual problem and data characteristics, the size of the sliding window w is determined, and the normalized sequence is divided into sliding windows with a step size of 1. For each window, the data of the voltage, current and temperature series at the same time step are combined to form a sample. Then the kth sample X k for: k=1,2,...,n-w+1;
[0044] (3) Constructing an LSTM feature extraction network: The input layer receives time series data of dimension (w, 3), that is, the shape of each sample is (w, 3). The input layer passes the data to the subsequent LSTM layer for processing. According to the complexity of the problem and the characteristics of the data, the appropriate number of LSTM units num is selected. At each time step, the LSTM unit calculates a new hidden state based on the current input and the hidden state of the previous time step. The output of the sequence is set, that is, the LSTM layer only outputs the hidden state of the last time step to obtain the global features of the entire sequence;
[0045] (4) Model training and parameter optimization: The constructed samples are divided into a training set and a validation set, usually in a ratio of 7:3. The training set is used for model training, and the validation set is used to evaluate the performance of the model and adjust the hyperparameters. The training is performed in a supervised manner, using the Adam optimizer, with a learning rate of 0.001, a batch size of 32, and a training cycle of 50.
[0046] (5) Time series feature extraction and output: For real-time or historical time series data, preprocess and sample construction are performed according to steps (1)-(2), and then the sample is input into the trained LSTM network. The LSTM network outputs the hidden state hw of the last time step. The dimension of the hidden state is d, and hw is the extracted time series feature.
[0047] The hidden state hw of the last time step of LSTM is concatenated with the manually extracted features to obtain the final input features, whose dimension is (d+12);
[0048] Perform the above feature extraction on all types of abnormal data, and then divide them into training set, test set, and validation set in a ratio of 7:2:1. The training set is used to train the random forest model and learn the mapping relationship between features and anomalies. The test set is used for hyperparameter optimization. The validation set is used for the final evaluation of the trained model.
[0049] The root optimization algorithm is used to optimize the hyperparameters of the random forest, namely the number of trees, maximum depth, and minimum number of sample splits. Each iteration generates a set of hyperparameters. The model is trained using the training set, and the fitness is calculated on the test set to select the optimal parameters. The test set is only used to evaluate the hyperparameters and does not participate in the update of the model training weights. It only serves as an "evaluation metric calculator."
[0050] Apply the optimal hyperparameters obtained by the root optimization algorithm to the random forest and retrain the model using the training set to ensure that all 70% of the data is used;
[0051] Use the trained random forest classifier to predict the validation set and calculate its accuracy. Since the validation set has never been involved in training or hyperparameter adjustment, its evaluation results reflect the model's generalization ability on completely unknown data and are a true reflection of the model's final performance.
[0052] For each type of anomaly, extensive data on the duration of each anomaly is collected. Taking the transformer oil temperature overheating anomaly as an example, the duration data of 10 transformer oil temperature overheating anomalies that occurred at the substation in the past year are collected. The duration sequence A = [2, 3, 5, 4, 3, 6, 4, 5, 3, 4] (unit: hours) is obtained. The minimum value min(A) = 2 and the maximum value max(A) = 6 in this sequence are found. [2, 6] is used as the duration interval of the transformer oil temperature overheating anomaly.
[0053] Different types of abnormalities are processed similarly to obtain duration intervals of different abnormalities. These intervals are then combined to determine the abnormalities corresponding to different duration intervals. For example, if the duration interval of abnormality 1 is [1, 4], such as transformer partial discharge, the duration interval of abnormality 2 is [2, 5], such as busbar overheating, and the duration interval of abnormality 3 is [6, 7], such as capacitor failure, then we can get: [1, 2] belongs to abnormality 1, [2, 4] belongs to abnormality 1 + abnormality 2, (4, 5] belongs to abnormality 2, and [6, 7] belongs to abnormality 3.
[0054] When multiple anomalies share a time interval, such as [2,4] above, which corresponds to anomaly one and anomaly two, the prior probability of each anomaly in the interval needs to be calculated based on historical data. For example, in the shared time interval [2,4], by collecting historical data, it is found that anomaly one occurred 30 times and anomaly two occurred 20 times, while the total number of times anomaly one occurred in each interval was 100 times, and the total number of times anomaly two occurred in each interval was 80 times. According to the formula P(anomaly i|t∈[a,b])=the number of times anomaly i occurred in the interval / the total number of times anomaly i occurred in each interval, it can be calculated that P(anomaly one|t∈[2,4])=30 / 100=0.3 and P(anomaly two|t∈[2,4])=20 / 80=0.25. The anomalies are sorted according to the size of P, that is, the probability of anomaly one is greater than that of anomaly two. Among them, the interval [a,b] represents the shared time interval;
[0055] Perform cluster analysis on each type of anomaly historical data. The required input features for clustering are the feature vectors that are a combination of the previously manually extracted features and the time series features extracted by LSTM. For example, cluster analysis is performed on the historical data of anomaly 1. The HDBSCAN clustering algorithm is used to group similar historical data of the same type into different clusters. The centroids of these clusters are calculated for use in subsequent anomaly assessment.
[0056] When a real-time anomaly occurs in substation equipment, the real-time voltage, current, and temperature series, as well as the duration of the anomaly, are collected. These three series are then processed through the same steps as the previous feature extraction, namely, calculating manually extracted features and extracting time series features using LSTM. The feature vectors are then substituted into the trained random forest classifier to obtain the probability of each anomaly occurring. Anomalies with probabilities greater than the specified threshold are placed in the first judgment anomaly set.
[0057] Locate the duration interval based on the anomaly duration. If the duration interval is a shared time interval, calculate the prior probability of the anomaly therein, sort the anomalies from large to small based on the prior probability, calculate the Mahalanobis distance between the feature vector of the real-time anomaly and each cluster of sorted anomalies, mark the clusters whose Mahalanobis distance is less than the specified threshold, and calculate the ratio of the number of marked clusters to the total number of clusters. If the ratio is greater than the specified ratio threshold, place the anomaly in the second judgment anomaly set. If not, continue to select the following types of anomalies according to the sorting and compare them with the real-time anomaly until a qualified anomaly is found;
[0058] The intersection of the first judgment anomaly set and the second judgment anomaly set is obtained, and the obtained anomaly is the diagnosed anomaly;
[0059] If the duration is not in the shared time interval but in the duration interval, and the duration interval corresponds to only one anomaly but does not meet the requirements, for example, the Mahalanobis distance calculation result does not meet the specified distance threshold, then the judgment is transferred to the shared time interval containing the anomaly. If it still does not meet the requirements, the distance between the duration and other duration intervals is calculated, and the duration intervals are located in ascending order of distance until a meeting the requirements is found;
[0060] Assume that a substation device experiences an abnormality at a certain moment and lasts for 3.5 hours. The real-time voltage sequence is [492, 495, 493, ...], the current sequence is [480, 485, 483, ...], and the temperature sequence is [33, 34, 35, ...];
[0061] We manually extract features from these three sequences and use LSTM to extract time series features to obtain feature vectors. We then substitute these feature vectors into the trained random forest classifier to determine the probability of each anomaly. We find that the probability of anomaly 1 is 0.7, and the probability of anomaly 2 is 0.4. With a probability threshold of 0.6, anomaly 1, whose probability is greater than the threshold, is placed in the first judgment anomaly set.
[0062] According to the duration of 3.5 hours, the duration interval is located. Since 3.5 hours is in [2,5], this interval corresponds to two anomalies, anomaly 1 and anomaly 2, which are shared time intervals. According to the anomaly sorting, the prior probability of anomaly 1 in this interval has been calculated to be greater than that of anomaly 2. First, the Mahalanobis distance between the real-time anomaly feature vector and each cluster of anomaly 1 is calculated. Assuming that anomaly 1 has 3 clusters, the Mahalanobis distance between the real-time anomaly feature vector and these 3 clusters is calculated. It is found that the Mahalanobis distance of 2 of the clusters is less than the specified distance threshold. The proportion of the total clusters is calculated to be 2 / 3≈0.67, which is greater than the specified proportion. The specified proportion is set to 0.5, so anomaly 1 meets the requirements and is placed in the second judgment anomaly set. At this time, there is no need to continue calculating the Mahalanobis distance between the real-time anomaly feature vector and each cluster of anomaly 2, because the ones ranked higher meet the requirements.
[0063] The intersection of the first and second abnormality judgment sets is calculated. Since there is only abnormality 1 in the first abnormality judgment set and abnormality 1 is also in the second abnormality judgment set, the final abnormality diagnosis is abnormality 1. If the intersection is an empty set, it is necessary to continue calculating the Mahalanobis distance between the real-time abnormality feature vector and the abnormal clusters after the sorting until a qualified one is found.
[0064] Some of the data in the above formulas are dimensionless and numerically calculated. Meanwhile, the contents not described in detail in this specification belong to the prior art known to those skilled in the art.
[0065] The above embodiments are only used to illustrate the technical method of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A method for diagnosing abnormal operation of substation equipment based on data analysis, characterized in that: include: S1. For each anomaly, collect the duration data of each anomaly to obtain a duration sequence. From this sequence, find the minimum and maximum values to form the duration interval corresponding to the anomaly. After several similar processes, obtain the duration intervals of different anomalies. These duration intervals are divided from small to large according to the lower limit of the interval to obtain the anomalies corresponding to different duration intervals. HDBSCAN clustering is performed on the historical data of each anomaly to calculate the centroid of these clusters. S2. When a real-time abnormality occurs in substation equipment, extract the feature vectors of the voltage sequence, current sequence, and temperature sequence, as well as the duration of the abnormality. Substitute the feature vectors into the random forest classifier to find abnormalities with a probability greater than a threshold. Place the abnormality in the first judgment abnormality set. Locate the abnormality corresponding to the duration interval based on the abnormality duration. Calculate the Mahalanobis distance between the feature vector of the real-time abnormality and the centroid of each located abnormal cluster. Mark clusters with a Mahalanobis distance less than a specified threshold distance. Calculate the ratio of the number of marked clusters to the total number of clusters. If the ratio is greater than the specified ratio threshold, place the abnormality in the second judgment abnormality set. Intersect the first judgment abnormality set with the second judgment abnormality set to obtain the final diagnosed abnormality.
2. The method for diagnosing abnormal operation of substation equipment based on data analysis according to claim 1, characterized in that: S0 is also included before S1. S0 obtains historical data of abnormalities in substation equipment, calculates manually extracted features of the voltage series, current series, and temperature series within the duration of the abnormality, and uses LSTM to extract the time series features of the three. The manually extracted features and the time series features are concatenated to obtain input features, and the input features are substituted into the random forest model for training to obtain a trained random forest classifier.
3. The method for diagnosing abnormal operation of substation equipment based on data analysis according to claim 2, characterized in that: The manually extracted features specifically include mean, variance, energy distribution entropy, and fractal box dimension.
4. The method for diagnosing abnormal operation of substation equipment based on data analysis according to claim 2, characterized in that: The specific steps of using LSTM to extract the three sequence time series features are: Collect the voltage series V, current series I, and temperature series T within each abnormal duration, check for missing values and abnormal values in each series, and perform Z-score standardization to obtain V', I', and T'; Determine the size of the sliding window w, and divide the normalized sequence into sliding windows with a step size of 1. For each window, combine the data of the voltage, current, and temperature series at the same time step to form the kth sample Xk: k=1,2,...,n-w+1, where n is the sequence length of voltage, current, and temperature; The samples are divided into training set and validation set in a ratio of 7:3, and substituted into the constructed LSTM to obtain the trained LSTM model; For real-time or historical time series data, the sample is input into the trained LSTM model, and the hidden state hw of the last time step output by the LSTM model is extracted. The dimension of the hidden state is d, and hw is the extracted time series feature.
5. The method for diagnosing abnormal operation of substation equipment based on data analysis according to claim 4 is characterized in that: The value range of the sliding window size w is [0.25n, 0.3n].
6. The method for diagnosing abnormal operation of substation equipment based on data analysis according to claim 4, characterized in that: In the constructed LSTM, the number of layers is set to 2, the number of units in each layer is set to 64, the optimizer is selected as Adam, the learning rate is 0.001, the batch size is 32, and the training cycle is 50.
7. The method for diagnosing abnormal operation of substation equipment based on data analysis according to claim 2, characterized in that: The root system optimization algorithm is used to optimize the number of trees, maximum depth, and minimum number of sample splits in the random forest.
8. The method for diagnosing abnormal operation of substation equipment based on data analysis according to claim 1, characterized in that: The duration interval can be divided into a shared time interval and a non-shared time interval. The shared time interval is the intersection interval obtained by calculating the duration intervals corresponding to two or more different anomalies, and the non-shared time interval is the duration interval corresponding to only one anomaly.
9. The method for diagnosing abnormal operation of substation equipment based on data analysis according to claim 1, characterized in that: The duration interval is located according to the real-time anomaly duration. If the duration interval is a shared time interval, the prior probability of each anomaly is calculated according to the formula P(anomaly i|t∈[a,b]) = the number of times anomaly i occurs in the interval / the total number of times anomaly i occurs in each interval. The anomalies are sorted from large to small according to the prior probability, and the feature vectors of the anomalies at the front of the sort are compared with the real-time anomaly one by one until the first matching anomaly is found and stopped. If the duration is not in the shared time interval but in the duration interval, and there is only one anomaly corresponding to the duration interval but it does not meet the requirements, then it is transferred to the shared time interval containing the anomaly for judgment. If it still does not meet the requirements, the distance between the duration and other duration intervals is calculated, and the duration interval is located according to the distance from small to large until a matching anomaly is found, where a and b are the upper and lower limits of the shared time interval.
10. The method for diagnosing abnormal operation of substation equipment based on data analysis according to claim 1, characterized in that: The intersection of the first abnormality set and the second abnormality set is calculated. If the intersection is an empty set, the Mahalanobis distance between the real-time abnormality feature vector and the abnormal clusters after the sorting needs to be calculated until a satisfactory one is found.