Substation equipment fault intelligent diagnosis method and system based on multi-source time series data
Patent Information
- Application Number
- CN202611015285.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]本发明针对现有变电站设备故障诊断依赖于固定阈值或人工经验,难以有效捕捉潜伏性故障的早期微弱特征,导致预警延迟与误报率高的技术问题,提供基于多源时序数据的变电站设备故障智能诊断方法及系统
本发明通过融合多源异构传感数据与深度时序分析,建立了一种自适应的设备健康状态动态基准与诊断机制。首先,采用一阶自回归模型与预训练的状态转移网络,能够从多参数时序中提取稳定性特征并量化状态演变的不确定性,克服了固定阈值难以适应动态工况的缺陷。其次,利用无监督聚类在转移概率空间构建正常状态参考簇,无需依赖故障样本即可自主学习设备健康基准,增强了方法的普适性与自适应性。最后,结合实时偏离度计算与动态分层判据规则,能够有效识别早期微弱异常,提高了对潜伏性故障的预警灵敏度与准确性,同时降低了误报率。
Smart Images

Figure CN122548581A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a method and system for intelligent fault diagnosis of substation equipment based on multi-source time-series data. Background Technology
[0002] With the deepening of smart grid construction, substations, as key nodes in the power system, require real-time monitoring of equipment operating status and fault early warning for ensuring the safe and stable operation of the power grid.
[0003] Existing technologies primarily rely on single-parameter alarm mechanisms based on fixed thresholds or comprehensive judgment methods based on human experience. Fixed-threshold methods cannot adapt to the dynamic changes in equipment status under different operating conditions and stages, and struggle to effectively capture the early, subtle characteristics of latent faults caused by long-term operation and slow degradation, resulting in critical information being buried under normal fluctuations. Methods based on human experience heavily depend on expert knowledge, are highly subjective, and are difficult to scale and continuously optimize, exhibiting low analysis efficiency when dealing with massive amounts of multi-source time-series data. Consequently, existing technologies generally suffer from long warning delays and high false alarm rates, failing to meet the urgent need for real-time, accurate, and early fault diagnosis of substation equipment. Summary of the Invention
[0004] This invention addresses the technical problem that existing substation equipment fault diagnosis relies on fixed thresholds or human experience, making it difficult to effectively capture the early and weak characteristics of latent faults, resulting in delayed warnings and high false alarm rates. It provides a method and system for intelligent fault diagnosis of substation equipment based on multi-source time-series data.
[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: In a first aspect, the present invention provides an intelligent fault diagnosis method for substation equipment based on multi-source time-series data, including: By deploying a heterogeneous sensor array on substation equipment, multi-source time-series monitoring data is collected, and the data is cleaned and time-aligned to obtain a pre-processed time-series dataset. For the historical sequence of each monitoring parameter in the time series dataset, a first-order autoregressive model is fitted, and the stability feature sequence of each monitoring parameter in the time series is extracted. Several stable feature sequences obtained by fusion are combined to form a multi-parameter joint feature time series, which is then input into a pre-trained temporal state transition learning model to obtain the transition probability matrix; Based on the transition probability matrix corresponding to historical time series data under normal operating conditions, an unsupervised clustering algorithm is used to establish a reference cluster for the normal state. The deviation between the real-time transition probability matrix and the normal state reference cluster is calculated, and combined with the dynamic hierarchical criterion rule, latent equipment faults are identified and warnings are issued.
[0006] Secondly, the present invention provides an intelligent fault diagnosis system for substation equipment based on multi-source time-series data, comprising: The data acquisition and preprocessing module is used to acquire multi-source time-series monitoring data through a heterogeneous sensor array deployed on substation equipment, and to perform data cleaning and time alignment to obtain a preprocessed time-series dataset. The stability feature extraction module is used to fit a first-order autoregressive model to the historical sequence of each monitoring parameter in the time series dataset and extract the stability feature sequence of each monitoring parameter in the time series. The transition probability modeling module is used to fuse several obtained stability feature sequences to form a multi-parameter joint feature time series, which is then input into a pre-trained time series state transition learning model to obtain the transition probability matrix. The normal state establishment module is used to establish a normal state reference cluster based on the transition probability matrix corresponding to historical time series data under normal operating conditions, using an unsupervised clustering algorithm. The fault diagnosis and early warning module is used to calculate the deviation between the real-time transition probability matrix and the normal state reference cluster, and combined with the dynamic hierarchical criterion rules, to identify latent faults in the equipment and issue early warnings.
[0007] The beneficial effects of this invention are: This invention establishes an adaptive dynamic benchmark and diagnostic mechanism for equipment health status by integrating multi-source heterogeneous sensor data with deep time-series analysis. First, by employing a first-order autoregressive model and a pre-trained state transition network, stability features can be extracted from multi-parameter time series data, and the uncertainty of state evolution can be quantified, overcoming the limitation of fixed thresholds in adapting to dynamic operating conditions. Second, unsupervised clustering is used to construct a normal state reference cluster in the transition probability space, enabling the method to autonomously learn the equipment health benchmark without relying on fault samples, thus enhancing the method's universality and adaptability. Finally, by combining real-time deviation calculation with dynamic hierarchical criterion rules, early subtle anomalies can be effectively identified, improving the sensitivity and accuracy of early warning of latent faults while reducing the false alarm rate. Attached Figure Description
[0008] Figure 1 A flowchart illustrating the intelligent fault diagnosis method for substation equipment based on multi-source time-series data provided by this invention; Figure 2 This is a schematic diagram of the intelligent fault diagnosis system for substation equipment based on multi-source time-series data provided by the present invention.
[0009] In the attached diagram, the components represented by each number are as follows: Data acquisition and preprocessing module 11, stability feature extraction module 12, transition probability modeling module 13, normal state establishment module 14, fault diagnosis and early warning module 15. Detailed Implementation
[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0011] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0012] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.
[0013] Example 1, as Figure 1 As shown, this embodiment of the invention provides an intelligent fault diagnosis method for substation equipment based on multi-source time-series data, including: S10: Collect multi-source time-series monitoring data through a heterogeneous sensor array deployed on substation equipment, and perform data cleaning and time alignment to obtain a preprocessed time-series dataset; The multi-source time-series monitoring data includes partial discharge signals, temperature data, and vibration data. The partial discharge signals are acquired through an ultra-high frequency partial discharge sensor, and the temperature data and vibration data are acquired through a temperature-vibration composite sensor.
[0014] First, multi-source time-series monitoring data is collected using a heterogeneous sensor array deployed on substation equipment. Substation equipment refers to key electrical devices in a power system used for voltage transformation, power distribution, control, and protection, such as power transformers, gas-insulated switchgear, or high-voltage circuit breakers. The heterogeneous sensor array is a collaborative monitoring network composed of sensing units with various physical principles and measurement functions, including electrical quantity sensors, non-electrical quantity sensors, and composite function sensors, such as ultra-high frequency partial discharge sensors and temperature / vibration composite sensors. This heterogeneous sensor array is used to collect multi-source time-series monitoring data, including partial discharge signals, temperature data, and vibration data.
[0015] Specifically, partial discharge signals are an important early sign of insulation degradation inside substation equipment. These partial discharge signals are acquired by an ultra-high frequency partial discharge sensor, which can effectively capture ultra-high frequency electromagnetic wave signals excited by insulation defects in power equipment. It has the characteristics of strong anti-interference ability and high sensitivity.
[0016] Temperature and vibration data reflect the thermal and mechanical states of substation equipment, respectively. A temperature and vibration composite sensor is used for integrated acquisition. This sensor can simultaneously measure the surface or internal temperature and three-dimensional vibration acceleration of key parts of substation equipment. This not only simplifies on-site installation and wiring, but also ensures that the temperature and vibration data have natural spatiotemporal synchronization at the source of acquisition, which facilitates subsequent multi-source feature fusion analysis.
[0017] The raw multi-source time-series monitoring data acquired must undergo data cleaning and time alignment before entering the analysis process due to differences in sensor sampling frequencies, clock deviations, and potential communication delays or data packet loss. Specifically, data cleaning includes removing obvious outliers caused by momentary sensor failures or strong electromagnetic interference, and appropriately interpolating to fill in short-term data gaps. Time alignment uses a unified timestamp benchmark, such as a high-precision network clock protocol, to resample or interpolate data streams from different sensors, ensuring that all monitoring parameters have a strictly corresponding time point at any analysis moment, thus forming a preprocessed, clean, and time-consistent time-series dataset. This time-series dataset serves as the foundation for subsequent feature extraction and model analysis, and can be used to characterize the continuous operational state evolution of substation equipment under multiple physical dimensions.
[0018] S20: For the historical sequence of each monitoring parameter in the time series dataset, fit a first-order autoregressive model and extract the stability feature sequence of each monitoring parameter in the time series. Specifically, for the historical sequence of each monitoring parameter in the time-series dataset, a first-order autoregressive model is fitted, and the stability feature sequence of each monitoring parameter over time is extracted, including: For each monitoring parameter's historical sequence, a sliding window method is used to divide it into multiple subsequences; For each subsequence, a first-order autoregressive model is fitted, and the model parameters are extracted, wherein the model parameters include autoregressive coefficients and residual variance; Calculate the statistical characteristics of each subsequence, wherein the statistical characteristics include the mean and standard deviation; The autoregressive coefficients, residual variances, means, and standard deviations corresponding to each subsequence are arranged in chronological order to form a stability characteristic sequence of each monitoring parameter.
[0019] First, for each monitoring parameter's historical sequence, a sliding window method is used for segmentation. Specifically, the sliding window method slides sequentially along the time axis with a fixed-length time window, sliding a fixed step size each time, thus dividing the continuous historical time series data into multiple overlapping or continuous subsequences. The size of the sliding window is determined comprehensively based on the rate of change of the physical characteristics of the specific monitoring parameter and the real-time requirements of the diagnostic task. For example, a shorter window can be used for rapidly changing vibration signals, while a longer window can be used for slowly changing temperature signals. This segmentation process aims to transform long-term, non-stationary time series data into a series of relatively stationary short-term data blocks to facilitate subsequent local model fitting and feature calculation.
[0020] Secondly, a first-order autoregressive model is fitted to each of the resulting subsequences. The first-order autoregressive model is a classic time series analysis method used to describe the linear dependence between the current observation and the previous observation. Fitting this model yields key model parameters. Specifically, these parameters include the autoregressive coefficients and residual variance. The autoregressive coefficients quantify the strength and direction of the linear relationship between the current and previous values; their absolute values directly reflect the autocorrelation or inertia strength of the sequence within that time window, making them a core indicator of time series dependence. Simultaneously, the residual variance represents the magnitude of the fluctuations that the first-order autoregressive model fails to explain, reflecting the noise level or intensity of random disturbances in the subsequence.
[0021] Furthermore, to further characterize the statistical properties of the subsequences, it is necessary to calculate the statistical features of each subsequence. Specifically, these statistical features include at least the arithmetic mean and standard deviation of all data points within the subsequence. The mean represents the overall level or baseline value of the monitored parameter during that time period, while the standard deviation quantifies the degree of dispersion of the data around the mean, i.e., the amplitude of fluctuation.
[0022] Finally, for each monitoring parameter, the autoregressive coefficients, residual variances, means, and standard deviations calculated for all its subsequences are arranged and combined according to the chronological order of the subsequences on the time axis. Each subsequence corresponds to a feature vector, which contains multiple feature values such as the autoregressive coefficients, residual variances, means, and standard deviations for that sub-window. Concatenating the feature vectors of all subsequences arranged in chronological order constitutes the time-series stability feature sequence of the monitoring parameter.
[0023] By performing the above processing on the historical sequence of each monitoring parameter, the final generated stability feature sequence comprehensively characterizes the parameter's dependence strength, fluctuation amplitude, baseline level and evolution trend in time series. Compared with the original time series dataset, it contains richer and more structured state information, thus providing a clearer data foundation for subsequent multi-source information fusion and other analyses.
[0024] S30: The obtained stable feature sequences are fused to form a multi-parameter joint feature time series, which is then input into a pre-trained time series state transition learning model to obtain the transition probability matrix; First, the aforementioned stability feature sequences are fused. Since each stability feature sequence originates from strictly time-aligned raw data and employs the same sliding window partitioning strategy, they naturally correspond in the time dimension. The fusion process is specifically as follows: at each identical time point or time window index, the stability feature vectors corresponding to all monitoring parameters at that moment are concatenated according to a preset parameter order. For example, first, the feature vector of the partial discharge signal is concatenated; second, the feature vector of the temperature data is concatenated; and finally, the feature vector of the vibration data is concatenated, thus forming a higher-dimensional joint feature vector. Arranging the joint feature vectors at all time points in chronological order constitutes a complete multi-parameter joint feature time series. This multi-parameter joint feature time series fully preserves the dynamic correlation information of each monitoring parameter during the time evolution process, providing a unified data structure for comprehensively describing the overall operating status of the equipment.
[0025] Secondly, the constructed multi-parameter joint feature time series is input into a pre-trained temporal state transition learning model. This temporal state transition learning model is a machine learning-based network architecture, such as a deep neural network employing attention mechanisms or improved causal self-attention mechanisms, specifically designed to capture complex long-term dependencies and dynamic patterns from long-term sequences. Specifically, this temporal state transition learning model takes a temporally arranged multi-dimensional joint feature vector as input and outputs a state transition probability matrix. Each row of this transition probability matrix represents the probability distribution of the substation equipment transitioning to various possible states at the next time step, given the state implied by the currently observed joint feature time series. Therefore, this transition probability matrix can be used to comprehensively and quantitatively characterize the evolution trend and uncertainty of the equipment operating state inferred from the multi-parameter joint feature time series.
[0026] Specifically, the pre-training steps of the temporal state transition learning model include: Based on the historical monitoring records of the substation equipment, a historical sample dataset is collected, wherein the historical sample dataset includes a time series of multi-parameter joint features organized in chronological order; The historical sample dataset is divided into states to obtain sample state labels corresponding to each time point, and a state transition probability matrix is calculated based on the sample state labels to obtain a set of sample state transition probability matrices. Based on machine learning, construct the network structure of a temporal state transition learning model; Using the multi-parameter joint feature time series as input and the corresponding sample state transition probability matrix as supervision signal, the time series state transition learning model is trained under supervision until the model evaluation index converges, thus completing the model pre-training.
[0027] First, based on the historical monitoring records of the substation equipment, a historical sample dataset is collected for training the time-series state transition learning model. Specifically, this historical sample dataset includes a multi-parameter joint feature time series strictly organized in chronological order. The construction method of this multi-parameter joint feature time series is the same as the aforementioned real-time processing flow, that is, it is generated by performing the same cleaning, alignment, feature extraction, and fusion steps on the historical monitoring data, ensuring the consistency of the training data and the future application data in the feature space.
[0028] Secondly, the historical sample dataset is divided into states and labeled. This process aims to assign discrete labels that reflect the operational state category to unlabeled time series data.
[0029] Specifically, the historical sample dataset is divided into states to obtain sample state labels corresponding to each time point, and a state transition probability matrix is calculated based on the sample state labels to obtain a set of sample state transition probability matrices, including: The time series of multi-parameter joint features in the historical sample dataset are divided using the sliding window method to obtain several feature subsequences; For all feature vectors within each feature subsequence, perform cluster analysis and use the cluster center category identifier as the sample state label for the time interval corresponding to the current subsequence; Based on the order of the sample state labels on the time axis, the frequency of transition from the current state label to the next state label is counted, and a state transition frequency matrix is constructed. Normalize each row of the state transition frequency matrix so that the sum of all elements in the same row is 1, and obtain the sample state transition probability matrix corresponding to the current historical time series data. Traverse all historical sample data segments to obtain a set of sample state transition probability matrices for model pre-training.
[0030] First, the multi-parameter joint feature time series already constructed in the historical sample dataset is divided using a sliding window method. Specifically, the sliding window setting can be consistent with the window setting in the aforementioned feature extraction stage or adjusted according to the granularity requirements of the state division. The sliding window cuts the long joint feature time series into a series of continuous, shorter feature subsequences. Each feature subsequence contains a high-dimensional feature vector sequence after the fusion of all monitoring parameters within a specific time interval, representing a segment of the comprehensive operating status of the substation equipment during that time period.
[0031] Secondly, each resulting feature subsequence is processed. Specifically, an unsupervised clustering algorithm is used to perform cluster analysis on the set of all high-dimensional feature vectors arranged chronologically within each subsequence. Examples of suitable clustering algorithms include K-Means clustering or density-based DBSCAN clustering. This clustering analysis aims to discover the main distribution patterns of feature vectors within the feature subsequence.
[0032] For example, taking the K-Means clustering algorithm, given a feature subsequence containing M high-dimensional feature vectors arranged in chronological order, and setting the number of clusters to be divided to K, the algorithm first randomly initializes K cluster centers. Then, iterative optimization is performed: in each iteration, the Euclidean distance from each feature vector in the subsequence to all K cluster centers is calculated, and the vector is assigned to the cluster represented by the nearest cluster center. After all vectors are assigned, for each cluster, the mean of all feature vectors within it is recalculated, and this mean is updated as the new cluster center. The above assignment and update steps are repeated until the positions of the cluster centers no longer change significantly or the preset number of iterations is reached, at which point the algorithm converges. Finally, the M feature vectors in the feature subsequence are divided into K clusters.
[0033] After clustering, the category identifier to which all feature vectors within the entire feature subsequence belong, or the cluster center closest to them, is used as the unified sample state label for the entire time interval corresponding to that feature subsequence. This sample state label is a discrete integer value, representing the dominant pattern category to which the joint operating state of the device's multiple parameters belongs in the feature space within that specific time window. Through the above clustering process, the continuous high-dimensional feature space is mapped onto discrete state labels, assigning each time interval a discrete symbol representing its dominant operating state.
[0034] Furthermore, based on the state label sequence formed by arranging the sample state labels of all feature subsequences in chronological order, state transition statistics are performed. Specifically, the entire state label sequence is traversed, and the frequency of transitioning from any state label i to state label j at the next time step is counted. Based on the statistical frequency of all possible state transition pairs i->j, a state transition frequency matrix is constructed, where the rows and columns of the matrix correspond to different state categories, and the matrix elements F ij It records the cumulative number of times the state transitions from state i to state j.
[0035] Next, the constructed state transition frequency matrix is subjected to probabilistic processing. Each row of the matrix is normalized by dividing the value of each element by the sum of all elements in that row. After this processing, the sum of the elements in each row becomes 1, and each element in the row represents the estimated probability of transitioning to state j in the next moment, given the current state i. Thus, a sample state transition probability matrix based on the statistical analysis of this historical time series data is obtained.
[0036] Finally, the above process is repeated, traversing all data segments or time intervals in the historical data that can be used for training. For each data segment, the entire process from sliding window partitioning, subsequence clustering and labeling to state transition probability matrix calculation is performed, thereby obtaining a set containing multiple sample state transition probability matrices. This set of sample state transition probability matrices comprehensively covers the state transition patterns in different periods and under different operating conditions in the historical data, and together constitutes a high-quality set of supervision signals for training the time-series state transition learning model.
[0037] Furthermore, based on machine learning principles, a temporal state transition learning model is constructed. Machine learning is a technological paradigm that enables computer systems to automatically learn from data and improve their performance without relying on explicit programming instructions. For example, sequence modeling methods based on deep neural networks, such as recurrent neural networks and long short-term memory networks, can be used. The aforementioned multi-parameter joint feature time series is used as the input data for the temporal state transition learning model, and the sample state transition probability matrix calculated for the corresponding time period is used as the training target, i.e., the supervision signal. Supervised training is performed on this temporal state transition learning model, and the training process continues until the model evaluation index converges, resulting in a trained temporal state transition learning model. This temporal state transition learning model possesses the ability to infer a transition probability matrix reflecting the current state evolution trend based on real-time multi-parameter joint feature time series. The convergence condition of the model evaluation index can be set according to the model's performance on an independent validation set. For example, within 10 consecutive training cycles, the average cross-entropy loss value between the model's predicted output and the true state transition probability matrix no longer decreases and the fluctuation range is less than a preset threshold of 0.001.
[0038] For example, since there are complex long-term dependencies and dynamic evolution patterns within the time series of multi-parameter joint features, and deep sequence models have significant advantages in capturing such nonlinear time series relationships, a sequence model based on deep neural networks is chosen as the basic architecture of the time series state transition learning model.
[0039] Specifically, this temporal state transition learning model mainly consists of a feature encoding layer, a sequence modeling layer, and a probability output layer. The feature encoding layer receives standardized multi-parameter joint feature temporal sequences. The sequence modeling layer employs a multi-layered stacked Transformer module with a causal attention mechanism. Each module contains a multi-head self-attention sub-layer and a feedforward neural network sub-layer, with a hidden layer dimension of 256 and 4 attention heads to effectively capture long-range dependencies across time steps. To avoid overfitting, residual connections and layer normalization operations are introduced after each attention sub-layer and feedforward sub-layer, and a Dropout layer with a dropout rate of 0.1 is added before the feedforward network output. The probability output layer consists of a fully connected layer and a Softmax activation function, mapping the final temporal representation output by the sequence modeling layer to a state transition probability matrix with dimensions N×N, where N is the preset total number of device state categories.
[0040] During training, key hyperparameters included an initial learning rate of 0.0005, 200 training epochs, and a batch size of 32. The learning rate was set using a warm-up and cosine annealing scheduling strategy to balance training stability and accuracy. The number of training epochs ensured the model fully learned the dynamic patterns of state transitions. The batch size balanced gradient update stability with computational resource constraints. Specifically, supervised learning was employed. Multi-parameter joint feature time series data were extracted from historical normal operating condition data, following the aforementioned steps, as the input sample set. Simultaneously, the corresponding sample state transition probability matrix was calculated to form a sample label set. The input sample set and the corresponding sample label set were divided into training, validation, and test sets in a 7:2:1 ratio.
[0041] Furthermore, the multi-parameter joint feature time series from the training set is used as the model input, and the corresponding sample state transition probability matrix is used as the supervision signal. The model parameters are iteratively optimized using the backpropagation algorithm and the Adam optimizer. The cross-entropy loss function is used to measure the distribution difference between the model's predicted probability matrix and the true probability matrix, and the training process is monitored using a validation set. Training is terminated when the average cross-entropy loss between the model's predicted output and the true state transition probability matrix no longer decreases and the fluctuation range is less than a preset threshold of 0.001 within 10 consecutive training epochs, resulting in a converged temporal state transition learning model. This temporal state transition learning model can effectively capture the complex mapping relationship between multi-parameter temporal features and state transition patterns, achieving accurate inference from real-time monitoring data to state evolution probabilities.
[0042] Finally, the obtained multi-parameter joint feature time series is input into the trained temporal state transition learning model, which calculates and outputs the transition probability matrix.
[0043] S40: Based on the transition probability matrix corresponding to historical time series data under normal operating conditions, an unsupervised clustering algorithm is used to establish a reference cluster for the normal state. Specifically, based on the transition probability matrix corresponding to historical time-series data under normal operating conditions, an unsupervised clustering algorithm is used to establish a reference cluster for the normal state, including: From the output layer of the time-series state transition learning model, extract the transition probability matrix corresponding to all time windows in the historical period of normal operating conditions of the substation equipment to form a set of normal operating condition transition probability matrices. Each transition probability matrix in the set of normal operating condition transition probability matrices is expanded into a one-dimensional feature vector in row-major order. Using all the expanded one-dimensional feature vectors as input samples, an unsupervised clustering algorithm is used for cluster analysis to obtain multiple cluster centers and the range of samples covered. Calculate the average distance from the feature vectors of all samples within each cluster to the cluster center, and use this distance as the intra-cluster dispersion of the cluster. Based on the intra-cluster dispersion, the clustering results are filtered to remove abnormal clusters, and the feature vectors corresponding to the remaining cluster centers are reconstructed into the form of transition probability matrices to jointly constitute the normal state reference cluster. The average intra-cluster dispersion of all clusters in the normal state reference cluster is calculated as a reliability index.
[0044] First, from the output layer of the already trained temporal state transition learning model, the transition probability matrices corresponding to all time windows within a clearly defined, manually or automatically confirmed period of normal operation of the substation equipment are extracted, forming a set of normal operation transition probability matrices. Each matrix in this set of normal operation transition probability matrices is a probability representation of state evolution inferred by the model based on real-time multi-source data of that time window under the equipment's healthy state, collectively defining the probability distribution space of state transitions during normal operation of the equipment.
[0045] Secondly, the data format of the normal operating condition transition probability matrix set is transformed to be suitable for cluster analysis. Specifically, for each N×N dimensional transition probability matrix in the normal operating condition transition probability matrix set, in row-major order (i.e., starting from the first row, all elements of each row are sequentially taken, concatenated, and expanded into a matrix of length N. 2 The one-dimensional eigenvectors. Through this operation, each matrix representing the state transition probability distribution is transformed into a point in a high-dimensional space.
[0046] Then, using all the expanded one-dimensional feature vectors as input sample points, an unsupervised clustering algorithm is employed for cluster analysis. This cluster analysis aims to discover the natural clustering patterns of state transition probabilities in the feature space under normal operating conditions.
[0047] Specifically, using all the expanded one-dimensional feature vectors as input samples, an unsupervised clustering algorithm is employed for cluster analysis to obtain multiple cluster centers and the range of samples they cover, including: Set an initial range for the number of clusters K, wherein the initial range is determined based on the total number of samples in the normal working condition transition probability matrix set and the desired clustering granularity, wherein K is a positive integer, and the lower limit of the range is 2, and the upper limit is a preset maximum value Kmax; Within the initial value range, different K values are selected sequentially for cluster analysis. For each selected K value, multiple independent clustering processes are performed, and K cluster centers are randomly initialized in each clustering process. Calculate the distance of each one-dimensional feature vector to all current cluster centers, and assign it to the cluster represented by the nearest cluster center; Based on the allocation results, the mean of all one-dimensional feature vectors within each cluster is recalculated and updated as the new cluster center for the corresponding cluster; After each iteration, the total distance sum of all clusters under the current clustering partition is calculated, where the total distance sum is the sum of the distances from samples within each cluster to their respective cluster centers; The result with the minimum total distance sum during multiple independent clustering processes is selected as the candidate clustering result and candidate total distance sum corresponding to the current K value; Analyze the candidate population distances corresponding to different K values and their changing trends as K increases, and select the K value corresponding to the inflection point in the changing trend as the final K value; The set of candidate cluster centers corresponding to the final K value is taken as the optimal clustering result, and the sample range covered by each cluster center is determined based on the optimal clustering result.
[0048] First, we define the initial range for the number of clusters, K. This initial range is determined based on the total number of samples in the normal operating condition transition probability matrix set and the desired clustering granularity, where K is a positive integer. The lower limit of this initial range is 2, indicating that there are at least two different typical state transition modes; the upper limit, Kmax, is a preset value, which needs to be set by comprehensively considering the total number of input samples and the data dimension N of the feature vectors. 2 This also includes prior knowledge of the complexity of the normal operating state patterns of the equipment. For example, for a set with a total sample size of M, the value of Kmax can be specifically set to min(15, √M) to ensure the feasibility and effectiveness of cluster analysis and avoid overfitting or excessive computation due to an excessive number of clusters. This initial value range provides a clear search space for the subsequent systematic exploration of the optimal number of clusters.
[0049] Secondly, within the set initial value range, each possible K value is selected sequentially for cluster analysis. For each selected specific K value, to overcome the sensitivity of clustering algorithms such as K-Means to the selection of initial centroids, multiple independent clustering processes are executed. In each independent clustering process, K cluster centers are randomly initialized.
[0050] Specifically, within each independent clustering process, standard iterative optimization is performed. The Euclidean distance from each one-dimensional feature vector to all K current cluster centers is calculated, and the feature vector is assigned to the cluster represented by its nearest neighbor. Based on this round of assignment, the arithmetic mean of all assigned one-dimensional feature vectors within each cluster is recalculated, and this mean is used to update the new cluster center for that cluster. After completing each round of iterative reassignment and center update, the total distance sum under the current cluster partition is calculated. This total distance sum is defined as the sum of intra-cluster distances across all clusters, i.e., the cumulative sum of Euclidean distances from all sample feature vectors within each cluster to their respective cluster centers. The total distance sum is a key indicator of the density of the current cluster partition; the smaller the value, the closer the sample points are to their respective centers, and the more compact the clustering effect.
[0051] After completing a preset number of iterations or reaching a convergence condition, the independent clustering process ends. The preset maximum number of iterations can be set according to the need for a balance between computational efficiency and accuracy, for example, 300 iterations. The convergence condition is set based on the stability of cluster center movement; for example, convergence is determined when the sum of the changes in the positions of all cluster centers in the current and subsequent iterations, measured by Euclidean distance, is less than a preset minimum threshold of 1e-5. In multiple independent clustering processes performed for the same K value, the result of the clustering process with the smallest final total distance sum is selected as the candidate clustering result for the current K value, and its corresponding total distance sum is the candidate total distance sum for that K value.
[0052] After iterating through all preset K values and obtaining the corresponding candidate total distance sums, the trend of the candidate total distance sums as the K value increases is analyzed. Typically, the total distance sums monotonically decrease as the K value increases. The K value corresponding to the point where a clear inflection point appears in this trend—that is, when the rate of decrease in the total distance sum resulting from increasing the K value significantly slows down—is selected as the final determined number of clusters. The inflection point represents the point where the marginal benefit of further increasing the number of clusters for improving sample compactness has become very low, and the increase in model complexity no longer brings significant improvement in clustering performance. Therefore, the K value corresponding to the inflection point achieves the best balance between representing the main data pattern and controlling model complexity, and is selected as the optimal number of clusters for constructing the normal-state reference cluster.
[0053] Finally, the optimal candidate clustering result corresponding to the final determined K value, i.e., its set of cluster centers, is taken as the optimal clustering result for the entire analysis process. Based on this optimal clustering result, the sample range covered by each cluster center can be clearly determined, that is, the set of all one-dimensional feature vectors assigned to the cluster represented by that cluster center in the entire input sample set. Determining this sample range provides a basis for subsequent calculation of intra-cluster dispersion and screening of anomalous clusters.
[0054] Furthermore, the average Euclidean distance from the feature vectors of all samples within each cluster to the center of its cluster is calculated, and this average value is defined as the intra-cluster dispersion of that cluster. Specifically, the intra-cluster dispersion measures the compactness of the samples within the cluster. The lower the intra-cluster dispersion, the more consistent and stable the state transition patterns within the cluster.
[0055] Based on the calculated intra-cluster dispersion and the number of samples in each cluster, the preliminary clustering results are screened to remove clusters that may represent noise, transient processes, or data anomalies. For example, the criteria for determining an anomalous cluster typically include: the number of samples in the cluster is less than a preset minimum sample size threshold, such as 1% of the total sample size; or the intra-cluster dispersion of the cluster is significantly higher than the average intra-cluster dispersion of all clusters, for example, exceeding twice the standard deviation of the overall average dispersion. Clusters that meet any of the above anomaly criteria are removed. The cluster center feature vectors of the remaining clusters, representing typical and stable normal patterns, are reconstructed into an N×N dimensional transition probability matrix by following the inverse process of expansion. The reconstructed matrix together constitutes a normal state reference cluster for fault diagnosis. This normal state reference cluster is a set of multiple typical normal pattern centers, which can more comprehensively characterize the reasonable fluctuation range that may exist in the normal state of the equipment.
[0056] Finally, the arithmetic mean of the intra-cluster dispersion of all ultimately retained clusters in the normal state reference cluster is calculated, and this mean is used as the confidence index for the entire normal state reference cluster. This confidence index reflects the inherent consistency of historical normal data; the higher the confidence, the more reliable the established normal baseline, and it can be used in subsequent dynamic threshold adjustments as a quantitative basis for assessing the baseline confidence level.
[0057] S50: Calculate the deviation between the real-time transition probability matrix and the normal state reference cluster, and combine it with the dynamic hierarchical criterion rules to identify latent equipment faults and issue warnings.
[0058] Specifically, calculating the deviation between the real-time transition probability matrix and the normal state reference cluster includes: The real-time collected multi-source time-series monitoring data is preprocessed, feature extracted and fused to generate a real-time multi-parameter joint feature time series, which is then input into the time-series state transition learning model to obtain the real-time transition probability matrix. Expand the real-time transition probability matrix in row-major order to obtain a real-time one-dimensional feature vector. Calculate the distance between the real-time one-dimensional feature vector and the feature vector corresponding to each cluster center in the normal state reference cluster, wherein the distance is calculated using Euclidean distance; The minimum value among several calculated distances is selected as the deviation between the real-time transition probability matrix and the normal state reference cluster.
[0059] First, a standardized process, identical to that used for historical data processing, is executed on the real-time multi-source time-series monitoring data. This process includes data cleaning and time alignment preprocessing steps, feature extraction based on a sliding window and first-order autoregressive model, and multi-parameter feature vector fusion, ultimately generating a real-time multi-parameter joint feature time series that matches the data structure used during model training. This real-time multi-parameter joint feature time series is then input into a pre-trained time-series state transition learning model, which outputs a real-time transition probability matrix representing the current state evolution trend of the device.
[0060] Secondly, to achieve distance measurement with each cluster center within the normal reference cluster, the real-time generated transition probability matrix needs to be converted into a feature vector representation with the same data structure as the reference cluster centers. Specifically, the N×N dimensional real-time transition probability matrix is sequentially concatenated according to its row-major order, from the first row to the Nth row, to expand and reconstruct a matrix of length N. 2 The one-dimensional feature vector. This one-dimensional feature vector is the real-time one-dimensional feature vector used for subsequent distance calculation.
[0061] Then, the distance between the aforementioned real-time one-dimensional feature vector and the feature vector corresponding to each cluster center in the normal state reference cluster is calculated. Preferably, Euclidean distance is a commonly used metric for measuring the overall difference between two high-dimensional vectors; the smaller the distance value, the more similar the state transition probability distributions represented by the two vectors are.
[0062] Finally, the minimum value among all calculated Euclidean distance values is selected as the final quantified deviation between the real-time transition probability matrix and the normal state reference cluster. Specifically, the normal state reference cluster consists of multiple cluster centers representing typical normal operating modes, and its overall structure defines the reasonable fluctuation range allowed by the state transition probability distribution when the equipment is in a healthy state. If the state corresponding to the real-time monitoring data is sufficiently similar to any typical normal mode in the reference cluster, the current equipment operating state can be determined to be within an acceptable normal fluctuation range. Therefore, the degree of deviation between the real-time state and the entire normal state reference cluster is determined by the distance between the real-time state and the center of the most similar typical normal mode within the reference cluster. The determined deviation is a non-negative scalar value, the magnitude of which directly quantifies the difference between the current state of the equipment inferred from the real-time monitoring data and the state baseline established under historical normal operating conditions. The larger the deviation value, the more significant the deviation of the current state transition mode from the typical normal mode, and the higher the probability of equipment anomalies or potential faults.
[0063] Furthermore, by combining dynamic hierarchical judgment rules, latent equipment faults can be identified and warnings issued.
[0064] The dynamic hierarchical criterion rules include: A baseline threshold matrix containing different warning levels is set, wherein each warning level corresponds to a set of threshold parameters, including a deviation threshold and a determination duration threshold. The mean of the absolute values of the autoregressive coefficients in the stability feature sequence is obtained in real time and used as a stability index. Obtain the current load rate of the substation equipment as an indicator of operating load. Based on the stability index and the operating load index, a basic dynamic adjustment coefficient is calculated, wherein the basic dynamic adjustment coefficient is positively correlated with the stability index and negatively correlated with the operating load index. By combining the credibility index of the normal state reference cluster, the basic dynamic adjustment coefficient is corrected to obtain the comprehensive dynamic adjustment coefficient; Using the comprehensive dynamic adjustment coefficient, the deviation thresholds of each level in the benchmark threshold matrix are numerically corrected, and the corresponding judgment duration thresholds are duration corrected to obtain a dynamic hierarchical criterion that takes effect in real time. When the deviation continuously exceeds the deviation threshold of a certain level in the dynamic hierarchical criterion, and the duration reaches the corrected duration threshold corresponding to the current level, an early warning for the corresponding level is triggered.
[0065] First, a baseline threshold matrix is established, defining different severity levels of warnings and their corresponding static judgment criteria. Each warning level is associated with a set of threshold parameters, including a deviation threshold and a judgment duration threshold. Specifically, the baseline threshold matrix is set based on statistical analysis of historical normal and abnormal data, performance parameters provided by equipment manufacturers, and relevant provisions in power grid operation and maintenance procedures. For example, the deviation threshold can be set to a specific high percentile of the deviation distribution under historical normal operating conditions, such as 95%; the judgment duration threshold is mainly set based on the physical time constant required for the development of different types of faults and the time required for maintenance personnel to arrive on-site and take initial measures. For example, the duration threshold for a level 3 warning can be set to 15 minutes, level 2 to 30 minutes, and level 1 to 45 minutes.
[0066] Secondly, from the currently calculated stability characteristic sequence, the mean of the absolute values of the autoregressive coefficients of all monitoring parameters is extracted. This value reflects the current overall autocorrelation strength of each parameter's time series and is defined as a stability index. Simultaneously, the real-time load rate of the equipment is obtained through the substation monitoring system as an indicator of operating load.
[0067] Furthermore, the basic dynamic adjustment coefficient is calculated based on the two real-time indicators mentioned above. Specifically, the ratio of the current stability indicator to its baseline stability indicator value is calculated as the relative value of the stability indicator, and the ratio of the current load indicator to its baseline load indicator value is calculated as the relative value of the load indicator. The baseline values of the stability indicator and the baseline values of the load indicator are reference values characterizing the operating status of the equipment under typical normal operating conditions. These values are obtained based on the statistical characteristics of long-term historical monitoring data during normal operating periods. For example, the baseline value of the stability indicator can be the long-term average or median of the stability indicator during historical normal periods, and the baseline value of the load indicator can be the rated load rate of the equipment or the average load rate during historical normal operation.
[0068] Specifically, the basic dynamic adjustment coefficient equals the relative value of the stability index divided by the relative value of the load index. When the equipment's operational stability is relatively high, it indicates that the time-series dependence of each monitoring parameter is strong, the system inertia is large, and short-term fluctuations have little impact on the overall state. Therefore, the tolerable deviation threshold can be relaxed accordingly. When the equipment is under high load, its electrical and mechanical stresses increase, state fluctuations may intensify, and even minor abnormalities may develop rapidly. Therefore, the early warning sensitivity needs to be improved accordingly, which is achieved by lowering the deviation threshold and the duration threshold. This basic dynamic adjustment coefficient dynamically reflects the degree of tightness or looseness of the current operating condition relative to the typical normal operating condition.
[0069] Furthermore, a reliability index of the normal state reference cluster is introduced to correct the basic dynamic adjustment coefficient, thus obtaining the comprehensive dynamic adjustment coefficient. The reliability index needs to be normalized first, converting it into a value between 0 and 1. Preferably, the normalized reliability is calculated as: (maximum intra-cluster dispersion - average intra-cluster dispersion) / (maximum intra-cluster dispersion - minimum intra-cluster dispersion). The formula for calculating the comprehensive dynamic adjustment coefficient is: Comprehensive dynamic adjustment coefficient = Basic dynamic adjustment coefficient × [1 + β] (1-Normalized Credibility). Wherein, β is the preset credibility correction weight coefficient, the value of which is greater than or equal to 0, and the specific value is determined based on the analysis of historical data quality and the prudence requirements of operation and maintenance strategies. The larger the value of β, the stronger the influence of the credibility of the normal state reference cluster on the threshold adjustment; when β=0, the comprehensive dynamic adjustment coefficient degenerates into the basic dynamic adjustment coefficient.
[0070] The calculated comprehensive dynamic adjustment coefficient is a composite adjustment factor that comprehensively considers the tightness of real-time operating conditions and the reliability of historical normal benchmarks. It represents the overall strictness calibration of the equipment status anomaly judgment at a given moment. This comprehensive dynamic adjustment coefficient can be used to make real-time, adaptive numerical and duration corrections to the preset benchmark warning thresholds, so that the final dynamic hierarchical criterion can simultaneously respond to changes in external operating conditions and the level of confidence of the internal model, achieving more accurate and robust fault warnings.
[0071] Furthermore, the calculated comprehensive dynamic adjustment coefficient is used to correct the thresholds in the baseline threshold matrix in real time. Specifically, the correction principle is as follows: when the comprehensive dynamic adjustment coefficient is greater than 1, it indicates that the current operating conditions are relatively lenient or the reliability of the normal state reference cluster is low. In this case, the warning criteria should be appropriately relaxed to reduce the risk of false alarms and improve the stability of the warning system. To this end, the baseline deviation threshold corresponding to each warning level in the baseline threshold matrix is divided by the comprehensive dynamic adjustment coefficient to obtain a larger corrected deviation threshold; at the same time, the baseline duration threshold corresponding to each warning level is also divided by the comprehensive dynamic adjustment coefficient to obtain a longer corrected duration threshold.
[0072] Conversely, when the comprehensive dynamic adjustment coefficient is less than 1, it indicates that the current operating conditions are relatively tense or the reliability of the normal state reference cluster is high. In this case, the early warning criteria should be appropriately tightened to improve the sensitivity to potential anomalies. To this end, the baseline deviation threshold corresponding to each early warning level in the baseline threshold matrix is divided by the comprehensive dynamic adjustment coefficient to obtain a smaller corrected deviation threshold. At the same time, the baseline duration threshold corresponding to each early warning level is also divided by the comprehensive dynamic adjustment coefficient to obtain a shorter corrected duration threshold.
[0073] Through the above calculation process, the static threshold parameters in the baseline threshold matrix are dynamically adjusted to values that are compatible with the current operating conditions and model confidence level, thereby generating a set of dynamic stratification criteria that take effect in real time.
[0074] Finally, the deviation of equipment status calculated from real-time monitoring data is continuously monitored. When the value of this deviation continuously exceeds the deviation threshold corresponding to a predefined warning level in the dynamic hierarchical criterion, after real-time correction, and the duration of this state exceeding the threshold reaches the judgment duration threshold corresponding to that warning level after real-time correction, then the current equipment operating status is determined to meet the conditions for triggering that level of warning, and the system automatically generates and issues the corresponding level of warning signal. In summary, the method of combining quantified status deviation, thresholds dynamically adjusted according to operating conditions and model reliability, and strict duration judgment can effectively distinguish between instantaneous fluctuations caused by brief external interference or measurement noise and gradual anomalies caused by the continuous development of potential defects in the equipment, thereby achieving accurate identification and graded early warning of latent faults in substation equipment.
[0075] In summary, the embodiments of this application have at least the following technical effects: Compared to existing technologies, this application constructs a high-quality, highly consistent multi-source time-series data foundation by deploying heterogeneous sensor arrays and performing rigorous spatiotemporal alignment and data cleaning, effectively solving the underlying technical obstacles to multi-source heterogeneous data fusion. Secondly, it utilizes a first-order autoregressive model and a pre-trained state transition network to extract stability features from the time-series data and quantify state evolution patterns. Furthermore, this application employs an unsupervised clustering algorithm to autonomously learn and establish a dynamic benchmark for the normal operating conditions of the equipment within the transition probability space, namely, a normal state reference cluster. Without relying on labeled fault samples, a robust health state model can be constructed using only the equipment's own historical normal data, possessing engineering practicality and adaptability.
[0076] Finally, a comprehensive diagnosis is performed by calculating the deviation between the real-time status and the normal reference cluster, combined with a hierarchical criterion rule that dynamically adjusts the real-time operating conditions and baseline reliability. This mechanism can sensitively capture early, subtle deviations from abnormal states, improving the early warning capability for latent faults, while effectively suppressing false alarms caused by normal fluctuations in operating conditions or transient interference. It achieves a closed-loop intelligent diagnosis of the entire process of substation equipment operating status, from multi-dimensional perception, intelligent feature extraction, adaptive modeling to accurate early warning, thereby improving the intelligence level and safety reliability of power grid operation and maintenance.
[0077] Example 2, as Figure 2As shown, based on the same inventive concept as the intelligent fault diagnosis method for substation equipment based on multi-source time-series data provided in Embodiment 1, this embodiment of the invention also provides an intelligent fault diagnosis system for substation equipment based on multi-source time-series data, including: The data acquisition and preprocessing module 11 is used to acquire multi-source time-series monitoring data through a heterogeneous sensor array deployed on substation equipment, and to perform data cleaning and time alignment to obtain a preprocessed time-series dataset. The stability feature extraction module 12 is used to fit a first-order autoregressive model to the historical sequence of each monitoring parameter in the time series dataset and extract the stability feature sequence of each monitoring parameter in the time series. The transition probability modeling module 13 is used to fuse several obtained stability feature sequences to form a multi-parameter joint feature time series, and input it into the pre-trained time series state transition learning model to obtain the transition probability matrix; The normal state establishment module 14 is used to establish a normal state reference cluster based on the transition probability matrix corresponding to the historical time series data under normal working conditions and using an unsupervised clustering algorithm. The fault diagnosis and early warning module 15 is used to calculate the deviation between the real-time transition probability matrix and the normal state reference cluster, and combined with the dynamic hierarchical criterion rules, to identify latent faults in the equipment and issue early warnings.
[0078] The data acquisition and preprocessing module 11 is specifically used for: By deploying heterogeneous sensor arrays on substation equipment, multi-source time-series monitoring data is collected, and the data is cleaned and time-aligned to obtain a preprocessed time-series dataset.
[0079] The multi-source time-series monitoring data includes partial discharge signals, temperature data, and vibration data. The partial discharge signals are acquired through an ultra-high frequency partial discharge sensor, and the temperature data and vibration data are acquired through a temperature-vibration composite sensor.
[0080] The stability feature extraction module 12 is specifically used for: For the historical sequence of each monitoring parameter in the time series dataset, a first-order autoregressive model is fitted, and the stability feature sequence of each monitoring parameter over time is extracted, including: For each monitoring parameter's historical sequence, a sliding window method is used to divide it into multiple subsequences; For each subsequence, a first-order autoregressive model is fitted, and the model parameters are extracted, wherein the model parameters include autoregressive coefficients and residual variance; Calculate the statistical characteristics of each subsequence, wherein the statistical characteristics include the mean and standard deviation; The autoregressive coefficients, residual variances, means, and standard deviations corresponding to each subsequence are arranged in chronological order to form a stability characteristic sequence of each monitoring parameter.
[0081] Specifically, the transition probability modeling module 13 is used for: The pre-training steps of the temporal state transition learning model include: Based on the historical monitoring records of the substation equipment, a historical sample dataset is collected, wherein the historical sample dataset includes a time series of multi-parameter joint features organized in chronological order; The historical sample dataset is divided into states to obtain sample state labels corresponding to each time point, and a state transition probability matrix is calculated based on the sample state labels to obtain a set of sample state transition probability matrices. Based on machine learning, construct the network structure of a temporal state transition learning model; Using the multi-parameter joint feature time series as input and the corresponding sample state transition probability matrix as supervision signal, the time series state transition learning model is trained under supervision until the model evaluation index converges, thus completing the model pre-training.
[0082] The historical sample dataset is divided into states to obtain sample state labels corresponding to each time point. Based on these sample state labels, a state transition probability matrix is calculated to obtain a set of sample state transition probability matrices, including: The time series of multi-parameter joint features in the historical sample dataset are divided using the sliding window method to obtain several feature subsequences; For all feature vectors within each feature subsequence, perform cluster analysis and use the cluster center category identifier as the sample state label for the time interval corresponding to the current subsequence; Based on the order of the sample state labels on the time axis, the frequency of transition from the current state label to the next state label is counted, and a state transition frequency matrix is constructed. Normalize each row of the state transition frequency matrix so that the sum of all elements in the same row is 1, and obtain the sample state transition probability matrix corresponding to the current historical time series data. Traverse all historical sample data segments to obtain a set of sample state transition probability matrices for model pre-training.
[0083] The normal state establishment module 14 is specifically used for: Based on the transition probability matrix corresponding to historical time-series data under normal operating conditions, an unsupervised clustering algorithm is used to establish a reference cluster for the normal state, including: From the output layer of the time-series state transition learning model, extract the transition probability matrix corresponding to all time windows in the historical period of normal operating conditions of the substation equipment to form a set of normal operating condition transition probability matrices. Each transition probability matrix in the set of normal operating condition transition probability matrices is expanded into a one-dimensional feature vector in row-major order. Using all the expanded one-dimensional feature vectors as input samples, an unsupervised clustering algorithm is used for cluster analysis to obtain multiple cluster centers and the range of samples covered. Calculate the average distance from the feature vectors of all samples within each cluster to the cluster center, and use this distance as the intra-cluster dispersion of the cluster. Based on the intra-cluster dispersion, the clustering results are filtered to remove abnormal clusters, and the feature vectors corresponding to the remaining cluster centers are reconstructed into the form of transition probability matrices to jointly constitute the normal state reference cluster. The average intra-cluster dispersion of all clusters in the normal state reference cluster is calculated as a reliability index.
[0084] Using all the expanded one-dimensional feature vectors as input samples, an unsupervised clustering algorithm is employed for cluster analysis to obtain multiple cluster centers and the sample ranges they cover, including: Set an initial range for the number of clusters K, wherein the initial range is determined based on the total number of samples in the normal working condition transition probability matrix set and the desired clustering granularity, wherein K is a positive integer, and the lower limit of the range is 2, and the upper limit is a preset maximum value Kmax; Within the initial value range, different K values are selected sequentially for cluster analysis. For each selected K value, multiple independent clustering processes are performed, and K cluster centers are randomly initialized in each clustering process. Calculate the distance of each one-dimensional feature vector to all current cluster centers, and assign it to the cluster represented by the nearest cluster center; Based on the allocation results, the mean of all one-dimensional feature vectors within each cluster is recalculated and updated as the new cluster center for the corresponding cluster; After each iteration, the total distance sum of all clusters under the current clustering partition is calculated, where the total distance sum is the sum of the distances from samples within each cluster to their respective cluster centers; The result with the minimum total distance sum during multiple independent clustering processes is selected as the candidate clustering result and candidate total distance sum corresponding to the current K value; Analyze the candidate population distances corresponding to different K values and their changing trends as K increases, and select the K value corresponding to the inflection point in the changing trend as the final K value; The set of candidate cluster centers corresponding to the final K value is taken as the optimal clustering result, and the sample range covered by each cluster center is determined based on the optimal clustering result.
[0085] The fault diagnosis and early warning module 15 is specifically used for: Calculating the deviation between the real-time transition probability matrix and the normal state reference cluster includes: The real-time collected multi-source time-series monitoring data is preprocessed, feature extracted and fused to generate a real-time multi-parameter joint feature time series, which is then input into the time-series state transition learning model to obtain the real-time transition probability matrix. Expand the real-time transition probability matrix in row-major order to obtain a real-time one-dimensional feature vector. Calculate the distance between the real-time one-dimensional feature vector and the feature vector corresponding to each cluster center in the normal state reference cluster, wherein the distance is calculated using Euclidean distance; The minimum value among several calculated distances is selected as the deviation between the real-time transition probability matrix and the normal state reference cluster.
[0086] The dynamic hierarchical criterion rules include: A baseline threshold matrix containing different warning levels is set, wherein each warning level corresponds to a set of threshold parameters, including a deviation threshold and a determination duration threshold. The mean of the absolute values of the autoregressive coefficients in the stability feature sequence is obtained in real time and used as a stability index. Obtain the current load rate of the substation equipment as an indicator of operating load. Based on the stability index and the operating load index, a basic dynamic adjustment coefficient is calculated, wherein the basic dynamic adjustment coefficient is positively correlated with the stability index and negatively correlated with the operating load index. By combining the credibility index of the normal state reference cluster, the basic dynamic adjustment coefficient is corrected to obtain the comprehensive dynamic adjustment coefficient; Using the comprehensive dynamic adjustment coefficient, the deviation thresholds of each level in the benchmark threshold matrix are numerically corrected, and the corresponding judgment duration thresholds are duration corrected to obtain a dynamic hierarchical criterion that takes effect in real time. When the deviation continuously exceeds the deviation threshold of a certain level in the dynamic hierarchical criterion, and the duration reaches the corrected duration threshold corresponding to the current level, an early warning for the corresponding level is triggered.
[0087] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0088] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0089] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.
Claims
1. A method for intelligent fault diagnosis of substation equipment based on multi-source time-series data, characterized in that, The method includes: By deploying a heterogeneous sensor array on substation equipment, multi-source time-series monitoring data is collected, and the data is cleaned and time-aligned to obtain a pre-processed time-series dataset. For the historical sequence of each monitoring parameter in the time series dataset, a first-order autoregressive model is fitted, and the stability feature sequence of each monitoring parameter in the time series is extracted. Several stable feature sequences obtained by fusion are combined to form a multi-parameter joint feature time series, which is then input into a pre-trained temporal state transition learning model to obtain the transition probability matrix; Based on the transition probability matrix corresponding to historical time series data under normal operating conditions, an unsupervised clustering algorithm is used to establish a reference cluster for the normal state. The deviation between the real-time transition probability matrix and the normal state reference cluster is calculated, and combined with the dynamic hierarchical criterion rule, latent equipment faults are identified and warnings are issued.
2. The intelligent fault diagnosis method for substation equipment based on multi-source time-series data according to claim 1, characterized in that, The multi-source time-series monitoring data includes partial discharge signals, temperature data, and vibration data. The partial discharge signals are acquired through an ultra-high frequency partial discharge sensor, and the temperature data and vibration data are acquired through a temperature-vibration composite sensor.
3. The intelligent fault diagnosis method for substation equipment based on multi-source time-series data according to claim 1, characterized in that, For the historical sequence of each monitoring parameter in the time series dataset, a first-order autoregressive model is fitted, and the stability feature sequence of each monitoring parameter over time is extracted, including: For each monitoring parameter's historical sequence, a sliding window method is used to divide it into multiple subsequences; For each subsequence, a first-order autoregressive model is fitted, and the model parameters are extracted, wherein the model parameters include autoregressive coefficients and residual variance; Calculate the statistical characteristics of each subsequence, wherein the statistical characteristics include the mean and standard deviation; The autoregressive coefficients, residual variances, means, and standard deviations corresponding to each subsequence are arranged in chronological order to form a stability characteristic sequence of each monitoring parameter.
4. The intelligent fault diagnosis method for substation equipment based on multi-source time-series data according to claim 1, characterized in that, The pre-training steps of the temporal state transition learning model include: Based on the historical monitoring records of the substation equipment, a historical sample dataset is collected, wherein the historical sample dataset includes a time series of multi-parameter joint features organized in chronological order; The historical sample dataset is divided into states to obtain sample state labels corresponding to each time point, and a state transition probability matrix is calculated based on the sample state labels to obtain a set of sample state transition probability matrices. Based on machine learning, construct the network structure of a temporal state transition learning model; Using the multi-parameter joint feature time series as input and the corresponding sample state transition probability matrix as supervision signal, the time series state transition learning model is trained under supervision until the model evaluation index converges, thus completing the model pre-training.
5. The intelligent fault diagnosis method for substation equipment based on multi-source time-series data according to claim 4, characterized in that, The historical sample dataset is divided into states to obtain sample state labels corresponding to each time point. Based on these sample state labels, a state transition probability matrix is calculated to obtain a set of sample state transition probability matrices, including: The time series of multi-parameter joint features in the historical sample dataset are divided using the sliding window method to obtain several feature subsequences; For all feature vectors within each feature subsequence, perform cluster analysis and use the cluster center category identifier as the sample state label for the time interval corresponding to the current subsequence; Based on the order of the sample state labels on the time axis, the frequency of transition from the current state label to the next state label is counted, and a state transition frequency matrix is constructed. Normalize each row of the state transition frequency matrix so that the sum of all elements in the same row is 1, and obtain the sample state transition probability matrix corresponding to the current historical time series data. Traverse all historical sample data segments to obtain a set of sample state transition probability matrices for model pre-training.
6. The intelligent fault diagnosis method for substation equipment based on multi-source time-series data according to claim 1, characterized in that, Based on the transition probability matrix corresponding to historical time-series data under normal operating conditions, an unsupervised clustering algorithm is used to establish a reference cluster for the normal state, including: From the output layer of the time-series state transition learning model, extract the transition probability matrix corresponding to all time windows in the historical period of normal operating conditions of the substation equipment to form a set of normal operating condition transition probability matrices. Each transition probability matrix in the set of normal operating condition transition probability matrices is expanded into a one-dimensional feature vector in row-major order. Using all the expanded one-dimensional feature vectors as input samples, an unsupervised clustering algorithm is used for cluster analysis to obtain multiple cluster centers and the range of samples covered. Calculate the average distance from the feature vectors of all samples within each cluster to the cluster center, and use this distance as the intra-cluster dispersion of the cluster. Based on the intra-cluster dispersion, the clustering results are filtered to remove abnormal clusters, and the feature vectors corresponding to the remaining cluster centers are reconstructed into the form of transition probability matrices to jointly constitute the normal state reference cluster. The average intra-cluster dispersion of all clusters in the normal state reference cluster is calculated as a reliability index.
7. The intelligent fault diagnosis method for substation equipment based on multi-source time-series data according to claim 6, characterized in that, Using all the expanded one-dimensional feature vectors as input samples, an unsupervised clustering algorithm is employed for cluster analysis to obtain multiple cluster centers and the sample ranges they cover, including: Set an initial range for the number of clusters K, wherein the initial range is determined based on the total number of samples in the normal working condition transition probability matrix set and the desired clustering granularity, wherein K is a positive integer, and the lower limit of the range is 2, and the upper limit is a preset maximum value Kmax; Within the initial value range, different K values are selected sequentially for cluster analysis. For each selected K value, multiple independent clustering processes are performed, and K cluster centers are randomly initialized in each clustering process. Calculate the distance of each one-dimensional feature vector to all current cluster centers, and assign it to the cluster represented by the nearest cluster center; Based on the allocation results, the mean of all one-dimensional feature vectors within each cluster is recalculated and updated as the new cluster center for the corresponding cluster; After each iteration, the total distance sum of all clusters under the current clustering partition is calculated, where the total distance sum is the sum of the distances from samples within each cluster to their respective cluster centers; The result with the minimum total distance sum during multiple independent clustering processes is selected as the candidate clustering result and candidate total distance sum corresponding to the current K value; Analyze the candidate population distances corresponding to different K values and their changing trends as K increases, and select the K value corresponding to the inflection point in the changing trend as the final K value; The set of candidate cluster centers corresponding to the final K value is taken as the optimal clustering result, and the sample range covered by each cluster center is determined based on the optimal clustering result.
8. The intelligent fault diagnosis method for substation equipment based on multi-source time-series data according to claim 1, characterized in that, Calculating the deviation between the real-time transition probability matrix and the normal state reference cluster includes: The real-time collected multi-source time-series monitoring data is preprocessed, feature extracted and fused to generate a real-time multi-parameter joint feature time series, which is then input into the time-series state transition learning model to obtain the real-time transition probability matrix. Expand the real-time transition probability matrix in row-major order to obtain a real-time one-dimensional feature vector. Calculate the distance between the real-time one-dimensional feature vector and the feature vector corresponding to each cluster center in the normal state reference cluster, wherein the distance is calculated using Euclidean distance; The minimum value among several calculated distances is selected as the deviation between the real-time transition probability matrix and the normal state reference cluster.
9. The intelligent fault diagnosis method for substation equipment based on multi-source time-series data according to claim 1, characterized in that, The dynamic hierarchical criterion rules include: A baseline threshold matrix containing different warning levels is set, wherein each warning level corresponds to a set of threshold parameters, including a deviation threshold and a determination duration threshold. The mean of the absolute values of the autoregressive coefficients in the stability feature sequence is obtained in real time and used as a stability index. Obtain the current load rate of the substation equipment as an indicator of operating load. Based on the stability index and the operating load index, a basic dynamic adjustment coefficient is calculated, wherein the basic dynamic adjustment coefficient is positively correlated with the stability index and negatively correlated with the operating load index. By combining the credibility index of the normal state reference cluster, the basic dynamic adjustment coefficient is corrected to obtain the comprehensive dynamic adjustment coefficient; Using the comprehensive dynamic adjustment coefficient, the deviation thresholds of each level in the benchmark threshold matrix are numerically corrected, and the corresponding judgment duration thresholds are duration corrected to obtain a dynamic hierarchical criterion that takes effect in real time. When the deviation continuously exceeds the deviation threshold of a certain level in the dynamic hierarchical criterion, and the duration reaches the corrected duration threshold corresponding to the current level, an early warning for the corresponding level is triggered.
10. A substation equipment fault intelligent diagnosis system based on multi-source time-series data, characterized in that, The method for intelligent fault diagnosis of substation equipment based on multi-source time-series data according to any one of claims 1-9 includes: The data acquisition and preprocessing module is used to acquire multi-source time-series monitoring data through a heterogeneous sensor array deployed on substation equipment, and to perform data cleaning and time alignment to obtain a preprocessed time-series dataset. The stability feature extraction module is used to fit a first-order autoregressive model to the historical sequence of each monitoring parameter in the time series dataset and extract the stability feature sequence of each monitoring parameter in the time series. The transition probability modeling module is used to fuse several obtained stability feature sequences to form a multi-parameter joint feature time series, which is then input into a pre-trained time series state transition learning model to obtain the transition probability matrix. The normal state establishment module is used to establish a normal state reference cluster based on the transition probability matrix corresponding to historical time series data under normal operating conditions, using an unsupervised clustering algorithm. The fault diagnosis and early warning module is used to calculate the deviation between the real-time transition probability matrix and the normal state reference cluster, and combined with the dynamic hierarchical criterion rules, to identify latent faults in the equipment and issue early warnings.