A power distribution equipment anomaly detection method based on time series transformer
Patent Information
- Application Number
- CN202610905263.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-15
Smart Images

Figure CN122763377A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power distribution anomaly detection technology, and in particular to a method for detecting power distribution equipment anomalies based on a time-series converter. Background Technology
[0002] Existing methods for detecting anomalies in power distribution equipment largely rely on limit-crossing alarms, empirical rules, and simple statistical analysis. These methods set fixed thresholds or empirical ranges for operating parameters such as voltage, current, active power, reactive power, power quality indicators, and temperature. Alarms are triggered by exceeding upper or lower limits or exceeding duration limits. Some solutions combine traditional time series models, support vector machines, or random forest methods to predict trends and classify states for single or a few monitored parameters. Current methods generally suffer from problems such as reliance on manually set features and thresholds, significant differences in parameter distribution across different types of power distribution equipment and operating conditions, difficulty in uniform adjustment, poor adaptability to complex scenarios involving operating condition switching, load surges, and power quality disturbances, and a high incidence of false alarms and missed alarms. They also struggle to identify early anomalies and gradual degradation processes in a timely manner, resulting in limited overall anomaly detection accuracy and early warning capabilities.
[0003] With the development of sensing and communication technologies, online monitoring of power distribution equipment is generating large-scale, multi-dimensional, and multi-sampling frequency runtime data. Some studies have begun to introduce time-series signal decomposition and deep learning models, such as using empirical mode decomposition, variational mode decomposition, or symplectic geometric mode decomposition to decompose current, voltage, or vibration signals into several scale components. These components are then combined with convolutional neural networks, long short-term memory networks, or attention-based converter models for prediction and fault identification. Other studies have introduced exponential smoothing time series models for trend and seasonality modeling. However, existing work is mostly focused on single signals or limited-dimensional data, lacking a unified modeling framework for multi-source runtime data of power distribution equipment. It also fails to adequately mine non-stationary, multi-scale, and strongly correlated operational characteristics. Coupling and aliasing still exist between modal components, and there is a lack of collaborative modeling of trends, seasonal structures, and residual behavior. Residuals are often used in the form of numerical bias, failing to fully utilize distribution characteristics. As a result, the robustness, generalization ability, and ability to identify complex anomaly patterns in online monitoring and anomaly detection remain significantly insufficient.
[0004] Therefore, how to provide a method for detecting anomalies in power distribution equipment based on time converters is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] One objective of this invention is to propose a method for detecting anomalies in power distribution equipment based on a time-series converter. This invention addresses multi-source operating time-series data of power distribution equipment and constructs a complete process from data acquisition and preprocessing, symplectic geometric mode decomposition, multi-scale feature construction, improved ETSformer modeling, anomaly index calculation, to adaptive alarm output. The method introduces a deformed symplectic mapping structure and a modal coupling interference cancellation structure in the signal decomposition stage to obtain mutually decoupled multi-scale modal components. In the time-series modeling stage, an improved ETSformer model is constructed, including a trend-seasonal module, a residual distribution module, and a component decoding module, achieving structured modeling of trend, seasonality, and residual distribution characteristics. An adaptive anomaly discrimination threshold is constructed based on point residual anomaly anomaly anomaly, interval residual anomaly anomaly, and distribution deviation anomaly ...
[0006] A method for detecting power distribution equipment anomalies based on a time-series converter according to an embodiment of the present invention includes: Collect multi-source runtime timing data of power distribution equipment, preprocess the multi-source runtime timing data to obtain preprocessed timing data; Symmetric geometric mode decomposition is performed on the target time series signal in the preprocessed time series data. Hamilton structure matrix is constructed and deformed symmetric mapping structure is applied to obtain symmetric geometric mode subspace. Mode coupling interference elimination structure is used to suppress and decouple the cross projection between modes, resulting in multiple mutually decoupled mode components. Modal components are filtered and grouped for reconstruction to form multi-scale subsequences, which are then combined with preprocessed time series data to obtain multi-scale input feature sequences. An improved ETSformer model is constructed, which uses the trend and seasonality module to extract trends and detrend multi-scale input feature sequences, uses the residual distribution module to perform window partitioning and statistical distribution feature mapping, and uses the component decoding module to calculate component importance coefficients to form decoded input features. The decoded input features are input into the temporal decoding unit of the decoding module of the component, and temporal decoding and output reconstruction are performed to obtain multi-step predicted values. The deviation is calculated based on the multi-step predicted values, and anomaly indices such as point residual anomaly, interval residual anomaly, and distribution deviation anomaly are constructed. Determine the adaptive threshold range for the anomaly index, determine that the corresponding power distribution equipment is in an abnormal operating state, and output anomaly alarm information.
[0007] Optionally, the multi-source runtime timing data includes voltage, current, active power, reactive power, power quality indicators, and temperature.
[0008] Optionally, obtaining the preprocessed time-series data includes: By using voltage transformers, current transformers, temperature sensors, and power quality monitoring devices installed on power distribution equipment, time-stamped voltage sampling sequences, current sampling sequences, active power sampling sequences, reactive power sampling sequences, power quality index sampling sequences, and temperature sampling sequences are collected at fixed sampling periods to obtain multi-source runtime sequence data. The multi-source runtime sequence data is aligned according to time stamps, the sampled data from different measurement points are mapped to a unified time axis, data points with duplicate time stamps are deleted, and data points with missing time stamps are interpolated to complete the data points with missing time stamps, resulting in a multi-source runtime sequence data matrix with a consistent time axis. Abnormal data in the multi-source runtime sequence data matrix is removed, each physical quantity is normalized according to its own numerical range, and the normalized multi-source runtime sequence data is divided into sample sequences of fixed length according to time order to form preprocessed time series data.
[0009] Optionally, obtaining multiple mutually decoupled modal components includes: Select the target time series signal from the preprocessed time series data, and reconstruct the phase space of the target time series signal according to the set embedding dimension and time delay to construct the trajectory matrix; The Hamilton structure matrix is constructed using the trajectory matrix. The Hamilton structure matrix consists of four sub-block matrices and satisfies the symplectic structure constraint. The Hamilton structure matrix and the standard symplectic matrix satisfy the symplectic preservation relationship. Construct a deformed symplectic mapping structure, define a set of symplectic preserving transformation operators with parameters, perform left and right multiplication transformation operations on the Hamilton structure matrix to obtain the deformed symplectic geometric modal feature matrix, and form a feature vector set containing multiple candidate modal feature vectors; Multiple initial modal subspaces are generated based on the feature vector set. The cross-projection energy between any two different modal subspaces is calculated, and a modal cross-projection matrix is constructed. In the modal coupling interference cancellation structure, the modal cross-projection matrix is used as input to perform orthogonal projection constraint operation and energy redistribution operation, reduce the cross-projection components between different modal subspaces, output a set of decoupled modal subspaces, and restore each modal subspace to the corresponding modal component time sequence signal according to the diagonal averaging principle, thus obtaining multiple mutually decoupled modal components.
[0010] Optionally, obtaining the multi-scale input feature sequence includes: For each modal component, calculate the energy value, mean, variance, correlation coefficient with the target time-series signal, and dominant frequency characteristics to obtain the modal feature set; Based on the energy threshold, correlation coefficient threshold, and effective range of dominant frequency, modal components with low energy and low correlation are removed from the modal feature set to obtain the effective modal component set; Based on the dominant frequency characteristics and energy distribution in the effective modal component set, the effective modal components are divided into multi-scale subsequences. The multi-scale subsequence set is generated by adding them point by point according to the time index and aligning them with the key monitoring quantities in the preprocessed time series data according to time. By splicing the multi-scale subsequence values and multi-source operating parameter values at each time step, a multi-scale input feature vector is constructed. The multi-scale input feature vector is arranged in chronological order to obtain the multi-scale input feature sequence.
[0011] Optionally, forming the decoding input features includes: Construct an improved ETSformer model, including a trend and seasonality module, a residual distribution module, and a component decoding module; The multi-scale input feature sequence is input into the trend and seasonal module. The trend extraction operation is performed on the multi-scale input feature sequence to generate the trend sequence. The trend sequence is subtracted from the multi-scale input feature sequence to obtain the detrended sequence. The seasonal modeling operation is performed on the detrended sequence to generate the seasonal sequence. The trend sequence and the seasonal sequence are then output. The residual sequence is obtained by subtracting the trend sequence and seasonal sequence from the multi-scale input feature sequence. The residual sequence is then input into the residual distribution module. The residual sequence is divided into windows according to the window length. The statistics of mean, variance, skewness, kurtosis and quantile are calculated for the residual samples in each window. The statistics are arranged in a fixed order to generate the residual distribution feature vector. The residual distribution feature vector is then arranged in time order to generate the residual distribution feature sequence. The trend sequence, seasonal sequence, and residual distribution feature sequence are input into the component decoding module. The component weight calculation unit in the component decoding module calculates the change magnitude index, stability index, and deviation index of the trend sequence, seasonal sequence, and residual distribution feature sequence respectively within the time window, and generates the importance coefficient of the trend component, the importance coefficient of the seasonal component, and the importance coefficient of the residual component. The trend sequence, seasonal sequence, and residual distribution feature sequence are multiplied by the corresponding component importance coefficients and concatenated in time order to generate the decoded input feature sequence.
[0012] Optionally, obtaining the multi-step predicted value includes: The decoding input feature sequence is input into the temporal decoding unit in the decoding module of the component, and the feature vectors of each time step in the decoding input feature sequence are read sequentially in the time dimension to generate the corresponding decoding hidden state sequence. Within the temporal decoding unit, temporal decoding operations are performed on the decoded hidden state sequence. For each prediction time step, the corresponding prediction feature vector sequence is output. The vectors of each time step in the prediction feature vector sequence are arranged in chronological order to form the prediction feature sequence. The predicted feature sequence is input into the output reconstruction unit. The output reconstruction unit performs linear transformation and dimension mapping on the predicted feature vector of each time step in the predicted feature sequence to obtain the predicted value sequence of the operating parameters of the power distribution equipment at multiple predicted time steps.
[0013] Optionally, the anomaly indexes for constructing point residual anomaly, interval residual anomaly, and distribution deviation anomaly include: At each time step, the actual measured value of the operating parameter is subtracted from the corresponding predicted value to obtain the residual time series. The absolute value of the residual at each time step in the residual time series is taken to form a point residual constant series. The residual time series is divided by a sliding time window. The residual mean and residual variance are calculated in each window to form an interval residual constant series. The probability distribution parameters of residual time series are statistically analyzed on the historical data of normal operation. In the prediction stage, the difference between the current residual time series and the normal residual distribution is calculated as the distribution deviation anomaly degree. This anomaly degree index is constructed by combining it with the point residual difference normality sequence and the interval residual difference normality sequence.
[0014] Optionally, the output of abnormal alarm information includes: During the normal operation phase in history, anomaly index sequence is recorded. For each anomaly index, mean, standard deviation and quantile are calculated on the historical time axis. Based on the mean, standard deviation and quantile, upper limit threshold and lower limit threshold are set for point residual anomaly, interval residual anomaly and distribution deviation anomaly, respectively, to form an adaptive discrimination threshold range. During the detection phase, the current anomaly index sequence is calculated in real time. Each anomaly value is compared with the corresponding adaptive discrimination threshold range. When any anomaly value exceeds the number of time steps, exceeds the upper threshold, or falls below the lower threshold, the corresponding power distribution equipment is marked as being in an abnormal operating state. An anomaly alarm message containing equipment identification, anomaly type, and time information is generated and sent to the operation and maintenance terminal.
[0015] The beneficial effects of this invention are: This invention significantly enhances the feature extraction and correlation modeling capabilities of multi-source operating time series data of power distribution equipment through the collaborative design of symplectic geometric mode decomposition and an improved ETSformer model. Symplectic geometric mode decomposition introduces a deformed symplectic mapping structure and a mode coupling interference cancellation structure, achieving multi-scale fine decomposition and mode decoupling of complex non-stationary time series signals during the signal decomposition stage, resulting in more clearly structured multi-scale sub-sequences. The improved ETSformer model, through trend and seasonal modules, residual distribution modules, and component decoding modules, performs structured modeling and dynamic weight allocation of trend, seasonality, and residual distribution features, thereby more fully exploring the long-term trends, periodic patterns, and abnormal disturbance characteristics in complex multi-source time series data. This improves the accuracy of abnormal pattern recognition and sensitivity to early potential problems, while reducing reliance on manual feature design and empirical thresholds.
[0016] This invention achieves refined and intelligent online anomaly detection for power distribution equipment through multi-scale input feature construction, joint measurement of point residual difference anomaly, interval residual difference anomaly, and distribution deviation anomaly, and an adaptive threshold discrimination mechanism. The deviation between multi-step predicted values and actual values is quantified at three levels: numerical, interval, and distribution, covering various anomaly forms such as sudden anomalies, gradual degradation, and statistical distribution shifts. The adaptive discrimination threshold is automatically generated based on historical normal operation data and updated with the operating status, ensuring good robustness and generalization ability under different equipment types, different load levels, and variable operating conditions. In summary, this invention effectively overcomes the problems of insufficient multi-source temporal correlation feature mining and reliance on fixed thresholds for anomaly detection with poor adaptability in existing technologies. It offers advantages such as high anomaly identification accuracy, strong early warning capability, ease of engineering deployment, and support for intelligent operation and maintenance. Attached Figure Description
[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a power distribution equipment anomaly detection method based on a time-series converter proposed in this invention; Figure 2 This is a structural block diagram of the symplectic geometric mode decomposition of a power distribution equipment anomaly detection method based on a time-series converter proposed in this invention; Figure 3 This is a functional diagram of the improved ETSformer model of the power distribution equipment anomaly detection method based on time converter proposed in this invention. Detailed Implementation
[0018] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0019] refer to Figure 1 , Figure 2 and Figure 3 A method for detecting anomalies in power distribution equipment based on a time-series converter, comprising: Collect multi-source runtime timing data of power distribution equipment, preprocess the multi-source runtime timing data to obtain preprocessed timing data; Symmetric geometric mode decomposition is performed on the target time series signal in the preprocessed time series data. Hamilton structure matrix is constructed and deformed symmetric mapping structure is applied to obtain symmetric geometric mode subspace. Mode coupling interference elimination structure is used to suppress and decouple the cross projection between modes, resulting in multiple mutually decoupled mode components. Modal components are filtered and grouped for reconstruction to form multi-scale subsequences, which are then combined with preprocessed time series data to obtain multi-scale input feature sequences. An improved ETSformer model is constructed, which uses the trend and seasonality module to extract trends and detrend multi-scale input feature sequences, uses the residual distribution module to perform window partitioning and statistical distribution feature mapping, and uses the component decoding module to calculate component importance coefficients to form decoded input features. The decoded input features are input into the temporal decoding unit of the decoding module of the component, and temporal decoding and output reconstruction are performed to obtain multi-step predicted values. The deviation is calculated based on the multi-step predicted values, and anomaly indices such as point residual anomaly, interval residual anomaly, and distribution deviation anomaly are constructed. Determine the adaptive threshold range for the anomaly index, determine that the corresponding power distribution equipment is in an abnormal operating state, and output anomaly alarm information.
[0020] In this embodiment, the multi-source runtime timing data includes voltage, current, active power, reactive power, power quality indicators, and temperature.
[0021] In this embodiment, obtaining the preprocessed time series data includes: By using voltage transformers, current transformers, temperature sensors, and power quality monitoring devices installed on power distribution equipment, time-stamped voltage sampling sequences, current sampling sequences, active power sampling sequences, reactive power sampling sequences, power quality index sampling sequences, and temperature sampling sequences are collected at fixed sampling periods to obtain multi-source runtime sequence data. The multi-source runtime sequence data is aligned according to time stamps, the sampled data from different measurement points are mapped to a unified time axis, data points with duplicate time stamps are deleted, and data points with missing time stamps are interpolated to complete the data points with missing time stamps, resulting in a multi-source runtime sequence data matrix with a consistent time axis. Abnormal data in the multi-source runtime sequence data matrix is removed, each physical quantity is normalized according to its own numerical range, and the normalized multi-source runtime sequence data is divided into sample sequences of fixed length according to time order to form preprocessed time series data.
[0022] In this embodiment, obtaining multiple mutually decoupled modal components includes: A target time-series signal is selected from the preprocessed time-series data. According to the set embedding dimension and time delay, the target time-series signal is reconstructed in phase space to construct a trajectory matrix. Specifically, the construction of the trajectory matrix involves: Select a target time series signal from the preprocessed time series data. Set the embedding dimension to the number of sampling points contained in each embedding vector. Set the time delay step to the number of interval sampling points of adjacent sampling points in the same embedding vector on the target time series signal. Take the first sampling point of the target time series signal as the starting position. Under the premise that several sampling points are selected in sequence according to the time delay step, the total number of selected sampling points is equal to the embedding dimension, and all sampling points fall within the length range of the target time series signal, move the starting position forward point by point and check whether the conditions are met one by one. The number of all starting positions that meet the conditions is determined as the number of valid time positions. Starting from each valid time position, sampling points of the embedding dimension are extracted sequentially from the target time-series signal in chronological order. The first sampling point is extracted from the starting point, and then one sampling point is extracted at each time delay step until the number of extracted sampling points equals the embedding dimension. All sampling points extracted at the same valid time position are arranged into a row of data according to the extraction order. The extraction and arrangement operation is repeated for all valid time positions. The resulting rows of data are stacked from top to bottom according to the order of the starting time positions to form a matrix. The number of rows in the matrix equals the number of valid time positions, and the number of columns in the matrix equals the embedding dimension. The current matrix is used as the trajectory matrix.
[0023] A Hamiltonian structure matrix is constructed using the trajectory matrix. The Hamiltonian structure matrix consists of four sub-block matrices and satisfies symplectic structure constraints. The Hamiltonian structure matrix and the standard symplectic matrix satisfy a symplectic preservation relationship, where: The method of constructing Hamiltonian structure matrix using trajectory matrix is as follows: perform matrix multiplication between trajectory matrix and transpose matrix to obtain a square matrix with the same number of rows and columns. Any element in the square matrix represents the degree of correlation or energy coupling between two rows or two columns in trajectory matrix. Through calculation, an energy correlation matrix suitable for constructing Hamiltonian structure can be extracted from the original trajectory data. The Hamiltonian structure matrix consists of four sub-block matrices and satisfies symplectic structure constraints, specifically: The Hamilton structure matrix is divided into four parts: the upper left sub-block, the upper right sub-block, the lower left sub-block, and the lower right sub-block. The four sub-blocks have the same number of rows and columns. The upper left sub-block is set as a zero matrix with all zeros. The lower right sub-block is also set as a zero matrix. The upper right sub-block directly uses the correlation matrix calculated from the trajectory matrix to characterize the coupling strength between different state components in the phase space. The lower left sub-block is set as the negative transpose of the upper right sub-block. That is, each element in the lower left sub-block is equal to the value obtained by taking the opposite of the corresponding element in the upper right sub-block and interchanging the rows and columns. Through the block structure design of the zero matrix plus the negative transpose matrix, the Hamilton structure matrix as a whole satisfies the symplectic structure constraint that the diagonal is zero and the two sides of the diagonal are negative transposes of each other. The Hamiltonian structure matrix and the standard symplectic matrix satisfy the symplectic preservation relation. Specifically, a fixed standard symplectic matrix is introduced, and the Hamiltonian structure matrix is multiplied by the standard symplectic matrix on the left. Then, the transpose of the Hamiltonian structure matrix is multiplied by the product on the right. If the result after the two matrix operations is still equal to the initial standard symplectic matrix, it indicates that the Hamiltonian structure matrix has a preservation effect on the standard symplectic matrix, that is, the symplectic preservation relation is satisfied. A deformed symplectic mapping structure is constructed, and a set of symplectic-preserving transformation operators with parameters are defined. Left and right multiplication transformations are performed on the Hamiltonian structure matrix to obtain the deformed symplectic geometric modal feature matrix, forming an feature vector set containing multiple candidate modal feature vectors. Specifically, the set of symplectic-preserving transformation operators with parameters is defined as follows: The parameter settings in the parameterized symplectic preserve transform operator include two types of real parameters: one is the diagonal control parameter, which controls the rotation and stretching intensity of each embedding dimension in the phase space; the other is the coupling control parameter, which controls the degree of coupling between different embedding dimensions. Based on the dimension of the Hamilton structure matrix, the number of diagonal control parameters is determined to be the same as the embedding dimension. Each diagonal control parameter corresponds to a state component in the phase space. At the same time, the number of coupling control parameters is determined to match the number of pairwise combinations between state components. Each coupling control parameter corresponds to the coupling relationship between a pair of state components. All diagonal control parameters are taken as real numbers between zero and one, limiting the rotation and stretching amplitudes to not exceed the preset range. All coupling control parameters are taken as real numbers between negative one and one, limiting the coupling strength to vary within the symmetric range. Multiple initial modal subspaces are generated based on the feature vector set. The cross-projection energy between any two different modal subspaces is calculated, and a modal cross-projection matrix is constructed, where: The generation of multiple initial modal subspaces specifically involves: extracting the complete set of eigenvectors from the deformed symplectic geometric modal feature matrix; dividing the set of eigenvectors into several groups based on the magnitude of the eigenvalues, the frequency range corresponding to the eigenvectors, or the oscillation characteristics; for each group of eigenvectors, linearly combining the current group of eigenvectors using the trajectory matrix; weighting and superimposing the rows or columns of the trajectory matrix according to the weights given by the eigenvectors to obtain a set of modal basis sequences that change on the time axis; and considering the modal basis sequences within the same group as the basis of a modal subspace to obtain multiple initial modal subspaces. The calculation of the cross-projection energy is specifically as follows: a modal subspace is selected as the projected subspace, and each modal basis sequence in the subspace is linearly represented on the basis of the other modal subspace. The weight coefficient of the linear representation is determined by the minimum mean square error. For each modal basis sequence, the energy of the linear representation part on the other modal subspace over the entire time axis is calculated, which is used as the projection energy of the current modal basis sequence on the other modal subspace. The projection energy of all modal basis sequences in the same modal subspace is averaged to obtain the cross-projection energy. The cross-projection energy is obtained by combining all modal subspaces pairwise. The cross-projection energy of each pair of modal subspaces is filled into a square matrix according to the row and column positions to form a modal cross-projection matrix. In the modal coupling interference cancellation structure, the modal cross-projection matrix is used as input to perform orthogonal projection constraint operations and energy redistribution operations, reducing the cross-projection components between different modal subspaces and outputting a set of decoupled modal subspaces. Then, each modal subspace is restored to its corresponding modal component time-series signal according to the diagonal averaging principle, resulting in multiple mutually decoupled modal components, where: The orthogonal projection constraint operation and energy redistribution operation are specifically as follows: Orthogonal projection constraint operation calculates the energy of the modal basis sequence in each modal subspace. The square root of the sum of the squares of the amplitudes at each time point is taken as the normalization factor. The normalization factor is used to divide the modal basis sequence point by point to obtain the normalized modal basis sequence. When two different modal subspaces are selected, the normalized modal basis sequence in one subspace is combined with the normalized modal basis sequence in the other subspace one by one. The inner product is the sum of the products of the amplitudes at the corresponding time points. The inner product result is used as the projection coefficient. The projection coefficient is multiplied by the reference modal basis sequence to obtain the projection component. The sum of all projection components is subtracted point by point from the current modal basis sequence. All modal basis sequences are updated in turn. Energy redistribution is based on the cross-projection energy between modal subspaces recorded in the modal cross-projection matrix. The cross-projection energy of each pair of subspaces is divided by the total energy of the modal basis sequences in the corresponding subspace to obtain the cross-energy ratio. Based on the cross-energy ratio, the subspace with a higher ratio in a certain frequency band is taken as the dominant subspace. The amplitude is increased according to the ratio in the modal basis sequences of the dominant subspace, and the amplitude is decreased according to the ratio in the modal basis sequences of the non-dominant subspace. The diagonal averaging principle is as follows: the modal trajectory matrix corresponding to each modal subspace is divided into multiple diagonal lines according to the fact that the sum of the row and column indices is a constant. Each diagonal line contains several matrix elements. For each diagonal line, all elements are added together and then divided by the number of elements to obtain the average value of the diagonal line. Then, the average values are taken out in order according to the order of the diagonal lines and arranged in order to form a one-dimensional time series. This time series is used as the modal component time series signal of the corresponding modal subspace.
[0024] In this embodiment, obtaining the multi-scale input feature sequence includes: For each modal component, calculate its energy value, mean, variance, correlation coefficient with the target time-series signal, and dominant frequency characteristics to obtain a set of modal features. Specifically, the correlation coefficient and dominant frequency characteristics of the target time-series signal are as follows: For each modal component and the target time series signal, first calculate the mean of both, then subtract the mean from the value at each time point, multiply the deviations of the two at each time point and sum them over the entire time range to obtain the co-variance. Then calculate the sum of the squares of the deviations of the two sequences and take the square root to obtain their respective standard deviations. Divide the co-variance by the product of the two standard deviations, and use the result as the correlation coefficient between the modal component and the target time series signal. For each modal component, frequency domain analysis is used to transform the time domain sequence of the modal component to the frequency domain to obtain the spectral amplitude corresponding to different frequency points. The frequency with the largest amplitude is found among all frequency points, and the current frequency is taken as the dominant frequency of the modal component. The corresponding spectral amplitude is taken as the dominant frequency amplitude. The two together constitute the dominant frequency feature of the modal component. Based on the energy threshold, correlation coefficient threshold, and effective range of the dominant frequency, modal components with low energy and low correlation are removed from the modal feature set to obtain the effective modal component set. Specifically, the energy threshold, correlation coefficient threshold, and effective range of the dominant frequency are: During a historical period when the power distribution equipment is in normal operation, the energy of each modal component obtained by symplectic geometric mode decomposition is calculated. All energy values are arranged in ascending order, and the energy value whose position in the arrangement is equal to 30% of the total number of modal components is taken as the energy threshold. The absolute value of the correlation coefficient between each modal component and the target time-series signal is taken, arranged in ascending order, and the value whose position in the arrangement is equal to 70% of the total number of modal components is taken as the correlation coefficient threshold. Based on the 50 Hz power frequency used in the power distribution system, the effective range of the main frequency is set to a power frequency base band of not less than 45 Hz and not greater than 55 Hz. Based on the dominant frequency characteristics and energy distribution in the effective modal component set, the effective modal components are divided into multi-scale subsequences. These subsequences are then added point-by-point according to their time indices to generate a multi-scale subsequence set, which is aligned with key monitoring quantities in the preprocessed time-series data. A multi-scale input feature vector is constructed by concatenating the multi-scale subsequence values and multi-source operating parameter values at each time step. The multi-scale input feature vectors are then arranged in chronological order to obtain the multi-scale input feature sequence, where: The process of dividing effective modal components into multi-scale subsequences involves: reading the dominant frequency and energy of each effective modal component, sorting them according to their dominant frequency from smallest to largest, assigning modal components with dominant frequencies not greater than a first frequency threshold to the low-frequency group, assigning modal components with dominant frequencies greater than the first frequency threshold but not greater than a second frequency threshold to the mid-frequency group, and assigning modal components with dominant frequencies greater than the second frequency threshold to the high-frequency group. The first frequency threshold is set to 10 Hz, and the second frequency threshold is set to 100 Hz. For each group of modal components, the values of all modal components in the current group are added at the same time index position to obtain a subsequence with the same number of time steps. The subsequences generated by the low-frequency group, mid-frequency group, and high-frequency group are recorded as different subsequences in the multi-scale subsequence set. The method involves constructing a multi-scale input feature vector by splicing multi-scale subsequence values and multi-source operating parameter values at each time step. Specifically, at each time step, the corresponding low-frequency subsequence values, mid-frequency subsequence values, and high-frequency subsequence values are extracted, and then the corresponding voltage values, current values, active power values, reactive power values, power quality index values, and temperature values are extracted in sequence. The multi-scale subsequence values and multi-source operating parameter values are arranged in the same vector in order. The splicing process is repeated for all time steps, and the multi-scale input feature vectors of each time step are arranged in chronological order to obtain the multi-scale input feature sequence.
[0025] In this embodiment, forming the decoding input feature includes: Construct an improved ETSformer model, including a trend and seasonality module, a residual distribution module, and a component decoding module, wherein: The trend and seasonal module is placed at the input of ETSformer. Trend extraction and detrending operations are performed before the multi-scale input feature sequence enters the encoder. The obtained trend sequence and detrending sequence are used as trend input and seasonal input. A residual distribution module is inserted between the encoder and decoder. The residual sequence is obtained by subtracting the trend sequence and seasonal sequence from the multi-scale input feature sequence. The residual sequence is calculated according to the time window, and the statistics of mean, variance, skewness, kurtosis and quantile are encoded into the residual distribution feature sequence. At the decoding end, a component decoding module replaces the original decoding structure. The component decoding module receives the trend sequence, seasonal sequence and residual distribution feature sequence respectively. The component weight calculation unit generates the importance coefficients of the trend component, seasonal component and residual component. The three types of component feature sequences are weighted and fused according to the coefficients and then sent to the time series decoding unit to output the multi-step operation parameter prediction sequence of the power distribution equipment. The multi-scale input feature sequence is input into the trend and seasonality module. A trend extraction operation is performed on the multi-scale input feature sequence to generate a trend sequence. The trend sequence is then subtracted from the multi-scale input feature sequence to obtain a detrended sequence. Seasonal modeling operations are then performed on the detrended sequence to generate a seasonal sequence. Both the trend sequence and the seasonal sequence are output. The process of performing trend extraction on the multi-scale input feature sequence to generate a trend sequence is as follows: the multi-scale input feature sequence is divided into overlapping sliding windows on the time axis according to the length of the trend window. For each time position, several sampling points before and after the time position are taken to form a window. The values of the same feature dimension within the window are weighted and summed, and then divided by the sum of the weights to obtain the trend value corresponding to the time position. The weighted average calculation is repeated for all time positions. The trend values of each time position are arranged in chronological order to form a trend sequence with the same length as the multi-scale input feature sequence. The process of generating a seasonal sequence by performing seasonal modeling operations on the detrended sequence is as follows: the seasonal cycle length is determined according to the cyclical characteristics of the power distribution equipment operation; the detrended sequence is divided into several complete cycle segments on the time axis according to the cycle length; for the same time position within each cycle, the detrended values at the current position in all cycle segments are collected; the arithmetic mean of these values is calculated to obtain the seasonal component corresponding to the cycle position; the seasonal components at each position are arranged in the order of the time positions within the cycle to obtain a seasonal template for one cycle; the seasonal template is then repeatedly expanded on the time axis according to the cycle length and aligned with the detrended sequence at each time position to form a complete seasonal sequence. The residual sequence is obtained by subtracting the trend sequence and seasonal sequence from the multi-scale input feature sequence. The residual sequence is then input into the residual distribution module. The residual sequence is divided into windows according to the window length. The statistics of mean, variance, skewness, kurtosis and quantile are calculated for the residual samples in each window. The statistics are arranged in a fixed order to generate the residual distribution feature vector. The residual distribution feature vector is then arranged in time order to generate the residual distribution feature sequence. The trend sequence, seasonal sequence, and residual distribution feature sequence are input into the component decoding module. Within the time window, the component weight calculation unit calculates the magnitude of change, stability, and deviation indices for the trend sequence, seasonal sequence, and residual distribution feature sequence, respectively, generating the importance coefficients for the trend component, seasonal component, and residual component. The trend sequence, seasonal sequence, and residual distribution feature sequence are then multiplied by their corresponding component importance coefficients and concatenated in chronological order to generate the decoded input feature sequence. The calculation of the variation magnitude index, stability index, and deviation index of the trend series, seasonal series, and residual distribution characteristic series respectively is as follows: For the trend series, seasonal series, and residual distribution characteristic series, the maximum value within the current window is subtracted from the minimum value, and the absolute average of the differences between adjacent time points within the window is calculated. The magnitude difference and the absolute average are weighted and added together as the magnitude of change index. The mean of the values within the window is calculated first, and then the square of the difference between the value at each time point and the mean is calculated and averaged. The variance obtained is transformed by monotonically decreasing to obtain the stability index. For trend series and seasonal series, the baseline mean is obtained at each time point during normal operation. The absolute value of the difference between the mean and the baseline mean is calculated point by point within the current window and the average is taken as the deviation. For residual distribution characteristic series, the mean, variance, skewness, kurtosis and quantile of the current window are subtracted from the corresponding statistics during normal operation, the absolute values are taken and summed according to weights as the deviation. The generation of the importance coefficients for the trend component, seasonal component, and residual component is specifically as follows: within each time window, the change magnitude index, stability index, and deviation index corresponding to the three sequences are added together according to their weights to obtain three comprehensive scores. The three comprehensive scores are then added together to obtain the total score. Each comprehensive score is then divided by the total score to obtain a normalization coefficient. These three normalization coefficients are defined as the importance coefficients for the trend component, seasonal component, and residual component, respectively.
[0026] In this embodiment, obtaining the multi-step predicted value includes: The decoding input feature sequence is input into the temporal decoding unit in the decoding module of the component, and the feature vectors of each time step in the decoding input feature sequence are read sequentially in the time dimension to generate the corresponding decoding hidden state sequence. Within the temporal decoding unit, temporal decoding operations are performed on the decoded hidden state sequence. For each prediction time step, a corresponding prediction feature vector sequence is output. The vectors of each time step in the prediction feature vector sequence are arranged in chronological order to form a prediction feature sequence. Specifically, the temporal decoding operation is performed as follows: Each hidden state vector in the decoded hidden state sequence is input sequentially in chronological order. The current hidden state vector is then summed with all hidden state vectors in the same sequence using a weighted summation. The weights are obtained by normalizing the similarity between the hidden states. The weighted summation result is added to the current hidden state vector and then fed into the feedforward transformation unit. In the feedforward transformation unit, linear transformation, nonlinear mapping, and another linear transformation are performed sequentially to output the predicted feature vector at the current time step. This process is repeated for all time steps in the decoded hidden state sequence. The predicted feature vectors obtained at each time step are then arranged in chronological order to form a predicted feature sequence. The predicted feature sequence is input into the output reconstruction unit. The output reconstruction unit performs linear transformation and dimension mapping on the predicted feature vector of each time step in the predicted feature sequence to obtain the predicted value sequence of the operating parameters of the power distribution equipment at multiple predicted time steps.
[0027] In this embodiment, the anomaly index for constructing point residual anomaly anomaly, interval residual anomaly, and distribution deviation anomaly includes: At each time step, the actual measured value of the operating parameter is subtracted from the corresponding predicted value to obtain the residual time series. The absolute value of the residual at each time step in the residual time series is taken to form a point residual constant series. The residual time series is divided by a sliding time window. The residual mean and residual variance are calculated in each window to form an interval residual constant series. The probability distribution parameters of residual time series are statistically analyzed on the historical data of normal operation. In the prediction stage, the difference between the current residual time series and the normal residual distribution is calculated as the distribution deviation anomaly degree. This anomaly degree index is constructed by combining it with the point residual difference normality sequence and the interval residual difference normality sequence.
[0028] In this embodiment, the output of abnormal alarm information includes: During the historical normal operation phase, an anomaly index sequence is recorded. For each anomaly index, the mean, standard deviation, and quantiles are calculated on the historical timeline. Based on the mean, standard deviation, and quantiles, upper and lower thresholds are set for point residual anomaly, interval residual anomaly, and distribution deviation anomaly, respectively, forming an adaptive discrimination threshold range. Specifically, the setting of the upper and lower thresholds is as follows: During the normal operation phase, the mean, standard deviation, first quantile, and ninety-ninth quantile of the point residual difference abnormality, interval residual difference abnormality, and distribution deviation anomaly were calculated for the entire time period. For each anomaly index, an initial upper limit value was obtained by adding three times the standard deviation to the mean, and an initial lower limit value was obtained by subtracting three times the standard deviation from the mean. The initial upper limit value was then compared with the ninety-ninth quantile, and the smaller value between the two was taken as the final upper limit threshold. The initial lower limit value was then compared with the first quantile, and the larger value between the two was taken as the final lower limit threshold. During the detection phase, the current anomaly index sequence is calculated in real time. Each anomaly value is compared with the corresponding adaptive discrimination threshold range. When any anomaly value exceeds the number of time steps, exceeds the upper threshold, or falls below the lower threshold, the corresponding power distribution equipment is marked as being in an abnormal operating state. An anomaly alarm message containing equipment identification, anomaly type, and time information is generated and sent to the operation and maintenance terminal.
[0029] Example 1: To verify the feasibility of this invention in practice, it was applied to a city power distribution company. A mixed-load distribution area containing ten distribution transformers, six ring main units, and four important feeders was selected within the company's jurisdiction. This area includes residential communities, commercial complexes, and small industrial users, resulting in significant load fluctuations. On-site, using existing distribution automation terminals and power quality monitoring devices, three months of operational data were continuously collected at a sampling period of one second, from 00:00 on January 1, 2025 to 24:00 on March 31, 2025. Monitored parameters included three-phase voltage, three-phase current, active power, reactive power, power quality indicators, and equipment temperature. Records of power outages and fault work orders from the same period were retrieved. Thirty-six confirmed equipment anomaly events and over three hundred maintenance and inspection records were used as a sample to evaluate the identification capabilities and early warning effects of different anomaly detection methods.
[0030] In the current distribution area scenario, the collected multi-source runtime time series data undergoes time alignment, denoising, interpolation, and normalization to construct preprocessed time series data. For the key monitoring quantities of each power distribution device, symplectic geometric mode decomposition is applied to decompose the original complex time series into multi-scale modal components. Multi-scale subsequences are obtained through modal filtering and group reconstruction, and aligned with key parameters of voltage, current, power quality, and temperature to construct multi-scale input feature sequences. These sequences are then fed into an improved ETSformer model, which includes a trend-seasonal module, a residual distribution module, and a component decoding module. On the same dataset, traditional fixed threshold methods, traditional time series prediction methods, machine learning-based classification methods, and deep learning-based prediction methods are run and compared with the method of this invention under the same alarm strategy. The alarm time, alarm level, and corresponding device are uniformly output for each detection.
[0031] In offline playback and online simulation of three consecutive months of data, the comparative results show that traditional fixed threshold methods generate more false alarms during periods of rapid load fluctuations and frequent switching of operating conditions, and are insensitive to slowly deteriorating anomalies. While schemes based on conventional time series modeling and deep learning have some predictive ability for single measurements, they are prone to instability in judging anomaly boundaries when dealing with complex time series data involving multi-source coupling, harmonic disturbances, and temperature changes. The method of this invention can provide early warnings in multiple anomaly events, and its response is relatively consistent across different equipment types, load levels, and operating periods. The changes in alarm curves within anomaly segments have a good correlation with the actual fault development process. Without increasing additional hardware investment, it provides maintenance personnel with clearer anomaly location data and reference points for handling opportunities.
[0032] Table 1. Performance Comparison of Different Methods in Power Distribution Equipment Anomaly Detection Tasks
[0033] As shown in Table 1, in terms of overall detection capability, both the detection accuracy and anomaly recall rate increase progressively with method complexity, with the method of this invention achieving the highest level in both categories. The detection accuracy improved from 90.1% for the fixed threshold method, 91.3% for the autoregressive method, and 92.7% for the random forest method, to 94.2% for the Long Short-Term Memory network and 95.1% for the original ETSformer method, with the method of this invention further improving to 97.2%. The anomaly recall rate increased sequentially from 78.3%, 82.5%, 85.9%, 89.6%, and 91.2%, reaching 94.8% for the method of this invention. The corresponding F1 score improved from 0.83, 0.86, 0.89, 0.92, and 0.94 to 0.96, and the area under the curve improved from 0.86, 0.88, 0.91, 0.94, and 0.96 to 0.985, indicating that the method of this invention outperforms the five comparative methods in terms of the overall balance between precision and recall.
[0034] From an error control perspective, the false alarm rate and false negative rate in the table reflect the stability of different methods in actual alarm scenarios. The false alarm rate gradually decreases from 7.9% for the fixed threshold method to 7.1% for the autoregressive method, 6.4% for the random forest method, 5.6% for the long short-term memory network method, and 4.8% for the original ETSformer method, with the method of this invention further reducing it to 3.1%. The false negative rate decreases sequentially from 21.7%, 17.5%, 14.1%, 10.4%, and 8.8%, with the method of this invention having the lowest rate at only 5.2%. It can be seen that under the same dataset and the same threshold strategy, this invention simultaneously reduces both false alarms and false negatives to the lowest levels among all methods, making it more suitable for the dual requirements of safety and reliability in power distribution scenarios.
[0035] Regarding timeliness and overhead, the average early warning time and average inference latency in the table reflect the trade-off between early warning effectiveness and computational burden. The average early warning time improved from 8.4 minutes for the fixed threshold method and 10.7 minutes for the autoregressive method to 15.3 minutes for random forest, 24.6 minutes for long short-term memory networks, and 27.8 minutes for the original ETSformer. The method of this invention achieves 32.7 minutes, the largest early warning time. Meanwhile, the average inference latency of each method is 3.2 ms, 4.5 ms, 6.8 ms, 8.9 ms, 9.1 ms, and 9.8 ms, respectively. The method of this invention only increases the latency by 0.7 ms compared to the original ETSformer, still keeping it within 10 ms. It achieves a longer early warning time while maintaining acceptable real-time response capability for online deployment.
[0036] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for detecting abnormality of power distribution equipment based on time-series transformer, characterized by, include: Collect multi-source runtime timing data of power distribution equipment, preprocess the multi-source runtime timing data to obtain preprocessed timing data; Symmetric geometric mode decomposition is performed on the target time series signal in the preprocessed time series data. Hamilton structure matrix is constructed and deformed symmetric mapping structure is applied to obtain symmetric geometric mode subspace. Mode coupling interference elimination structure is used to suppress and decouple the cross projection between modes, resulting in multiple mutually decoupled mode components. Modal components are filtered and grouped for reconstruction to form multi-scale subsequences, which are then combined with preprocessed time series data to obtain multi-scale input feature sequences. An improved ETSformer model is constructed, which uses the trend and seasonality module to extract trends and detrend multi-scale input feature sequences, uses the residual distribution module to perform window partitioning and statistical distribution feature mapping, and uses the component decoding module to calculate component importance coefficients to form decoded input features. The decoded input features are input into the temporal decoding unit of the decoding module of the component, and temporal decoding and output reconstruction are performed to obtain multi-step predicted values. The deviation is calculated based on the multi-step predicted values, and anomaly indices such as point residual anomaly, interval residual anomaly, and distribution deviation anomaly are constructed. Determine the adaptive threshold range for the anomaly index, determine that the corresponding power distribution equipment is in an abnormal operating state, and output anomaly alarm information.
2. The method of claim 1, wherein, The multi-source runtime sequence data includes voltage, current, active power, reactive power, power quality indicators, and temperature.
3. The method of claim 1, wherein, The obtained preprocessed time series data includes: By using voltage transformers, current transformers, temperature sensors, and power quality monitoring devices installed on power distribution equipment, time-stamped voltage sampling sequences, current sampling sequences, active power sampling sequences, reactive power sampling sequences, power quality index sampling sequences, and temperature sampling sequences are collected at fixed sampling periods to obtain multi-source runtime sequence data. The multi-source runtime sequence data is aligned according to time stamps, the sampled data from different measurement points are mapped to a unified time axis, data points with duplicate time stamps are deleted, and data points with missing time stamps are interpolated to complete the data points with missing time stamps, resulting in a multi-source runtime sequence data matrix with a consistent time axis. Abnormal data in the multi-source runtime sequence data matrix is removed, each physical quantity is normalized according to its own numerical range, and the normalized multi-source runtime sequence data is divided into sample sequences of fixed length according to time order to form preprocessed time series data.
4. The method of claim 1, wherein, The process of obtaining multiple mutually decoupled modal components includes: Select the target time series signal from the preprocessed time series data, and reconstruct the phase space of the target time series signal according to the set embedding dimension and time delay to construct the trajectory matrix; The Hamilton structure matrix is constructed using the trajectory matrix. The Hamilton structure matrix consists of four sub-block matrices and satisfies the symplectic structure constraint. The Hamilton structure matrix and the standard symplectic matrix satisfy the symplectic preservation relationship. Construct a deformed symplectic mapping structure, define a set of symplectic preserving transformation operators with parameters, perform left and right multiplication transformation operations on the Hamilton structure matrix to obtain the deformed symplectic geometric modal feature matrix, and form a feature vector set containing multiple candidate modal feature vectors; Multiple initial modal subspaces are generated based on the feature vector set. The cross-projection energy between any two different modal subspaces is calculated, and a modal cross-projection matrix is constructed. In the modal coupling interference cancellation structure, the modal cross-projection matrix is used as input to perform orthogonal projection constraint operation and energy redistribution operation, reduce the cross-projection components between different modal subspaces, output a set of decoupled modal subspaces, and restore each modal subspace to the corresponding modal component time sequence signal according to the diagonal averaging principle, thus obtaining multiple mutually decoupled modal components.
5. The method of claim 1, wherein, The process of obtaining the multi-scale input feature sequence includes: For each modal component, calculate the energy value, mean, variance, correlation coefficient with the target time-series signal, and dominant frequency characteristics to obtain the modal feature set; Based on the energy threshold, correlation coefficient threshold, and effective range of the dominant frequency, modal components with low energy and low correlation are removed from the modal feature set to obtain the effective modal component set; Based on the dominant frequency characteristics and energy distribution in the effective modal component set, the effective modal components are divided into multi-scale subsequences. The multi-scale subsequence set is generated by adding them point by point according to the time index and aligning them with the key monitoring quantities in the preprocessed time series data according to time. By splicing the multi-scale subsequence values and multi-source operating parameter values at each time step, a multi-scale input feature vector is constructed. The multi-scale input feature vector is arranged in chronological order to obtain the multi-scale input feature sequence.
6. The method of claim 1, wherein, The formation of the decoding input features includes: Construct an improved ETSformer model, including a trend and seasonality module, a residual distribution module, and a component decoding module; The multi-scale input feature sequence is input into the trend and seasonal module. The trend extraction operation is performed on the multi-scale input feature sequence to generate the trend sequence. The trend sequence is subtracted from the multi-scale input feature sequence to obtain the detrended sequence. The seasonal modeling operation is performed on the detrended sequence to generate the seasonal sequence. The trend sequence and the seasonal sequence are then output. The residual sequence is obtained by subtracting the trend sequence and seasonal sequence from the multi-scale input feature sequence. The residual sequence is then input into the residual distribution module. The residual sequence is divided into windows according to the window length. The statistics of mean, variance, skewness, kurtosis and quantile are calculated for the residual samples in each window. The statistics are arranged in a fixed order to generate the residual distribution feature vector. The residual distribution feature vector is then arranged in time order to generate the residual distribution feature sequence. The trend sequence, seasonal sequence, and residual distribution feature sequence are input into the component decoding module. The component weight calculation unit in the component decoding module calculates the change magnitude index, stability index, and deviation index of the trend sequence, seasonal sequence, and residual distribution feature sequence respectively within the time window, and generates the importance coefficient of the trend component, the importance coefficient of the seasonal component, and the importance coefficient of the residual component. The trend sequence, seasonal sequence, and residual distribution feature sequence are multiplied by the corresponding component importance coefficients and concatenated in time order to generate the decoded input feature sequence.
7. The method of claim 1, wherein, The process of obtaining multi-step prediction values includes: The decoding input feature sequence is input into the temporal decoding unit in the decoding module of the component, and the feature vectors of each time step in the decoding input feature sequence are read sequentially in the time dimension to generate the corresponding decoding hidden state sequence. Within the temporal decoding unit, temporal decoding operations are performed on the decoded hidden state sequence. For each prediction time step, the corresponding prediction feature vector sequence is output. The vectors of each time step in the prediction feature vector sequence are arranged in chronological order to form the prediction feature sequence. The predicted feature sequence is input into the output reconstruction unit. The output reconstruction unit performs linear transformation and dimension mapping on the predicted feature vector of each time step in the predicted feature sequence to obtain the predicted value sequence of the operating parameters of the power distribution equipment at multiple predicted time steps.
8. The method of claim 1, wherein, The anomaly indices for constructing point residual anomaly anomaly, interval residual anomaly, and distribution deviation anomaly include: At each time step, the actual measured value of the operating parameter is subtracted from the corresponding predicted value to obtain the residual time series. The absolute value of the residual at each time step in the residual time series is taken to form a point residual constant series. The residual time series is divided by a sliding time window. The residual mean and residual variance are calculated in each window to form an interval residual constant series. The probability distribution parameters of residual time series are statistically analyzed on the historical data of normal operation. In the prediction stage, the difference between the current residual time series and the normal residual distribution is calculated as the distribution deviation anomaly degree. This anomaly degree index is constructed by combining it with the point residual difference normality sequence and the interval residual difference normality sequence.
9. The method of claim 1, wherein, The output of abnormal alarm information includes: During the normal operation phase in history, anomaly index sequence is recorded. For each anomaly index, mean, standard deviation and quantile are calculated on the historical time axis. Based on the mean, standard deviation and quantile, upper limit threshold and lower limit threshold are set for point residual anomaly, interval residual anomaly and distribution deviation anomaly, respectively, to form an adaptive discrimination threshold range. During the detection phase, the current anomaly index sequence is calculated in real time. Each anomaly value is compared with the corresponding adaptive discrimination threshold range. When any anomaly value exceeds the number of time steps, exceeds the upper threshold, or falls below the lower threshold, the corresponding power distribution equipment is marked as being in an abnormal operating state. An anomaly alarm message containing equipment identification, anomaly type, and time information is generated and sent to the operation and maintenance terminal.