Transformer fault prediction method based on time-space-spectrum multi-source information fusion

Through time-space-spectral multi-source information fusion technology, combined with TV-ARMA, AWPT and WOA, DBN models are trained, and the problem of insufficient accuracy and generalization capabilities of transformer fault prediction in the existing technology is solved, and the fault prediction effect with high accuracy and dynamic adaptability is achieved.

CN120180276AActive Publication Date: 2025-06-20NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Patent Information

Application Number
CN202510488832.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-06-20
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The existing multi-source information fusion technology has significant defects in data preprocessing, feature extraction, model optimization and judgment rules in transformer fault prediction, and it is difficult to meet the needs of high precision, strong generalization and dynamic adaptability.

Method used

The method based on time-space-spectral multi-source information fusion is adopted to dynamically capture the changes of electrical parameters through time-varying autoregressive sliding averaging model (TV-ARMA), optimize spectrum feature extraction with adaptive wavelet packet transformation (AWPT), and use whale optimization algorithm (WOA) to train the deep confidence network (DBN) to achieve high-precision prediction of transformer failures.

Benefits of technology

It significantly improves the accuracy and reliability of identification of faults such as winding short circuit, core grounding and insulation aging, and meets the demand for high accuracy, strong generalization and dynamic adaptability of transformer fault prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180276A_ABST
    Figure CN120180276A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of transformer fault prediction, and discloses a transformer fault prediction method based on time-space-spectrum multi-source information fusion, and the key points of the technical scheme are as follows: data acquisition, data preprocessing, feature extraction, model construction and training, and fault prediction. Through fusion of time-space-spectrum multi-source information, a time-varying autoregressive moving average model (TV-ARMA) is adopted to dynamically capture electrical parameter changes, adaptive wavelet packet transform (AWPT) is combined to optimize spectrum feature extraction, and a whale optimization algorithm (WOA) is utilized to train a deep belief network (DBN), so that high-precision prediction of a transformer fault is realized. And the identification accuracy and reliability of faults such as winding short circuit, iron core grounding and insulation aging are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of transformer fault prediction, and more specifically, it relates to a transformer fault prediction method based on time-space-spectrum multi-source information fusion. Background Art

[0002] Although certain progress has been made in the application of multi-source information fusion technology in the field of transformer fault prediction, the existing multi-source information fusion technology has significant defects in aspects such as data preprocessing, feature extraction, model optimization, and decision rules, and it is difficult to meet the requirements of transformer fault prediction for high precision, strong generalization, and dynamic adaptability.

[0003] Therefore, the present invention provides a transformer fault prediction method based on time-space-spectrum multi-source information fusion, which improves the above technical problems. Summary of the Invention

[0004] The embodiments of the present disclosure aim to address the deficiencies of the prior art and provide a transformer fault prediction method based on time-space-spectrum multi-source information fusion. The present invention dynamically captures the changes in electrical parameters by fusing time-space-spectrum multi-source information and using a time-varying autoregressive moving average model (TV-ARMA), optimizes the extraction of spectral features by combining with adaptive wavelet packet transform (AWPT), and trains a deep belief network (DBN) using the whale optimization algorithm (WOA), achieving high-precision prediction of transformer faults and significantly improving the recognition accuracy and reliability of faults such as winding short circuits, core grounding, and insulation aging.

[0005] The above technical objectives of the present invention are achieved through the following technical solutions: A transformer fault prediction method based on time-space-spectrum multi-source information fusion includes the following steps: S1. Data collection, which includes: time information collection, space information collection, and spectrum information collection; S2. Data preprocessing, which includes: normalization processing and denoising operation; S3. Feature extraction, which includes: time feature extraction, space feature extraction, and spectrum feature extraction; S4. Model construction and training, which includes: constructing a deep belief network (DBN) and using the whale optimization algorithm (WOA) to optimize the weight parameters of the DBN; S5. Fault prediction, which includes: inputting test data into the trained DBN model, outputting the fault probability, and comprehensively judging winding short circuit, core grounding, and insulation aging faults in combination with the threshold decision rules of time, space, and spectrum features.

[0006] As a preferred technical solution of the present invention, the process of collecting time information is as follows: The voltage signal V(t), current signal I(t), and power factor PF(t) during the operation of the transformer are collected in real time through high-precision sensors, and the timestamps are recorded; The time-varying autoregressive moving average model (TV-ARMA) is used to dynamically model the voltage and current signals, and the power factor is calculated through an improved power factor formula; The process of collecting spatial information is as follows: Sensors are arranged on the transformer windings, iron cores, and tank walls to obtain the spatial distribution data of temperature, vibration, and partial discharge parameters; The spatial distribution characteristics are analyzed through the spatial autocorrelation function I(T) and the variogram γ(d); Through the probability density function, the spatial distribution probability of the partial discharge amount can be understood, and the high partial discharge area can be determined; The process of collecting spectral information is as follows: The spectral analysis is performed on the neutral point current, winding current, and partial discharge signal, and the variational mode decomposition (VMD) combined with the adaptive wavelet packet transform (AWPT) algorithm is used to extract spectral features. The extracted spectral features include: the proportion of band energy P n,m and the central frequency f n,m .

[0007] As a preferred technical solution of the present invention, the calculation formula of the spatial autocorrelation function I(T) is:

[0008]

[0009] where is the average value of all temperature data; The spatial weight matrix d ij is the sensor spacing, and σ is the weight decay parameter; The TV-ARMA model solves the time-varying coefficients a i (t) and c i (t) through the alternating direction multiplier method (ADMM), and is updated in real time to reflect the dynamic changes caused by load mutations.

[0010] As a preferred technical solution of the present invention, the process of normalization is as follows: Z-score normalization is performed on time, space, and spectral data respectively, and mapped to the interval [-1,1]; The process of denoising is as follows: The multi-scale morphological filtering algorithm is used, and the data is subjected to opening and closing operations using structural elements of different scales to remove noise.

[0011] As a preferred technical solution of the present invention, the multi-scale morphological filtering uses two-dimensional structural elements to process spatial data, and the corrosion and dilation operation formulas are:

[0012]

[0013] As a preferred technical solution of the present invention, the process of time feature extraction is as follows: extracting statistical features, autocorrelation coefficients, cross-correlation coefficients and power spectral densities of voltage, current and power factor; the statistical features include: mean value, variance, kurtosis and skewness; the process of spatial feature extraction is as follows: calculating temperature gradient Spatial difference σ of vibration intensity A , local discharge spatial positioning coordinates (x, y); the process of spectrum feature extraction is as follows: extracting characteristic frequency f char , band energy E i , center frequency Bandwidth B i and spectrum entropy value S.

[0014] As a preferred technical solution of the present invention, the search conditions for characteristic frequency f char are as follows:

[0015] and and X(f) > τ

[0016] wherein, and are the first derivative and the second derivative of spectrum data X(f) with respect to frequency f respectively, and τ is a preset threshold;

[0017] The calculation formula for band energy E i is as follows:

[0018]

[0019] wherein, k i and k i+1 are the discrete frequency point indexes corresponding to the frequency band [f i , f i+1 , and Δf is the frequency resolution;

[0020] The calculation formula for center frequency is as follows:

[0021]

[0022] The calculation formula for bandwidth B i is: B i = f high - f low

[0023] The calculation formula for spectrum entropy value S is:

[0024]

[0025] wherein, M is the number of discrete points of spectrum data, It represents the proportion of the energy at the i-th frequency point in the total energy.

[0026] As a preferred technical solution of the present invention, the process of constructing the deep belief network (DBN) is as follows: It is composed of multiple stacked restricted Boltzmann machine (RBM) layers. The number of nodes in the input layer matches the number of features, and the output layer corresponds to the fault type or probability. The process of using the whale optimization algorithm (WOA) to optimize the DBN weight parameters includes: initializing the population, fitness evaluation, simulating the hunting behavior, and iterative optimization until convergence.

[0027] As a preferred technical solution of the present invention, the output layer of the DBN uses the softmax function for multi-classification, and the formula is:

[0028]

[0029] Among them, o k is the output value of the k-th node in the output layer, representing the probability that the sample belongs to the k-th type of fault; is the weight connecting the j-th hidden layer unit of the last RBM to the k-th output layer unit, is the bias of the k-th output layer unit, and C is the total number of fault types;

[0030] The parameter update formula of the whale optimization algorithm (WOA) includes: Encircling behavior: X i (t + 1) = X * (t) - A · |C · X * (t) - X i (t)|; Spiral behavior: X i (t + 1) = X * (t) · e bl · cos(2πl) + X * (t); Random search behavior: X i (t + 1) = X rand - A · |C · X rand - X i (t)|.

[0031] As a preferred technical solution of the present invention, the fault determination rules in the fault prediction include: Winding short-circuit fault: The DBN output probability P w > 0.7, and the temperature gradient The energy of the frequency band E w > E w-thresh ; Core grounding fault: The DBN output probability P c > 0.6, and the power factor skewness |S PF | > S PF-thresh , the amplitude of the spectral characteristic frequency Ach ar-c >A c-th resh ; Insulation aging fault: the output probability P of DBN a >0.5, and the local discharge spectrum entropy value S p >S p-th resh , the change rate of power spectral density

[0032] In summary, the present invention has the following beneficial effects:

[0033] First, multi-dimensional information collaborative enhancement of diagnostic ability: The time-varying autoregressive moving average model (TV-ARMA) is used to dynamically capture the transient characteristics of voltage and current, solve the problem of poor adaptability of traditional fixed-parameter models to load mutations, and improve the sensitivity of abnormal state detection. By quantitatively analyzing the temperature and vibration distribution laws through the spatial autocorrelation function and variogram, the fault area is accurately located, avoiding misjudgment caused by sparse sensors in traditional methods. Based on variational mode decomposition (VMD) and adaptive wavelet packet transform (AWPT), features such as band energy and spectrum entropy are extracted to improve the spectral analysis accuracy of non-stationary signals and effectively suppress noise interference.

[0034] Second, efficient data preprocessing and feature fusion: Through Z-score normalization and multi-scale morphological filtering, the heterogeneity of multi-source data is eliminated, the feature fusion efficiency is improved, and the model training convergence speed is accelerated. Combining time-space-spectrum statistical features (such as kurtosis, temperature gradient, band energy) and dynamic features (such as time-varying coefficients, autocorrelation coefficients) comprehensively characterizes complex fault patterns and enhances the feature representation ability.

[0035] Third, intelligent model optimization and high-precision prediction: The whale optimization algorithm (WOA) is used to train the deep belief network (DBN), which improves the parameter tuning efficiency compared with the genetic algorithm, enhances the model generalization ability, and reduces the false alarm rate under complex working conditions such as load fluctuations and environmental temperature changes. Combining multi-dimensional threshold criteria (such as DBN output probability, temperature gradient, spectrum entropy) to comprehensively determine the fault type. Brief Description of the Drawings

[0036] Figure 1 It is a flowchart of the transformer fault prediction method based on time-space-spectrum multi-source information fusion provided by the embodiment of the present invention. Detailed Embodiments

[0037] The embodiments of the present disclosure aim at the problem that existing multi-source information fusion technologies have significant defects in aspects such as data preprocessing, feature extraction, model optimization, and decision rules, and it is difficult to meet the requirements of high precision, strong generalization, and dynamic adaptability for transformer fault prediction. In view of this, the embodiments of the present disclosure propose a transformer fault prediction method based on spatio-temporal-spectral multi-source information fusion. Through deep fusion of multi-source information, dynamic modeling optimization, and intelligent criterion design, it solves the technical bottlenecks of traditional methods in data heterogeneity processing, feature one-sidedness, and insufficient model generalization ability, and provides a highly reliable and low-cost transformer fault prediction solution for the power system.

[0038] Please refer to Figure 1 , Figure 1 which shows the flowchart of the transformer fault prediction method based on spatio-temporal-spectral multi-source information fusion according to the embodiments of the present disclosure. The overall process mainly includes the following 5 steps:

[0039] S1. Data acquisition.

[0040] S1.1 Acquisition of time information: Use high-precision sensors to collect voltage, current, and power factor during the operation of the transformer in real time, and accurately record the time stamps.

[0041] In an actual power system, the variation of these electrical parameters over time contains rich information and can reflect the operating state of the transformer.

[0042] For the collected voltage signal V(t) and current signal I(t), considering the harmonic interference and signal fluctuation characteristics existing in the actual power grid, a time-varying autoregressive moving average model (TV-ARMA) is used to describe their variation laws over time. The traditional autoregressive moving average model (ARMA) assumes that the model parameters are fixed. However, during the operation of the transformer, due to the influence of load changes, environmental factors, etc., the variation laws of electrical parameters are time-varying. The expression of the TV-ARMA model is:

[0043]

[0044] where a i (t), b j (t), c i (t), d j (t) are the time-varying autoregressive coefficients and moving average coefficients respectively, and their variation with time t reflects the dynamic characteristics of the variation laws of voltage and current signals; p and q are the autoregressive order and moving average order respectively; ∈(t) is a white noise sequence with a mean of 0, representing unpredictable random interference.

[0045] The power factor PF(t) can be calculated through the phase difference between voltage and current. Considering the errors in actual measurements and the time-varying characteristics of the signal, the following improved formula is adopted:

[0046]

[0047] where, is the phase difference between the voltage V(t) and the current I(t); ΔV(t) and ΔI(t) are the differences in voltage and current between the current moment and the previous moment, respectively, reflecting the change rates of voltage and current; and are the average values of voltage and current over a period of time, respectively; k and m are weighting coefficients adjusted according to the actual situation, used to correct measurement errors and consider the influence of signal changes on the power factor.

[0048] Through these formulas, the variation relationships of voltage, current, and power factor with time can be described more precisely. When collecting these time-series data, the accurate recording of timestamps is crucial. Using a high-precision clock synchronization system to accurately capture the minute changes in electrical parameters in a short time provides strong support for analyzing the transient changes in the operating state of the transformer.

[0049] When a load mutation occurs in the transformer, through the TV-ARMA model, the time-varying coefficients a i (t), c i (t), etc. of voltage and current can be observed to change rapidly, reflecting the adjustment of the signal change law; at the same time, using the improved power factor formula, the change of the power factor during load mutation can be calculated more accurately, providing a more accurate basis for judging the operating state of the transformer.

[0050] S1.2 Spatial information acquisition: Sensors are arranged on the windings, iron cores, and tank walls of the transformer to obtain the spatial distribution information of temperature, vibration, and partial discharge parameters, and determine the state differences at different positions.

[0051] To accurately describe the spatial distribution characteristics of these parameters, the concepts of Spatial Autocorrelation Function (SAF) and Variogram are introduced.

[0052] Taking the data collected by temperature sensors arranged at different positions of the transformer as an example for the temperature parameter. Let the position coordinates of the temperature sensor be (x i , y i ), i = 1, 2,..., n, and the temperature values at the corresponding positions be T(x i , y i ). The spatial autocorrelation function is used to measure the spatial correlation of temperature, and its formula is:

[0053]

[0054] wherein, is the average value of all temperature data; the spatial weight matrix which defines the spatial relationship between positions (x i , y i ) and (x j , y j ), σ is a parameter controlling the attenuation rate of the spatial weight, is the distance between two points. The value of I(T) ranges from -1 to 1. The closer the value is to 1, the stronger the positive correlation of temperature in space, that is, the more similar the temperature values at adjacent positions; the closer the value is to -1, the stronger the negative correlation; the value close to 0 indicates no correlation in space. By calculating the spatial autocorrelation function, it can be judged whether there is an aggregation phenomenon in the transformer temperature distribution. For example, if the value of I(T) is high in a certain area, it means that the temperature change in this area is relatively consistent, and there may be a risk of local overheating.

[0055] The variogram is used to describe the degree of temperature change with spatial distance, and the formula is:

[0056]

[0057] i+d , y i+d ) is the position at a distance d from (x i , y i ). The variogram γ(d) changes with the change of d. When d is small, if γ(d) grows slowly, it means that the temperature changes little in a short distance and the spatial continuity is good; on the contrary, if γ(d) grows rapidly, it means that the temperature changes greatly in a short distance and there may be local temperature anomalies. By analyzing the curve of the variogram, the spatial scale of temperature change can be determined, that is, to determine within what spatial range the temperature has similarity or difference, which is very important for locating the fault area.

[0058] For vibration parameters, a similar method can be adopted. Let the vibration amplitudes collected by vibration sensors at different positions (x m , y m ) be A(x m , y m ). In order to consider the propagation characteristics of vibration in different directions, an anisotropic variogram is introduced:

[0059]

[0060] where θ represents the direction; (x i+d,θ , y i+d,θ ) is the position at a distance d and in the direction θ from (x i , y i ); N θ (d) is the number of sample point pairs at a distance d in the direction θ. By calculating the anisotropic variogram in different directions, the propagation law and attenuation of vibration in different directions can be analyzed. For example, if γ θ (d) grows rapidly in a certain direction, it indicates that the change of vibration in this direction is large, and there may be structural defects or faults related to this direction.

[0061] For partial discharge parameters, due to their strong randomness and locality, a spatial distribution description based on probability density is adopted. Let the partial discharge quantity detected by the partial discharge sensor at the position (x k , y k ) be Q(x k , y k ), and the probability density function p(Q) of the partial discharge quantity is defined as:

[0062]

[0063] where δ is the Dirac function. When Q = Q(x k , y k ), δ(Q - Q(x k , y k )) = 1, otherwise it is 0; n is the total number of detected partial discharge events. Through the probability density function, the spatial distribution probability of the partial discharge quantity can be intuitively understood, and the high-incidence area of partial discharge can be determined. At the same time, combined with the analysis methods of spatial autocorrelation function and variogram, the spatial correlation and variation law of partial discharge can be further studied, providing a basis for judging the type and severity of partial discharge faults.

[0064] Through these formulas and methods, the spatial distribution information of transformer temperature, vibration, and partial discharge parameters can be deeply analyzed, and the operating state of the transformer can be more comprehensively understood from the spatial dimension, providing a richer basis for fault prediction.

[0065] S1.3 Spectrum information acquisition: Collect the neutral point current, winding current, and partial discharge signals of the transformer and perform spectrum analysis, and use the variational mode decomposition combined with the adaptive wavelet packet transform algorithm to obtain the spectrum characteristics of the signals in different frequency bands.

[0066] First, perform variational mode decomposition (VMD) on the collected signal x(t). VMD decomposes the signal into K modal components u k (t), and each modal component has its corresponding central frequency ωk . Its core is to solve a variational problem, and the Lagrangian expression of this variational problem is:

[0067]

[0068] where α is a quadratic penalty factor used to balance the signal reconstruction accuracy and the modal bandwidth; λ(t) is a Lagrange multiplier used to ensure the satisfaction of the constraint conditions; δ(t) is the Dirac function; * represents the convolution operation; ‖·‖2 represents the L2 norm.

[0069] By solving the above variational problem through the Alternating Direction Method of Multipliers (ADMM) (optimizing the time-varying coefficients through the ADMM algorithm), each modal component u k (t) can be obtained. The original signal x(t) is decomposed into K modal components with different center frequencies, realizing the preliminary separation of different frequency components of the signal.

[0070] Perform Adaptive Wavelet Packet Transform (AWPT) on each modal component u k (t). In the wavelet packet transform, the signal is recursively decomposed and reconstructed. Let be the scaling function and ψ(t) be the wavelet function. For the wavelet packet function W n,m (t) of the nth layer and the mth node, there is:

[0071]

[0072] where h(l) and g(l) are the coefficients of the low-pass filter and the high-pass filter respectively, and Z represents the set of integers.

[0073] In the Adaptive Wavelet Packet Transform, the optimal wavelet packet decomposition tree structure is selected according to the characteristics of the signal. Here, the information entropy criterion is introduced to select the optimal node. For a signal segment y(t) corresponding to a node, its information entropy E is defined as:

[0074]

[0075] During the wavelet packet decomposition process, calculate the information entropy of each node, and select the node with the minimum information entropy to continue the decomposition until the preset decomposition depth is reached or other stopping conditions are met. This can make the result of the wavelet packet decomposition more in line with the frequency characteristics of the signal and more accurately capture the characteristics of the signal in different frequency bands.

[0076] After the Adaptive Wavelet Packet Transform, the wavelet packet coefficients c n,m of each frequency band are obtained. In order to extract more representative spectral features, calculate the energy proportion P n,m of each frequency band, and the formula is:

[0077]

[0078] Energy proportion P n,m It reflects the proportion of each frequency band in the total signal energy and can highlight the main frequency components of the signal. At the same time, calculate the center frequency f of the frequency band n,m :

[0079]

[0080] where f(i) is the frequency value corresponding to the frequency sampling point. The center frequency f n,m can more accurately reflect the frequency position of each frequency band and help analyze the changes in the signal frequency components.

[0081] The spectral features calculated by the above variational mode decomposition combined with adaptive wavelet packet transform algorithm and related formulas, such as energy proportion and center frequency, can more comprehensively and accurately describe the characteristics of transformer neutral point current, winding current and partial discharge signals in different frequency bands. After these spectral features are fused with time information and space information, they can provide richer and more accurate basis for transformer fault prediction, improving the accuracy and reliability of fault prediction.

[0082] S2. Data preprocessing.

[0083] S2.1 Normalization processing: Adopt the Zscore normalization method to uniformly map time, space and spectral data with different ranges and dimensions to the interval [-1, 1], thus eliminating the influence of dimensions.

[0084] For the voltage, current and power factor data in time information, let the original data be X = {x1, x2, …, x n}, and its mean is and the standard deviation is σ x . The Z-score normalization formula is:

[0085]

[0086] where z i is the normalized data value. Through this formula, the original data is standardized to have zero mean and unit variance. Transform z i to map it to the interval [-1, 1], and the transformation formula is:

[0087]

[0088] Here, min(z) and max(z) are the minimum and maximum values in the data z after preliminary Z-score normalization respectively. z i' is the time information data finally mapped to the interval [-1, 1]. This processing method not only eliminates the dimensionality influence of the time information data, but also makes different data on the same comparable scale, facilitating subsequent analysis. For example, the voltage data may vary between several hundred volts and several thousand volts, and the current data varies between several amperes and several hundred amperes. After this normalization processing, they are numerically comparable and can better participate in subsequent model operations.

[0089] For the temperature, vibration, and partial discharge parameter data in the spatial information, similar processing is also carried out. Taking the temperature data as an example, let the original temperature data set be T = {t1, t2, …, t m}, first calculate its mean value and standard deviation σ t :

[0090]

[0091] Then perform preliminary Z-score normalization:

[0092]

[0093] Then map it to the interval [-1, 1]:

[0094]

[0095] Among them, is the temperature value after preliminary Z-score normalization, min(t z ) and max(t z ) are the minimum and maximum values of the temperature data after preliminary normalization, is the temperature data finally normalized to the interval [-1, 1]. In this way, the temperature data distributed in space is unified to the same scale, enabling the temperature data at different positions to be directly compared and analyzed, which is beneficial to exploring the spatial characteristics of the temperature data.

[0096] For the spectrum information, let the wavelet packet coefficient data of a certain frequency band obtained after variational mode decomposition and adaptive wavelet packet transform be C = {c1, c2, …, c k}, first perform Z-score normalization on it:

[0097]

[0098] Among them, is the mean value of the wavelet packet coefficients, σ c is the standard deviation, which are calculated respectively through the following formulas:

[0099]

[0100] Next, to map it to the interval [-1, 1], the transformation is used:

[0101]

[0102] Here, is the wavelet packet coefficient value after preliminary Z - score normalization, min(c z ) and max(c z ) are the minimum and maximum values of the wavelet packet coefficient data after preliminary normalization, is the spectral information data finally normalized to the interval [-1, 1]. After such processing, the spectral feature data of different frequency bands are expressed on the same scale, enhancing the comparability between spectral features and helping to more accurately analyze the relationship between spectral information and transformer faults.

[0103] Through the above - mentioned detailed Z - score normalization and mapping processing, the dimensional differences of time, space, and spectral data can be effectively eliminated, providing a standardized and comparable data basis for the subsequent transformer fault prediction model based on multi - source information fusion, and improving the performance of the model and the accuracy of fault prediction.

[0104] S2.2 Denoising operation: Adopt the multi - scale morphological filtering algorithm to remove the noise and interference in the data and improve the purity of the data.

[0105] Multi - scale morphological filtering is based on the principle of mathematical morphology. By using structural elements of different scales to process the data, the noise of different scales can be effectively removed. Mathematical morphology mainly has two basic operations: erosion and dilation, and on this basis, opening operation and closing operation are derived.

[0106] For the discrete one - dimensional data sequence x(n), n = 1, 2, …, N (here taking the voltage data in time information as an example, and other data processing methods are similar), the formula for the erosion operation is:

[0107]

[0108] where b(m) is the structural element and Z is the set of integers. The erosion operation can remove the small protrusions in the signal and suppress the positive pulse noise.

[0109] The formula for the dilation operation is:

[0110]

[0111] The dilation operation is the opposite of erosion and is used to fill the small holes in the signal and suppress the negative pulse noise.

[0112] The opening operation is to erode first and then dilate, and its formula is:

[0113]

[0114] Opening operation can remove isolated noise points and tiny interferences in the signal, smooth the contour of the signal, and retain the main features of the signal.

[0115] Closing operation is dilation followed by erosion, and the formula is:

[0116]

[0117] Closing operation can fill tiny gaps in the signal, connect broken parts, and also play a role in smoothing the signal.

[0118] In multi-scale morphological filtering, data is processed using multiple structuring elements of different scales. Let the scale of the structuring element be s, and for structuring elements b s (m) of different scales, multi-scale opening and closing operations are performed on the data x(n). For example, first perform a small-scale opening operation on the data:

[0119]

[0120] where s1 represents a smaller scale, is the data after the small-scale opening operation. The small-scale opening operation can remove noise and detail interferences at a smaller scale and retain the basic features of the signal.

[0121] Then perform a large-scale closing operation:

[0122]

[0123] Here s2 > s1, and s2 represents a larger scale, is the data after the large-scale closing operation. The large-scale closing operation can further smooth the signal, remove noise and interferences at a larger scale, and at the same time retain the overall trend of the signal.

[0124] By combining such multi-scale opening and closing operations, different-scale noises in the data can be comprehensively processed. Small-scale operations focus on removing detail noises, while large-scale operations are used to handle more macroscopic interferences, thereby improving the purity of the data.

[0125] For two-dimensional spatial information data, such as the distribution data T(x i ,y i ) of temperature at different positions of a transformer, the structuring element becomes a two-dimensional form b(x m ,y m ), and the corresponding formulas for erosion, dilation, opening operation, and closing operation are extended in the two-dimensional space. For example, the formula for two-dimensional erosion operation is:

[0126]

[0127] Among them, S is the area covered by the structural element b. The formulas for two-dimensional dilation, opening operation, and closing operation are also extended according to a similar logic to meet the processing requirements of two-dimensional data:

[0128]

[0129] Process the spatial data with two-dimensional structural elements of different scales to remove noise and outliers in the spatial distribution and better display the spatial characteristics of the data.

[0130] For spectral information, due to the different data characteristics from time and spatial information, when performing multi-scale morphological filtering, the spectral data is regarded as a discrete frequency sequence X(f j ), j = 1, 2, …, M. The above formulas for erosion, dilation, opening operation, and closing operation are also applied. However, the design of the structural element b(f k ) needs to be adjusted according to the frequency characteristics of the spectrum. For example, a Gaussian-type structural element can be designed, and its expression in the frequency domain is:

[0131]

[0132] Among them, σ controls the scale of the structural element. The larger σ is, the wider the coverage range of the structural element in the frequency domain. Perform multi-scale morphological filtering on the spectral data with such structural elements of different scales, which can remove the noise frequency bands in the spectrum, highlight the useful spectral characteristics, and provide more accurate data for subsequent analysis based on spectral characteristics.

[0133] Through the above multi-scale morphological filtering algorithm, use structural elements of different scales to specifically process time, space, and spectral information data, effectively remove various noises and interferences, improve the data quality, and provide more reliable data support for subsequent transformer fault prediction.

[0134] S3. Feature extraction.

[0135] S3.1 Time feature extraction: Extract statistical features, correlation features, and frequency domain features from the time series data. Here, the statistical features include mean, variance, kurtosis, and skewness, the correlation features include autocorrelation coefficient and cross-correlation coefficient, and the frequency domain features include power spectral density, so as to obtain the time variation law of the data.

[0136] 1) Statistical features

[0137] <1> Mean: For the voltage sequence V(n), n = 1, 2, …, N, its mean The calculation formula is:

[0138]

[0139] The mean value reflects the average voltage level during transformer operation and can be used as a benchmark to measure voltage stability. For example, if the mean voltage value deviates greatly from the rated voltage over a period of time, it may indicate that the transformer is operating abnormally.

[0140] For the current sequence I(n), the mean for:

[0141]

[0142] The current mean reflects the average size of the transformer load current, which helps to understand the load condition of the transformer.

[0143] The mean value of the power factor series PF(n)

[0144]

[0145] The mean power factor reflects the average level of the transformer's power utilization efficiency and is of great significance for evaluating the transformer's operating economy.

[0146] <2> Variance: Voltage variance It is used to measure the degree of voltage fluctuation around the mean value. The formula is:

[0147]

[0148] The larger the voltage variance, the more severe the voltage fluctuation, which may cause greater impact on the insulation and other components of the transformer, increasing the risk of failure.

[0149] Current variance

[0150]

[0151] The current variance reflects the stability of the current. Abnormal current fluctuations may be related to factors such as internal winding faults in the transformer and sudden load changes.

[0152] Power Factor

[0153]

[0154] The variance of the power factor can reflect its stability. An unstable power factor will affect the power transmission efficiency of the power grid and may be a signal of problems in the internal or external circuits of the transformer.

[0155] <3> Kurtosis: Voltage Kurtosis K VUsed to describe the sharpness of the voltage distribution, the formula is:

[0156]

[0157] If the voltage kurtosis value significantly deviates from the kurtosis value of the normal distribution (usually 3), it may indicate the presence of abnormal spikes or pulse signals in the voltage, which may be caused by harmonic interference in the power grid, partial discharge inside the transformer, etc.

[0158] Current kurtosis K I :

[0159]

[0160] Abnormal changes in current kurtosis can reflect abnormal conditions in the current signal, such as current mutations, increased harmonic components, etc., which may all be related to transformer faults.

[0161] Power factor kurtosis K PF :

[0162]

[0163] Changes in power factor kurtosis can help judge the stability of the power factor and whether there are abnormal fluctuations. An abnormal kurtosis value may imply abnormal changes in the power conversion efficiency of the transformer.

[0164] <4>Skewness: Voltage skewness S V Used to measure the degree of asymmetry of the voltage distribution, the formula is:

[0165]

[0166] Voltage skewness can reflect whether there are asymmetric fluctuations in the voltage signal. For example, if the voltage skewness is not zero, it may indicate the presence of a certain DC component or other asymmetric interference factors in the voltage.

[0167] Current skewness S I :

[0168]

[0169] Current skewness helps to detect the asymmetric characteristics of the current distribution, and this asymmetry may be related to factors such as the asymmetry of the transformer winding and the imbalance of the load.

[0170] Power factor skewness S PF :

[0171]

[0172] The variation of power factor skewness can reflect the asymmetry of power factor distribution, which has reference value for analyzing the power conversion characteristics and potential faults of transformers.

[0173] 2) Correlation characteristics

[0174] <1> Autocorrelation coefficient: The autocorrelation coefficient r of voltage V (k) is used to measure the correlation between the voltage sequence and itself at different time delays k, and the formula is:

[0175]

[0176] By analyzing the variation of the voltage autocorrelation coefficient with the time delay k, the periodicity and memory of voltage fluctuations can be understood. For example, if the voltage autocorrelation coefficient is high at a certain specific time delay k, it indicates that the voltage has certain similarity and periodicity within this time interval, which may be related to the periodic load changes of the power grid or the inherent characteristics of the transformer itself.

[0177] The autocorrelation coefficient of current r I (k):

[0178]

[0179] The autocorrelation coefficient of current can reflect the time correlation of the current signal and help judge the law and stability of current fluctuations.

[0180] The autocorrelation coefficient of power factor r PF (k):

[0181]

[0182] The autocorrelation coefficient of power factor helps analyze the change trend and stability of the power factor, and has certain reference value for predicting the future change of the power factor.

[0183] <2> Cross-correlation coefficient: The cross-correlation coefficient r of voltage and current VI (k) is used to measure the correlation between voltage V(n) and current I(n), and the formula is:

[0184]

[0185] The cross-correlation coefficient can help analyze the phase relationship and mutual influence between voltage and current. For example, under normal operating conditions, there is a certain phase difference between voltage and current, and the cross-correlation coefficient can quantify this relationship. When the cross-correlation coefficient shows abnormal changes, it may indicate that the load characteristics of the transformer have changed or there are faults.

[0186] The cross-correlation coefficient of voltage and power factor r VPF (k):

[0187]

[0188] This cross - correlation coefficient can reflect the influence of voltage changes on the power factor and the potential relationship between the two, which is of great significance for analyzing the power conversion efficiency and operating status of the transformer.

[0189] The cross - correlation coefficient r between current and power factor IPF (k):

[0190]

[0191] By analyzing this cross - correlation coefficient, the relationship between current changes and the power factor can be understood, and whether the load characteristics and power conversion efficiency of the transformer are normal can be judged.

[0192] 3) Frequency - domain characteristics - Power spectral density: The voltage power spectral density P V (f) is used to describe the distribution of the power of the voltage signal at different frequencies and is calculated by the Welch method. The formula is:

[0193]

[0194] where M is the number of segments, and V i (f) is the Fourier transform of the i - th segment of the voltage sequence V(n). The voltage power spectral density can clearly show the energy distribution of the voltage in different frequency bands. If there are abnormal peaks in the power spectral density in a certain specific frequency band, it may indicate the existence of voltage problems related to that frequency, such as harmonic interference, etc.

[0195] The current power spectral density P I (f):

[0196]

[0197] where I i (f) is the Fourier transform of the i - th segment of the current sequence I(n). The current power spectral density helps to detect the harmonic components and frequency characteristics in the current and plays an important role in judging the load nature and operating status of the transformer.

[0198] The power - factor power spectral density P PF (f):

[0199]

[0200] where PF i(f) is the Fourier transform of the i-th segment of the power factor sequence PF(n). The power factor power spectrum density can reflect the energy distribution of the power factor at different frequencies, help analyze the frequency characteristics and stability of the power factor, and provide valuable information for transformer operation monitoring and fault prediction.

[0201] Through the above-mentioned specific calculation formulas for voltage, current, and power factor data, the time variation patterns in time series data can be more comprehensively and deeply explored, providing powerful time dimension information support for transformer fault prediction.

[0202] S3.2 Spatial feature extraction: Based on the spatially distributed sensor data, the temperature gradient, spatial differences in vibration intensity, and spatial positioning of partial discharge are calculated to obtain the spatial distribution characteristics that reflect the internal state of the transformer.

[0203] 1) Temperature gradient calculation: The temperature gradient is used to measure the severity of temperature changes between different positions of the transformer, reflecting the direction and rate of heat transfer. i ,y i ) and (x j ,y j ) of two temperature sensors, whose temperature values ​​are T(x i ,y i ) and T(x j ,y j ), temperature gradient vector The formula for calculating the components in two-dimensional space is:

[0204] In the x-direction: (When x j -x i ≠0 hours)

[0205] In the y direction: (When y j -y i ≠0 hours)

[0206] Combining the x and y directions, the magnitude of the temperature gradient is:

[0207] Temperature gradient direction: (The gradient direction is determined by the inverse tangent function)

[0208] The calculation of temperature gradient helps to find the heat concentration area and heat propagation path inside the transformer. For example, in the transformer winding, if the temperature gradient in a certain area is large, it means that the temperature changes drastically in this area, and there may be poor heat dissipation or local overheating problems, which may lead to accelerated insulation aging and increase the risk of failure.

[0209] 2) Calculation of spatial differences in vibration intensity: The spatial differences in vibration intensity can be measured by calculating the standard deviation of vibration amplitudes at different positions. Let the vibration amplitudes collected by vibration sensors at n different positions (x m , y m ) be A(x m , y m ), where m = 1, 2, …, n. First, calculate the mean value of vibration amplitudes at all positions

[0210]

[0211] Then calculate the spatial difference (standard deviation) σ of vibration intensity A :

[0212]

[0213] A larger spatial difference in vibration intensity means that the vibration conditions at different positions of the transformer vary greatly, which may imply inhomogeneity or local defects in the internal structure. For example, when the spatial difference in vibration intensity in a certain area is significantly higher than that in other areas, it may indicate that there are loose components or structural damages in that area, which will affect the mechanical stability of the transformer and may lead to more serious faults in the long run.

[0214] 3) Spatial localization calculation of partial discharge: Hyperbolic localization method can be used for spatial localization with multiple partial discharge sensors. Assume there are three partial discharge sensors with positions (x1, y1), (x2, y2) and (x3, y3) respectively, and the times when they detect partial discharge signals are t1, t2 and t3. Given that the propagation speed of the partial discharge signal is v, then the signal propagation distances are r1 = v(t1 - t0), r2 = v(t2 - t0), r3 = v(t3 - t0) (t0 is the starting time of partial discharge, which can be estimated by signal processing methods).

[0215] According to the definition of a hyperbola, for two sensors, such as sensor 1 and sensor 2, there is:

[0216] |r2 - r1| = 2a 12 (2a 12 is the length of the real axis of the hyperbola)

[0217] |r3 - r1| = 2a 13 (2a 13 is the length of the real axis of another hyperbola)

[0218] By establishing the hyperbola equation:

[0219]

[0220] By simultaneously solving these two hyperbola equations, the coordinates (x, y) of the local power source can be obtained. In actual calculations, numerical methods such as iterative algorithms or least squares methods can be used to solve, so as to improve the positioning accuracy. Accurate spatial positioning of partial discharges is crucial for judging the internal insulation condition of the transformer and determining the fault location, which helps to take targeted measures in a timely manner to prevent the further expansion of the fault.

[0221] Through the above calculations of the spatial differences in temperature gradient, vibration intensity, and the spatial positioning of partial discharges, the spatial distribution characteristics of the internal state of the transformer can be comprehensively obtained, providing key spatial dimension information for transformer fault prediction.

[0222] S3.3 Spectrum feature extraction: Extract characteristic frequencies, band energies, center frequencies, bandwidths, and spectral entropy values from the spectrum data, so as to analyze the relationship between the spectral characteristics of the signal and the operating state of the transformer.

[0223] 1) Characteristic frequency extraction: The characteristic frequency refers to the frequency component that has special significance in the spectrum and can characterize the operating state of the transformer. Through the spectrum analysis of the neutral point current, winding current, and partial discharge signal of the transformer, a peak search algorithm is used to determine the characteristic frequency. Let the spectrum data of a certain frequency band obtained after variational mode decomposition and adaptive wavelet packet transform be X(f). In the frequency range [f min , f max , the search condition for the characteristic frequency f ch ar is defined as:

[0224] and and X(f) > τ

[0225] where and are the first derivative and second derivative of the spectrum data X(f) with respect to the frequency f respectively, which are used to determine the extreme points of the spectrum curve; τ is a preset threshold, which is used to screen out the peak frequencies with larger amplitudes and practical significance. The frequency f that satisfies the above conditions is the characteristic frequency f ch ar .

[0226] For example, when a core fault occurs in the transformer, harmonics of specific frequencies will be generated, and these harmonic frequencies appear as peaks with higher amplitudes in the spectrum. Through the above characteristic frequency extraction method, these characteristic frequencies related to the fault can be accurately identified, providing an important basis for fault diagnosis.

[0227] 2) Band energy calculation: The band energy reflects the energy size contained in the signal within a specific frequency interval, and is of great significance for analyzing the frequency component distribution of the transformer signal. The entire spectrum is divided into multiple frequency bands [f i , f i+1, where \(i = 1, 2, \ldots, N\), and the band energy \(E\) i is calculated by the formula:

[0228]

[0229] In actual calculation, since the spectrum data is usually discrete, the trapezoidal integration method is used for approximate calculation:

[0230]

[0231] where \(k\) i and \(k\) i+1 are respectively the discrete frequency point indices corresponding to the frequency band \([f\) i , \(f\) i+1 , and \(\Delta f\) is the frequency resolution.

[0232] Different fault types will cause changes in the energy distribution of the signal in different frequency bands. For example, a winding short - circuit fault may cause a significant increase in the energy of certain specific frequency bands. By monitoring the changes in the band energy, the abnormality of the transformer operation state can be detected in a timely manner.

[0233] 3) Center frequency calculation: The center frequency is used to describe the central position of the frequency distribution within the frequency band and can more accurately reflect the frequency characteristics of the frequency band. For each frequency band \([f\) i , \(f\) i+1 , the center frequency is calculated by the formula:

[0234]

[0235] Similarly, in the case of discrete data, a numerical calculation method is adopted:

[0236]

[0237] The change in the center frequency can reflect the shift of the signal frequency components. When a fault occurs inside the transformer, the frequency components of the signal will change, thereby causing a change in the center frequency. By monitoring the change trend of the center frequency, it can be effectively judged whether the operation state of the transformer is normal.

[0238] 4) Bandwidth calculation: The bandwidth is used to measure the width of the frequency band and reflects the degree of dispersion of the signal frequency components. For the frequency band \([f\) i , \(f\) i+1 , the bandwidth \(B\) i can be calculated using the half - power bandwidth method, that is, finding the two frequencies \(f\) low and \(f\) high (satisfying \(f\) low \(\leq f\) high),then the bandwidth B i is:

[0239] B i = f high - f low

[0240] The change in bandwidth is closely related to the type and severity of transformer faults. For example, partial discharge faults may cause the bandwidth of the signal to widen because partial discharges generate multiple frequency components, making the frequency distribution more dispersed. By monitoring the change in bandwidth, valuable information can be provided for fault diagnosis.

[0241] 5) Spectrum entropy value calculation: The spectrum entropy value is an index that measures the complexity and uncertainty of the spectrum and can reflect the degree of uniformity of the distribution of signal frequency components. For the spectrum data X(f), the calculation formula for the spectrum entropy value S is based on the information entropy theory and uses the following formula:

[0242]

[0243] where M is the number of discrete points of the spectrum data, represents the proportion of the energy at the i-th frequency point to the total energy.

[0244] Under normal operating conditions, the spectrum entropy value of the transformer is relatively stable. When a transformer fails, the spectrum components become more complex, resulting in a change in the spectrum entropy value. For example, during the insulation aging process, the spectrum entropy value may gradually increase. By monitoring the change in the spectrum entropy value, potential fault hazards can be detected in advance.

[0245] By extracting the above spectrum features, namely characteristic frequency, band energy, center frequency, bandwidth, and spectrum entropy value, and combining time information and space information for comprehensive analysis, the operating state of the transformer can be understood more comprehensively and deeply, providing strong support for transformer fault prediction.

[0246] S4. Model construction and training.

[0247] S4.1 Construct a deep belief network (DBN): It is composed of multiple stacked restricted Boltzmann machines (RBMs), plus an output layer. The number of nodes in the input layer is determined according to the number of extracted features, and the number of nodes in the output layer corresponds to different fault types or fault probability levels of the transformer.

[0248] The restricted Boltzmann machine (RBM) is an energy-based undirected graph model, consisting of a visible layer and a hidden layer. There are no connections between the nodes within the layer, only connections between the nodes of different layers. For an RBM with n v visible units and n h hidden units, its energy function is defined as:

[0249]

[0250] where v i represents the state of the i-th unit in the visible layer (usually taking values of 0 or 1), and h j represents the state of the j-th unit in the hidden layer (usually taking values of 0 or 1), a i is the bias of the visible layer unit i, and b j is the bias of the hidden layer unit j, and w ij is the weight connecting the visible layer unit i and the hidden layer unit j.

[0251] Based on the energy function, the joint probability distribution of the RBM is:

[0252]

[0253] where Z = ∑ v ∑ h exp(-E(v, h)) is the partition function, which is used to normalize the probability distribution. Since the exact calculation of the partition function is very difficult in high-dimensional cases, the Contrastive Divergence (CD) algorithm is usually used for approximate training.

[0254] In the DBN, the hidden layer output of the previous RBM serves as the visible layer input of the next RBM. For the l-th RBM (l = 1, 2,..., L - 1, where L is the total number of RBMs), let its visible layer input be v (l) , and the hidden layer output be h (l) .

[0255] The activation probability from the visible layer to the hidden layer is calculated as follows:

[0256]

[0257] where is the sigmoid activation function, is the number of nodes in the visible layer of the l-th RBM, and are the bias of the hidden layer unit j and the weight connecting the visible layer unit i and the hidden layer unit j in the l-th RBM, respectively. This formula represents the probability that the hidden layer unit j is activated (taking the value of 1) given the visible layer state v (l) , which depends on the bias of the hidden layer unit, the connection weight, and the state of the visible layer unit.

[0258] The reconstruction probability from the hidden layer to the visible layer is calculated as:

[0259]

[0260] wherein, is the bias of the i-th visible layer unit of the l-th RBM, is the number of nodes in the hidden layer of the l-th RBM. This formula is used to calculate the probability that the visible layer unit i is reconstructed as 1 when the hidden layer state h (l) is known.

[0261] During the training process, the parameters (weights w and biases a, b) of the RBM are adjusted by minimizing the reconstruction error. Taking the l-th RBM as an example, the contrastive divergence algorithm is adopted, and its goal is to minimize the following reconstruction error function:

[0262]

[0263] wherein, V is the set of all visible layer samples in the training dataset. The contrastive divergence algorithm approximately calculates the gradient by performing k steps (usually k = 1 or k = 2) of Gibbs sampling on the training samples, and then updates the parameters. The specific update formulas are as follows:

[0264]

[0265] wherein, ∈ is the learning rate, which controls the step size of parameter update; <·> data represents the expectation calculated based on the training data, and <·> recon represents the expectation calculated based on the reconstructed data. By continuously iteratively updating these parameters, the RBM can better learn the feature representation of the data.

[0266] After training all the RBMs layer by layer, the output of the hidden layer of the last RBM is connected to the output layer. Assuming that the output layer is a fully connected layer, for the multi-classification problem (corresponding to different transformer fault types), the softmax activation function is adopted in the output layer, and its calculation formula is:

[0267]

[0268] wherein, o k is the output value of the k-th node in the output layer, representing the probability that the sample belongs to the k-th type of fault; is the weight connecting the j-th hidden layer unit of the last RBM to the k-th output layer unit, is the bias of the k-th output layer unit, and C is the total number of fault types.

[0269] For the regression problem (corresponding to the fault probability level), a linear activation function can be adopted for the output layer, and the output value directly represents the predicted fault probability level. Through such a network structure and training method, DBN can effectively learn the complex relationship between the input data (data after time, space, and spectrum feature fusion) and the transformer fault type or fault probability level, providing strong model support for fault prediction.

[0270] S4.2 Optimize the DBN weight parameters using the Whale Optimization Algorithm (WOA); divide the data after preprocessing and feature extraction into a training set and a test set. Use the training set to train the DBN. During the training process, use the Whale Optimization Algorithm (WOA) to optimize the DBN weight parameters. The specific steps are as follows:

[0271] 1) Initialization: Randomly generate a group of whale individuals (solutions), and each individual represents a set of weight parameters of the DBN. Set the parameters of the algorithm, including the population size, the maximum number of iterations, and the search space range.

[0272] Let the whale population size be N. In this patent, to ensure that the Whale Optimization Algorithm (WOA) can search in a diverse initial solution space and increase the probability of finding the global optimal solution, randomly initialize the weights and biases of the restricted Boltzmann machines (RBMs) in each layer of the DBN.

[0273] For the l-th layer RBM in the DBN (l = 1, 2, …, L - 1, where L is the total number of RBM layers), the weight matrix W (l) of its elements ( is the number of visible layer nodes in the l-th layer RBM, is the number of hidden layer nodes), and the initialization formula is:

[0274]

[0275] where, The reason for setting the search space range in this way is that the magnitude of the weight values has a crucial impact on the training process and performance of the neural network. Smaller weight values help prevent overfitting in the initial stage of model training, making the updates of neurons more stable and facilitating the model to robustly learn data features; while larger weight values may lead to overly large neuron outputs, making it difficult for the model to converge during the training process. Setting the value range of the weights to [-1, 1] can not only ensure that the weights have sufficient diversity, giving the model the opportunity to explore different parameter combinations, but also limit the variation range of the weights, avoiding the adverse effects of overly large or small weight values on model training. rand() is a random number generation function uniformly distributed in the interval [0, 1]. The random numbers generated by this function enable the weights of each whale individual to be randomly initialized within the specified range, thus ensuring the diversity of the initial population and providing a rich initial solution space for the WOA to search for the global optimal solution in the subsequent search process.

[0276] For the bias vector a of the visible layer of the l-th layer RBM (l) of the elements and the bias vector b of the hidden layer (l) of the elements The initialization formulas are respectively:

[0277]

[0278] Among them, Bias provides additional learnable parameters for neurons in the neural network, helping the model better fit the data. Setting the search space of the bias to [-1, 1] can avoid the excessive influence of overly large or small bias values on neuron outputs, ensuring that the model can reasonably adjust the bias during the training process, and thus optimizing the performance of the model. Randomly generating bias values within the corresponding range through the rand() function further increases the diversity of the initial solutions, enabling the algorithm to search from different starting points in the initial stage of training and increasing the probability of finding the global optimal solution.

[0279] Set the maximum number of iterations to T = 300 times. The maximum number of iterations is a key parameter in WOA, which determines the number of iteration rounds in the process of the algorithm searching for the optimal solution. If the maximum number of iterations is set too small, the algorithm may not be able to fully explore the solution space, making it difficult to find the global optimal solution or a parameter combination close to the global optimal solution, resulting in a low prediction accuracy of the model. Since the transformer fault prediction problem involves multi-source information fusion and the data features are complex, and the DBN contains multiple RBM layers, the parameters of each layer need to be finely tuned. Too few iteration times will cause the algorithm to fail to fully mine the potential laws in the data. For example, if the number of iterations is only a few dozen times, WOA may stop searching before fully exploring the complex relationship between the data features and transformer faults, making the trained DBN model unable to fit the data well, thereby affecting fault prediction.

[0280] 2) Fitness evaluation: For each whale individual, apply the corresponding weight parameters to the DBN, train it using the training data, and calculate the mean square error on the validation data. Take the reciprocal of the mean square error as the fitness value. The higher the fitness value, the better the performance of the DBN corresponding to this set of weight parameters.

[0281] Let the training set be where x m is the input sample containing time, space, and spectrum fusion features, and y m is the corresponding true label (which can be the encoding of the fault type or the fault probability level). The validation set is

[0282] After applying a set of weight parameters represented by a whale individual to the DBN, the predicted output of the DBN for the input sample x n is The mean square error (MSE) is used to measure the error degree between the predicted value and the true value, and its calculation formula is:

[0283]

[0284] The meaning of this formula is to square the difference between the predicted value and the true value of each sample in the validation set, and then find the average of these squared differences. The squaring operation is to avoid the cancellation of positive and negative errors, making the measurement of errors more reasonable. The smaller the value of the mean square error, the closer the predicted value of the model is to the true value, and the better the prediction performance of the model.

[0285] Take the reciprocal of the mean square error as the fitness value Fitness, that is:

[0286]

[0287] The reciprocal of the mean squared error is used as the fitness value because in the optimization process, we hope to find the combination of weight parameters that minimizes the model prediction error. The smaller the mean squared error, the larger its reciprocal. In the whale optimization algorithm, a larger fitness value represents a better solution. In this way, the prediction performance of the model on the validation set is quantified as a fitness value, which facilitates the whale optimization algorithm to judge the quality of each whale individual (i.e., each group of weight parameters) according to the size of the fitness value during the iteration process, guiding the algorithm to search in the direction of finding better weight parameters, continuously improving the prediction accuracy of the DBN model, and thus better serving the transformer fault prediction task.

[0288] 3) Hunting behavior simulation: Simulate the behaviors of humpback whales such as surrounding prey, spiraling upward, and random searching, and update the position of the whale individual (i.e., the weight parameters of the DBN). By adjusting the position, the whale individual gradually approaches the optimal solution.

[0289] In the whale optimization algorithm, when hunting for prey, humpback whales will surround the prey. Assume that the position vector of the current whale individual (corresponding to the weight parameter vector of the DBN) is X i and the optimal position vector found in the current iteration (corresponding to the optimal combination of weight parameters) is X * . First, introduce the parameter a, which linearly decreases from 2 to 0 as the iteration number t increases. The formula is:

[0290]

[0291] where t is the current iteration number and T is the set maximum iteration number. a plays a key role in controlling the search range and convergence speed in the algorithm. As the iteration progresses, a gradually decreases, causing the search range of the whales to gradually shrink, and the algorithm gradually converges to the optimal solution.

[0292] Calculate the coefficient vectors A and C. The formulas are respectively:

[0293] A = 2a·r1 - a

[0294] C = 2·r2

[0295] where r1 and r2 are random vectors in the range [0, 1], and are regenerated each time of calculation. A and C are used to determine the way and direction of updating the position of the whale individual. The A vector, combined with the value of a, in the early stage of the algorithm, due to a being larger, the value range of A is larger, enabling the whale individual to search within a larger range and explore different solution space regions; in the later stage of the algorithm, a decreases, the value range of A also becomes smaller, and the search range of the whale individual gradually focuses near the optimal solution. The C vector, by taking random values, increases the randomness of the search process and prevents the algorithm from falling into a local optimal solution.

[0296] According to the relationship between the absolute value of A and 1, the whale has two different behavioral patterns. When |A| < 1, the whale executes the behavior of surrounding the prey, and the formula for updating the position is:

[0297] X i (t + 1) = X * (t) - A·|C·X * (t) - X i (t)|

[0298] This formula indicates that the individual whale approaches the current optimal position, and the degree and direction of approach are jointly determined by A, C, and the difference between the current individual position and the optimal position. The whale takes the current optimal position as a reference point and adjusts its own position according to the calculation results of A and C, gradually reducing the distance from the optimal solution.

[0299] When |A| ≥ 1, the whale executes the random search behavior, and the formula is:

[0300] X i (t + 1) = X rand -A·|C·X rand -X i (t)|

[0301] Among them, X rand is the position of a whale individual randomly selected from the population. In this case, the whale individual no longer searches around the current optimal position, but randomly selects a position in the solution space as a reference for position update. This helps the algorithm to widely explore the solution space in the initial stage of the search, increasing the possibility of finding the global optimal solution and avoiding the algorithm from converging to the local optimal solution prematurely.

[0302] In addition to the above two behaviors, the humpback whale also has the behavior of spiraling up to approach the prey. Introduce a parameter l, which follows a normal distribution to simulate the spiral shape. The position update formula for the spiral up behavior is:

[0303] X i (t + 1) = X * (t)·e bl ·cos(2πl) + X * (t)

[0304] Among them, b is a constant used to control the shape of the spiral. In this patent, according to actual tests and experience, b = 1 is set. This formula indicates that the individual whale gradually approaches the current optimal position in a spiral manner. This behavioral pattern combines the characteristics of global search and local search, while maintaining a fine search in the area near the current optimal solution, and can also jump out of the local optimal to a certain extent to explore better solutions.

[0305] During the actual execution of the algorithm, for each whale individual, a certain probability p (set p = 0.5 in this patent) is used to determine whether to perform the spiral upward behavior or the above-mentioned behavior of surrounding the prey or random search. Through this comprehensive simulation of hunting behavior, the positions of whale individuals (i.e., the weight parameters of DBN) are continuously updated, enabling the whale population to continuously explore and optimize in the solution space, gradually approaching the global optimal solution, thereby finding a more suitable combination of DBN weight parameters for the transformer fault prediction task and improving the prediction performance of the model.

[0306] 4) Iterative optimization: Repeat the fitness evaluation and hunting behavior simulation steps until the maximum number of iterations is reached or the termination condition is satisfied to obtain the optimal weight parameters. When performing iterative optimization, in addition to repeating the fitness evaluation and hunting behavior simulation steps, a dynamic learning rate adjustment strategy and a population diversity preservation mechanism are also introduced to further improve the algorithm performance.

[0307] In each iteration, dynamically adjusting the learning rate helps balance the convergence speed and search accuracy of the algorithm. If the learning rate is too large, the algorithm may skip the optimal solution during the search process, while if the learning rate is too small, the convergence speed will be too slow. Therefore, the following dynamic learning rate adjustment formula is adopted:

[0308]

[0309] where η(t) represents the learning rate at the t-th iteration, η0 is the initial learning rate, λ is the parameter controlling the learning rate decay speed, and T is the maximum number of iterations. As the number of iterations t increases, the learning rate η(t) gradually decreases, maintaining a larger learning rate at the beginning of the algorithm to accelerate the search speed and decreasing the learning rate when approaching the optimal solution, enabling the algorithm to converge more precisely to the optimal solution.

[0310] To maintain the diversity of the population and avoid the algorithm prematurely falling into a local optimum, a diversity evaluation method based on population entropy is introduced. The calculation formula for population entropy H is:

[0311]

[0312] where N is the population size, p i is the proportion of the i-th whale individual in the population. When the population entropy H is lower than a certain threshold H thresh , it indicates that the population diversity is insufficient. At this time, some whale individuals are randomly perturbed to increase the population diversity. The random perturbation formula is:

[0313]

[0314] where, is the position of the perturbed whale individual (corresponding to the weight parameter vector of DBN), X iis the original position, α is the disturbance intensity coefficient, and X max and X min are the maximum and minimum values of the search space respectively, and randn() is a random number obeying the standard normal distribution.

[0315] During the iteration process, when the maximum iteration number T is reached or a specific termination condition is met (such as the change in the optimal solution is less than a certain minimum value ∈ after consecutive iterations), the iteration stops and the optimal weight parameters are obtained.

[0316] S5. Fault prediction: Input the test set data into the trained DBN model to obtain the fault prediction results of the transformer, and then judge whether the transformer has faults and the type or severity of the faults. For different types of faults, the following prediction and determination schemes are adopted:

[0317] (1) Winding short - circuit fault determination

[0318] 1) Based on the output probability of the DBN model: When the probability P w corresponding to the winding short - circuit fault in the output result of the DBN model exceeds the set threshold P w-th resh , it is preliminarily determined that the transformer may have a winding short - circuit fault. P w-th resh is an empirical value obtained through statistical analysis of a large amount of historical fault data and normal operation data, and is used to distinguish between normal operation and winding short - circuit fault states. For example, through the analysis of long - term monitoring data of multiple transformers, it is found that when P w is greater than 0.7, the probability of an actual winding short - circuit fault is relatively high, and at this time, P w-th resh can be set to 0.7.

[0319] 2) Combining time characteristics: Use the time - varying autoregressive moving average model (TV - ARMA) to analyze the variation of the time - varying coefficients of voltage V(t) and current I(t). If within a certain period of time, the change rates of the voltage autoregressive coefficient a i (t) and the current autoregressive coefficient c i (t) exceed their respective thresholds S a-th resh and S c-th resh , then the basis for judging the possibility of a winding short - circuit fault is increased. The change rate calculation formulas are respectively:

[0320]

[0321] where Δt is the time interval. S a-th resh and S c-th resh are obtained through statistical analysis of the change rates of time - varying coefficients under normal operation and known winding short - circuit faults, and are used to measure the abnormality of the coefficient change.

[0322] 3) Combine spatial features: Calculate the temperature gradient in the transformer winding area If indicates that the temperature change in this area is drastic, and there may be a local overheating problem caused by winding short - circuit. T g-th resh is the threshold determined according to the historical data of the temperature gradient in the winding area during normal operation of the transformer. Exceeding this value means abnormal temperature distribution. At the same time, if the spatial difference (standard deviation) σ A of the vibration intensity is greater than the vibration threshold σ A-th resh , it further supports the judgment of winding short - circuit fault. σ A-th resh is obtained based on the statistics of a large number of transformer vibration data under normal operation and fault conditions, and is used to define the normal range of vibration differences.

[0323] 4) Combine spectral features: Analyze the spectrum of the winding current. If the band energy E w in a specific frequency band (such as f1 - f2) satisfies E w >E w-th resh , and the center frequency f c of this frequency band deviates from the normal range by more than Δf c-th resh , it indicates that the spectrum of the winding current is abnormal and is related to the winding short - circuit fault. The formula for calculating the band energy is The formula for calculating the center frequency is

[0324]

[0325] E w-th resh , Δf c-th resh are thresholds determined through spectral analysis of a large number of winding short - circuit fault samples and normal operation samples, and are used to judge the abnormality degree of the band energy and the center frequency.

[0326] (2) Determination of iron - core grounding fault

[0327] 1) Based on the output probability of the DBN model: If in the prediction result of the DBN model, the probability P c of the iron - core grounding fault is greater than the threshold P c-th resh , it is preliminarily judged that the transformer has a risk of iron - core grounding fault. P c-th resh is determined based on historical fault data and normal operation data. For example, through the analysis of a large number of transformer data, it is found that when P c is greater than 0.6, the probability of actual iron - core grounding fault is relatively high, and P c-thresh can be set to 0.6.

[0328] 2) Combine time features: Calculate the skewness S PF and kurtosis K PF of the power factor PF(t). If |S PF |>S PF-th resh and |KPF |>K PF-th resh It indicates that the power factor distribution is abnormal and may be related to the iron core grounding fault. S PF-thresh and K PF-th resh are thresholds determined based on the skewness and kurtosis statistical data of the power factor during normal operation, and are used to measure the abnormality degree of the power factor distribution.

[0329] 3) Combining spatial characteristics: Use the spatial autocorrelation function I(T) to analyze the spatial correlation of the temperature in the transformer iron core area. If in the iron core area, I(T) is greater than the threshold I T-thresh , it indicates that the positive correlation of the temperature in this area is too strong, and there may be a phenomenon of local heat accumulation caused by iron core grounding. I T-thresh is obtained through the statistical analysis of the spatial autocorrelation function values of the temperature in the iron core area during normal operation and iron core grounding fault, and is used to judge the abnormal situation of the temperature correlation.

[0330] 4) Combining spectral characteristics: In the spectrum of the transformer neutral point current, if the characteristic frequency f ch ar-c (such as around 100 Hz, which is a characteristic frequency often appearing in iron core grounding faults) is detected, and the amplitude A ch ar-c corresponding to this characteristic frequency is greater than the threshold A c-th resh , and at the same time, the spectral entropy value S c at this frequency differs from the spectral entropy value S c-normal in the normal state by more than ΔS c-th resh , then the possibility of iron core grounding fault is further confirmed. A c-th resh and ΔS c-th resh are thresholds determined through the comparative analysis of a large amount of neutral point current spectrum data during iron core grounding faults and normal operation, and are used to judge the abnormal degree of the characteristic frequency amplitude and spectral entropy value.

[0331] (3) Judgment of insulation aging fault

[0332] 1) Based on the output probability of the DBN model: When the insulation aging fault probability P a output by the DBN model is greater than the threshold P a-th resh , it is considered that there is a possibility of insulation aging fault in the transformer. P a-th resh is determined according to the historical data and experience related to the insulation aging of the transformer. For example, through the long-term monitoring and analysis of multiple transformers, it is found that when P a is greater than 0.5, the probability of the transformer showing insulation aging signs is relatively high, and P a-thresh can be set to 0.5.

[0333] 2) Combining time characteristics: Analyze the power spectral density P V (f), P I(f). If within a specific frequency range (such as f3 - f4), the rate of change of the voltage power spectral density and the rate of change of the current power spectral density respectively exceed the thresholds and then it indicates that the frequency characteristics of the voltage and current have changed, which may be related to insulation aging.

[0334] f3 and f4 are boundary parameters of the frequency interval used to define the analysis of the rate of change of the voltage and current power spectral densities. f3 is the starting frequency of this frequency interval, f4 is the ending frequency, and f3 < f4. The values of these two parameters are related to the electrical characteristics during the normal operation of the transformer and the frequency change characteristics under fault conditions. They are determined through the spectral analysis of a large amount of transformer operation data and combined with the theoretical research on frequency changes during transformer faults. In practical applications, faults such as transformer insulation aging will cause changes in the frequency components of the voltage and current signals, and these changes will be concentrated in certain specific frequency ranges. For example, in some research and actual monitoring, it has been found that when a transformer has an insulation aging fault, within the frequency range of 1000 Hz - 5000 Hz, the changes in the power spectral densities of the voltage and current are more obvious. At this time, f3 can be set to 1000 Hz and f4 can be set to 5000 Hz. For different types of transformers and different operating environments, the values of f3 and f4 may vary, and need to be adjusted and optimized according to the characteristics and historical data of the specific transformer to ensure that frequency change information related to faults can be effectively captured within this frequency interval and provide an accurate basis for fault determination.

[0335] The calculation formulas for the rate of change of the power spectral density are respectively:

[0336]

[0337] and are obtained through the statistical analysis of the rate of change of the voltage and current power spectral densities during normal operation and insulation aging, and are used to measure the abnormality of the change in the power spectral density.

[0338] 3) Combining spatial characteristics: Analyze the probability density function p(Q) of partial discharge. If in some areas inside the transformer, the probability density p(Q) of the partial discharge amount is greater than the threshold p Q-th resh , and the spatial autocorrelation function I(Q) of the partial discharge differs from that in the normal state by more than ΔI Q-th resh , then it indicates that the partial discharge activity is abnormal, which may be caused by the decline of the insulation performance due to insulation aging. p Q-th resh , ΔI Q-th reshIt is a threshold determined based on the statistical analysis of partial discharge data during normal operation and insulation aging, and is used to judge the abnormal degree of the partial discharge probability density and the spatial autocorrelation function.

[0339] 4) Combine spectral features: Calculate the spectral entropy value S of the transformer partial discharge signal p , if S p > S p-th resh , it indicates that the spectral complexity increases and the frequency component distribution becomes more uneven, which is in line with the change characteristics of the partial discharge spectrum during the insulation aging process. S p-th resh is obtained through the statistical analysis of the spectral entropy values of partial discharge signals during normal operation and insulation aging, and is used to define the normal range of the spectral entropy value.

[0340] Through the above comprehensive judgment scheme based on the output results of the DBN model, combined with the multi-source information characteristics of time, space, and spectrum and the corresponding thresholds, it is possible to more accurately predict the short-circuit fault of the transformer winding, the iron core grounding fault, and the insulation aging fault, provide strong support for the operation and maintenance of the transformer, timely discover potential fault hazards, and ensure the safe and stable operation of the power system.

Claims

1. A transformer fault prediction method based on time-space-spectrum multi-source information fusion, characterized in that: The method comprises the following steps: S1. Data collection, wherein the data collection includes: time information collection, space information collection and spectrum information collection; S2, data preprocessing, the data preprocessing includes: normalization processing and denoising operation; S3, feature extraction, the feature extraction includes: time feature extraction, spatial feature extraction and spectrum feature extraction; S4, model construction and training, the model construction and training include: building a deep belief network (DBN) and optimizing DBN weight parameters using a whale optimization algorithm (WOA); S5. Fault prediction, which includes: inputting the test data into the trained DBN model, outputting the fault probability, and combining the threshold judgment rules of time, space, and spectrum characteristics to comprehensively judge the winding short circuit, core grounding, and insulation aging faults.

2. The transformer fault prediction method based on time-space-spectrum multi-source information fusion according to claim 1 is characterized in that: The process of time information collection is as follows: using high-precision sensors to collect the voltage signal V(t), current signal I(t), and power factor PF(t) of the transformer in real time, and recording the timestamp; using a time-varying autoregressive moving average model (TV-ARMA) to dynamically model the voltage and current signals, and using an improved power factor formula to calculate the power factor; The process of spatial information collection is as follows: arranging sensors on transformer windings, cores and oil tank walls to obtain spatial distribution data of temperature, vibration and partial discharge parameters; analyzing spatial distribution characteristics through spatial autocorrelation function I(T) and variation function γ(d); Through the probability density function, we can understand the spatial distribution probability of partial discharge and determine the areas with high incidence of partial discharge; The process of spectrum information collection is: perform spectrum analysis on neutral point current, winding current and partial discharge signal, and extract spectrum features by using variational mode decomposition (VMD) combined with adaptive wavelet packet transform (AWPT) algorithm. The extracted spectrum features include: frequency band energy proportion P n,m and center frequency f n,m .

3. The transformer fault prediction method based on time-space-spectrum multi-source information fusion according to claim 2 is characterized in that: The calculation formula of the spatial autocorrelation function I(T) is: in, is the average value of all temperature data; the spatial weight matrix d ij is the sensor spacing, σ is the weight decay parameter; The TV-ARMA model uses the alternating direction method of multipliers (ADMM) to solve the time-varying coefficients a i (t) and c i (t) and is updated in real time to reflect the dynamic changes caused by load mutations.

4. The transformer fault prediction method based on time-space-spectrum multi-source information fusion according to claim 1 is characterized in that: The normalization process is as follows: Z-score normalization is performed on the time, space and spectrum data respectively, and mapped to the interval [-1, 1]; The process of the denoising operation is: adopting a multi-scale morphological filtering algorithm, using structural elements of different scales to perform opening and closing operations on the data to remove noise.

5. The transformer fault prediction method based on time-space-spectrum multi-source information fusion according to claim 4 is characterized in that: The multi-scale morphological filtering uses a two-dimensional structural element to process spatial data, and the corrosion and expansion operation formulas are:

6. The transformer fault prediction method based on time-space-spectrum multi-source information fusion according to claim 1 is characterized in that: The process of extracting the time features is as follows: extracting the statistical features, autocorrelation coefficient, mutual correlation coefficient and power spectrum density of voltage, current and power factor; the statistical features include mean, variance, kurtosis and skewness; The process of spatial feature extraction is: calculating the temperature gradient Spatial difference of vibration intensityσ A , partial discharge spatial location coordinates (x, y); The process of spectrum feature extraction is: extracting the characteristic frequency f char , frequency band energy E i , center frequency Bandwidth B i And the spectrum entropy value S.

7. The transformer fault prediction method based on time-space-spectrum multi-source information fusion according to claim 6 is characterized in that: Characteristic frequency f char The search criteria are: and And X(f)>τ in, and are the first-order derivative and second-order derivative of the spectrum data X(f) with respect to frequency f, respectively, and τ is a preset threshold; Band Energy E i The calculation formula is: Among them, k i and k i+1 They are the frequency bands [f i ,f i+1 ] is the discrete frequency point index corresponding to the pixel, Δf is the frequency resolution; Center frequency The calculation formula is: Bandwidth B i The calculation formula is: i =f high -f low The calculation formula of spectrum entropy value S is: Where M is the number of discrete points of the spectrum data, Indicates the ratio of the energy of the i-th frequency point to the total energy.

8. The transformer fault prediction method based on time-space-spectrum multi-source information fusion according to claim 1 is characterized in that: The process of constructing a deep belief network (DBN) is as follows: it is composed of a stack of multiple restricted Boltzmann machine (RBM) layers, the number of input layer nodes matches the number of features, and the output layer corresponds to the fault type or probability; The process of optimizing DBN weight parameters using the whale optimization algorithm (WOA) includes: initializing the population, fitness evaluation, hunting behavior simulation and iterative optimization until convergence.

9. The transformer fault prediction method based on time-space-spectrum multi-source information fusion according to claim 8 is characterized in that: The output layer of the DBN uses the softmax function for multi-classification, and the formula is: Among them, k is the output value of the kth node in the output layer, indicating the probability that the sample belongs to the kth type of fault; is the weight connecting the last RBM hidden layer unit j with the output layer unit k, is the bias of output layer unit k, C is the total number of fault types; The parameter update formula of the whale optimization algorithm (WOA) includes: Surrounding behavior: X i (t+1)=X * (t)-A·|C·X * (t)-X i (t)|; Spiral Behavior: X i (t+1)=X * (t)·e bl ·cos(2πl)+X * (t); Random search behavior: X i (t+1)=X rand -A·|C·X rand -X i (t)|.

10. The transformer fault prediction method based on time-space-spectrum multi-source information fusion according to claim 1 is characterized in that: The fault judgment rules in the fault prediction include: Winding short circuit fault: DBN output probability P w >0.7, and the temperature gradient Band Energy E w >E w-th resh ; Core grounding fault: DBN output probability P c >0.6, and the power factor skewness |S PF |>S PF-thresh , spectrum characteristic frequency amplitude A char-c >A c-th resh ; Insulation aging failure: DBN output probability P a >0.5, and the partial discharge spectrum entropy value S p >S p-thresh , power spectral density change rate

Citation Information

Patent Citations

  • Transformer fault maintenance prediction method based on real-time monitoring information

    CN104914327A

  • Remote sensing inversion method of ocean internal thermohaline structure considering spatial nonstationarity

    CN109543356A

  • Rolling bearing fault trend prediction method based on PSO-LSTM

    CN115096592A

  • Power transformer fault diagnosis method based on deep learning

    CN116562169A

  • Parking space utilization rate evaluation method and system for intelligent parking

    CN118824047A

Cited By

  • Surface crack detection method and system based on array surface waves

    CN120275498A

  • Surface crack detection method and system based on array surface wave

    CN120275498B

  • Transformer load remote monitoring and detecting system

    CN120610204A

  • Closed-loop adaptive motor turn-to-turn short circuit dynamic evolution simulation system

    CN120974956A

  • Tunnel lighting cable bridge fault detection method based on improved positioning

    CN121208504A