Transformer fault diagnosis and prediction method and storage medium

By improving the Ant Lion Optimization Algorithm to optimize the deep learning model and combining it with the CNN-BiLSTM-Attention hybrid structure, the problem of insufficient prediction accuracy in transformer fault diagnosis is solved, efficient and stable fault prediction is achieved, and the safety and reliability of the power grid system are improved.

CN120611745AInactive Publication Date: 2025-09-09NANTONG INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510717948.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing grid transformer fault diagnosis methods lack prediction accuracy when processing high-dimensional, nonlinear characteristic data, making it difficult to meet the needs of real-time monitoring and efficient management of the grid. Traditional methods are prone to falling into local optimal solutions and have insufficient convergence speed, resulting in difficulty in timely detection of grid equipment failures, affecting grid stability and security.

Method used

The improved Ant Lion Optimization (IALO) algorithm was used to optimize the deep learning model. Combined with the CNN-BiLSTM-Attention hybrid model, a transformer fault prediction model was constructed through data preprocessing and feature extraction. The Kaggle public dataset was used for training and validation, and hyperparameters were optimized to improve model accuracy and stability.

Benefits of technology

The accuracy and stability of transformer fault prediction have been significantly improved, with the test accuracy reaching 98.28% and the test loss reduced to 0.1851. The model has shown unique advantages in processing temporal and spatial characteristics, improving the safety and reliability of the smart grid system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611745A_ABST
    Figure CN120611745A_ABST
Patent Text Reader

Abstract

The invention discloses a transformer fault diagnosis and prediction method and a storage medium, and the fault prediction method comprises the steps: introducing elite retention, adaptive step length, diversity protection and a local search strategy into an original ant lion optimization algorithm, and obtaining an improved ant lion optimization algorithm; optimizing the deep learning model by adopting an improved ant lion optimization algorithm to obtain an improved deep learning model; training the improved deep learning model by using the preprocessed transformer fault data set as input data to obtain a fault prediction model; obtaining and preprocessing transformer operation data; and inputting the preprocessed transformer operation data based on the fault prediction model, and outputting a transformer fault diagnosis and prediction result. The method has unique advantages in the aspects of processing time sequence data and spatial features, the parameters of the deep learning model or the hybrid model thereof are further optimized through the improved ant lion optimization algorithm, and the safety and reliability of the intelligent power grid system are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a transformer fault diagnosis and prediction method, and in particular to a transformer fault diagnosis and prediction method, device and storage medium based on a hybrid neural network. Background Art

[0002] As the global energy transition continues, the penetration of renewable energy in power grids is rapidly increasing. While this trend improves energy efficiency, it also significantly increases the complexity and uncertainty of grid operations. Compared with traditional power generation methods, renewable energy sources are characterized by unstable supply and significant intermittent fluctuations. These characteristics lead to ongoing challenges in the balance of grid operations. This imbalance can easily cause grid equipment, especially transformers, to fail or even be damaged, further threatening the safe and stable operation of the entire grid system. Therefore, in this context, how to accurately predict and diagnose transformer failures and improve grid stability has become a key research topic in the energy sector.

[0003] Faced with the increasingly complex operating environment of power grids, traditional grid stability prediction methods, relying on empirical experience and fixed rules, are no longer able to meet the current demands for real-time monitoring and efficient management of power grids. These methods often exhibit significant shortcomings when dealing with large amounts of complex data, particularly when processing high-dimensional, nonlinear data, where their prediction accuracy significantly decreases. This lack of accuracy makes it difficult for grid managers to promptly identify and resolve potential equipment failures, potentially leading to serious grid accidents. Therefore, grid management urgently needs more effective prediction methods to improve the accuracy and real-time nature of transformer failure predictions and grid stability.

[0004] To address these issues, hybrid approaches combining optimization algorithms with deep learning models have become a hot topic of research. Hybrid models automatically search for optimal parameter configurations through optimization algorithms, thereby improving the model's predictive performance. The ALO algorithm (Ant Lion Optimization) is widely used for deep learning model optimization tasks due to its outstanding search capabilities. However, in practical applications, the ALO algorithm has also been exposed to problems such as susceptibility to local optimal solutions and insufficient convergence speed. Therefore, research on improved ALO algorithms (IALO) has gradually attracted attention. By adjusting the optimization strategy, IALO further enhances the algorithm's global search performance and convergence speed, making it more suitable for deep learning model optimization tasks in complex environments.

[0005] Research in the field of transformer fault diagnosis and prediction is increasingly showing a trend toward multi-method integration and multi-index integrated evaluation to address the increasingly complex operational conditions of power grids and practical engineering needs. In recent years, addressing the sample imbalance problem inherent in real-world data, some existing technologies have employed integrated sampling methods and advanced machine learning models to improve transformer fault prediction performance. Addressing the limited accuracy of traditional prediction methods when processing small sample sizes, the fusion of grey theory and relevance vector machines (RVMs) has gained increasing attention. Dynamic prediction of transformer failure probability for effective early warning is also gaining attention. Xu Cong employed a vector autoregression (VAR) model to construct prediction models for both accidental and aging failures. By combining historical operating data with real-time health indices, he achieved accurate predictions of failure probability. Research results demonstrate that this method demonstrates strong practicality in accurately predicting transformer failure probability, effectively improving the reliability and stability of power systems.

[0006] Transformer stability refers to the ability of a transformer to maintain key parameters such as voltage and current within safe ranges in the face of load fluctuations, external disturbances, and environmental changes during grid operation, and to quickly recover to steady state after a disturbance. In the context of smart grids, transformers not only play a pivotal role in connecting energy supply and user demand, but their stability is also directly related to the reliability and security of the entire grid system. Theoretical research on grid dynamic stability has shown that the response time, nominal power, and elasticity coefficient of grid nodes jointly determine the system's ability to regulate disturbances, and these theories are equally applicable to explaining transformer stability.

[0007] Response time reflects the transformer's ability to react to electricity price fluctuations and sudden load changes. This adjustment rate plays a crucial role in maintaining grid balance. If the response time is too short, the system may experience oscillations due to overly rapid adjustments. If the response time is too long, load changes may not be compensated in time, leading to power imbalance. Rated power directly relates to the distribution of energy within the grid. The discrepancy between power generation and load causes voltage fluctuations, placing significant stress on the transformer. A chronic imbalance in power distribution between grid nodes can not only cause voltage collapse but also damage the transformer's internal insulation. Meanwhile, the elasticity coefficient reflects the sensitivity of each node to electricity price fluctuations. Its function is to adjust to load changes. However, excessive sensitivity can amplify grid disturbances, posing a potential threat to system stability. These factors are both independent and intertwined, forming a complex network that influences transformer stability and failure risk.

[0008] The goal of transformer fault diagnosis is to promptly identify transformer anomalies and provide early warning of potential failures, thereby ensuring the safe and stable operation of smart grids. To achieve this goal, the diagnostic system must possess robust data processing and feature extraction capabilities. Transformer operating data typically comes from multiple sensors and monitoring systems, and this data may be affected by noise, missing values, and measurement errors. Therefore, data standardization is essential to ensure the accuracy of subsequent analysis. Effective preprocessing not only improves data quality but also provides uniform and clear input for feature engineering.

[0009] Feature extraction is a key step in diagnostic system design. Analyzing transformer operating data and extracting features relevant to fault occurrence can effectively improve fault diagnosis accuracy. The diagnostic system must accurately identify parameters associated with fault risk, such as load fluctuations, temperature variations, and response time. By properly selecting and constructing features, the system can more accurately detect abnormal changes in complex power grid environments and effectively predict the likelihood of faults.

[0010] Model selection and hyperparameter optimization are also important requirements. Fault diagnosis systems need to select appropriate machine learning or deep learning algorithms based on data characteristics and application requirements. The selected algorithm must not only ensure accuracy but also possess strong real-time responsiveness to ensure rapid and accurate fault predictions in dynamic power grid environments. Hyperparameter optimization can also further enhance the model's stability and generalization capabilities, ensuring stable operation despite data noise, missing values, or abnormal fluctuations. Hyperparameter optimization is particularly effective in ensuring the stability and accuracy of prediction results when grid load fluctuates significantly. Summary of the Invention

[0011] Purpose of the invention: The purpose of the present invention is to provide a transformer fault diagnosis and prediction method to solve how to improve the accuracy and stability of transformer fault diagnosis and prediction.

[0012] Technical solution: The transformer fault diagnosis and prediction method described in the present invention includes the following steps:

[0013] By introducing elite preservation, adaptive step size, diversity protection and local search strategy into the original Ant Lion Optimization Algorithm, an improved Ant Lion Optimization Algorithm is obtained.

[0014] An improved deep learning model is obtained by optimizing the deep learning model using an improved ant lion optimization algorithm;

[0015] The preprocessed transformer fault dataset is used as input data to train the improved deep learning model and obtain the fault prediction model;

[0016] Acquire and preprocess transformer operation data;

[0017] Based on the fault prediction model, the preprocessed transformer operation data is input and the transformer fault diagnosis and prediction results are output.

[0018] This paper aims to construct a CNN-BiLSTM-Attention hybrid model optimized based on the IALO algorithm to achieve accurate diagnosis of transformer faults and efficient prediction of smart grid stability. This prediction method focuses on multiple key links, including data analysis, model construction, parameter optimization, and model evaluation. The paper utilizes a grid stability simulation dataset publicly available on Kaggle, which contains 60,000 observations, each of which includes 12 dynamic feature variables and a binary stability label. Based on this dataset, the paper conducts in-depth exploratory data analysis to determine the degree of influence of each feature on grid stability, thereby providing a data foundation for subsequent model training.

[0019] Based on data analysis, the study constructed three basic deep learning models: CNN, BiLSTM and CNN-BiLSTM-Attention, to preliminarily explore the applicability of different network structures in the problem of power grid stability prediction. These three basic models have become important research objects for subsequent optimization due to their advantages in spatial feature extraction, time series processing and feature attention allocation. In order to further improve the prediction performance of the model, the present invention uses ALO (Ant Lion Optimization Algorithm) and improved IALO (Improved Ant Lion Optimization Algorithm) to optimize the hyperparameters of the basic model, which significantly improves the prediction accuracy of the model on the power grid stability dataset. Among them, the IALO algorithm, due to its stronger global search capability, effectively overcomes the defect of the traditional ALO algorithm that is prone to falling into local optimal solutions, and demonstrates better optimization effects.

[0020] Preferably, the elite retention strategy includes:

[0021] Assume that the population size is N and the elite retention ratio is ρ e , then the number of elite individuals N e for:

[0022]

[0023] in represents the rounding-up operation, which ensures that at least one optimal solution is retained when the population is small or the elite ratio is low;

[0024] Let all candidate solutions in the population be {X1,X2,…,X N} and its corresponding fitness is {f(X1),f(X2),…,f(X N )}; Sort these fitness values ​​in ascending order, and the elite individual set is:

[0025] E={X (1) ,X (2) ,…,X (Ne)}

[0026] where X (i) represents the i-th optimal solution.

[0027] Preferably, the adaptive step size strategy includes:

[0028] Let the maximum number of iterations be K max ,At the kth iteration, the adaptive step size factor α is defined as:

[0029]

[0030] When k is small, α is close to 1, allowing a larger step size for global search. As k increases, α gradually decreases, and the minimum step size remains at least 0.1;

[0031] Based on the adaptive step size factor α, the position update of non-elite individuals is expressed as:

[0032] x new =x current +(ba)×0.1×(rand-0.5)×α

[0033] Where [a, b] is the allowed interval of the parameter, and rand represents a random number generated from a uniform distribution [0, 1], x current Represents the current position, x new Represents the updated position.

[0034] Preferably, the diversity protection strategy includes:

[0035] Let the diversity ratio be p d , the number of worst individuals in the population N d Expressed as:

[0036]

[0037] Where N is the total population, Represents the rounding-up operation. For these worst individuals, new candidate solutions will be obtained by random initialization. The formula for generating new candidate solutions is:

[0038] x new =a+rand×(ba).

[0039] Preferably, the local search strategy includes:

[0040] Assume that the elite individual parameter is x_elite and the allowed interval is [a, b]. The ratio of the neighborhood perturbation step size to the search space is defined as β. Then the perturbation of the local search is expressed as:

[0041] Δ=(ba)×β×(rand-0.5)

[0042] Where rand represents a random number generated from a uniform distribution [0,1];

[0043] The update formula of the elite individual in the neighborhood is:

[0044] x neig hbor=x elite +Δ.

[0045] Preferably, the process of the improved ant lion optimization algorithm includes:

[0046] Initialize the population and set the parameters search boundary, population size, maximum number of iterations, elite ratio, diversity ratio and number of local searches;

[0047] Randomly generate N solutions and calculate the population fitness;

[0048] Sort in ascending order of fitness, record the elite set and the worst set;

[0049] Record the current optimal solution;

[0050] Calculate the adaptive step size factor;

[0051] Update the position of non-elite individuals and evaluate their fitness, and replace them if they are better;

[0052] Perform local search on each elite and replace it if a better solution is found;

[0053] Reinitialize the worst individual to maintain diversity and output iterative information;

[0054] After iteration, the optimal solution and its fitness are output.

[0055] Preferably, the original ant lion optimization algorithm includes:

[0056] Assume that the random walk step length is t, the maximum number of iterations is n, and the random function r(t) is defined as:

[0057]

[0058] Where rand represents a random number generated from a uniform distribution [0,1];

[0059] Use the accumulation and cumsum operation to construct the ant's random walk path X(t):

[0060] X(t)=cumsum(2r(t)-1)

[0061] The position of the i-th variable in the t-th generation The update formula is:

[0062]

[0063] where a i with d i are the minimum and maximum allowed values ​​of the ith variable, respectively, and and They represent the minimum and maximum values ​​of the i-th variable in the t-th iteration respectively;

[0064] The update process is affected by the location of the Ant Lion Trap, which is mapped to:

[0065]

[0066] where c t and d t represent the minimum and maximum values ​​of all variables in the current iteration, respectively, and represents the position information of the j-th ant lion at iteration t;

[0067] When an ant is updated, its fitness is Exceeds the current antlion's fitness When , the ant lion will capture the ant and replace its position with the new ant lion position, the formula is as follows:

[0068]

[0069] Preferably, the deep learning model includes one of a convolutional neural network, a bidirectional long short-term memory network, a hybrid model of a convolutional neural network, a bidirectional long short-term memory network, and an attention mechanism;

[0070] The preprocessing method is normalization or standardization;

[0071] The method for obtaining the transformer operation data is as follows:

[0072] Use sensors to collect real-time operating status data of transformers under different working conditions on site;

[0073] After data collection is completed, the original records are transferred to the backend for centralized preprocessing.

[0074] Preferably, the training process of the improved deep learning model is:

[0075] The standardized transformer fault dataset is divided into a training set and a validation set;

[0076] Set training parameters and choose optimizer and loss function;

[0077] Initialize the improved deep learning model, input the training set, calculate the training loss and update the weights, and evaluate the model performance on the validation set;

[0078] Adjust hyperparameters, repeat training, and output the optimal model after reaching the training round.

[0079] Another aspect of the present invention discloses a computer storage medium, wherein the computer storage medium stores instructions, and when the instructions are executed on a computer, the computer is caused to execute the above method.

[0080] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:

[0081] The fault prediction model in the present invention has unique advantages in processing time series data and spatial features. The parameters of the deep learning model or its hybrid model are further optimized through the improved ant lion optimization algorithm, effectively improving the safety and reliability of the smart grid system. The CNN-BiLSTM-Attention hybrid model optimized by the improved ant lion optimization algorithm (IALO) achieved the highest prediction accuracy (98.28%) and the lowest test loss (0.1851) in the test set. This significantly improved prediction performance fully demonstrates the effectiveness of the IALO algorithm in the hyperparameter optimization process, and the unique advantages of the CNN-BiLSTM-Attention hybrid structure in capturing the characteristics of power grid stability data. In addition, the present invention further analyzes the confusion matrix and ROC curve to intuitively demonstrate the classification performance and stability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 This is a data distribution diagram;

[0083] Figure 2 Box plots for data outlier detection;

[0084] Figure 3 is the correlation heat map;

[0085] Figure 4 is a relationship diagram between variables;

[0086] Figure 5 Basic flow chart for model training and evaluation;

[0087] Figure 6 This is the structure diagram of the convolutional neural network;

[0088] Figure 7 This is a diagram of the neural structure of long-term and short-term memory;

[0089] Figure 8 This is the flowchart of the improved antlion optimization algorithm;

[0090] Figure 9 Flowchart of the model training process;

[0091] Figure 10 Flowchart for transformer fault prediction;

[0092] Figure 11 This is a diagram of the distribution of some data;

[0093] Figure 12 This is a comparison chart of data boxes after normalization and standardization;

[0094] Figure 13 The ROC curves of each model are shown in Figure 2. DETAILED DESCRIPTION

[0095] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0096] A transformer fault diagnosis and prediction method comprises the following steps:

[0097] 1. Dataset Description and Exploratory Analysis

[0098] 1.1 Dataset Source and Feature Description

[0099] The dataset used in this paper comes from the "Electrical Grid Stability Simulated Dataset" in the UCI Machine Learning Repository. This dataset contains multiple monitoring features of power grid operation and aims to predict grid stability using machine learning methods. The dataset contains 60,000 observations and has been processed through data augmentation. During data preprocessing, the target variable (stabf) is mapped to a binary classification label, where stable is 1 and unstable is 0. A description of the main features in the dataset is shown in Table 1.

[0100] Table 1 Feature Description

[0101]

[0102] Response time refers to how quickly a transformer reacts to external disturbances, such as changes in voltage, current, or load. A shorter response time means the transformer can adapt more quickly to environmental changes, reflecting its regulation capabilities and operating status to a certain extent.

[0103] Nominal power describes the rated power output of a transformer under normal operating conditions. Positive values ​​typically represent the rated output of the transformer when supplying power to the grid, while negative values ​​may indicate load operation in some cases. This parameter is crucial for evaluating a transformer's performance under varying load conditions.

[0104] The elasticity factor reflects the transformer's sensitivity to load fluctuations or other changes in operating conditions. Transformers with a higher elasticity factor exhibit a more pronounced response to changing operating conditions. This metric helps assess the equipment's adaptability and potential risks.

[0105] Grid stability is a binary target variable that represents the stability of the power grid, where 1 represents stability and 0 represents instability. This target variable is the prediction target of the present invention.

[0106] 1.2 Exploratory Data Analysis

[0107] Exploratory data analysis (EDA) is a preliminary step in data analysis. Its purpose is to discover the structure, distribution, and potential relationships of the data through visualization and statistical description, providing a basis for subsequent model construction and feature selection. The main task of EDA is to use various graphs to display the characteristic distribution of transformer operating data, the relationship between variables, and any outliers or data imbalances. As shown in Table 2, the dataset is complete, with no missing values.

[0108] Table 2 Data missing table

[0109]

[0110] By drawing a feature distribution graph, you can observe the overall distribution of the data. For example, if p1 is a bell-shaped curve, it means that the feature values ​​are concentrated within a certain range. The feature distribution graph helps to understand the feature distribution and provides a basis for subsequent standardization and normalization processing. The data distribution is as follows: Figure 1 shown.

[0111] Figure 2 The box plots of the raw data are shown. Each box plot represents the data distribution of a different feature, with the horizontal axis representing the feature index (from 1 to 13) and the vertical axis representing the feature's numerical range. It can be observed that the value distribution of some features is very concentrated, with most data points concentrated in a range close to 0, while other features are more discrete, with some features having a wide range of values ​​and even extreme values ​​(such as feature 5). Feature 5, in particular, has a box plot with significant outliers, which fall outside the upper whisker, indicating the presence of strong outliers in this feature.

[0112] Correlation heatmaps reveal the relationships between features. Figure 3As shown, stabf and stab have a strong correlation, indicating that the stability of the target variable is directly related to other features. Furthermore, the strong negative correlation between p1 and p3 suggests that the power and energy consumption of a node may have opposite trends. These correlation analyses can clarify which features are likely to have a significant impact on fault diagnosis, thus giving them more attention during feature selection.

[0113] exist Figure 4 The relationship diagram between the variables clearly shows the relationships between some features in the data. There is a clear relationship between stab and stabf. The values ​​of stab are concentrated between 0 and 1 and exhibit a clear distribution pattern. This indicates a direct correlation between stab and the target variable stabf. Therefore, stab is a promising feature that can serve as a key input in fault diagnosis models.

[0114] The distribution of stabf reveals a severe class imbalance in the data. The vast majority of samples have a stabf value of 0 (unstable), while the number of samples with a stabf value of 1 (stable) is far lower than that of class 0. This can affect subsequent model training, especially in classification tasks, where it may cause the model to favor the majority class. To address this imbalance, oversampling or undersampling techniques may be necessary during model training.

[0115] 2. Improved Deep Reinforcement Learning Algorithm

[0116] 2.1 Basic Algorithm Principles

[0117] Machine learning aims to use training data to learn a model with generalization ability and make accurate predictions on new samples, such as Figure 5 The figure shows the complete process from training data to model evaluation. A large number of samples from the training set are used to train the model, and the model continuously adjusts its internal parameters during training to make the predictions closer to the true labels. After training, the features of the test set are input into the model, and the model's performance is evaluated by comparing them with the true labels of the test set. If the model performs as expected, it can be put into use. If the evaluation results are still insufficient, the model needs to be revised or reconstructed, and the above process repeated. This iterative approach allows the model to gradually improve its generalization ability as it continuously absorbs new data information, and it also provides a viable approach for deep learning in more complex data scenarios.

[0118] 2.2 CNN Model Structure and Principle

[0119] Convolutional neural networks (CNNs) process input data through hierarchical feature extraction and have shown good performance in areas such as image recognition and time series analysis. Figure 6 This network structure is usually composed of a stack of convolutional layers, pooling layers, and fully connected layers. The convolutional and pooling layers are mainly responsible for feature extraction and dimensionality reduction, while the fully connected layers are used for the final classification or regression tasks. When extracting local features, the convolutional layer uses a learnable convolution kernel and the input matrix to perform a sliding convolution operation. The convolution operation process can be expressed as shown in formula (3-1):

[0120]

[0121] where X (l-1) is the output of the (l-1)th layer, is the i-th convolution kernel of the l-th layer, B (l) is the bias term, * represents the discrete convolution operation, and σ(·) is the activation function (such as ReLU or Tanh). When the convolution kernel slides on the input matrix, the local information is integrated into each position of the output feature map through weighted summation, so that the network can capture representative local patterns. The pooling layer follows the convolution layer. Common maximum pooling and average pooling will downsample the feature map according to the set sliding window to reduce the data dimension and the number of parameters. The specific process can be understood as retaining the maximum value or average value in each sliding window. The downsampling operation can alleviate the risk of overfitting while retaining important features because a large amount of redundant information is discarded. The fully connected layer maps the pooled multi-dimensional features to a one-dimensional vector, and the Softmax function is usually used in the last layer to output the classification probability. The calculation method of the Softmax function can be expressed as shown in formula (3-2):

[0122]

[0123] where z i is the linear transformation result of the i-th neuron, and M is the number of output categories. Cross-entropy is often used as a loss function for classification tasks, measuring the difference between the model's prediction and the true label. A smaller loss function indicates that the model's prediction is closer to the true distribution. The backpropagation algorithm updates the convolution kernel parameters and fully connected layer weights based on the loss function, enabling the model to continuously improve its ability to capture data patterns during iteration.

[0124] Through the stacking of multiple layers of convolution and pooling and the comprehensive decision-making of fully connected layers, CNN shows high efficiency and accuracy in extracting complex features.

[0125] 2.3 BiLSTM Model Structure and Principle

[0126] The BiLSTM model captures past and future information in the sequence by simultaneously utilizing forward and backward LSTMs, thereby fully reflecting the inherent laws of time series data. This model has obvious advantages in processing tasks with strong time dependence because it can fuse information from both ends of the sequence to achieve a global understanding of the context. In practical applications, the BiLSTM architecture is widely used in natural language processing, speech recognition, and power grid fault diagnosis, because it performs well in capturing complex time series patterns. The core of this model is that each LSTM unit realizes the selective transmission of information through a gating mechanism, thereby effectively solving the problem of gradient disappearance or explosion in long sequences, such as Figure 7 As shown in Figure 1, the internal structure of a single LSTM unit is shown, including key parts such as the input gate, forget gate, output gate, and cell state.

[0127] The input x at each moment t Will be the same as the hidden state h at the previous moment t-1 and cell state c t-1 Entering the gating mechanism, the input gate determines the importance of the current input information, the forget gate determines how much of the previous memory is retained, and the output gate controls how the new hidden state is transmitted downward. The red dashed boxes indicate the internal interactions between different gates and states. The gating signal, combined with the cell state, dynamically filters information, effectively alleviating the vanishing or exploding gradient problem.

[0128] The overall architecture of the bidirectional LSTM is that the forward and reverse LSTMs process sequence data separately on the input layer and concatenate and fuse the outputs to fully utilize the context information.

[0129] The forward module starts from x in chronological order t-1 to x t+1 The reverse module processes the input, reading the sequence information backward from the future. Both generate their own hidden states and cell states in the hidden layer. Finally, the outputs from these two directions are concatenated or fused in the output layer, preserving both past and future context. This bidirectional processing enables the model to capture complex temporal dependencies with greater flexibility and accuracy, providing rich feature representations for subsequent fault diagnosis or situation prediction.

[0130] Each LSTM unit realizes the selective transmission of information through gate control at the current moment. The input gate controls the retention of the current input information. Its calculation formula is shown in formula (3-3):

[0131] i t =σ(W i x t +U i h t-1 +bi ) (3-3)

[0132] Where σ represents the Sigmoid activation function, W i with U i Represent the input weight and hidden state weight respectively, b i is the bias term, x t With h t-1 Representing the hidden state of the previous moment and the input at the current moment respectively. The role of the input gate is to ensure that important information is transmitted, thus laying the foundation for subsequent state updates.

[0133] Another key part of the gate control mechanism is the forget gate, which is used to determine how much previous information to retain. The calculation formula of the forget gate is shown in formula (3-4):

[0134] f t =σ(W f x t +U f h t-1 +b f ) (3-4)

[0135] In formula (3-4), W f 、U f with b f Represent the input weight matrix, hidden state weight matrix and bias term respectively. The forget gate also generates a control signal between 0 and 1 through the Sigmoid function, thereby achieving selective forgetting of historical memories. The candidate memory unit generates new memory candidates through nonlinear transformation, as shown in formula (3-5):

[0136]

[0137] Tanh introduces nonlinear characteristics, ensuring that the value of the candidate memory cell varies between [-1, 1]. This process provides the basis for fine-grained information updates. Wc, Uc, and bc are the parameters of the candidate memory cell, responsible for generating potential new information at the current moment. Next, the update of the cell state integrates the control of the forget gate and the input gate. Formula (3-6) gives the update mechanism:

[0138]

[0139] where ⊙ represents element-wise multiplication, c t-1 It is the state of the LSTM memory unit at the previous time step and carries the long-term dependency information of the sequence. This formula ensures that new information is balanced with old memory, providing a stable memory transfer mechanism for the LSTM unit.

[0140] The function of the output gate is to determine the hidden state generation at the current moment, as shown in formula (3-7):

[0141] o t =σ(W o x t +U o h t-1 +b o ) (3-7)

[0142] Among them, Wo, Uo and bo are the output gate parameters, which determine the explicit output of the memory cell state to the current hidden state. Then the hidden state is determined by formula (3-8):

[0143] h t =o t ⊙tanh(c t ) (3-8)

[0144] where h t 、c t Represent the hidden state and memory cell state at the current time step, respectively. These formulas constitute the complete internal mechanism of the LSTM unit, ensuring that time series information is effectively captured and transmitted at every moment. Based on this, the bidirectional LSTM fuses the outputs of the forward and reverse LSTMs, as shown in formula (3-9):

[0145]

[0146] in Indicates bidirectional LSMT, h t Represents the output of a unidirectional LSMT. This structure not only enhances the model’s ability to capture temporal dependencies, but also provides a more comprehensive information representation for complex fault diagnosis tasks, thereby significantly improving prediction accuracy and robustness.

[0147] 2.4 Attention Mechanism

[0148] The Attention mechanism allocates learnable weights between different time steps or feature dimensions, enabling the model to focus on the most valuable input information, thereby improving the modeling ability of long sequences or high-dimensional data.

[0149] Assume that the hidden state of the input sequence is {h1,h2,…,h T}, the model needs to determine the importance of each hidden state for a certain query vector q. The scoring function first compares the query vector q with each hidden state h t Map them to the same space, and then use the learnable parameters to measure the similarity, and we get formula (3-10):

[0150]

[0151] Where W q and W h They are mapping matrices, v is a learnable vector, b e is the bias term. The scoring result is e t Used to represent the hidden state h t The degree of matching with the query vector q, the larger the value, the better h t More important for the current task. Next, the attention weight α t Formula (3-11) is obtained by performing Softmax normalization on the scoring results:

[0152]

[0153] where α t The value of is in the interval (0,1), and the sum of the weights at all moments is 1. τ is the index variable in the summation, which represents the time step index involved in the attention mechanism, that is, traversing and weighting all time steps. T is the total number of time steps in the input sequence. In this way, the model can focus on the time steps or feature dimensions with larger weights when predicting, while weakening the rest accordingly, thereby focusing on key information in long sequences or high-dimensional inputs. Finally, as shown in formula (3-12), the attention-weighted context vector c is obtained by the weighted sum of each hidden state:

[0154]

[0155] The Attention mechanism uses this dynamic weighting approach to enable the model to process long sequence data without having to rely on a single fixed-length vector to express all information, thereby effectively highlighting the features of key moments while retaining the global context.

[0156] 2.5 Multilayer Perceptron (MLP) Model Structure and Principle

[0157] Multilayer Perceptron (MLP) maps input data layer by layer to the output space through multiple layers of linear transformation and nonlinear activation function, forming an abstract process from simple features to high-dimensional representation. The model structure usually includes an input layer, several hidden layers and an output layer. The neurons in each layer are fully connected to the neurons in the previous layer, and nonlinear mapping is achieved under the action of the activation function. Assume that the input vector of the lth layer is h (l-1) , the weight matrix is ​​W (l) , the bias term is b (l) , then the forward propagation can be expressed by formulas (3-13) and (3-14):

[0158] z (l) =W(l) h (l-1) +b (l) (3-13)

[0159] h (l) =σ(z (l) ) (3-14)

[0160] Where σ(·) represents a nonlinear activation function, such as ReLU, Sigmoid, or Tanh. The output h of the hidden layer (l) It will be iterated as the input of the next layer until the final output layer generates the target result. For classification tasks, the output layer often uses the Softmax function to normalize the multi-class probability distribution and uses the cross entropy loss function to measure the gap between the prediction and the true label.

[0161] Model parameter updates are typically implemented using the backpropagation algorithm. Backpropagation calculates the gradient of each layer's weights and biases with respect to the loss function using the chain rule. During the gradient descent process, the parameters are corrected layer by layer, gradually improving the network's ability to fit the training data. Multilayer perceptrons capture high-level features of data through stacked linear transformations and activation functions, and are therefore widely used in fields such as fault diagnosis, image recognition, and natural language processing.

[0162] 2.6 Principles of Logistic Regression (LR) Model

[0163] Logistic Regression (LR) is a classic linear classification model. Its basic idea is to map the linear combination of input features to the probability space to achieve binary classification tasks.

[0164] The model constructs a linear equation To integrate the input vector x with the weight parameter w and bias b; this linear combination provides the basis for subsequent nonlinear mapping. Then, the Sigmoid function is used to convert z into a probability value, as shown in formula (3-15):

[0165]

[0166] Here, σ(z) compresses any real number into the range (0, 1), which can be interpreted as the probability that the sample belongs to the positive class. This probability output provides a quantitative basis for the model's decision-making, and predictions are usually made based on a preset threshold (such as 0.5). During model training, the cross-entropy loss function is used to measure the difference between the predicted probability and the true label. Its formula is shown in Formula (3-16):

[0167]

[0168] Where N represents the total number of samples, y(i) is the true label of the i-th sample, p (i) is the predicted probability. Gradient descent updates the parameters based on this loss function, gradually bringing the model output closer to the true distribution. This mechanism is both simple and efficient, while also offering good interpretability, ensuring the model's stable performance in practical applications such as fault diagnosis.

[0169] 2.7 Improved Ant Lion Optimization Algorithm

[0170] 2.7.1 Principle of the Original Antlion Optimization Algorithm

[0171] The original Ant Lion Optimizer (ALO) is a continuous, single-objective optimization algorithm. Its core concept is derived from the predatory behavior of ant lions. By simulating the random walks, position updates, and captures of ants within the search space, it achieves a dynamic balance between global search and local exploitation. This process is constructed using three main operators, corresponding to the ant's random walk, ant's position update, and ant lion's capture process.

[0172] The ant random walk algorithm provides a variety of candidate solutions for the initial stage of the algorithm. Assume that the random walk step length is t, the maximum number of iterations is n, and the random function r(t) is defined as shown in formula (3-17):

[0173]

[0174] Where rand represents a random number generated from a uniform distribution [0,1]. Therefore, using the cumulative sum (cumsum) operation to construct the ant's random walk path, we can obtain formula (3-18):

[0175] X(t)=cumsum(2r(t)-1) (3-18)

[0176] This formula ensures that ants fully explore the solution space, laying a diversified foundation for subsequent searches.

[0177] Next, the ant position update mechanism maps the solution obtained by random walk to the allowed interval, realizing the organic combination of local search and global development. Formula (3-19) defines the position update formula of the i-th variable in the t-th generation:

[0178]

[0179] where a i with d i are the minimum and maximum allowed values ​​of the ith variable, respectively, and and Represent the minimum and maximum values ​​of the i-th variable in the t-th iteration. The update process is also affected by the position of the ant lion trap, and its position mapping is described by formula (3-20):

[0180]

[0181] where c t and d t represent the minimum and maximum values ​​of all variables in the current iteration, respectively, and represents the position information of the jth ant lion at iteration t. This process makes the updated position of the ant closely related to the position of the ant lion, thus guiding the search to converge to the local optimal area.

[0182] Finally, during the ant lion capture phase, the survival of the fittest is achieved through fitness comparison. Exceeding the current ant lion When the fitness of the ant lion reaches , the ant lion will "capture" the ant and replace its position with the new ant lion position, as shown in formula (3-21):

[0183]

[0184] This replacement mechanism ensures the continuous updating of the optimal solution in the group and drives the search process to converge to the global optimal solution.

[0185] Overall, the original ALO algorithm achieves comprehensive exploration and local development of the solution space through the above three steps. Its random walk, position update, and capture mechanisms form a complete optimization closed loop, providing an effective solution strategy for solving complex optimization problems.

[0186] 2.7.2 Improved Ant Lion Optimization Algorithm

[0187] The improved ant lion optimization algorithm (IALO) is based on the original ALO. By introducing strategies such as elite retention, adaptive step size, diversity protection, and local search, it effectively balances the global search and local development processes, and achieves a better compromise between convergence speed and optimal solution quality, making the algorithm more robust and efficient in complex high-dimensional optimization problems.

[0188] (1) Elite retention strategy

[0189] The elite retention strategy is the key mechanism to ensure the inheritance of excellent solutions in the improved Ant Lion Optimization Algorithm. Its core is to retain individuals with high fitness (i.e., low loss) during each iteration, so that these excellent solutions can play a leading role in subsequent searches. This strategy ranks the fitness of each candidate solution in the population and selects the top several as elite individuals to avoid losing the optimal solution during random updates or reinitialization. To achieve this goal, let the population size be N and the elite retention ratio be ρ e , then the number of elite individuals N e It can be determined by formula (3-22):

[0190]

[0191] in The formula ensures that at least one optimal solution is retained when the population is small or the proportion of elites is low. The basic idea of ​​this strategy is that after each round of iteration, the fitness of all individuals f(·) is calculated and sorted to select the N with the best fitness. e This process can be described by the following steps: Let all candidate solutions in the population be {X1,X2,…,X N} and its corresponding fitness is {f(X1),f(X2),…,f(X N )}; Sort these fitness values ​​in ascending order, and the elite individual set is as shown in formula (3-23):

[0192]

[0193] where X (i) represents the i-th optimal solution. This sorting operation ensures that the elite solution always remains in the population, thus providing a reference for other candidate solutions in subsequent iterations and playing a role in guiding the global search. Elite retention not only improves the stability of the algorithm, but also accelerates the convergence process, because excellent solutions will not be easily overwritten by random perturbations. Overall, the elite retention strategy enables the improved algorithm to effectively inherit and utilize historical optimal information through the precise description of formulas (3-22) and (3-23). ​​This feature is particularly important when optimizing complex high-dimensional problems. As described in the article, this mechanism achieves stable global convergence while maintaining diversity.

[0194] (2) Adaptive step size factor

[0195] The adaptive step size factor is designed to enable the algorithm to have a larger search step size in the early stage so as to fully explore the global solution space, and gradually reduce the step size during the iteration process, so as to achieve a more detailed search during local development. The core of this mechanism is to dynamically adjust the step size so that the search area can be greatly disturbed in the initial stage, and then finely converge to the optimal solution in the later stage. To this end, the maximum number of iterations is set to K max , at the kth iteration, the adaptive step size factor α is defined as shown in formula (3-23):

[0196]

[0197] This formula shows that when k is small, α is close to 1, allowing a larger step size for global search. As k increases, α gradually decreases, maintaining a minimum step size of at least 0.1, thus preventing search stagnation. Based on this factor, the position update of non-elite individuals can be expressed as formula (3-24):

[0198] x new =x current +(ba)×0.1×(rand-0.5)×α (3-25)

[0199] Where [a, b] is the permissible interval for the parameter, and rand represents a random number generated from the uniform distribution [0, 1]. This update formula ensures a high degree of exploration in the early stages while gradually focusing on the local optimal solution as the iterations proceed, thus achieving a dynamic balance between global search and local exploitation. The introduction of an adaptive step size factor not only alleviates the problem of excessive randomness in the later stages of the original algorithm, but also provides more refined control over the overall optimization process. This improvement significantly improves the algorithm's convergence speed and solution quality.

[0200] (3) Local search enhancement

[0201] Local search aims to further fine-tune elite individuals to achieve higher accuracy near the optimal solution. Specifically, each elite individual is subjected to multiple small random perturbations within its neighborhood, and the fitness of the new solutions is calculated. If a superior solution is found, it is immediately replaced. Setting the neighborhood perturbation step size to a small coefficient relative to the search space allows for fine-tuning of excellent solutions without disrupting the overall search. Combining local search with elite retention allows the algorithm to delve deeper into promising areas in later iterations.

[0202] The strategy achieves in-depth exploration of the local search space by applying small random perturbations in the neighborhood of the elite individual. Let the elite individual parameter be x_elite and the allowed interval be [a, b]. The ratio of the neighborhood perturbation step size to the search space is defined as β. Then the perturbation of the local search can be expressed as formula (3-26):

[0203] Δ=(ba)×β×(rand-0.5) (3-26)

[0204] Where rand represents a random number generated from a uniform distribution [0,1]. Furthermore, the update formula for the elite individual in the neighborhood is shown in (3-27):

[0205] x neighbor =x elite +Δ (3-27)

[0206] This ensures that the perturbation amplitude is small enough so that the search does not deviate from the global optimal region, while also fine-tuning the local optimal solution. When the newly generated neighborhood solution is better than the original elite solution, it can replace the original solution, further improving the search accuracy.

[0207] (4) Diversity protection

[0208] The diversity preservation strategy prevents the population from overconverging and becoming trapped in local optima during evolution by reinitializing the individuals with the worst fitness. In each iteration, the algorithm selects a certain proportion of the worst individuals, randomly generates new candidate solutions, and replaces them, thus maintaining population diversity.

[0209] The diversity protection strategy is to prevent the population from converging to the local optimum too early and maintain the population diversity by reinitializing individuals with poor fitness. Let the diversity ratio be p d , the number of worst individuals in the population N d It can be expressed as formula (3-28):

[0210]

[0211] Where N is the total population, Indicates the rounding operation. For these worst individuals, new candidate solutions will be obtained by random initialization, and the generation formula is as shown in (3-29):

[0212] x new =a+rand×(ba) (3-29)

[0213] The formula ensures that each reinitialized individual is evenly distributed in the search space, thereby injecting new vitality into the algorithm and avoiding falling into local optimality.

[0214] Through these improvements, IALO effectively overcomes the limitations of the original algorithm, such as slow or premature convergence in the late stages, in complex optimization scenarios. Elite retention strengthens the inheritance and utilization of excellent solutions, adaptive step size and local search enhance the algorithm's precision development capabilities, and diversity protection prevents the population from stagnating or reaching local extremes, making the entire optimization process more robust and efficient.

[0215] 2.7.3 Improved Antlion Optimization Algorithm Process

[0216] The overall process of the improved ant lion optimization algorithm is as follows Figure 8 As shown in the figure, at the beginning of the process, the algorithm initializes the population and sets key parameters, including the search boundary, elite ratio, and adaptive step size factor. Randomly generated ants initially form a diverse set of candidate solutions, and the individuals with the lowest loss are selected as antlions through fitness evaluation. The random walk process allows the ants to fully explore the solution space, and the fitness comparison between the ants and antlions is calculated in each iteration.

[0217] During the search process, the algorithm will mark high-performing individuals as elites based on the elite retention strategy and avoid random updates or reinitialization of these elites in subsequent iterations. The adaptive step size factor plays a role in this stage, achieving a dynamic balance between global and local optimization by adjusting the search stride. If an ant's fitness exceeds that of an antlion, the antlion's position will be replaced by that of the ant, thereby strengthening the group's inheritance of excellent solutions. The local search mechanism performs small perturbations around elite individuals in order to achieve higher accuracy near the optimal solution, while diversity protection reinitializes the worst individuals at the end of each iteration, providing the algorithm with the possibility of escaping local extremes.

[0218] The entire process ends when the termination condition is met, and the individual with the best current fitness is output as the final solution. Four improved strategies—elite retention, adaptive step size, local search, and diversity protection—are integrated throughout the algorithm's main steps. Through the orderly connection of linear expressions, they form an optimization closed loop, achieving faster convergence and better search quality in complex high-dimensional problems. The algorithm flow is shown in Table 3:

[0219] Table 3 Pseudocode of the improved ant lion algorithm

[0220]

[0221]

[0222] 3. Case test

[0223] 3.1 Test plan and indicators

[0224] This paper first compared three different data processing methods using MLP and LR. The data processing methods mainly include raw data processing, standardized data processing, and post-optimization data processing. For both raw and standardized data processing, standardized data showed the best training results, manifested in faster convergence and lower loss.

[0225] In terms of model evaluation, the present invention uses a variety of common evaluation indicators, including loss function, accuracy, confusion matrix and receiver operating characteristic curve (ROC curve), etc. These indicators together provide a comprehensive evaluation of model performance. The loss function is used to measure the difference between the model prediction value and the true value. The smaller the value, the better the model. The accuracy rate represents the proportion of correct predictions, which can directly reflect the classification effect of the model. The confusion matrix further reveals the classification performance of the model in each category by displaying data such as the model's true positive (TP), false positive (FP), true negative (TN) and false negative (FN). Finally, the ROC curve and the area under the curve (AUC) are important indicators for measuring the performance of the classification model. The larger the AUC, the better the performance of the model at different thresholds, and the better it can distinguish between positive and negative categories.

[0226] The fault determination criterion of the present invention is based on the variable stab in the data set. Specifically, a constant threshold is used for classification, wherein when stab>0, the transformer state is determined to be "faulty", and when stab<0, the transformer state is determined to be "normal".

[0227] In the subsequent optimization process, all adjustments and improvements to the model were based on standardized data, further verifying the key role of standardization in improving model performance.

[0228] 3.2 Performance Comparison Results

[0229] In the experiment, we first compared three different data processing methods (original data, normalized data, and standardized data) using MLP and LR. The specific accuracy comparison results are shown in Table 4:

[0230] Table 4 Comparison of machine learning algorithm accuracy

[0231]

[0232] As can be seen in the table above, the effect of normalizing the data on the LR model is significantly improved (0.85), while the performance of the MLP model on both normalized and standardized data is similar, both at 0.71. Therefore, data normalization is particularly effective for the LR model.

[0233] The experiment further compared the performance of the CNN model based on standardized data, the BiLSTM model, and the hybrid model of CNN and BiLSTM combined with the Attention mechanism on the original ALO and IALO algorithms.

[0234] During the model training process, the standardized transformer fault dataset was selected as the input data to ensure that the features of the data are at the same scale and reduce the impact of different row features on model training. By setting a series of appropriate training parameters, the accuracy and generalization ability of the model are improved. The training process is as follows: Figure 9 shown.

[0235] During model training, CNN, BiLSTM, and CNN-BiLSTM-Attention models were used. Training rounds were set to 50 to ensure sufficient model training. After training, the model performance was evaluated on a validation set using metrics such as accuracy, confusion matrix, AUC, and ROC curve. These metrics provide a comprehensive view of the model's classification performance. Table 5 further illustrates the training parameters used in the experiment.

[0236] Table 5 Training parameter settings

[0237]

[0238] The results show that IALO performs well on all models, especially in terms of accuracy. As the complexity of the model increases, the advantage of IALO becomes more obvious. The specific performance comparison results are shown in Table 6:

[0239] Table 6. Comparison of the accuracy of deep learning algorithms

[0240]

[0241] As can be seen from Table 6, IALO generally shows lower loss values ​​and higher accuracy compared to the traditional ALO algorithm. Especially on complex models such as CNN-BiLSTM-Attention, IALO's optimization effect is most significant.

[0242] After data normalization, the accuracy of the traditional LR model increased by approximately 32.81%. Among deep learning models, the CNN model improved by approximately 6.07% after using ALO optimization, and by approximately 3.41% through improved IALO. The BiLSTM model improved by approximately 6.31% and 2.40%, respectively. The complex CNN-BiLSTM-Attention model improved by approximately 3.86% and 2.95%, respectively.

[0243] These results show that the advantages of IALO become more pronounced as the model complexity increases, which also confirms the close connection between data preprocessing and optimization algorithm improvements. Through these optimization methods, the model can effectively improve classification performance when faced with more complex data.

[0244] 4. Practical application and engineering practice

[0245] 4.1 Actual Application Construction and Deployment

[0246] 4.1.1 Overall Architecture Design

[0247] The present invention takes the deep learning model as the core and completes fault prediction by acquiring transformer operation data, performing preprocessing operations, and calling the trained IALO-CNN-BiLSTM-Attention model. Figure 10 The prediction flowchart shown shows the main connection relationships between these modules and provides an overall framework reference for subsequent practical deployment.

[0248] The data acquisition link of the present invention assumes the responsibility of information collection. In practical applications, sensors are required to collect key indicators of the transformer in real time on site, including values ​​such as response time, nominal power and elastic coefficient. These indicators reflect the operating status of the equipment under different working conditions.

[0249] After data collection is complete, the original records are transmitted to the backend for centralized processing. This process often relies on network communication protocols and data storage middleware to ensure transmission efficiency and data integrity. Because the original data often contains features with different dimensions or significant distribution differences, these differences can affect the training stability and prediction accuracy of the model. To address this problem, the present invention uses a standardized method to transform the values ​​of each feature so that different features are within a comparable numerical range.

[0250] Understanding the distribution of data is crucial during data preprocessing and modeling. Figure 11 The distribution of stab is shown. The visualization of these data helps to understand the relationship between different features and target variables, especially how they affect the performance of the model. Figure 11 As can be seen from the figure, the distribution of the stab feature appears more dispersed, especially with a higher concentration of samples in the low-value range. This observation helps identify which features have a greater impact on model training and further supports the importance of data standardization and normalization, ensuring that the model can work efficiently in a relatively balanced data environment.

[0251] In deep learning, data preprocessing methods commonly used are data normalization and standardization. Data normalization adjusts the data scale to bring all eigenvalues ​​within the range [0, 1]. Standardization adjusts the data mean to 0 and the standard deviation to 1, giving the data a uniform distribution.

[0252] In order to achieve better prediction results, all input data are normalized to ensure that the value of each feature is compressed between 0 and 1. The normalization operation is implemented by formula (4-1):

[0253]

[0254] Among them, X represents the original data, X norm Represents the normalized data, where min(X) and max(X) are the minimum and maximum values ​​of feature X, respectively. Normalization can significantly accelerate the training process and avoid the gradient descent instability problem caused by differences in feature ranges. Standardization converts the data values ​​of each feature to a form with a mean of 0 and a standard deviation of 1. The formula is as follows:

[0255]

[0256] Where μ is the mean of the feature and σ is the standard deviation of the feature. Standardization can make the data have zero mean and unit variance, thereby reducing the impact of different features on model training to a certain extent and making the contribution of each feature to the final output more balanced.

[0257] Figure 12 The normalized data and the box plot of the standardized data are shown in . Figure 2 The box plot of the raw data shows a relatively dispersed distribution of features, particularly feature 5, which exhibits multiple outliers. The boxes for features 1 through 4 are smaller, indicating a relatively concentrated distribution of these features. However, feature 5 exhibits a wide range, with extreme values ​​at its upper edge, indicating its instability. Overall, most features in the raw data exhibit significant dispersion, which may affect subsequent modeling and training.

[0258] In the normalized data box plot, the number of outliers for feature 5 has decreased, while the distribution of other features is more even, with smaller boxes, indicating that the data has become more concentrated. Normalization scales the data to a relatively fixed range [0, 1], helping to reduce dimensional differences between features and thus improving model training. Standardized data further compresses the differences between features, making their distribution more compact. Normalization ensures consistent scaling of all features by adjusting the data to a distribution with a mean of 0 and a standard deviation of 1. As can be seen in the figure, the outlier for feature 5 has largely disappeared, and the overall data has become more balanced, with no significant skewness. This makes the data training and modeling process more stable, helping to improve model performance.

[0259] By comparing these three data processing methods, we can see that standardization plays an important role in eliminating the impact of extreme values ​​in the data and improving the balance of features.

[0260] Preprocessed data can better meet the feature distribution requirements of deep learning models and reduce the adverse effects of outliers or noise. The model prediction module performs fault diagnosis tasks based on the trained IALO-CNN-BiLSTM-Attention model. The present invention obtains data input from the preprocessing stage and loads it into the IALO-CNN-BiLSTM-Attention model. The model then outputs prediction results based on its internal parameters and network structure.

[0261] The combination of multiple technologies ensures both accuracy and robustness in the prediction process, while providing users with a reliable reference for fault warnings. The output phase presents the model's predictions to users in an intuitive manner. Users can choose to store the results in a database and display them through a visual interface or alarm module. Decision makers can use this information to assess the transformer's health and implement targeted maintenance strategies.

[0262] 4.1.2 Model deployment and operating environment configuration

[0263] This section focuses on deploying trained deep learning models into real-world applications and ensuring that the models can deliver real-world predictions. This process begins with model encapsulation, where the model is saved in a standard format, enabling seamless loading and callability across different platforms. Next, the preprocessing module converts transformer operation data, acquired in real time or batches, into a format with a uniform numerical range, ensuring that the input data meets the model's feature distribution requirements.

[0264] The deployed model will receive preprocessed data, perform fault prediction, and output prediction results. The experimental environment and platform setup are crucial for model training and application. A good environment not only speeds up experiments but also potentially improves model performance. Table 7 shows the hardware environment and deep learning framework configuration used in this paper.

[0265] Table 7 Operating platform and parameter configuration table

[0266]

[0267]

[0268] The operating environment and platform configuration provide a stable and efficient computing foundation for this invention, ensuring smooth operation and providing support for subsequent research on optimization models. Next, we will verify the actual prediction performance of the model on a large-scale dataset and evaluate its reliability in fault diagnosis tasks.

[0269] 4.2 Model Prediction

[0270] Based on the aforementioned system deployment and environment configuration, this section evaluates fault prediction using 15,000 real-world transformer data sets to validate the model's applicability in large-scale scenarios. After normalization, the data is fed into the trained IALO-CNN-BiLSTM-Attention model.

[0271] The model receives preprocessed data and automatically calculates the failure probability or classification result corresponding to each sample based on internal parameters and network structure. This step ensures the efficient conversion of data to prediction results.

[0272] This paper uses 15,000 data points to comprehensively test the model. Here, some of the prediction results are shown. The first few columns represent the input features (including response time tau, nominal power p, and elastic coefficient g), and the last column "stabf" is the prediction label given by the model, which is based on the variable stab.

[0273] Here, a constant threshold is used for stab classification: when stab > 0, the transformer is considered "faulty" and the stabf label is set to "0." When stab < 0, the transformer is considered "normal" and the stabf label is set to "1." In actual testing, comparisons of the confusion matrix, ROC curve, and AUC metric constructed from the predicted label (stabf) and the true state (stab) demonstrate the model's high classification performance and accuracy under large-scale data, further validating its reliability and superiority in practical applications. The model's prediction results are shown in Table 8.

[0274] Table 8 Model prediction results

[0275]

[0276] Table 8 shows the specific confusion matrix results. Analyzing the values ​​in the confusion matrix provides a deeper understanding of the model's performance in real-world applications. The confusion matrix shows the model's prediction accuracy for both stable and unstable states, with predictions reflected in four different regions.

[0277] Table 9 Confusion matrix results

[0278]

[0279]

[0280] As can be seen from the matrix, the fault prediction model demonstrates high accuracy in predicting stable states. Of the samples predicted as stable, 7,374 were correctly predicted as stable, while 147 were misclassified as unstable. The relatively small number of misclassified samples indicates that the fault prediction model accurately identifies stable states. On the other hand, the fault prediction model also demonstrates excellent classification capabilities in predicting unstable states. 7,326 samples were correctly predicted as unstable, while only 153 were misclassified as stable. The fault prediction model has a low misclassification rate.

[0281] Overall, the fault prediction model performed very well and exhibited strong stability. By comparing the number of false positives and correct classifications, the model's accuracy can be further quantified and improved using metrics such as accuracy, precision, recall, and F1 score. While the accuracy rate was 98.06%, demonstrating high accuracy, there is still room for improvement, particularly by optimizing the model architecture and enhancing feature selection to further enhance prediction robustness.

[0282] Figure 13 The ROC curves of various models are compared, with the Area Under the Circumference (AUC) value used as a metric to measure model classification performance. The CNN-BiLSTM-Attention-iALO model (CBATT iALO in the figure) achieves an AUC value of 0.98, demonstrating its superior ability to distinguish between positive and negative samples, far exceeding other models such as MLP and LR. This significant improvement in AUC demonstrates the model's superiority in processing complex data, particularly high-dimensional time series data, where the CBATT iALO model exhibits greater adaptability and generalization capabilities.

[0283] Practice has proven that each step in the prediction process is tightly integrated, forming a complete closed loop of data collection, preprocessing, model deployment, and result feedback, providing users with stable and rapid technical support. This closed-loop model not only helps improve the timeliness of operational decisions but also lays a solid foundation for subsequent promotion and optimization of the model in more complex environments.

Claims

1. A transformer fault diagnosis and prediction method, characterized in that: The steps include: By introducing elite preservation, adaptive step size, diversity protection and local search strategy into the original Ant Lion Optimization Algorithm, an improved Ant Lion Optimization Algorithm is obtained. An improved deep learning model is obtained by optimizing the deep learning model using an improved ant lion optimization algorithm; The preprocessed transformer fault dataset is used as input data to train the improved deep learning model and obtain the fault prediction model; Acquire and preprocess transformer operation data; Based on the fault prediction model, the preprocessed transformer operation data is input and the transformer fault diagnosis and prediction results are output.

2. The transformer fault diagnosis and prediction method according to claim 1, characterized in that: The elite retention strategies include: Assume that the population size is N and the elite retention ratio is ρ e , then the number of elite individuals N e for: in represents the rounding-up operation, which ensures that at least one optimal solution is retained when the population is small or the elite ratio is low; Let all candidate solutions in the population be {X1,X2,…,X N } and its corresponding fitness is {f(X1),f(X2),…,f(X N )}; Sort these fitness values ​​in ascending order, and the elite individual set is: where X (i) represents the i-th optimal solution.

3. The transformer fault diagnosis and prediction method according to claim 1 is characterized in that ,The adaptive step size strategy includes: Let the maximum number of iterations be K max ,At the kth iteration, the adaptive step size factor α is defined as: When k is small, α is close to 1, allowing a larger step size for global search. As k increases, α gradually decreases, and the minimum step size remains at least 0.1; Based on the adaptive step size factor α, the position update of non-elite individuals is expressed as: x new =x current +(b-a)×0.1×(rand-0.5)×α Where [a, b] is the allowed interval of the parameter, and rand represents a random number generated from a uniform distribution [0, 1], x current Represents the current position, x new Represents the updated position.

4. The transformer fault diagnosis and prediction method according to claim 1, characterized in that: The diversity protection strategies include: Let the diversity ratio be p d , the number of worst individuals in the population N d Expressed as: Where N is the total population, Represents the rounding-up operation. For these worst individuals, new candidate solutions will be obtained by random initialization. The formula for generating new candidate solutions is: x new =a+rand×(b-a)。 5. The transformer fault diagnosis and prediction method according to claim 1, characterized in that: The local search strategy includes: Assume that the elite individual parameter is x_elite and the allowed interval is [a, b]. The ratio of the neighborhood perturbation step size to the search space is defined as β. Then the perturbation of the local search is expressed as: Δ=(ba)×β×(rand-0.5) Where rand represents a random number generated from a uniform distribution [0,1]; The update formula of the elite individual in the neighborhood is: x neig hbor=x elite +Δ。 6. The transformer fault diagnosis and prediction method according to claim 1, characterized in that: The process of the improved ant lion optimization algorithm includes: Initialize the population and set the parameters search boundary, population size, maximum number of iterations, elite ratio, diversity ratio and number of local searches; Randomly generate N solutions and calculate the population fitness; Sort in ascending order of fitness, record the elite set and the worst set; Record the current optimal solution; Calculate the adaptive step size factor; Update the position of non-elite individuals and evaluate their fitness, and replace them if they are better; Perform local search on each elite and replace it if a better solution is found; Reinitialize the worst individual to maintain diversity and output iterative information; After iteration, the optimal solution and its fitness are output.

7. The transformer fault diagnosis and prediction method according to claim 1, characterized in that: The original ant lion optimization algorithm includes: Assume that the random walk step length is t, the maximum number of iterations is n, and the random function r(t) is defined as: Where rand represents a random number generated from a uniform distribution [0,1]; Use the accumulation and cumsum operation to construct the ant's random walk path X(t): X(t)=cumsum(2r(t)-1) The position of the i-th variable in the t-th generation The update formula is: where a i with d i are the minimum and maximum allowed values ​​of the random walk of the i-th variable, respectively, and and They represent the minimum and maximum values ​​of the i-th variable in the t-th iteration respectively; The update process is affected by the location of the Ant Lion Trap, which is mapped to: where c t and d t represent the minimum and maximum values ​​of all variables in the current iteration, respectively, and represents the position information of the j-th ant lion at iteration t; When an ant is updated, its fitness is Exceeds the current antlion's fitness When , the ant lion will capture the ant and replace its position with the new ant lion position, the formula is as follows:

8. The transformer fault diagnosis and prediction method according to claim 1, characterized in that: The deep learning model includes one of a convolutional neural network, a bidirectional long short-term memory network, and a hybrid model of a convolutional neural network, a bidirectional long short-term memory network, and an attention mechanism; The preprocessing method is normalization or standardization; The method for obtaining the transformer operation data is as follows: Use sensors to collect real-time operating status data of transformers under different working conditions on site; After data collection is completed, the original records are transferred to the backend for centralized preprocessing.

9. The transformer fault diagnosis and prediction method according to claim 1, characterized in that: The training process of the improved deep learning model is as follows: The standardized transformer fault dataset is divided into a training set and a validation set; Set training parameters and choose optimizer and loss function; Initialize the improved deep learning model, input the training set, calculate the training loss and update the weights, and evaluate the model performance on the validation set; Adjust hyperparameters, repeat training, and output the optimal model after reaching the training round.

10. A computer storage medium, characterized in that The computer storage medium stores instructions, which, when executed on a computer, enable the computer to perform the method according to any one of claims 1 to 9.