Transformer fault diagnosis method based on ANOVA-IBKA-CatBoost

By adopting the ANOVA-IBKA-CatBoost-based method in transformer fault diagnosis, the problem of insufficient modeling capabilities for complex nonlinear data in the prior art is solved, and higher diagnostic accuracy and stability are achieved.

CN120145209APending Publication Date: 2025-06-13CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510197808.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing transformer fault diagnosis methods lack the ability to model complex nonlinear relationships between data, and the diagnostic results are not stable enough when processing boundary data or noise data.

Method used

The transformer fault diagnosis method based on ANOVA-IBKA-CatBoost is adopted, and the classification gradient enhancement algorithm is optimized through ANOVA-IBKA-CatBoost algorithm, combined with lens imaging reverse learning strategy and adaptive t distribution strategy, the hyperparameters of the CatBoost model are optimized to improve diagnostic accuracy.

Benefits of technology

It improves the accuracy and robustness of transformer fault diagnosis, can process complex nonlinear data more effectively, enhances the processing ability of boundary data and noise data, and improves the stability of diagnostic results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145209A_ABST
    Figure CN120145209A_ABST
Patent Text Reader

Abstract

The invention discloses a transformer fault diagnosis method based on ANOVA-IBKA-CatBoost, and the method comprises the steps: collecting data of dissolved gas in transformer oil, and carrying out the expansion of original gas characteristics based on a gas ratio; carrying out F value calculation by adopting variance analysis, and measuring the relevance between each feature and the transformer fault; sorting is carried out according to F values of the features, the features are input to a CatBoost model for training, and an optimal feature subset is selected; introducing a lens imaging reverse learning strategy and a self-adaptive t distribution strategy to improve a black wing plinary algorithm; setting a hyper-parameter initial range of the CatBoost model, wherein the hyper-parameter initial range comprises the number of trees, the depth of the trees, the learning rate and the random intensity; the method comprises the following steps: taking a minimum error rate as a target function, performing hyper-parameter optimization by applying an improved black-wing plinergic algorithm, automatically adjusting hyper-parameters of a CatBoost model through multi-round iterative optimization of the improved black-wing plinergic algorithm, and ending circulation when the maximum number of iterations is reached. According to the method, statistical analysis, an intelligent optimization algorithm and a machine learning technology are fused, and the accuracy and robustness of transformer fault diagnosis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of transformer fault diagnosis, and particularly relates to a transformer fault diagnosis method based on ANOVA-IBKA-CatBoost. Background Technique

[0002] As a key device in the power system, the operating state of a transformer is directly related to the safety and stability of the power system. However, due to the long-term operation of the transformer in a complex electromagnetic and thermal environment, it is easily affected by load fluctuations, environmental factors, and aging effects, and various faults may occur, such as winding short circuits, insulation aging, and overheating faults. If these faults are not detected and processed in time, it will lead to serious economic losses and even safety accidents. Therefore, real-time monitoring of the health status of the transformer and timely diagnosis and discovery of potential faults have become an important means to ensure the stable operation of the power system.

[0003] At present, Dissolved Gas Analysis (DGA) is the main technical means for transformer fault diagnosis. It infers potential fault types by analyzing the gas components and their change rules in transformer oil. Traditional DGA methods include the IEC three-ratio method, Rogers four-ratio method, David triangle method, Duval method, etc. However, most of the above methods are based on empirical formulas or preset rules, lack the ability to model complex non-linear relationships between data, and the diagnostic results are not stable enough when dealing with boundary data or noisy data.

[0004] With the continuous development of artificial intelligence technology, machine learning has gradually become an important means to improve the accuracy and efficiency of transformer fault diagnosis. By introducing machine learning models, the limitations of traditional methods can be effectively overcome, and their powerful data processing capabilities can be utilized to provide more accurate and efficient fault diagnosis methods. Summary of the Invention

[0005] The present invention proposes a transformer fault diagnosis method based on Analysis of Variance (ANOVA) and an Improved Black Kite Algorithm Optimized Classification Gradient Boosting Algorithm (IBKA-CatBoost). This method combines statistical analysis, intelligent optimization algorithms, and machine learning technologies, aiming to improve the accuracy and robustness of transformer fault diagnosis and is applicable to the condition monitoring and fault diagnosis of transformers in power systems.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A transformer fault diagnosis method based on ANOVA-IBKA-CatBoost includes the following steps:

[0008] Step 1: Collect the dissolved gas data in transformer oil, expand the original gas characteristics based on gas ratios, and further explore potential fault information;

[0009] Step 2: Calculate the F-value of the expanded characteristic data using ANOVA (Analysis of Variance) to measure the correlation between each characteristic and transformer faults;

[0010] Step 3: Sort the characteristics according to their F-values, and gradually input the characteristics into the CatBoost model for training in the order of the sorting results. Evaluate the performance of each feature subset using the classification accuracy of the test set, and finally select the feature combination that maximizes the test set accuracy as the optimal feature subset;

[0011] Step 4: Introduce the lens imaging reverse learning strategy and the adaptive t-distribution strategy to improve the traditional Black Kite Algorithm (BKA);

[0012] Step 5: Set the initial range of the hyperparameters of the CatBoost model, mainly including the number of trees (iterations), the depth of the trees (depth), the learning rate (learning_rate), and the random strength (random_strength);

[0013] Step 6: Use the improved Black Kite Algorithm (IBKA) in Step 4 to optimize the hyperparameters with the objective of minimizing the error rate. Automatically adjust the hyperparameters of the CatBoost model through multiple rounds of iteration optimization of the improved Black Kite Algorithm (IBKA), and end the loop when the maximum number of iterations is reached;

[0014] Step 7: Divide the training set and the test set, train the CatBoost model with the best parameters, and comprehensively evaluate the performance of the ANOVA-IBKA-CatBoost model according to evaluation metrics such as accuracy, recall rate, and F1 score.

[0015] In the above Step 1, the expanded gas characteristics include: H 2 , CH 4 , C 2 H 6 , C 2 H 4 , C 2 H 2 , TH, C 2 H 2 / C 2 H 4 , CH 4 / H 2 , C 2 H 4 / C 2 H 6 H2 / CH 4 、H 2 / C 2 H 2 、H 2 / C 2 H 4 、H 2 / C 2 H 6 、H 2 / TH、CH 4 / TH、C 2 H 6 / TH、C 2 H4 / TH、C 2 H 2 / TH、(H 2 +CH 4 ) / TH、(H 2 +C2H 6 ) / TH、(H 2 +C 2 H 2 ) / TH、(CH 4 +C 2 H 6 ) / TH、(H 2 +C 2 H 2 +C 2 H 4 ) / TH, a total of 23-dimensional feature quantities, where TH represents the sum of four hydrocarbon gases dissolved in oil.

[0016] The dissolved gas data in transformer oil is shown in Table 1. This dataset covers six different fault types, namely medium and low temperature overheating, high temperature overheating, partial discharge, low energy discharge, high energy discharge, and normal state. The data in Table 1 provides a solid foundation for studying the correlation between different fault types and gas characteristics, helps to extract key features highly relevant to faults, thereby optimizing the fault diagnosis model and improving the accuracy and robustness of diagnosis.

[0017] Table 1 Partial data of dissolved gases in transformer oil

[0018]

[0019] In the second step described above, let the expanded feature data be: X = {x 1 , x 2 , …, x 23}, where each feature x i corresponds to n samples, and the label data is y, including six types of transformer fault types: medium and low temperature overheating fault, high temperature overheating fault, partial discharge, low energy discharge, high energy discharge, and normal. For the i-th feature xi , the F value is calculated by the following formula:

[0020]

[0021] In the formula: F i is a statistic that measures the ability of feature i to distinguish different transformer fault types in the dissolved gas data in oil. The larger its value, the higher the importance of this feature in fault diagnosis; MSB i is the mean square between groups of feature x i , representing the degree of change of the feature values between different fault categories; MSW i is the mean square within groups of feature x i , representing the degree of change of the feature values within the same fault category; c - 1 is the degree of freedom of MSB; n - c is the degree of freedom of MSW; n j is the number of samples in the j - th class; is the mean of x i in the j - th class; is the overall mean of x i ; x i,j,k is the feature value of the k - th sample in the j - th class. In step three, the F values of all feature data are sorted in descending order to obtain a feature sorted list {x (1) , x (2) ,..., x (23)}, where x (1) represents the feature with the strongest correlation with the fault; x (2) represents the feature with the second - strongest correlation with the fault; x (23) represents the feature with the lowest correlation with the fault.

[0022] In step three, the CatBoost model is a machine - learning algorithm based on gradient - boosted decision trees (GBDT), which can efficiently handle categorical data and the non - linear relationship between features and performs excellently in classification tasks. According to the gas data in step one and the method described in step two, the F values of 23 features are obtained, and the specific results are: 4.23, 2.03, 12.07, 11.28, 17.93, 1.95, 44.96, 30.18, 5.19, 7.06, 4.56, 10.02, 4.35, 6.88, 36.31, 68.29, 87.95, 20.17, 7.08, 6.88, 6.80.

[0023] The features sorted from largest to smallest F value are: C 2 H 2 / TH, C 2 H 4 / TH, (CH 4 +C 2 H6 ) / TH, C 2 H 2 / C 2 H 4 , C 2 H 6 / TH, CH 4 / H 2 , CH 4 / TH, C 2 H 2 , C 2 H 6 , C 2 H 4 , H 2 / C 2 H 4 , (H 2 +CH 4 ) / TH, H 2 / CH 4 , H 2 / TH, (H 2 +C 2 H 6 ) / TH, (H 2 +C 2 H 2 ) / TH, (H 2 +C 2 H 2 +C 2 H 4 ) / TH, C 2 H 4 / C 2 H 6 , H 2 / C 2 H 2 , H 2 / C 2 H 6 , H 2 , CH 4 , TH.

[0024] By incrementally sorting the features and using them as input features for the CatBoost model, and combining the validation metrics of the CatBoost model on the test set, the test set accuracy reaches its maximum when the number of features is 11. The optimal feature subset selected is: C 2 H 2 / TH, C 2 H 4 / TH, (CH 4 +C 2 H 6 ) / TH, C 2 H 2 / C 2 H4 , C 2 H 6 / TH, CH 4 / H 2 , CH 4 / TH, C 2 H 2 , C 2 H 6 , C 2 H 4 , H 2 / C 2 H 4 .

[0025] In the fourth step, an adaptive t - distribution strategy and a lens imaging reverse imaging strategy are introduced to improve the Black - winged Kite Algorithm (BKA). The specific process is as follows:

[0026] As a common probability distribution, the adaptive t - distribution has a long tail. Compared with the standard normal distribution, it can generate larger perturbations. This characteristic enables BKA to effectively jump out of local minima or maxima when searching for the global optimal solution. This strategy dynamically adjusts the degrees of freedom and gradually reduces the perturbation amplitude as the iteration process progresses to achieve a smooth transition from global search to local optimization. The probability density function of the adaptive t - distribution is:

[0027]

[0028] In Equation (4): v is the degrees of freedom, and its dynamic adjustment formula is Γ(·) is the gamma function.

[0029] The lens imaging reverse imaging strategy generates new solutions based on the principle of lens imaging, enabling the optimization process to search a broader search space, effectively avoiding the premature convergence problem of the algorithm, and further improving the accuracy of the solution and the global convergence ability. The core idea of this strategy is to dynamically generate a solution with reverse learning characteristics according to the relationship between the position of the current solution and the boundary of the search space; for an individual in the t - th iteration, its new position is calculated by the following formula:

[0030]

[0031] In Equation (5): New_BKA represents the newly generated individual, k is the dynamically updated factor, which gradually decreases as the iteration number t increases, specifically X(i) is the position of the current individual; ub represents the upper bound of the interval, and lb represents the lower bound of the interval. Step 4 also includes performing a performance test on IBKA and conducting a comparative analysis with BKA, PSO, and WOA. Through experiments, the performance of each algorithm in terms of optimization efficiency, convergence speed, accuracy, etc. is evaluated to verify the advantages of IBKA in transformer fault diagnosis and ensure the reliability and efficiency of its optimization process.

[0032] The test functions used are:

[0033]

[0034] In Equation (6): F 1 (x) represents a unimodal test function, and x j represents the j-th component in the vector x, and n represents the dimension of the variable.

[0035]

[0036] In Equation (7): F 2 (x) represents a multimodal test function, and x i represents the i-th element in the vector x.

[0037]

[0038] In Equation (8): F 3 (x) represents a multimodal test function, and e is the natural constant.

[0039]

[0040] In Equation (9): F 4 (x) represents a composite benchmark test function, and a i represents the i-th number in a set of known data, usually the target value, and b i represents the i-th number in a set of known data, usually a specific parameter, and x 1 、x 2 、x 3 、x 4 represent the input independent variables.

[0041] In Step 5, the initial parameter ranges are set as: iterations ∈ [50, 300], deep ∈ [4, 10], learning_rate ∈ [0.001, 1], random_strength ∈ [1, 6].

[0042] In Step 6, with minimizing the error rate as the objective function, the error rate expression is:

[0043]

[0044] where: N is the total number of samples, y i is the actual label of the i-th sample, is the predicted label of the i-th sample, is the indicator function; θ is the set hyperparameter, specifically iterations, depth, learning_rate, and random_strength; represents the predicted label of the i-th sample under the condition of the given parameter θ.

[0045] Apply the improved black-winged kite algorithm (IBKA) in step four for hyperparameter optimization:

[0046] The traditional black-winged kite algorithm includes three stages: initializing the population, the attack stage, and the migration stage.

[0047] First, set the population size, attack probability, and number of iterations of the improved black-winged kite algorithm (IBKA), and define the search range of hyperparameters for the CatBoost model. Take equation (10) as the objective function, and after initialization, the black-winged kite individuals enter the attack stage and the migration stage;

[0048] Next, according to equation (4), adopt the adaptive t-distribution strategy to dynamically adjust the black-winged kite individuals;

[0049] Subsequently, apply the lens imaging reverse imaging strategy according to equation (5) to obtain the best black-winged kite individual.

[0050] In each round of iteration, evaluate the performance of the CatBoost model under the current hyperparameter combination. Take the error rate as the evaluation index. If the current error rate is less than the error rate of the previous round of iteration, update the hyperparameters. Through multiple rounds of iteration optimization of the improved black-winged kite algorithm (IBKA), dynamically adjust the hyperparameters of the CatBoost model. When the maximum number of iterations is reached, end the loop.

[0051] In step seven, the formulas for each evaluation index are:

[0052]

[0053]

[0054] where: Accuracy represents the accuracy rate, Recall represents the recall rate, and Precision represents the precision rate. TP represents true positive, that is, the number of samples predicted as faulty and actually faulty; TN represents true negative, that is, the number of samples predicted as normal and actually normal; FP represents false positive, that is, the number of samples predicted as faulty but actually normal; FN represents false negative, that is, the number of samples predicted as normal but actually faulty.

[0055] A transformer fault diagnosis method based on ANOVA - IBKA - CatBoost has the following technical effects:

[0056] 1) In step 1 of the present invention, the original features are expanded through gas ratios, which can mine more potential information from the original data and enhance the expression ability of the features. The expanded features can capture more potential associations related to faults, providing more powerful inputs for subsequent analysis and model training.

[0057] 2) In step 2 of the present invention, ANOVA is used to analyze the variances of each feature to quantify the association strength between the features and the fault categories. Through F - value evaluation, features with higher discrimination ability can be effectively identified, redundant or irrelevant features can be effectively reduced, and the training efficiency and accuracy can be improved.

[0058] 3) In step 3 of the present invention, the features are gradually input into the CatBoost model according to the F - value ranking, which can accurately identify the features crucial for fault classification. Evaluating the performance of the feature subsets through the test - set accuracy helps to select the optimal feature combination and avoid overfitting or underfitting.

[0059] 4) In step 4 of the present invention, the lens imaging reverse learning strategy and the adaptive t - distribution strategy are introduced, which is beneficial to improving the search process of the Black - winged Kite Algorithm (IBKA) and avoiding falling into local optima. It enhances the global search ability of the optimization algorithm and quickly finds more suitable hyperparameters.

[0060] 5) In step 5 of the present invention, by setting the hyperparameters of the CatBoost model, the complexity and training efficiency of the CatBoost model can be controlled, the model performance can be improved, and overfitting can be avoided.

[0061] 6) In step 6 of the present invention, the IBKA algorithm is used to dynamically adjust the hyperparameters of the CatBoost model, reducing manual intervention and improving the optimization efficiency. By minimizing the error rate as the objective function, it ensures that the hyperparameter adjustment always focuses on enhancing the model's prediction ability. The multiple rounds of iteration of IBKA improve the accuracy of hyperparameter selection and reduce the uncertainty in the training process. Description of the Drawings

[0062] Figure 1 is the flowchart of IBKA.

[0063] Figure 2 is F 1 the iteration curve of the test function;

[0064] Figure 3 is F 2 the iteration curve of the test function;

[0065] Figure 4 is F 3Iteration curve of the test function;

[0066] Figure 5 is F 4 Iteration curve of the test function.

[0067] Figure 6 is the overall flowchart of the ANOVA-IBKA-CatBoost model. Specific implementation manner

[0068] A transformer fault diagnosis method based on ANOVA-IBKA-CatBoost comprehensively uses analysis of variance, intelligent optimization algorithms, and machine learning techniques. It conducts significant screening of gas characteristics based on analysis of variance, optimizes the hyperparameters of the CatBoost model by combining the improved black-winged kite algorithm, and obtains the CatBoost model with optimal parameters, thereby achieving high-precision diagnosis of transformer faults.

[0069] As Figure 1 shown, it is the flowchart of the improved black-winged kite algorithm IBKA, and the specific steps are as follows:

[0070] S1: Set parameters such as the attack probability, random seed, and maximum number of iterations;

[0071] S2: Initialize the population;

[0072] S3: Enter the attack stage;

[0073] S4: Enter the migration stage;

[0074] S5: Introduce the reverse learning strategy of lens imaging;

[0075] S6: Introduce the adaptive t-distribution strategy;

[0076] The initialization stage in S2 aims to create an initial population, which consists of a group of randomly generated solutions and is evenly distributed in the solution space. The initial position of each individual (i.e., the black-winged kite) is determined by the following formula:

[0077]

[0078] In formula (15): represents the position of the i-th individual in the j-th dimension; and are the lower and upper bounds of the j-th dimension respectively; rand is a random number in the interval [0, 1].

[0079] After initialization, select the individual with the optimal fitness value in the initial population as the leader, and its position X L is expressed as:

[0080]

[0081] In Equation (16): f(X i ) is the fitness function, used to evaluate the quality of individuals; i ∈ [1, pop] indicates that the solution index i is within the interval [1, pop], where pop is the population size.

[0082] The attack phase in S3 refers to the behavior of the black-winged kite using hovering observation and rapid diving during predation to achieve a dynamic balance between global exploration and local search. The mathematical model of the attack behavior is described as follows:

[0083]

[0084] In Equation (17): t is the number of iterations, and respectively represent the positions of the i-th individual in the j-th dimension at iteration times t and t + 1; r is a random number within [0, 1]; p is the attack probability, with a value of 0.9; n is a dynamic adjustment factor used to control the step size, defined as where T is the total number of iterations.

[0085] Through this model, BKA can effectively adjust the search range during the attack behavior to achieve fine search and rapid convergence of the solution space.

[0086] The migration phase in S4 refers to simulating the group migration pattern of the black-winged kite adapting to environmental changes. Its core idea is to dynamically select the optimal leader to guide the population to approach the optimal region of the solution. The mathematical model of the migration behavior is as follows:

[0087]

[0088] In Equation (18), represents the position of the i-th individual in the j-th dimension at iteration time t + 1. represents the current leader position in the j-th dimension; F i and F ri are the fitness values of the i-th individual and the randomly generated individual respectively; C(0, 1) is a perturbation factor introduced by the Cauchy distribution to enhance population diversity; m = 2sin(r + π / 2), where r is a random number.

[0089] In addition, the probability density function of the Cauchy distribution is expressed as:

[0090]

[0091] In Equation (19), δ represents the scale parameter of the distribution, controlling the width of the distribution, x represents the independent variable, μ represents the position parameter of the distribution, and C(0, 1) is f(x, 1, 0).

[0092] The migration behavior ensures the balance between the global and local population by dynamically adjusting the positional relationship between individuals and leaders, effectively avoiding being trapped in local optima.

[0093] Figures 2 to 5 Curves of the average fitness varying with the number of iterations for BKA, GWO, PSO, and IBKA under complex functions such as unimodal and multimodal functions after 30 independent experiments. It can be seen from the curve results that the improved black-winged kite algorithm is superior to the other three algorithms in terms of convergence performance and optimization ability.

[0094] To further apply the improved black-winged kite algorithm, Figure 6 For the ANOVA-IBKA-CatBoost model, the overall model steps are as follows:

[0095] Step (1): Collect data on dissolved gases in transformer oil on-site.

[0096] Step (2): Effectively expand the features according to the gas ratio, deriving 23 feature dimensions.

[0097] Step (3): Calculate the F-value of each feature using ANOVA.

[0098] Step (4): Arrange the features in descending order of the F-value and input them into the CatBoost model in sequence. Select the optimal features according to the accuracy index.

[0099] Step (5): With the minimization of the error rate as the objective function, set the search range for each parameter of CatBoost, and use IBKA to dynamically optimize each parameter to find the best hyperparameter combination.

[0100] Step (6): Divide the training set and the test set, set the best hyperparameters, train the model and conduct fault diagnosis, and finally evaluate the model effect according to various performance indicators.

[0101] The error rate expression in step (5) is:

[0102]

[0103] In Equation (10): N is the total number of samples, y i The actual label of the i-th sample, is the predicted label of the i-th sample, is the indicator function; θ is the set hyperparameter, specifically iterations, depth, learning_rate, and random_strength; represents the predicted label of the i-th sample under the condition of the given parameter θ.

[0104] In summary, the present invention proposes a transformer fault diagnosis method based on ANOVA-IBKA-CatBoost. By implementing key steps such as the collection of dissolved gas data in oil, feature expansion and screening, and hyperparameter optimization, it can effectively improve the accuracy and reliability of fault diagnosis and provide effective guiding opinions for on-site fault detection.

Claims

1. Transformer fault diagnosis method based on ANOVA-IBKA-CatBoost, characterized by The following steps are involved: Step 1: Collect dissolved gas data in transformer oil and expand the original gas features based on gas ratios; Step 2: Use variance analysis to calculate the F value of the expanded feature data to measure the correlation between each feature and the transformer fault; Step 3: Sort the features according to their F values, and input the features into the CatBoost model for training step by step according to the sorting results, and select the optimal feature subset; Step 4: Introduce the lens imaging reverse learning strategy and adaptive t-distribution strategy to improve the Black Kite algorithm; Step 5: Set the initial range of hyperparameters of the CatBoost model, including the number of trees, tree depth, learning rate, and random intensity; Step 6: Taking minimizing the error rate as the objective function, apply the improved Black Kite algorithm in step 4 to optimize the hyperparameters. Through multiple rounds of iterative optimization of the improved Black Kite algorithm, the hyperparameters of the CatBoost model are automatically adjusted, and the loop ends when the maximum number of iterations is reached.

2. The transformer fault diagnosis method based on ANOVA-IBKA-CatBoost according to claim 1 is characterized in that: It also includes step seven: dividing the training set and test set, setting the CatBoost model with the best parameters for training, and comprehensively evaluating the model performance based on evaluation indicators such as accuracy, recall, and F1 score.

3. The transformer fault diagnosis method based on ANOVA-IBKA-CatBoost according to claim 1 is characterized in that: In the step one, the expanded gas features include: H2, CH4, C2H6, C2H4, C2H2, TH, C2H2 / C2H4, CH4 / H2, C2H4 / C2H6, H2 / CH4, H2 / C2H2, H2 / C2H4, H2 / C2H6, H2 / TH, CH4 / TH, C2H6 / TH, C2H4 / TH, C2H2 / TH, (H2+CH4) / TH, (H2+C2H6) / TH, (H2+C2H2) / TH, (CH4+C2H6) / TH, (H2+C2H2+C2H4) / TH, a total of 23 dimensions of features, wherein TH represents the sum of the four hydrocarbon gases dissolved in the oil.

4. The transformer fault diagnosis method based on ANOVA-IBKA-CatBoost according to claim 1 is characterized in that: In the step 2, the expanded feature data is assumed to be: X = {x1, x2, ..., x 23 }, Among them, each feature xi corresponds to n samples, and the label data is y, including medium and low temperature overheating fault, high temperature overheating fault, partial discharge, low energy discharge, high energy discharge and normal, a total of 6 types of transformer fault types; for the i-th feature x i , the F value is calculated by the following formula: Where: F i It is a statistic that measures the ability of feature i to distinguish different transformer fault types in the dissolved gas in oil data. The larger its value, the more important the feature is in fault diagnosis. i is feature x i The mean square error between groups indicates the degree of variation of the characteristic value between different fault categories; MSW i is feature x i The within-group mean square error indicates the degree of variation of the characteristic value within the same fault category; c-1 is the degree of freedom of MSB; nc is the degree of freedom of MSW; n j is the number of samples in the jth class; is x in the jth class i The mean of For x i The overall mean of x i,j,k is the eigenvalue of the kth sample in the jth class.

5. The transformer fault diagnosis method based on ANOVA-IBKA-CatBoost according to claim 1 is characterized in that: In step 3, the F values ​​of all feature data are sorted in descending order to obtain a feature sorting list {x (1) ,x (2) ,...,x (23) }, where x (1) Indicates the feature most associated with the fault; x (2) Indicates the feature with the second strongest fault correlation; x (23) Indicates the feature with the lowest fault correlation.

6. The transformer fault diagnosis method based on ANOVA-IBKA-CatBoost according to claim 3 is characterized in that: In step 3, according to the F values ​​of the 23 features obtained, the features after sorting the F values ​​from large to small are: C2H2 / TH, C2H4 / TH, (CH4+C2H6) / TH, C2H2 / C2H4, C2H6 / TH, CH4 / H2, CH4 / TH, C2H2, C2H6, C2H4, H2 / C2H4, (H2+CH4) / TH, H2 / CH4, H2 / TH, (H2+C2H6) / TH, (H2+C2H2) / TH, (H2+C2H2+C2H4) / TH, C2H4 / C2H6, H2 / C2H2, H2 / C2H6, H2, CH4, TH; By sorting the features in ascending order and using them as input features of the CatBoost model, combined with the validation indicators of the CatBoost model on the test set, the test set accuracy reaches the maximum when the number of features is 11, and the optimal feature subsets are screened out: C2H2 / TH, C2H4 / TH, (CH4+C2H6) / TH, C2H2 / C2H4, C2H6 / TH, CH4 / H2, CH4 / TH, C2H2, C2H6, C2H4, and H2 / C2H4.

7. The transformer fault diagnosis method based on ANOVA-IBKA-CatBoost according to claim 1 is characterized in that: In step 4, the adaptive t distribution strategy and the lens imaging reverse imaging strategy are introduced to improve the black kite algorithm. The specific process is as follows: The probability density function of the adaptive t distribution is: In formula (4), v is the degree of freedom, and its dynamic adjustment formula is: Γ(·) is the gamma function; The lens imaging reverse imaging strategy dynamically generates a solution with reverse learning characteristics according to the relationship between the position of the current solution and the boundary of the search space; for an individual in the tth iteration, its new position is calculated by the following formula: In formula (5), New_BKA represents the newly generated individual, k is the dynamic update factor, which gradually decreases with the increase of the number of iterations t; X(i) is the position of the current individual; ub represents the upper bound of the interval, and lb represents the lower bound of the interval.

8. The transformer fault diagnosis method based on ANOVA-IBKA-CatBoost according to claim 1 is characterized in that: The step 4 also includes a performance test of the improved black kite algorithm, and the test function used is: In formula (6): F1(x) represents the single-peak test function, x j represents the jth component in vector x, and n represents the dimension of the variable; In formula (7): F2(x) represents the multi-peak test function, x i represents the i-th element in vector x; In formula (8): F3(x) represents the multi-peak test function, e is a natural constant; In formula (9): F4(x) represents the composite benchmark function, a i Represents the i-th number of a set of known data, b i represents the i-th number in a set of known data, and x1, x2, x3, and x4 represent the input independent variables.

9. The transformer fault diagnosis method based on ANOVA-IBKA-CatBoost according to claim 7 is characterized in that: In step 6, the objective function is to minimize the error rate, and the error rate expression is: Where: N is the total number of samples, y i The actual label of the i-th sample, is the predicted label of the i-th sample, is the indicator function; θ is the set hyperparameter; It represents the predicted label of the i-th sample under the given parameter θ.

10. The transformer fault diagnosis method based on ANOVA-IBKA-CatBoost according to claim 9 is characterized in that: The improved black kite algorithm in step 4 is used to optimize hyperparameters, as follows: The black kite algorithm consists of three stages: initialization of the population, attack stage and migration stage. First, the population size, attack probability and number of iterations of the improved black kite algorithm are set, and the optimization range of hyperparameters is defined for the CatBoost model. Formula (10) is used as the objective function, and after initialization, the black kite individuals enter the attack stage and migration stage. Then, according to formula (4), the adaptive t distribution strategy is adopted to dynamically adjust the black-winged kite individuals; Then, the lens imaging reverse imaging strategy is applied according to formula (5) to obtain the best black-winged kite individual; In each round of iteration, the performance of the CatBoost model under the current hyperparameter combination is evaluated; the error rate is used as the evaluation indicator. If the current error rate is less than the error rate of the previous round of iteration, the hyperparameters are updated; the hyperparameters of the CatBoost model are dynamically adjusted through multiple rounds of iterative optimization of the improved Black Kite algorithm; when the maximum number of iterations is reached, the loop ends.