Oil-immersed transformer fault diagnosis method, system and readable storage medium
The SMOTE-XGBoost algorithm is used to solve the sample imbalance problem in oil-immersed transformer fault diagnosis, improve the diagnostic accuracy and reliability, reduce operation and maintenance costs, and achieve efficient fault diagnosis.
Patent Information
- Application Number
- CN202111414926.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-25
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2041-11-25
AI Technical Summary
Existing fault diagnosis methods for oil-immersed transformers suffer from sample imbalance, resulting in low accuracy of traditional artificial intelligence algorithms, unreliable diagnostic results, lack of high-quality historical data sets, and reliance on the experience of maintenance personnel.
SMOTE oversampling technology and XGBoost ensemble learning algorithm are used for transformer fault diagnosis. By synthesizing minority class samples and gradient boosting decision trees, the reliability and generalization ability of model training results are improved.
It improves the accuracy of oil-immersed transformer fault diagnosis, reduces operation and maintenance costs, saves manpower and material resources, and improves power supply reliability.
Smart Images

Figure CN114358116B_ABST
Abstract
Description
Technical field
[0001] The present invention relates to the technical field of transformer fault diagnosis, which is implemented based on a SMOTE-XGBoost ensemble learning algorithm and specifically provides a method, system and readable storage medium for diagnosing faults of oil-immersed transformers. [Background Technology]
[0002] Transformer fault diagnosis is a crucial component of daily substation maintenance. It impacts the safety and reliability of power supply and is a key component in ensuring the safe operation of the power grid. Transformer fault diagnosis involves monitoring the condition of substation equipment, analyzing and mining key transformer status data, and using experience or big data analysis algorithms to determine the likely type of transformer fault and schedule maintenance. With the application of big data technology and advanced sensors, transformer fault diagnosis strategies have evolved from passive maintenance, such as corrective and scheduled maintenance, to proactive maintenance based on condition monitoring. Proactive maintenance based on oil-immersed transformer condition monitoring involves obtaining transformer status information through sensors and other instruments while the equipment is operating, without requiring equipment downtime. Using relevant big data mining algorithms, the transformer's actual operating conditions are analyzed to identify and predict potential faults.
[0003] Currently, the fault diagnosis methods for oil-immersed power transformers are basically developed based on dissolved gas analysis (DGA). The analysis methods based on dissolved gas in oil can be divided into traditional qualitative empirical methods and data-driven artificial intelligence methods.
[0004] Qualitative empirical methods are based on extensive practical experience and summarize data to identify transformer fault types. These methods can be categorized into the characteristic gas method and the three-ratio method. The characteristic gas method separates the volatile gases generated by the decomposition of oil-immersed transformer insulating oil into primary and secondary gases. Based on extensive practical experience, the concentrations of these gases are then used to directly determine the transformer fault type. While this method is simple and easy to implement, it lacks analysis of the relationships between characteristic gases, is highly subjective, and relies heavily on empirical experience, making it difficult to apply on a large scale. The three-ratio method uses the ratios of individual characteristic gases dissolved in the oil-immersed transformer insulating oil to diagnose transformer faults. Characteristic gases are selected, three ratios of their concentrations are generated, and these ratios are encoded. The combination of these ratios is then used to determine the transformer fault type. This method considers some simple coupling between gases and uses the ratios of the characteristic gases as a basis for judgment. However, it lacks a rigorous mathematical formula and is essentially still a qualitative judgment based on empirical experience, thus remaining highly subjective.
[0005] Data-driven AI-based diagnostic methods utilize extensive historical transformer operating data to train models. These models are capable of analyzing the complex coupling relationships between different characteristic gases. These models possess rigorous mathematical derivations, avoid subjectivity, and offer high accuracy in diagnostic results. Furthermore, they enable online diagnosis, making them a focus of research in transformer fault diagnosis both domestically and internationally. Data-driven AI-based transformer fault diagnosis methods include: Multilayer Perceptron (MLP), K-Nearest Neighbor (KNN), Auto-Encoder Network (AE), Fuzzy Logic (FL), Support Vector Machine (SVM), Decision Tree (DT), Gradient Boosting Decision Tree (GBDT), and Random Forest. While traditional AI-based transformer fault diagnosis methods can achieve good results, they still face numerous challenges. Transformer failures are inherently minority events. Therefore, the majority of transformer operating data we obtain often corresponds to normal operation, with only a very small number of faulty samples. This makes data-driven transformer fault diagnosis an unbalanced learning process. Transformer fault types are categorized into overheating and discharge faults, each with varying degrees of severity. Therefore, actual transformer fault diagnosis is a multi-classification problem based on an unbalanced sample distribution. Unbalanced samples cause traditional artificial intelligence algorithms to prioritize the majority class samples during model learning, ignoring the minority class samples. This significantly impairs model learning, reduces generalization, and ultimately leads to significant bias in diagnostic and classification results. Furthermore, for multi-classification problems, especially those with unbalanced samples, simply selecting accuracy is not an appropriate evaluation metric. While high accuracy can be achieved by assuming all results are normal, such results are inaccurate. Therefore, selecting an appropriate model algorithm is crucial for achieving effective transformer fault diagnosis.
[0006] The transformer fault diagnosis method based on the SMOTE-XGBoost algorithm is specifically designed to address the sample imbalance problem in transformer data. It effectively handles unbalanced learning, resulting in more reliable model training results, higher fault diagnosis efficiency, and stronger generalization capabilities. The synthetic minority over-sampling technique (SMOTE) improves the overfitting problem often associated with random oversampling algorithms. Its basic idea is to artificially synthesize new samples based on the k-nearest neighbors of minority class samples, thereby achieving a balanced number of samples across all classes in the dataset. XGBoost (eXtreme Gradient Boosting) is an improved algorithm based on the gradient boosting decision tree (GBDT), aiming to maximize the algorithm's speed and efficiency. Compared to GBDT, XGBoost's greatest improvement lies in the setting of its objective function, which is composed of the sum of a loss function and a regularization term. Among them, the expansion of the second-order derivative of Taylor's formula makes the gradient descent of the loss function faster and more accurate, and the results are more precise; the use of regularization terms controls the complexity of the model, avoids model overfitting to the greatest extent, and improves the model's generalization ability.
[0007] The factors that cause transformer failures are complex and intertwined. Furthermore, actual substations contain numerous transformers, making manual maintenance a massive workload that relies heavily on the maintenance personnel's experience and is subject to significant subjectivity. While traditional AI-based fault diagnosis algorithms overcome this subjectivity and significantly reduce the workload, they also face challenges such as a lack of high-quality historical datasets and a severely unbalanced dataset distribution. This results in low accuracy and unreliable diagnostic results based on these algorithms.
[0008] Therefore, it is necessary to study a method, system and readable storage medium for oil-immersed transformer fault diagnosis to address the shortcomings of the existing technology and to solve or alleviate one or more of the above problems. [Summary of the invention]
[0009] In view of this, the purpose of the present invention is to make up for the shortcomings of the existing oil-immersed transformer fault diagnosis algorithm based on traditional artificial intelligence algorithm. The present invention performs transformer fault diagnosis based on SMOTE oversampling technology and XGBoost ensemble learning algorithm, and provides an oil-immersed transformer fault diagnosis method, system and readable storage medium to improve the accuracy of oil-immersed transformer fault diagnosis and improve power supply reliability.
[0010] In one aspect, the present invention provides a method for diagnosing a fault of an oil-immersed transformer, the method comprising:
[0011] S1: Obtain transformer fault samples and normal samples from the historical database, and obtain the corresponding gas content under different fault conditions and normal conditions from the samples through gas analysis method;
[0012] S2: Filter the gases in S1, select gas data with concentration and integrity as the characteristic gas input model, and encode the transformer state corresponding to the characteristic gas into integers as the output result of model training;
[0013] S3: Preprocess the model’s input data and output results, and retain valid data as the first training set;
[0014] S4: performing stratification, sample synthesis balance and standardization on the first training set to obtain the second training set;
[0015] S5: Input the second training set into the XGBoost algorithm for training and testing;
[0016] S6: Transformer fault diagnosis is performed based on the macro-averaged precision, recall, and F1 score of the training and testing results in S5.
[0017] According to the aspects and any possible implementations described above, an implementation is further provided, wherein the types of gases obtained in different fault states and normal states in S1 include but are not limited to H2, CH4, C2H6, C2H4, C2H2, CO and CO2.
[0018] According to the above aspects and any possible implementation, an implementation is further provided, wherein the characteristic gases in S2 include H2, CH4, C2H6, C2H4 and C2H2.
[0019] According to the above aspects and any possible implementation, an implementation is further provided, wherein the transformer status in S2 includes normal, partial discharge, low energy discharge, high energy discharge, medium and low temperature overheating, and high temperature overheating.
[0020] According to the above aspects and any possible implementation, an implementation is further provided, wherein the data preprocessing in S3 is specifically: eliminating the results with missing values, extreme values and / or negative values in the data, and retaining valid data.
[0021] According to the aspects and any possible implementation methods described above, an implementation method is further provided, wherein the hierarchical division in S4 is specifically as follows: the ratio of transformer fault samples and normal samples in the historical database is used as a preset ratio condition, and the first training set is hierarchically divided in the same proportion.
[0022] According to the above aspects and any possible implementation, an implementation is further provided, wherein the sample synthesis balance in S4 is specifically: using the synthetic minority sample method SMOTE to identify and analyze the minority samples and synthesize new minority samples.
[0023] According to the above aspects and any possible implementation, an implementation is further provided, wherein the standardization processing in S4 specifically comprises: performing standardization processing on the data set, and the standardization formula adopts:
[0024]
[0025] Among them, x * represents the standardized data, x is the original data, and x std are the mean and standard deviation respectively.
[0026] According to the above aspects and any possible implementation, a fault diagnosis system for an oil-immersed transformer is further provided, for completing any one of the fault diagnosis methods described above, wherein the fault diagnosis system comprises:
[0027] A data acquisition module is used to obtain the corresponding gas content under different fault conditions and normal conditions;
[0028] The data screening module is used to select gases with concentration and integrity as characteristic gases to be input into the model, and integer encode the transformer states corresponding to the characteristic gases as the output results of the model training;
[0029] The data preprocessing module is used to preprocess the input model and output results, and retain valid data as the first training set;
[0030] A stratified balancing and standardization module is used to perform stratification, sample synthesis balancing and standardization on the first training set to obtain a second training set;
[0031] The model training and testing module is used to input the second training set into the XGBoost algorithm for training and testing;
[0032] The fault diagnosis module diagnoses transformer faults based on the results of the model training and testing module through the macro-averaged accuracy, recall rate and F1 score.
[0033] According to the above aspects and any possible implementation, a readable storage medium is further provided, including: a memory storing a program; a processor, which implements any one of the fault diagnosis methods when executing the program.
[0034] Compared with the prior art, the present invention can achieve the following technical effects:
[0035] 1): The present invention adopts SMOTE oversampling technology to solve the serious imbalance problem of classification samples in transformer fault diagnosis, especially the serious scarcity of fault samples, and improves the reliability of subsequent artificial intelligence algorithm training model.
[0036] 2) This paper uses the XGBoost ensemble learning algorithm for transformer fault diagnosis. This model is efficient, flexible, robust, and has strong generalization capabilities. Compared with traditional artificial intelligence algorithms, this model can make transformer fault diagnosis results more accurate, efficient, and reliable, effectively saving significant manpower and material resources and reducing operation and maintenance costs.
[0037] Of course, any product implementing the present invention does not necessarily need to achieve all of the above-mentioned technical effects at the same time.
Brief Description of the Drawings
[0038] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0039] Figure 1 This is a distribution diagram of various types of transformer sample data in a data set provided by an embodiment of the present invention;
[0040] Figure 2 1 is a logarithmic loss curve diagram of a transformer fault diagnosis model provided by an embodiment of the present invention during each round of iteration;
[0041] Figure 3 is a classification error rate curve diagram of the transformer fault diagnosis model provided by one embodiment of the present invention during each round of iteration;
[0042] Figure 4 This is a confusion matrix diagram for transformer fault diagnosis provided by one embodiment of the present invention;
[0043] Figure 5 The figure is a flowchart of the steps of a fault diagnosis method provided by an embodiment of the present invention. [Specific implementation method]
[0044] In order to better understand the technical solution of the present invention, the embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0045] It should be understood that the embodiments described are only a portion of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by persons of ordinary skill in the art without creative work are within the scope of protection of the present invention.
[0046] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "an", "the" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.
[0047] The present invention performs transformer fault diagnosis based on SMOTE oversampling technology and XGBoost ensemble learning algorithm, and provides an oil-immersed transformer fault diagnosis method, system and readable storage medium to improve the accuracy of oil-immersed transformer fault diagnosis and improve power supply reliability.
[0048] A method for diagnosing faults of an oil-immersed transformer, the method comprising:
[0049] S1: Obtain transformer fault samples and normal samples from the historical database, and obtain the corresponding gas content under different fault conditions and normal conditions from the samples through gas analysis method;
[0050] S2: Filter the gases in S1, select gas data with concentration and integrity as the characteristic gas input model, and encode the transformer state corresponding to the characteristic gas into integers as the output result of model training;
[0051] S3: Preprocess the model’s input data and output results, and retain valid data as the first training set;
[0052] S4: performing stratification, sample synthesis balance and standardization on the first training set to obtain the second training set;
[0053] S5: Input the second training set into the XGBoost algorithm for training and testing;
[0054] S6: Transformer fault diagnosis is performed based on the macro-averaged precision, recall, and F1 score of the training and testing results in S5.
[0055] The gas types obtained in S1 under different fault states and normal states include but are not limited to H2, CH4, C2H6, C2H4, C2H2, CO and CO2. The characteristic gases in S2 include H2, CH4, C2H6, C2H4 and C2H2.
[0056] The transformer states in S2 include normal, partial discharge, low energy discharge, high energy discharge, medium and low temperature overheating, and high temperature overheating.
[0057] The data preprocessing in S3 specifically includes: eliminating the results with missing values, extreme values and / or negative values in the data, and retaining valid data.
[0058] The stratification in S4 is specifically as follows: the ratio of transformer fault samples to normal samples in the historical database is used as a preset ratio condition, and the first training set is stratified in the same ratio. The sample synthesis balance is specifically as follows: the minority class sample synthesis method SMOTE is used to identify and analyze the minority class samples and synthesize new minority class samples. The standardization process is specifically as follows: the data set is standardized, and the standardization formula is:
[0059]
[0060] Among them, x * represents the standardized data, x is the original data, and x std are the mean and standard deviation respectively.
[0061] The present invention also provides an oil-immersed transformer fault diagnosis system for implementing any one of the above-mentioned fault diagnosis methods, wherein the fault diagnosis system comprises:
[0062] A data acquisition module is used to obtain the corresponding gas content under different fault conditions and normal conditions;
[0063] The data screening module is used to select gases with concentration and integrity as characteristic gases to be input into the model, and integer encode the transformer states corresponding to the characteristic gases as the output results of the model training;
[0064] The data preprocessing module is used to preprocess the input model and output results, and retain valid data as the first training set;
[0065] A stratified balancing and standardization module is used to perform stratification, sample synthesis balancing and standardization on the first training set to obtain a second training set;
[0066] The model training and testing module is used to input the second training set into the XGBoost algorithm for training and testing;
[0067] The fault diagnosis module diagnoses transformer faults based on the results of the model training and testing module through the macro-averaged accuracy, recall rate and F1 score.
[0068] The present invention also includes a readable storage medium, comprising: a memory storing a program; and a processor, which implements any one of the fault diagnosis methods when executing the program.
[0069] The actual operation of the oil-immersed transformer fault diagnosis method based on SMOTE-XGBoost of the present invention is as follows:
[0070] S1: Use the transformer oil chromatography online monitoring system to collect the gas content in the oil of the oil-immersed transformer during operation, and select as much oil-immersed transformer oil gas content data under normal operating conditions and fault conditions as possible, namely DGA data.
[0071] S2: Determine five characteristic gases, H2, CH4, C2H6, C2H4, and C2H2, as model inputs. Six transformer status types, including normal, partial discharge, low-energy discharge, high-energy discharge, medium- and low-temperature overheating, and high-temperature overheating, are encoded with integers from 0 to 5 as model outputs.
[0072] S3: Preprocess the data. Simply delete missing data and remove some extreme values or negative values.
[0073] S4: Then the preprocessed dataset is divided into training set and test set in a stratified manner with a ratio of 8:2.
[0074] S5: Use SMOTE oversampling technology to synthesize minority class samples in the training set data to achieve the purpose of balancing the data set.
[0075] S6: Perform 0-1 normalization on the input data of the training set and test set to avoid numerical calculation problems that may occur during model training.
[0076] S7: Use the XGBoost ensemble learning algorithm to train the model on the training set data after SMOTE balancing, and then use the trained XGBoost model to predict the test set data to obtain the predicted diagnostic results.
[0077] S8: Use the evaluation indicators of multi-classification problems to evaluate the performance of the model on the diagnostic results of the test set. Based on the evaluation results, decide whether it is necessary to reversely adjust the model parameters and conduct multiple experiments to obtain the desired model performance.
[0078] The working principle of the present invention is as follows:
[0079] This paper discloses a method for oil-immersed transformer fault diagnosis based on the SMOTE-XGBoost ensemble learning algorithm. Its working principle is achieved through the following detailed steps:
[0080] S1: Based on prior knowledge, the oil gas analysis method, which is commonly used for oil-immersed transformer fault diagnosis, is used to select as many transformer fault samples and normal samples as possible. Each sample contains the various gas contents in the transformer oil and the corresponding fault status.
[0081] S2: The cracking of insulating oil in oil-immersed transformers produces a variety of gases, including H2, CH4, C2H6, C2H4, C2H2, CO, and CO2. Since CO and CO2 are very rare, dispersed, and incompletely collected on-site, five characteristic gases, H2, CH4, C2H6, C2H4, and C2H2, were selected as model inputs. The concentrations of these five characteristic gases in the oil vary depending on the transformer's operating state. According to my country's "Guidelines for Analysis and Judgment of Dissolved Gases in Transformer Oil" (DL / T722-2000), transformer fault types are classified into six states: normal, partial discharge, low-energy discharge, high-energy discharge, medium-low temperature overheating, and high temperature overheating. These six states are each assigned an integer code ranging from 0 to 5, which serves as the output of model training.
[0082] S3: Because transformer failures are inherently rare events, the collected data cannot be time series data. Actual data collected may contain missing values, extreme values, and even negative values, so preprocessing is required. Since it is not time series data, there is no correlation between previous and subsequent data, so missing values, extreme values, and even negative values are simply removed.
[0083] S4: Because transformer fault samples have a severely unbalanced distribution—a large number of samples correspond to normal conditions, while a small number correspond to faulty conditions—a simple random partitioning strategy cannot be used when dividing the dataset into training and test sets. This will lead to a significant imbalance between normal and faulty conditions in the training and test sets, disrupting subsequent model training and causing significant errors. Therefore, a stratified partitioning principle should be adopted to ensure that the proportion of each type of sample in the training and test sets is consistent with that in the original dataset.
[0084] S5: After stratification, the proportions of each class of samples in the training set are the same as those in the metadata set, indicating a significant imbalance in the distribution of samples. This imbalance significantly impairs the training and learning process of AI algorithms, as the model learning process tends to favor majority class samples and neglect minority class samples. This results in the learned model having excellent recognition performance for majority class samples but poor performance for minority class samples, significantly reducing the model's generalization ability. Therefore, the sample distribution of the training set needs to be balanced. The test set is used for model evaluation, so the test set samples do not need to be balanced.
[0085] S6: Using the synthetic minority sample technology SMOTE, we can identify and analyze minority samples and synthesize new minority samples to achieve the purpose of balancing the data set and reduce the overfitting problem that may be caused by random oversampling. The algorithm flow is as follows:
[0086] (1) For the minority class sample set P m , from the set P m Randomly select one of the minority class samples x i , x i ∈P m , calculate the Euclidean distance between other minority class samples and x i distance, select distance x i The K most recent minority class samples Get x i K nearest neighbors;
[0087] (2) For each selected K nearest neighbor sample, it is different from the original sample x i New samples are randomly generated on the connection line of , and the formula is:
[0088]
[0089] Among them, x new is a new synthesized sample, ε is a random number between the interval (0,1), and |·| represents the Euclidean distance. Step 6:
[0090] Since the variations of various characteristic gases vary greatly, the data set must be standardized to ensure the stability of the numerical calculation.
[0091] Processing, standardization formula is:
[0092]
[0093] Among them, x * represents the standardized data, x is the original data, and x std are the mean and standard deviation respectively.
[0094] S7: The SMOTE-balanced and standardized dataset is fed into the XGBoost algorithm for training and testing. The XGBoost algorithm is essentially a gradient-based algorithm, which is an improvement on the GBDT algorithm. The biggest difference lies in the different settings of the objective function. XGBoost is an additive model composed of multiple Classification and Regression Trees (CART) models.
[0095]
[0096] in, represents the i-th prediction result, K represents the number of CART trees, f k A model representing a specific CART tree.
[0097] The objective function of XGBoost is defined as:
[0098]
[0099] Where Obj xgb represents the objective function of the XGBoost model, Represents a custom loss function. For example, depending on the task, the mean square error function can be used as the loss function for regression problems, and the logarithmic loss function can be used for classification problems. Ω(f k ) represents the regularization term of the CART tree, and n is the number of samples.
[0100] The training process of the XGBoost model is an iterative process based on Boost gradient boosting, with the goal of minimizing the objective function. After the gradient boosting iterative process of each CART tree and the expansion of the second-order derivative of the Taylor formula, the final loss function expression can be obtained by simplification:
[0101]
[0102] in, They represent the cumulative sum of the first-order partial derivatives and the cumulative sum of the second-order partial derivatives of the samples contained in the leaf node j, and both quantities are constants. j ={i|q(x i )=j} means that all samples x belonging to the jth leaf node i It is classified into a sample set of a leaf node. is the prediction result of the i-th variable at the t-1-th iteration, is the prediction result of the i-th variable after the t-th iteration. T represents the number of leaves, and γ represents the leaf coefficient. A larger γ indicates a greater penalty, and a more desirable tree structure. This parameter can effectively prevent an excessive number of leaves. λ represents the weight vector coefficient of the leaf nodes, which is actually the step size adjustment in the regularization method. A larger λ indicates a more desirable tree structure.
[0103] The above formula gives the final result of the objective function. In the actual solution, λ and γ are custom parameters. The first-order derivative s of each sample at each node should be calculated first. i and the second-order derivative p i , then sum the samples contained in each node to get S j and P j , and finally traverse the nodes of the decision tree to obtain the objective function.
[0104] In the actual training process, due to the uncertainty of the tree structure, XGBoost uses a greedy algorithm to find the optimal split point of the leaf node and calculate the benefit of the split node. For the objective function, the benefit after the split is
[0105]
[0106] Where L represents the left subtree and R represents the right subtree. The larger the gain, the greater the difference between the pre-split and post-split values, which means the more obvious the gradient descent.
[0107] S8: For transformer fault diagnosis, a multi-classification problem based on unbalanced data, accuracy should not be used as an evaluation metric, as this can lead to significant bias in the evaluation results. Because all transformer fault types are equally important, macro-averaged accuracy, recall, and F1 score should be used. These three metrics are essentially calculated based on the confusion matrix, which is defined in Table 1 below:
[0108]
[0109] Table 1
[0110] For a classification problem with L categories, the macro-averaged precision, recall, and F1 score are defined as:
[0111]
[0112]
[0113] The oil-immersed transformer fault data set collected by this invention contains a total of 6910 data, each of which contains five characteristic gases: H2, CH4, C2H6, C2H4, and C2H2, and their corresponding labels. The distribution of various transformer sample data in the data set is as follows: Figure 1 The logarithmic loss curve of the transformer fault diagnosis model based on the SMOTE-XGBoost algorithm in each round of iteration is shown in Figure 2 As shown in the figure, the x-axis represents the number of model iterations, and the y-axis represents the model's logarithmic loss. The dashed line and the dotted line represent the logarithmic loss curves for the training set and test set, respectively. As can be seen from the figure, as the number of iterations increases, the training set has a lower logarithmic loss than the test set, indicating that the model does not suffer from significant overfitting and performs as expected.
[0114] At the same time, the classification error rate curve of the transformer fault diagnosis model based on the SMOTE-XGBoost algorithm in each round of iteration is as follows: Figure 3As shown in the figure, the x-axis represents the number of model iterations, and the y-axis represents the model's classification error rate. The dashed line and the dotted line represent the classification error rate curves for the training set and test set, respectively. The figure shows that as the number of iterations increases, the training set has a lower classification error rate than the test set, which is consistent with the changes in the logarithmic loss curve.
[0115] The final transformer fault diagnosis classification results based on the SMOTE-XGBoost algorithm are as follows: Figure 4 The confusion matrix is shown in Table 2. Compared with other algorithms, the performance comparison results of the transformer fault diagnosis model are shown in Table 2.
[0116]
[0117]
[0118] Table 2
[0119] From the simulation comparison results, it can be seen that compared with GBDT, XGBoost has improved model accuracy and training speed. The use of SMOTE oversampling will greatly increase the number of training samples, so the training time of the SMOTE-XGBoost model is prolonged. However, at the same time, the proposed SMOTE-XGBoost algorithm has the highest macro-average F1 score and can achieve more accurate oil-immersed transformer fault diagnosis and classification results.
[0120] The above describes in detail the oil-immersed transformer fault diagnosis method, system, and readable storage medium provided by the embodiments of the present application. The description of the above embodiments is only intended to help understand the method and core concept of the present application. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present application. In summary, the contents of this specification should not be construed as limiting the present application.
[0121] For example, certain words are used in the specification and claims to refer to specific components. Those skilled in the art should understand that hardware manufacturers may use different nouns to refer to the same component. This specification and claims do not use differences in names as a way to distinguish components, but use differences in the functions of components as the criteria for distinction. For example, "including" and "comprising" mentioned throughout the specification and claims are open-ended terms, so they should be interpreted as "including / including but not limited to". "Approximately" means that within an acceptable error range, those skilled in the art can solve the technical problems within a certain error range and basically achieve the technical effects. The subsequent description in the specification is a preferred embodiment of the present application, but the description is for the purpose of illustrating the general principles of the present application, and is not used to limit the scope of the present application. The scope of protection of the present application shall be as defined in the attached claims.
[0122] It should also be noted that the terms "include," "comprises," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a product or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such product or system. In the absence of further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the product or system comprising the element.
[0123] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.
[0124] The above description shows and describes several preferred embodiments of the present application. However, as previously mentioned, it should be understood that the present application is not limited to the form disclosed herein and should not be construed as excluding other embodiments. Instead, the present application can be used in various other combinations, modifications, and environments and can be modified within the scope of the application concept described herein through the above teachings or technology or knowledge in the relevant field. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present application should be protected by the claims appended hereto.
Claims
1. A method for diagnosing faults of an oil-immersed transformer, characterized in that: The fault diagnosis method comprises: S1: Obtain transformer fault samples and normal samples from the historical database, and obtain the corresponding gas content under different fault conditions and normal conditions from the samples through gas analysis method; S2: Filter the gases in S1, select gas data with concentration and integrity as the characteristic gas input model, and encode the transformer state corresponding to the characteristic gas into integers as the output result of model training; S3: Preprocess the model’s input data and output results, and retain valid data as the first training set; S4: performing stratification, sample synthesis balance and standardization on the first training set to obtain the second training set; S5: Input the second training set into the XGBoost algorithm for training and testing; S6: Transformer fault diagnosis is performed based on the macro-averaged precision, recall, and F1 score of the training and testing results in S5; The sample synthesis balance in S4 is specifically: using the synthetic minority sample method SMOTE to identify and analyze the minority samples and synthesize new minority samples; The standardization process in S4 is specifically: standardizing the data set, and the standardization formula is: Among them, x * represents the standardized data, x is the original data, and x std are the mean and standard deviation, respectively; The SMOTE-balanced and standardized dataset in S5 is fed into the XGBoost algorithm for training and testing as follows: XGBoost is an additive model composed of multiple Classification and Regression Trees (CART) models. in, represents the i-th prediction result, K represents the number of CART trees, f k Represents a specific CART tree model; The objective function of XGBoost is defined as: Where Obj xgb represents the objective function of the XGBoost model, represents a custom loss function. Depending on the task, the mean square error function can be used as the loss function for regression problems, and the logarithmic loss function can be used for classification problems. k ) represents the regularization term of the CART tree, and n is the number of samples; The training process of the XGBoost model is an iterative process based on Boost gradient improvement. The purpose is to minimize the objective function. After the gradient improvement iterative process of each CART tree and the expansion of the second-order derivative of Taylor's formula, the final loss function expression can be obtained by simplification: in, They represent the cumulative sum of the first-order partial derivatives and the cumulative sum of the second-order partial derivatives of the samples contained in the leaf node j, and both quantities are constants; I j ={i|q(x i )=j} means that all samples x belonging to the jth leaf node i Enter into the sample set of a leaf node; is the prediction result of the i-th variable at the t-1-th iteration, is the prediction result of the i-th variable after the t-th iteration; T represents the number of leaves, γ represents the leaf number coefficient, the larger the γ is, the greater the penalty imposed, and the more hope to obtain a tree with a simple structure. This parameter can effectively prevent an excessive number of leaves; λ represents the weight vector coefficient of the leaf node, which is actually the step size adjustment in the regularization method. The larger λ is, the more hope to obtain a tree with a simple structure. The above formula gives the final result of the objective function. In the actual solution, λ and γ are custom parameters. The first-order derivative s of each sample at each node should be calculated first. i and the second-order derivative p i , then sum the samples contained in each node to get S j and P j , and finally traverse the nodes of the decision tree to obtain the objective function; In the actual training process, due to the uncertainty of the tree structure, XGBoost uses a greedy algorithm to find the optimal split point of the leaf node and calculate the benefit of the split node; then for the objective function, the benefit after splitting is Among them, L represents the left subtree and R represents the right subtree; the larger the gain, the greater the difference between before and after the split, which means the more obvious the gradient descent.
2. The fault diagnosis method according to claim 1, characterized in that: The gas types obtained in S1 under different fault states and normal states include but are not limited to H2, CH4, C2H6, C2H4, C2H2, CO and CO2.
3. The fault diagnosis method according to claim 2, characterized in that: The characteristic gases in S2 include H2, CH4, C2H6, C2H4 and C2H2.
4. The fault diagnosis method according to claim 1, characterized in that: The transformer states in S2 include normal, partial discharge, low energy discharge, high energy discharge, medium and low temperature overheating, and high temperature overheating.
5. The fault diagnosis method according to claim 1, characterized in that: The data preprocessing in S3 specifically includes: eliminating the results with missing values, extreme values and / or negative values in the data, and retaining valid data.
6. The fault diagnosis method according to claim 1, characterized in that: The hierarchical division in S4 is specifically as follows: using the ratio of transformer fault samples to normal samples in the historical database as a preset ratio condition, the first training set is hierarchically divided in the same ratio.
7. A fault diagnosis system for an oil-immersed transformer, used to implement any one of the fault diagnosis methods according to claims 1-6, characterized in that: The fault diagnosis system comprises: A data acquisition module is used to obtain the corresponding gas content under different fault conditions and normal conditions; The data screening module is used to select gases with concentration and integrity as characteristic gases to be input into the model, and integer encode the transformer states corresponding to the characteristic gases as the output results of the model training; The data preprocessing module is used to preprocess the input model and output results, and retain valid data as the first training set; A stratified balancing and standardization module is used to perform stratification, sample synthesis balancing and standardization on the first training set to obtain a second training set; The model training and testing module is used to input the second training set into the XGBoost algorithm for training and testing; The fault diagnosis module diagnoses transformer faults based on the results of the model training and testing module through the macro-averaged accuracy, recall rate and F1 score.
8. A readable storage medium comprising: a memory storing a program; A processor, wherein when executing the program, the processor implements the fault diagnosis method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Transformer fault diagnosis method based on OneclassSVM algorithm
CN112183590A