Power Transformer Fault Detection Method for Detecting Dissolved Gases Based on Data Clustering
Through a random forest classifier based on data clustering, the relative concentration and fault point position information of the five gases in the power transformer are solved, and the problem of low fault recognition accuracy in the traditional Duval pentagonal method is achieved, achieving higher fault recognition accuracy.
Patent Information
- Application Number
- CN202310204158.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-03-06
AI Technical Summary
The traditional Duval pentagonal method has problems with data overlap and fault boundary rigidity when identifying power transformer failures, making it difficult to accurately identify the fault type.
A random forest classifier based on data clustering is used to construct a regular pentagonal analysis graph, extract the Euclidean distance between the fault point and the center point of the feature and the relative concentration of the five gases as input features, and a classification model is constructed to identify the fault type of the power transformer.
It significantly improves the accuracy of power transformer fault identification, reduces data overlapping problems, and provides more accurate fault type discrimination.
Smart Images

Figure CN116183880B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power equipment condition monitoring and fault diagnosis, and particularly relates to a method for accurately detecting dissolved gas data by using a random forest classifier based on data clustering Background Art
[0002] Power transformers are one of the most closely monitored units in the entire power system because their failures can bring disasters to the power transmission network. Among various condition monitoring techniques used for continuous health assessment of power transformers, dissolved gas analysis (DGA) is one of the earliest and most widely used techniques. This method detects the evolution of different types of gases inside oil-immersed transformers because oil-immersed transformers may fail due to continuous operation and service aging. One of the main advantages of this method is that only a small amount of oil sample is required, which can be taken from the drain valve of the transformer without disconnecting the transformer from operation. In this way, the opportunity for on-line monitoring is provided. The most important gases generated due to transformer failures are H 2 , CH 4 , C 2 H 6 , C 2 H 4 and C 2 H 2 , and these gases remain dissolved in the oil. In DGA analysis, according to the concentrations of these gases in the oil, it can be known whether the transformer has an early failure. Therefore, correctly detecting faults based on DGA data is a long-term research field, and different DGA interpretation methods have been developed to obtain accurate power transformer faults
[0003] The Duval pentagon proposed by Duval and Lamarre is a method for identifying faults occurring inside power transformers based on the visualization of DGA data. As Figure 1 shown, the fault identification method based on the Duval pentagon divides the pentagon into several regions with fixed fault boundaries. Each region corresponds to a fault type respectively; when the DGA data of the insulating oil of the power transformer falls into a specific region, it is identified as the fault type corresponding to that region. However, due to the existence of fixed boundaries and the misinterpretation of those faults near the boundaries, it is one of the defects of the traditional Duval pentagon. Because the fault distributions themselves are not strictly separated from each other, they overlap in many regions of the pentagon, especially near the band boundaries. Therefore, it becomes difficult to accurately identify the DGA fault points near the boundaries of the two fault bands. Therefore, it is necessary to improve the traditional Duval pentagon to eliminate the rigid nature of the fault band boundaries, so as to accurately detect the early fault types Summary of the Invention
[0004] The object of the present invention is to provide a power transformer fault detection method for detecting dissolved gases based on data clustering.
[0005] The power transformer fault detection method includes the following steps:
[0006] Step 1: Collect the insulating oil in multiple power transformers with known fault types, and detect the relative concentrations of H 2 , CH 4 , C 2 H 6 , C 2 H 4 , and C 2 H 2 in the insulating oil as a training data set.
[0007] Step 2: Construct an analysis diagram of a regular pentagon; use the lines connecting the centroid of the regular pentagon to its five endpoints as five coordinate axes respectively. The five coordinate axes correspond to H 2 , CH 4 , C 2 H 6 , C 2 H 4 , and C 2 H 2 respectively. For each sample data, determine the position of the fault point in the analysis diagram as follows: Mark the corresponding identification points in the regular pentagon according to the relative concentrations of H 2 , CH 4 , C 2 H 6 , C 2 H 4 , and C 2 H 2 in each sample data. Connect the five identification points in sequence to obtain an identification pentagon. Take the centroid of the identification pentagon as the fault point. The fault points corresponding to each fault type form a fault point set respectively, and m fault point sets are obtained; m is the number of fault types.
[0008] Step 3: Each of the m fault point sets forms a characteristic center point P i , i = 1, 2,..., m. The process of forming the characteristic center point P i for the i-th fault type is as follows:
[0009] 3-1. Traverse all the fault points in the fault point set corresponding to the i-th fault type, and screen out the core points. The core points are the fault points where the number of fault points of the same type within the feature neighborhood is greater than or equal to the preset number. All the fault points of the same type as the core point within the feature neighborhood of each core point are called the adjacent points of the core point. Each core point and its corresponding adjacent points form a feature point set; the feature neighborhood of a fault point represents: a circular area centered on the fault point with a preset value ε as the radius.
[0010] 3-2. Merge the feature point sets with a distance less than or equal to the preset value ε until the distance between any two feature point sets is greater than the preset value ε. The distance between two feature point sets represents the distance between the two closest fault points in the two feature point sets.
[0011] 3-3. Generate convex polygons for each merged feature point set respectively; the generated convex polygons use some or all of the fault points in the feature point set as endpoints and cover all the fault points in the feature point set.
[0012] 3-4. Calculate the centroid coordinates of each convex polygon obtained in step 3-3 respectively.
[0013] 3-5. Take the average coordinate value of the centroid coordinates of the convex polygon as the feature center point P of the i-th fault type i .
[0014] Step 4. Extract 11 input features F1 to F11 of the random forest algorithm for each sample data respectively as follows: F1 is the Euclidean distance between the fault point and the feature center point P1; F2 is the Euclidean distance between the fault point and the feature center point P2; F3 is the Euclidean distance between the fault point and the feature center point P3; F4 is the Euclidean distance between the fault point and the feature center point P4; F5 is the Euclidean distance between the fault point and the feature center point P5; F6 is the Euclidean distance between the fault point and the feature center point P6; F7 is the relative concentration of H 2 gas; F8 is the relative concentration of C 2 H 6 gas; F9 is the relative concentration of CH 4 gas; F10 is the relative concentration of C 2 H 4 gas; F11 is the relative concentration of C 2 H 2 gas.
[0015] Step 5. Construct a classification model based on the random forest algorithm, and use the training data set to train the classification model. The classification model takes the 11 input features described in step 4 as input and the fault type as output.
[0016] Step 6: Collect the insulating oil in the power transformer under test and detect the relative concentrations of H 2 , CH 4 , C 2 H 6 , C 2 H 4 and C 2 H 2 in it; according to the relative concentrations of the five gases, determine the 11 input features F1 to F11 corresponding to the power transformer under test; input the obtained input features F1 to F11 into the classification model obtained in Step 5; the classification model outputs the fault type of the power transformer under test.
[0017] Preferably, the detected fault types include the following six categories, namely partial discharge fault, low-energy discharge fault, high-energy discharge fault, low-temperature overheating fault, medium-temperature overheating fault, and high-temperature overheating fault.
[0018] Preferably, the scale range of each axis is from 0% to 100%.
[0019] Preferably, take the geometric center of the analysis graph as the coordinate origin to establish a plane rectangular coordinate system; sort the five gases; the coordinates (x i , y i ) of the identification point corresponding to the i-th gas in the plane rectangular coordinate system are expressed as shown in Equation (1):
[0020]
[0021] where r i is the relative concentration of the i-th gas, and α i is the angle between the x-axis of the plane rectangular coordinate system and the axis corresponding to the i-th gas. i = 1, 2, 3, 4, 5.
[0022] Preferably, the coordinate (C x , C y ) of the centroid position of the polygon is expressed as shown in Equation (2):
[0023]
[0024] where x i , y i are the horizontal and vertical coordinates of the i-th endpoint of the polygon respectively; n is the number of endpoints of the polygon; A is a characteristic parameter, and its expression is as shown in Equation (3):
[0025]
[0026] Preferably, the method for forming the feature region in step 3-3 is as follows: for each set of feature points, a convex polygon is generated; the convex polygon has some or all of the fault points in the set of feature points as endpoints and covers all the fault points in the set of feature points. Then, the endpoints of the convex polygon are smoothed to obtain the feature region.
[0027] Preferably, the process of generating the convex polygon is as follows:
[0028] (1) For a set of feature points with the number of fault points less than or equal to 3, connect the fault point with the smallest abscissa and the fault point with the largest abscissa in the set of feature points as the endpoints of the convex hull to form a side of the convex hull.
[0029] (2) Take the fault point outside the current convex hull and farthest from the current convex hull as the new endpoint of the convex hull.
[0030] (3) Connect all the endpoints of the convex hull in sequence to obtain a new convex hull.
[0031] (4) Repeat steps (2) and (3) until all the sets of feature points are on or inside the edge of the convex hull, and the final convex polygon is obtained.
[0032] Preferably, the smoothing process of the endpoints of the convex polygon is realized by k-order spline interpolation.
[0033] Preferably, the order k = 4.
[0034] Preferably, the classification model described in step five includes m decision trees; the m decision trees correspond to m fault types respectively.
[0035] The beneficial effects of the present invention are as follows:
[0036] 1. The present invention utilizes the information obtained from the unique spatial distribution pattern of the fault layer in the gas concentration graph plane and the numerical pattern of the relative concentration of a single gas into the machine learning algorithm. Compared with the traditional Duval pentagon method, the accuracy of identifying power transformer faults by the present invention is significantly improved.
[0037] 2. The present invention forms different clusters inside the analysis graph of the pentagon by clustering, reduces the data overlap problem of the traditional Duval pentagon recognition method, and improves the accuracy of power transformer fault recognition based on dissolved gas detection.
[0038] 3. The present invention uses a distance parameter in combination with the relative concentration of dissolved gases as input features for fault identification in a decision tree model to improve the accuracy of fault detection. At the same time, the present invention uses Dissolved Gas Analysis (DGA) technology to detect the evolution of different types of gases inside a power transformer, only requiring a small amount of oil sample to be taken from the oil drain valve of the transformer without disconnecting the operation of the transformer, providing an opportunity for on-line monitoring. Description of the Drawings
[0039] Figure 1 is a schematic diagram of the Duval pentagon used in the traditional power transformer identification code.
[0040] Figure 2 is a flow chart of the present invention.
[0041] Figure 3 is an analysis diagram with the fault point location obtained in Step 2 of the present invention.
[0042] Figure 4 is a schematic diagram of forming a convex polygon in Step 3-3 of the present invention. Detailed Embodiments
[0043] The present invention will be further described below with reference to the accompanying drawings.
[0044] As Figure 2 shown, a power transformer fault detection method based on data clustering for detecting dissolved gases constructs a fault case sample library by collecting dissolved gas sample data from various fault cases collected in actual engineering.
[0045] The detected fault types include the following six categories:
[0046] PD: Defined as a partial discharge fault.
[0047] D1: Defined as a low-energy discharge fault.
[0048] D2: Defined as a high-energy discharge fault.
[0049] T1: Defined as a low-temperature overheat fault, T < 300 °C.
[0050] T2: Defined as a medium-temperature overheat fault, 300 °C - 700 °C.
[0051] T3: Defined as a high-temperature overheat fault, T > 700 °C.
[0052] This power transformer fault detection method determines the fault type of the oil-immersed transformer by respectively detecting the relative concentration of the proportions of five dissolved gases in the insulating oil of the oil-immersed transformer. The five dissolved gases are H 2 , CH 4 , C 2H 6 , C 2 H 4 and C 2 H 2 .
[0053] The power transformer fault detection method comprises the following steps:
[0054] Step 1: Collect the insulating oil from multiple power transformers with known fault types and perform DGA (Dissolved Gas Analysis) tests on them to obtain the H 2 , CH 4 , C 2 H 6 , C 2 H 4 and C 2 H 2 The relative concentration of is used as the training data set. The training data set contains samples corresponding to all fault types, and the number of samples for each fault type is greater than or equal to 475.
[0055] Step 2: Figure 3 As shown, construct an analysis diagram of a regular pentagon; the lines connecting the geometric center of the regular pentagon and the five endpoints are used as five axes. The scale range of each axis is 0% to 100%, and the geometric center of the regular pentagon is the 0 scale point. The five axes correspond to H 2 , CH 4 , C 2 H 6 , C 2 H 4 and C 2 H 2 .
[0056] For each sample data, the location of the fault point is determined in the analysis diagram. The specific method is: according to the H in each sample data 2 , CH 4 , C 2 H 6 , C 2 H 4 and C 2 H 2 The relative concentrations of the corresponding identification points are marked in the regular pentagon. The five identification points are connected in sequence to obtain an irregular identification pentagon. The centroid of the identification pentagon is taken as the fault point. The position of the fault point corresponds to a specific fault type. The six fault types form six fault point sets respectively.
[0057] The specific process of extracting the centroid of the identification pentagon as the fault point is as follows:
[0058] 2-1. Take the geometric center of the analysis diagram as the coordinate origin and establish a plane rectangular coordinate system; sort the five soluble gases in the order of H 2 , CH 4 , C 2 H 6 , C 2 H 4 and C 2 H 2 ; the coordinates (x i , y i ) of the identification point corresponding to the i-th gas in the plane rectangular coordinate system are expressed as shown in Equation (1):
[0059]
[0060] where r i is the relative concentration (%) of the i-th gas, and α i is the angle between the x-axis of the plane rectangular coordinate system and the axis corresponding to the i-th gas (measured counterclockwise, and the angle range is 0° to 360°). i = 1, 2, 3, 4, 5.
[0061] 2-2. Connect the five coordinates (x i , y i ) in sequence to obtain an identification pentagon; the fault point is the centroid of the identification pentagon, and its coordinates (C x , C y ) are expressed as shown in Equation (2):
[0062]
[0063] where A is the area of the polygon, and its expression is as shown in Equation (3):
[0064]
[0065] It can be seen from Equation (2) that if the relative concentration of a specific gas is 100% and the relative concentrations of all other gases are 0%, the fault point is located at the point corresponding to the "40%" concentration on the axis corresponding to the gas with a relative concentration of 100%, which means that the fault point is restricted within the regular pentagon formed by the 40% points on the five axes of the analysis diagram.
[0066] Perform DGA tests on multiple insulating oils of power transformers with different fault types according to the above process and extract the fault points. The positions of the obtained fault points in the analysis diagram are as Figure 2 shown; it can be seen from Figure 2 that the positions of many fault points do not fall into Figure 1Rather than falling into the corresponding specific area in the provided Duval pentagon, it falls into the fault areas corresponding to other fault types. Especially at the junction of the two fault areas in the conventional Duval pentagon, there is a phenomenon where multiple fault points of different types are distributed overlappingly, indicating that the current method of directly dividing the Duval pentagon artificially is difficult to accurately judge the faults of power transformers. In addition, the density of fault points within the area covered by a specific fault is not uniform. In some areas, the fault points are densely distributed, while in other areas, the fault points are sparsely distributed. If these areas can be clearly identified and used to design fault detection methods, the efficiency of fault detection can be improved.
[0067] Step 3: Perform density-based clustering (DBSCAN) on each of the six sets of fault points, so that each of the six fault types forms a characteristic center point P i . The process of the i-th fault type forming the characteristic center point P i is as follows:
[0068] 3-1. Traverse all the fault points in the set of fault points corresponding to the i-th fault type, and screen out the core points. The core points are the fault points whose number of fault points of the same type (i.e., belonging to the same set of fault points) within the characteristic neighborhood is greater than or equal to the preset number Minpts. All the fault points of the same type as the core point within the characteristic neighborhood of each core point are called the adjacent points of the core point. Each core point and its corresponding adjacent points form a set of characteristic points; the characteristic neighborhood of a fault point means: a circular area centered on this fault point with a preset value ε as the radius.
[0069] 3-2. Calculate the distance between each two sets of characteristic points respectively, and merge the two sets of characteristic points with a distance less than or equal to the preset value ε into the same set of characteristic points until the distance between any two sets of characteristic points is greater than the preset value ε. The distance between two sets of characteristic points represents the distance between the two closest fault points in these two sets of characteristic points (if the two sets of characteristic points include the same fault point, the distance between the two sets of characteristic points is 0).
[0070] 3-3. As Figure 4As shown in the figure, a convex polygon is generated for each set of feature points after merging; the generated convex polygon has some or all of the fault points in the set of feature points as endpoints and covers all the fault points in the set of feature points (coverage means that all fault points are inside or on the boundary of the convex polygon). In this embodiment, the convex polygon is generated by the convex hull algorithm. The fault points that serve as endpoints on any convex polygon are called boundary points; the fault points outside all convex polygons are called noise points. The basic idea is to construct the convex hull by recursively dividing the point set into several subsets. Specifically, it first finds the smallest convex hull that contains all the points in the point set and divides the point set into two subsets, one inside the convex hull and the other outside the convex hull. Then, for these two subsets, it recursively constructs the convex hulls and merges the obtained convex hulls into the convex hull of the entire point set. The specific steps are as follows:
[0071] (1) Judge the number in the set of feature points; if the number of fault points in the set of feature points is less than 3, then delete the set of feature points; if the number is greater than or equal to 3, then go to step two.
[0072] (2) Connect the points with the smallest and largest abscissas in the set of feature points to form an edge of the convex hull; if there are points with the same abscissa, then take the point with the smallest ordinate among the points with the same abscissa.
[0073] (3) Find the point in the set of feature points that is farthest from the current convex hull projection.
[0074] (4) Sort the points in the current convex hull and the point farthest from the current convex hull according to the polar angle.
[0075] (5) Connect these points in order of the polar angle size to form a new convex hull.
[0076] (6) Divide the set of feature points into two subsets, one inside the convex hull and the other outside the convex hull.
[0077] (7) Re-execute steps three to six for the subset outside the convex hull. Each time, the subsets inside and outside the convex hull will be re-divided until the subset outside the convex hull is an empty set.
[0078] Further, the vertices of the convex polygon are smoothly connected to avoid mutations at the connection points, resulting in non-differentiability. The specific steps are as follows:
[0079] a. Add the vertex with the smallest abscissa in the smallest convex polygon as the starting point of the closed curve; if there are multiple vertices with the smallest abscissa, then take the vertex with the smallest ordinate among them as the starting point of the closed curve.
[0080] b. Add the remaining vertices to the set in clockwise order.
[0081] c. Check whether the last point of the curve is equal to the starting point; if it is not equal to the starting point, add the starting point to the end of the set again.
[0082] d. Perform spline interpolation; use the "Generate Interpolation Spline" function to create a spline interpolation object, and set the order k of the spline interpolation using k = 4. The size of k can be adjusted as needed to control the smoothness of the generated curve.
[0083] e. Use the "Linear Equally Divide" function to generate a floating-point number array t containing 10 times the number of vertices of the convex polygon, and perform spline interpolation on the curve in it; among them, the multiple can be adjusted as needed to control the smoothness of the generated curve.
[0084] f. Store the generated curve points in the floating-point number array t.
[0085] 3-4. Calculate the centroid coordinates of each convex polygon obtained in step 3-3 respectively.
[0086] 3-5. Take the average coordinate value of the centroid coordinates of the convex polygon as the characteristic center point P of the i-th fault type. i 。
[0087] Step 4: Extract 11 input features F1 to F11 of the random forest algorithm for each sample data as follows: F1 is the Euclidean distance between the fault point and the characteristic center point P1; F2 is the Euclidean distance between the fault point and the characteristic center point P2; F3 is the Euclidean distance between the fault point and the characteristic center point P3; F4 is the Euclidean distance between the fault point and the characteristic center point P4; F5 is the Euclidean distance between the fault point and the characteristic center point P5; F6 is the Euclidean distance between the fault point and the characteristic center point P6; F7 is the relative concentration (%) of H 2 gas; F8 is the relative concentration (%) of C 2 H 6 gas; F9 is the relative concentration (%) of CH 4 gas; F10 is the relative concentration (%) of C 2 H 4 gas; F11 is the relative concentration (%) of C 2 H 2 gas.
[0088] Step 5: Construct a classification model based on the random forest algorithm. This classification model takes the 11 input features described in step 4 as input and the fault type as output.
[0089] The classification model based on the random forest algorithm contains decision trees; the number of decision trees is set manually. Each decision tree needs to be trained to act as a base classifier. The training data for each base classifier is prepared by randomly selecting from the entire training dataset with a certain degree of duplication, and the duplicated data is randomly generated from the training dataset, accounting for approximately one-third of the training dataset. This process is called "bagging", and the training dataset for each base classifier is called "in-bag data". This method can ensure that each decision tree is trained on approximately two-thirds of the original data. The remaining data is called "out-of-bag" data, which is used to verify the accuracy of the constructed decision tree training.
[0090] When constructing a decision tree, an attribute (feature) is selected at the root node first. Therefore, a parameter is needed to quantify the proximity or distance of the fault point relative to different convex polygons, and then a distance-based feature parameter is proposed for accurate classification of faults in transformers. In the present invention, the Euclidean distance between the fault point and the feature center point Pi is used as a new parameter, which can be used as a distinguishing feature for different types of faults; at the same time, the relative concentrations of five gases are also used as input features.
[0091] In each decision tree, each root node represents the possible values of the corresponding training data features, that is, each possible value of the training data features will generate a branch. The generation and selection of branches are completed according to the information gain of the possible values of the features (select the attribute with the highest information gain at the root node).
[0092] 5-1. Taking the construction of the PD fault type decision tree as an example, the results can be divided into two types: PD faults and other faults. Starting from the first attribute (feature) F1, the interval between the theoretical maximum value and the theoretical minimum value of this feature in the training data is divided into six equal parts to obtain six data subsets, and each subset contains several data points of PD faults and other faults.
[0093] 5-2. Substituting the data in each data subset into the information entropy calculation formula respectively, the information entropy under different data subsets can be obtained, and its expression is shown in Equation (4).
[0094]
[0095] Among them, E(Y i ) is the information entropy of the data subset Y i divided from the training dataset; N is the number of fault types (i.e., the number of classifications N = 6), p j is the feature ratio in the data subset Y i , and when p j takes the value of 1, the information entropy is 0 (obviously, in this case, the probability 1 represents a certain event without any uncertainty).
[0096] 5-3. Calculate the information gain IG under the input features F1 to F11 using Equation (5), and its expression is as follows
[0097]
[0098] where the operator |·| is the size of the set, that is, the number of elements in the set; Y is the training data; Y i is the data subset split from the training data Y.
[0099] 5-4. Calculate the information gain IG of the remaining features, and select the feature with the largest information gain IG as the first branch node of the decision tree. If there are features with the same information gain IG, any one of them can be selected as the branch node, and the features that have been used as branch nodes will no longer participate in the subsequent process of constructing the decision tree. Therefore, the first branch node derives six branches in total. There is a data subset under each branch, and there are several data points of PD faults and other faults under each data subset.
[0100] 5-5. Starting from the first data subset, use the data in the data subset to calculate the information gain IG of the remaining features, and obtain the feature with the largest information gain IG as the new node under this branch. By analogy, calculate the features with the largest information gain for the remaining 5 data subsets to generate new branch nodes (for the range of information entropy of 0, no subsequent branches will be created).
[0101] 5-6. In the new branch node, divide the interval between the theoretical maximum value and the theoretical minimum value of the data in the data subset corresponding to this feature into six equal parts to obtain six data subsets, and each subset contains several data points of different faults. Repeat the above process to complete the establishment of the PD fault decision tree.
[0102] 5-7. Repeat the above process to complete the establishment of decision trees for the remaining fault types. All decision trees are integrated into the classification model based on the random forest algorithm. When performing fault type discrimination, all decision trees give decisions. The decision with the highest number of votes is the decision of the classification model.
[0103] Step 6: Collect the insulating oil in the measured power transformer and conduct DGA tests on the obtained insulating oil to obtain H 2 、CH 4 、C 2 H 6 、C 2 H 4 and C 2 H 2The relative concentration. Calculate the 11 input features F1 to F11 corresponding to the power transformer under test; input the obtained input features F1 to F11 into the classification model obtained in Step Five; the classification model outputs the probabilities of the power transformer under test belonging to various fault types respectively. Take the fault class corresponding to the maximum probability as the fault type of the power transformer under test.
Claims
1. A power transformer fault detection method for detecting dissolved gases based on data clustering, characterized in that: It includes the following steps: Step 1: Collect insulating oil from multiple power transformers with known fault types and detect the relative concentrations of H 2 , CH 4 , C 2 H 6 , C 2 H 4 and C 2 H 2 in the insulating oil as the training data set; Step 2: Construct an analysis diagram of a regular pentagon; use the lines connecting the centroid of the regular pentagon to its five endpoints as five coordinate axes respectively; the five coordinate axes respectively correspond to H 2 、CH 4 、C 2 H 6 、C 2 H 4 and C 2 H 2 ; For each sample data, determine the position of the fault point in the analysis diagram. The process is as follows: Mark the corresponding identification points in the regular pentagon according to the relative concentrations of H 2 、CH 4 、C 2 H 6 、C 2 H 4 and C 2 H 2 in each sample data respectively; Connect the five identification points in sequence to obtain an identification pentagon; Take the centroid of the identification pentagon as the fault point; The fault points corresponding to each fault type respectively form a fault point set, and m fault point sets are obtained; m is the number of fault types; Step 3: Each of the m sets of fault points forms a characteristic center point P i , where i = 1, 2,..., m; the process of the i-th fault type forming the characteristic center point P i is as follows: 3-1. Traverse all the fault points in the fault point set corresponding to the i-th fault type, and screen out the core points; the core points are the fault points whose number of fault points of the same type within the characteristic neighborhood is greater than or equal to the preset number; all the fault points of the same type as the core point within the characteristic neighborhood of each core point are called the adjacent points of the core point; each core point and its corresponding adjacent points form a characteristic point set; the characteristic neighborhood of a fault point means: a circular area centered on the fault point with a preset value ε as the radius; 3-2. Merge the characteristic point sets with a distance less than or equal to the preset value ε until the distance between any two characteristic point sets is greater than the preset value ε; The distance between two characteristic point sets represents the distance between the two closest fault points in the two characteristic point sets; 3-3. Generate characteristic regions for each of the merged characteristic point sets respectively; The characteristic region covers all the fault points in the characteristic point set; 3-4. Calculate the centroid coordinates of the characteristic regions obtained in step 3-3 respectively; Take the average coordinate value of the centroid coordinates of the characteristic region as the characteristic center point P of the i-th fault type i ; Step 4. Extract the following 11 input features F1 to F11 of the random forest algorithm for each sample data respectively: F1 is the Euclidean distance between the fault point and the feature center point P1; F2 is the Euclidean distance between the fault point and the feature center point P2; F3 is the Euclidean distance between the fault point and the feature center point P3; F4 is the Euclidean distance between the fault point and the feature center point P4; F5 is the Euclidean distance between the fault point and the feature center point P5; F6 is the Euclidean distance between the fault point and the feature center point P6; F7 is the relative concentration of H 2 gas; F8 is the relative concentration of C 2 H 6 gas; F9 is the relative concentration of CH 4 gas; F10 is the relative concentration of C 2 H 4 gas; F11 is the relative concentration of C 2 H 2 gas; Step Five: Construct a classification model based on the random forest algorithm, and use the training data set to train the classification model; the classification model takes the 11 input features described in step four as inputs and the fault type as the output; Step 6: Collect the insulating oil in the power transformer under test and detect the relative concentrations of H 2 , CH 4 , C 2 H 6 , C 2 H 4 and C 2 H 2 ; According to the relative concentrations of the five gases, determine the 11 input features F1 to F11 corresponding to the power transformer under test; Input the obtained input features F1 to F11 into the classification model obtained in Step 5; The classification model outputs the fault type of the power transformer under test.
2. A power transformer fault detection method for detecting dissolved gases based on data clustering according to claim 1, characterized in that: The fault types to be tested and detected include the following six categories, namely partial discharge fault, low-energy discharge fault, high-energy discharge fault, low-temperature overheat fault, medium-temperature overheat fault, and high-temperature overheat fault.
3. A power transformer fault detection method for detecting dissolved gases based on data clustering according to claim 1, characterized in that: The scale range of each axis is from 0% to 100%.
4. A power transformer fault detection method for detecting dissolved gases based on data clustering according to claim 1, characterized in that: Taking the geometric center of the analysis diagram as the coordinate origin, a plane rectangular coordinate system is established; the five gases are sorted; the coordinates (x i , y i ) of the identification point corresponding to the i-th gas are expressed as shown in Equation (1): where r i is the relative concentration of the i-th gas, and α i is the angle between the x-axis of the rectangular coordinate system and the axis corresponding to the i-th gas; i = 1, 2, 3, 4, 5.
5. A power transformer fault detection method for detecting dissolved gases based on data clustering according to claim 1, characterized in that: The centroid position coordinates (C x , C y ) of the polygon are expressed as shown in Equation (2): where x i , y i are the abscissa and ordinate of the i-th endpoint of the polygon respectively; n is the number of endpoints of the polygon; A is a characteristic parameter, and its expression is as shown in Equation (3):
6. A power transformer fault detection method for detecting dissolved gases based on data clustering according to claim 1, characterized in that: The method for forming the characteristic region in step 3-3 is: generate a convex polygon for each characteristic point set respectively; the convex polygon takes some or all of the fault points in the characteristic point set as endpoints and covers all the fault points in the characteristic point set; then smooth the endpoints of the convex polygon to obtain the characteristic region.
7. A power transformer fault detection method for detecting dissolved gases based on data clustering according to claim 6, characterized in that: The process of generating the convex polygon is as follows: (1) For a characteristic point set with the number of fault points less than or equal to 3, connect the fault point with the smallest abscissa and the fault point with the largest abscissa in the characteristic point set as the endpoints of the convex hull to form a side of the convex hull; (2) Select the fault point that is farthest from the current convex hull and outside the current convex hull as the new endpoint of the convex hull; (3) Connect all the endpoints of the convex hull in sequence to obtain a new convex hull; (4) Repeat steps (2) and (3) until all the feature point sets are on the edge or inside the convex hull, and the final convex polygon is obtained.
8. A power transformer fault detection method for detecting dissolved gases based on data clustering according to claim 6, characterized in that: The smoothing process of the endpoints of the convex polygon is realized by k-order spline interpolation.
9. A power transformer fault detection method for detecting dissolved gases based on data clustering according to claim 8, characterized in that: The order k = 4.
10. A power transformer fault detection method for detecting dissolved gases based on data clustering according to claim 1, characterized in that: The classification model described in step five includes m decision trees; the m decision trees correspond to m fault types respectively.
Citation Information
Patent Citations
Random-forest-model-based power transformer fault diagnosis method
CN102221655A
Marine diesel engine fault location method based on union belief rule base and ant colony algorithm
CN110132603A