Distribution transformer equipment state evaluation method and system based on machine learning

Through machine learning-based methods, the improved DBSCAN and CatBoost algorithms are used to achieve accurate evaluation of equipment status of distribution transformers, solving the problem of inaccurate equipment status evaluation in the prior art, and improving the reliability and maintenance efficiency of equipment operation.

CN119989018APending Publication Date: 2025-05-13CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311499162.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to achieve accurate assessment of the equipment status of distribution transformers, resulting in losses and safety hazards caused by equipment failure and shutdown.

Method used

Using a machine learning-based method, the operating data and environmental data of the distribution transformer are obtained, and the cluster analysis and interval value method are used to normalize the distribution transformer by using the improved DBSCAN algorithm. Combined with the pre-constructed state evaluation model, the CatBoost algorithm is used to evaluate the operating status of the distribution transformer.

Benefits of technology

It effectively improves the intelligence level of equipment status evaluation, reduces errors and uncertainties caused by human subjective judgments, improves the reliability and maintenance efficiency of equipment operation, and reduces operation and maintenance costs and failure risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989018A_ABST
    Figure CN119989018A_ABST
Patent Text Reader

Abstract

The invention discloses a distribution transformer equipment state evaluation method and system based on machine learning. The method comprises the following steps: acquiring operation data and environment data of a distribution transformer; based on the operation data and the environment data, utilizing an improved DBSCAN algorithm to perform clustering analysis, and adopting an interval value method to perform normalization processing to obtain a data set; calculating based on the data set in combination with a pre-constructed state evaluation model to obtain an operation state of the distribution transformer; wherein the state evaluation model is obtained by taking the operation state of the distribution transformer as output, and training by utilizing the input data and combining with a CatBoost algorithm; according to the invention, the improved DBSCAN algorithm is utilized to optimize the data set, so that the realization effect of the state evaluation model can be effectively improved; according to the method, the state evaluation model is established by using the machine learning algorithm, so that errors and uncertainty caused by human subjective judgment are reduced, the intelligent level of equipment state evaluation is improved, loss caused by equipment failure and shutdown is effectively avoided, and accurate evaluation of the transformer equipment state is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power equipment status assessment, and in particular to a distribution transformer equipment status assessment method and system based on machine learning. Background Art

[0002] At present, the technology of smart substation is developing rapidly, and the stability, safety and other performance of distribution transformers have been significantly improved. However, due to the numerous and unstable factors reflecting the operating status information of distribution transformers, the evaluation of transformer equipment status is complicated. In order to better put transformers into engineering use, improve the reliability of power supply of distribution networks, and strengthen the management of the entire life cycle of assets, it is necessary to conduct a good evaluation of their operating reliability.

[0003] For a long time, the judgment of the health level and operating status of distribution transformers has been mainly based on regular maintenance, which often leads to "over-maintenance" and "under-maintenance", resulting in huge waste of manpower and material resources, and also reduces the reliability of power supply. In previous studies, some literatures have extracted the features of transformer operating status data, constructed the health index of equipment operating status, and evaluated the transformer status, but they are limited to using partial information for judgment, which is relatively one-sided. There are also literatures that consider the complex relationship between transformer evaluation indicators, propose a method for determining the importance of indicators that combines hierarchical analysis method and entropy weight method, and use the minimum variance theory to achieve the selection of optimal weights. The optimal weight method proposed in this literature combines the subjectivity of the hierarchical analysis method with the objectivity of the entropy weight method, overcoming the adverse effects of the excessive subjective factors in status evaluation for a long time, but failing to distinguish the influence of the weight of objective factors. Fuzzy diagnosis and back propagation neural network were subsequently used for transformer state assessment and decision making. Some researchers used fuzzy learning vector quantization network for transformer state assessment. By using fuzzy classifiers to divide dissolved gas analysis (DGA) data into different subclasses, and training fuzzy learning vector quantization networks for each class, the correctness of the assessment was improved. This is better than the previous fuzzy diagnosis and back propagation neural network methods, but relying solely on DGA data reduces the effectiveness of equipment state assessment and diagnosis. Some literature proposes that distribution transformer state assessment is a multi-attribute decision problem, and applies evidence theory to fuse the information of transformer diagnostic data to evaluate the transformer state category, but does not select a reasonable state assessment indicator. In addition, object meta-theory, fuzzy comprehensive evaluation, Bayesian network and grey target theory are also used to evaluate transformer state. It can be seen that the current distribution transformer operation state evaluation still faces a series of problems. How to achieve accurate evaluation of the distribution transformer operation state to support the safe and stable operation of the distribution network is of great significance. Summary of the invention

[0004] In order to solve the problem of how to accurately evaluate the status of transformer equipment in the prior art, the present invention proposes a distribution transformer equipment status evaluation method based on machine learning, comprising:

[0005] Obtain operating and environmental data of distribution transformers;

[0006] Based on the operation data and the environmental data, cluster analysis is performed using an improved DBSCAN algorithm and normalization is performed using an interval value method to obtain a data set;

[0007] Calculating based on the data set combined with a pre-built state evaluation model to obtain the operating state of the distribution transformer;

[0008] The state evaluation model is constructed by using the data set as input and utilizing the improved CatBoost algorithm to evaluate the operating state of the distribution transformer.

[0009] Preferably, the method further includes a process of constructing a state evaluation model, wherein the process of constructing the state evaluation model includes:

[0010] Taking the data set as input and the operating status of the distribution transformer as output;

[0011] Generate multiple feature rankings based on the data set for learning, and calculate the conversion value of the classification feature;

[0012] Based on the conversion values ​​of the classification features, a combination of classification features is established using a greedy strategy and a tree structure is selected as a weak learning classification tree;

[0013] An ordered Boosting algorithm is used to calculate the gradient, and the gradient is used to train the weak learning classification tree;

[0014] Among them, the operating status of the distribution transformer includes: healthy operation, low temperature overheating, medium temperature overheating, high temperature overheating, partial discharge, low energy discharge and high energy discharge.

[0015] Preferably, the cluster analysis based on the operation data and the environmental data is performed using an improved DBSCAN algorithm and normalized using an interval value method to obtain a data set, including:

[0016] Based on the operation data and the environmental data, cluster analysis is performed using an improved DBSCAN algorithm and noise data is eliminated to obtain a clustering result;

[0017] Based on the clustering results, normalization processing is performed using an interval value method to obtain normalized data to construct a data set.

[0018] Preferably, the clustering analysis is performed based on the operation data and the environmental data using an improved DBSCAN algorithm and noise data is eliminated to obtain a clustering result, including:

[0019] Based on the operation data and the environment data, a neighborhood radius and a minimum number of points are calculated using a genetic algorithm;

[0020] Based on the operation data and the environment data, any data object is selected to determine whether it has been classified into a certain cluster or is determined to be noise data. If so, the data object is reselected; otherwise, the number of data points within the neighborhood radius of the data object as the center point is counted;

[0021] A determination is made based on whether the number of data points within the neighborhood radius of the data object as the center point is less than the minimum number of points. If so, the data object is marked as a boundary point or a center point; otherwise, a new cluster is established with the data object as the core point;

[0022] Based on the unlabeled data objects within the neighborhood radius of the data object as the center point, the number of data points within the neighborhood radius is selected to determine whether it is less than the minimum number of points. If so, the unlabeled data objects are reselected; otherwise, the data objects within the neighborhood radius that are not included in other clusters are added to the new cluster to obtain the clustering result.

[0023] Preferably, the operating data and environmental data of the distribution transformer include: key parameter data and insulation state data in the operating conditions of the dry-type distribution transformer; load information of the oil-immersed distribution transformer, gas dissolved concentration data in the transformer oil, ambient temperature, humidity monitoring quantities and corresponding operating status of the equipment.

[0024] Based on the same inventive concept, the present invention also proposes a distribution transformer equipment status evaluation system based on machine learning, comprising:

[0025] A data acquisition module, used to acquire operating data and environmental data of the distribution transformer;

[0026] A data pre-training module, used to perform cluster analysis based on the operation data and the environmental data using an improved DBSCAN algorithm and perform normalization processing using an interval value method to obtain a data set;

[0027] A model solving module, used to calculate based on the data set in combination with a pre-built state evaluation model to obtain the operating state of the distribution transformer;

[0028] The state evaluation model is constructed by using the data set as input and utilizing the improved CatBoost algorithm to evaluate the operating state of the distribution transformer.

[0029] Preferably, it also includes a model building module, and the model building module is specifically used to:

[0030] Taking the data set as input and the operating status of the distribution transformer as output;

[0031] Generate multiple feature rankings based on the data set for learning, and calculate the conversion value of the classification feature;

[0032] Based on the conversion values ​​of the classification features, a combination of classification features is established using a greedy strategy and a tree structure is selected as a weak learning classification tree;

[0033] An ordered Boosting algorithm is used to calculate the gradient, and the gradient is used to train the weak learning classification tree;

[0034] Among them, the operating status of the distribution transformer includes: healthy operation, low temperature overheating, medium temperature overheating, high temperature overheating, partial discharge, low energy discharge and high energy discharge.

[0035] Preferably, the data pre-training module includes:

[0036] A cluster analysis submodule, used to perform cluster analysis based on the operation data and the environmental data using an improved DBSCAN algorithm and remove noise data to obtain a clustering result;

[0037] The normalization processing submodule is used to perform normalization processing based on the clustering results using the interval value method to obtain normalized data to construct a data set.

[0038] Preferably, the cluster analysis submodule is specifically used for:

[0039] Based on the operation data and the environment data, a neighborhood radius and a minimum number of points are calculated using a genetic algorithm;

[0040] Based on the operation data and the environment data, any data object is selected to determine whether it has been classified into a certain cluster or is determined to be noise data. If so, the data object is reselected; otherwise, the number of data points within the neighborhood radius of the data object as the center point is counted;

[0041] A determination is made based on whether the number of data points within the neighborhood radius of the data object as the center point is less than the minimum number of points. If so, the data object is marked as a boundary point or a center point; otherwise, a new cluster is established with the data object as the core point;

[0042] Based on the unlabeled data objects within the neighborhood radius of the data object as the center point, the number of data points within the neighborhood radius is selected to determine whether it is less than the minimum number of points. If so, the unlabeled data objects are reselected; otherwise, the data objects within the neighborhood radius that are not included in other clusters are added to the new cluster to obtain the clustering result.

[0043] Preferably, the operating data and environmental data of the distribution transformer in the data acquisition module include: key parameter data and insulation state quantity data in the operating conditions of the dry-type distribution transformer; load information of the oil-immersed distribution transformer, gas dissolved concentration data in the transformer oil, ambient temperature, humidity monitoring quantities and corresponding operating status of the equipment.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] A distribution transformer equipment state evaluation method and system based on machine learning, comprising: obtaining operation data and environmental data of the distribution transformer; performing cluster analysis based on the operation data and environmental data using an improved DBSCAN algorithm and performing normalization processing using an interval value method to obtain a data set; performing calculation based on the data set combined with a pre-built state evaluation model to obtain the operation state of the distribution transformer; wherein the state evaluation model is obtained by training the input data combined with a CatBoost algorithm with the operation state of the distribution transformer as output; the present invention optimizes the data set using an improved DBSCAN algorithm, which can effectively improve the implementation effect of the state evaluation model; the present invention uses a machine learning algorithm to build a state evaluation model to reduce errors and uncertainties caused by human subjective judgment, improves the intelligent level of equipment state evaluation, effectively avoids losses caused by equipment failures and downtime, and realizes accurate evaluation of the transformer equipment state; and uses an improved CatBoost algorithm to evaluate the transformer equipment state, solves the gradient offset problem in the model training process, improves the generalization ability, reduces the possibility of overfitting, and enhances the robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 A flow chart of a method for evaluating the state of a distribution transformer based on machine learning according to the present invention;

[0047] Figure 2 It is the overall flow chart of the distribution transformer status evaluation of the present invention;

[0048] Figure 3 The data sample clustering analysis flow chart of the present invention. DETAILED DESCRIPTION

[0049] The present invention proposes a distribution transformer equipment status evaluation method based on machine learning to improve the diagnostic capability and intelligence level of distribution transformer equipment status evaluation and realize accurate evaluation of transformer equipment status. The method can effectively evaluate the overall health status of the transformer, help operation and maintenance personnel to timely discover equipment failures and hidden dangers, and issue alarms for abnormal equipment status. It has high feasibility and accuracy, and can further improve the operation reliability and maintenance efficiency of the equipment, reduce operation and maintenance costs and failure risks, and improve the safety and stability of network operation. In order to better understand the present invention, the content of the present invention is further described below in conjunction with the accompanying drawings and embodiments of the specification.

[0050] Embodiment 1:

[0051] A distribution transformer equipment status evaluation method based on machine learning, the specific process is as follows Figure 1 As shown, including:

[0052] Step 1, obtaining operation data and environmental data of the distribution transformer;

[0053] Step 2, performing cluster analysis based on the operation data and the environmental data using an improved DBSCAN algorithm and performing normalization processing using an interval value method to obtain a data set;

[0054] Step 3, calculating based on the data set and a pre-built state evaluation model to obtain the operating state of the distribution transformer;

[0055] The state evaluation model is constructed by using the data set as input and utilizing the improved CatBoost algorithm to evaluate the operating state of the distribution transformer.

[0056] This embodiment provides a distribution transformer equipment status evaluation method based on machine learning. First, the data used for transformer status evaluation is obtained; second, the collected data is clustered and analyzed, and different status categories are divided according to data characteristics, and noise data is removed to obtain the data set required for model training; finally, a classification model is constructed based on the previously calculated clustering results and known status labels to achieve accurate evaluation of the transformer status, such as Figure 2 shown.

[0057] Before step 1, a state evaluation model construction process is also included, and the construction process specifically includes:

[0058] Based on the calculated clustering results and known status labels, a classification model is constructed to achieve accurate evaluation of the transformer status.

[0059] With the help of CatBoost algorithm, the transformer condition evaluation model is built and trained:

[0060] Considering the problem of sample imbalance, a CatBoost model suitable for multi-feature problems is established. The model input features are transformer load information, as well as the collected data on gas dissolved concentration in transformer oil and operating environment data, and the output label is the equipment status P.

[0061] Compared with other GBDT algorithms, Catboost is optimized in many aspects. First, CatBoost adopts the ordering principle to avoid the conditional displacement problem inherent in the iteration of the GBDT algorithm, while making it possible to use the entire data set for training and learning. Secondly, CatBoost converts the traditional gradient boosting algorithm into an ordered Boosting algorithm, thereby solving the inevitable gradient shift problem in iteration, improving the generalization ability, reducing the possibility of overfitting, and enhancing the robustness of the model. Finally, CatBoost constructs a combination of classification features through a greedy strategy and uses the above combination as an additional feature, which makes it easier for the model to capture high-order dependencies and more significantly improves the prediction accuracy. In addition, CatBoost chooses the forgetting decision tree as the basic prediction cycle, which reduces the possibility of overfitting and increases the execution speed of the model.

[0062] In the present invention, the standardized data set is first set as:

[0063]

[0064] In the formula, D is the standardized data set; X i For the sample group, is the m-th eigenvector of the i-th group of samples; i is the order of the sample group; Y i is the tag value.

[0065] The steps of CatBoost algorithm include:

[0066] First, randomly generate multiple feature rankings for learning, find samples of the same class under each feature, and calculate the classification feature conversion value:

[0067]

[0068] In the formula, is the conversion value of the classification feature; j is the sample group ranking; i is the sample group ranking; n is the total number of sample groups; is the indicator function, when When , the indicator function is 1, otherwise it is 0; is the k-th eigenvector of the j-th group sample; is the k-th eigenvector of the i-th group sample; Y j is the label value; α is the prior weight; P is a prior value.

[0069] Secondly, a combination of classification features is established according to the greedy strategy, and a tree structure is selected.

[0070] Use the ordered Boosting algorithm to calculate the sample group X i The gradient of the weak learning classification tree is used to train the weak learning classification tree. In addition, the final model is obtained by weighting. For model training, a cluster in the data set uses the same label. In the present invention, the ratio of the training set, the validation set and the test set is 3:1:1.

[0071] Two groups of models are built, and the information of the two types of transformers to be evaluated are respectively input into the training model and the parameters are adjusted. After the training is completed, the model can directly obtain the corresponding transformer status evaluation results.

[0072] In step 1, the operating data and environmental data of the distribution transformer are obtained, including:

[0073] Obtain data used for transformer condition evaluation;

[0074] In terms of model input, for dry-type distribution transformers, the required data include key parameter data and insulation state data in the operating conditions of dry-type transformers. The key parameters include factory design life, service life and operating load level. The insulation state includes hot spot temperature and electrical indicators of the winding. The electrical indicators include absorption ratio, core grounding current and direct resistance unbalance coefficient. For oil-immersed distribution transformers, according to GB-T 7252-2016 "Guidelines for Analysis and Judgment of Transformer Oil-soluble Gases", the required data include transformer load information, gas dissolved concentration data in transformer oil, ambient temperature, humidity monitoring, and corresponding operating status of the equipment. Obtain at least 300 sets of operating data of the two types of transformers under various working conditions, and then train the corresponding models of the two types of equipment separately.

[0075] In step 2, cluster analysis is performed based on the operation data and the environmental data using an improved DBSCAN algorithm and normalization is performed using an interval value method to obtain a data set, which specifically includes:

[0076] Perform cluster analysis on the collected data, divide it into different status categories according to the data characteristics, remove noise data, and organize the data set required for model training;

[0077] Since any single feature in the data cannot accurately determine the state type of the transformer, and there are some coupling relationships between the characteristic attributes of the data, it is necessary to analyze the data features and organize them to obtain the final data set required for model training. The implementation process is as follows Figure 3 shown.

[0078] Cluster the existing data samples, classify them according to the state type, and remove the noise data.

[0079] Based on multi-dimensional features, the improved DBSCAN algorithm is used to cluster data samples. The advantage of using this algorithm is that there is no need to pre-set the number of cluster centers, but it is preliminarily assumed that the features of each cluster are similar, and different clustering results are obtained after feature clustering. DBSCAN is an information clustering method based on data density, and its brief process is shown in the following table:

[0080]

[0081] Table 1

[0082] To evaluate the distribution density of data points, two parameters, neighborhood radius ε and minimum points MinPts(μ), are used to form clusters. The algorithm starts with an optional point and calculates the points with a radius less than ε near the point. If the number of points is greater than the parameter μ, the points form a cluster; otherwise, the point is considered an outlier. In the next step, the point can be identified as part of a cluster. The advantage of this method is that it can distinguish outliers from other data.

[0083] Introduce the DB index Davies Bouldin Index, DBI; evaluate the DBSCAN clustering effect.

[0084] The present invention introduces DBI to evaluate the clustering results calculated by DBSCAN. It calculates the distance within the cluster and the distance between clusters. The best selection of clusters will be made because DBI is minimized and the index is calculated as follows:

[0085]

[0086] In the formula, DBI is the clustering effect index; N is the number of clusters; a is the rank of the clusters; b is the rank of the clusters; S a Cluster C a The average distance within S b Cluster C b The average distance of ab Cluster C a and cluster C b Average linkage between distances.

[0087]

[0088] Where, d ab Cluster C a and cluster C b Average linkage between distances; P r is point r; C a is the ath cluster; Cb is the bth cluster; P s is point s; ‖C a ‖ is the Euclidean norm of cluster a; ‖C b ‖ is the Euclidean norm of cluster b.

[0089]

[0090] In the formula, S a Cluster C a The average distance within P r is point r; C a is the ath cluster; C b is the bth cluster; P s is point s; ‖C a ‖ is the Euclidean norm of cluster a. Where ‖·‖ is the Euclidean norm, and P r ∈C a This means that point r belongs to cluster a.

[0091] By optimizing the parameter selection method, the computational performance of the original algorithm is improved.

[0092] In this algorithm, the most important role is to find the appropriate values ​​of the neighborhood radius ε and the minimum point μ. The values ​​can be calculated by general statistics and classical algorithms, but in most cases the calculated results allow the algorithm to run with high precision. Therefore, in the present invention, the genetic algorithm Genetic Algorithm, GA, is used as a heuristic algorithm to estimate the exact values ​​of these parameters, and the verification results show that the proposed method has achieved significant improvements over the original algorithm.

[0093] In order to design an improved DBSCAN, a GA algorithm is used to find the optimal values ​​of data points P, neighborhood radius ε, and minimum point μ. In the adaptive GA algorithm for improving DBCSAN for a dataset with U objects and M attributes, each chromosome is an M+2 dimensional array, as shown in the following formula:

[0094]

[0095] min

[0096] dis min =1≤u,r≤U‖P u ,P r ‖

[0097] u≠r

[0098] max

[0099] dis max =1≤u,r≤U‖P u ,P r ‖

[0100] u≠r

[0101] The first M elements represent the initial point P for the DBSCAN algorithm. The M+1th element represents the neighborhood radius ε, and the last element represents the value μ of the minimum point.

[0102] In the formula, chro u is the uth basic operation unit; chro uv is the gene combination of u and v; min is the minimum distance between data points; max is the maximum distance between data points; u represents the object; v is the attribute; U is the number of objects; M is the number of attributes; P u is point u.

[0103] The input of the genetic algorithm is the population size pop_size, the crossover rate p c , mutation rate p m , the maximum number of iterations MaxItr and / or other termination criteria. In Algorithm 2, a brief process of the genetic algorithm based on initialization, crossover, mutation and selection is given. Regarding the chromosome structure, pop_size chromosomes are randomly generated as the initial population in the zero generation P(0). To evaluate each chromosome, Algorithm 1 is executed with DBI as its fitness function.

[0104] The improved DBSCAN algorithm is shown in Table 2. c The roulette cycle is considered to select a new population P(t+1) from P(t)∪C(t) as a selection rule. The optimal determination of mutation and crossover rates is a challenge for genetic algorithms. These parameters are determined empirically and have a significant impact on the efficiency, accuracy, and speed of the algorithm.

[0105]

[0106]

[0107] Table 2

[0108] Dataset standardization:

[0109] Considering that the input data features vary greatly, which affects the processing speed of the model, it is necessary to normalize the data. The present invention uses the interval value method to normalize the data, scaling the data to a specific interval in proportion to avoid interactions between values. Here, the extreme value method is selected for linear function transformation, and the formula is as follows:

[0110]

[0111] In the formula, X(d) is the normalized data with a mapping interval of [-1, 1]; X is the original data; maxX is the maximum value in the data sample; minX is the minimum value in the data sample. The normalized data after dimensionality reduction can be used as input to train and test the model.

[0112] In step 3, the operation status of the distribution transformer is obtained by performing calculation based on the data set combined with a pre-built status evaluation model, which specifically includes:

[0113] Inputting the data set into the state evaluation model and outputting the operating state of the distribution transformer;

[0114] According to the common fault types of transformers, healthy operation, low-temperature overheating, medium-temperature overheating, high-temperature overheating, partial discharge, low-energy discharge and high-energy discharge are used as the output features of the transformer status evaluation model.

[0115] This embodiment uses machine learning technology to build a machine learning model that accurately diagnoses the working status of the equipment by analyzing the operating data and environmental data of the distribution transformer equipment, and can identify the equipment operating status and potential faults;

[0116] The machine learning algorithm used in this embodiment can automatically learn and optimize the diagnostic model based on historical data and equipment characteristics, making it more accurate and intelligent, helping operation and maintenance personnel to formulate more scientific and reasonable maintenance and repair plans, and reducing the errors and uncertainties caused by human subjective judgment compared with traditional manual judgment;

[0117] By analyzing abnormal patterns and regularities in the data, we can determine the operating status of the equipment, issue early warning information, and take corresponding repair and maintenance measures in a timely manner, effectively avoiding losses caused by equipment failure and downtime.

[0118] Embodiment 2:

[0119] A distribution transformer equipment status evaluation system based on machine learning, comprising:

[0120] A data acquisition module, used to acquire operating data and environmental data of the distribution transformer;

[0121] A data pre-training module, used to perform cluster analysis based on the operation data and the environmental data using an improved DBSCAN algorithm and perform normalization processing using an interval value method to obtain a data set;

[0122] A model solving module, used to calculate based on the data set in combination with a pre-built state evaluation model to obtain the operating state of the distribution transformer;

[0123] The state evaluation model is constructed by using the data set as input and utilizing the improved CatBoost algorithm to evaluate the operating state of the distribution transformer.

[0124] It also includes a model building module, which is specifically used to:

[0125] Taking the data set as input and the operating status of the distribution transformer as output;

[0126] Generate multiple feature rankings based on the data set for learning, and calculate the conversion value of the classification feature;

[0127] Based on the conversion values ​​of the classification features, a combination of classification features is established using a greedy strategy and a tree structure is selected as a weak learning classification tree;

[0128] An ordered Boosting algorithm is used to calculate the gradient, and the gradient is used to train the weak learning classification tree;

[0129] Among them, the operating status of the distribution transformer includes: healthy operation, low temperature overheating, medium temperature overheating, high temperature overheating, partial discharge, low energy discharge and high energy discharge.

[0130] The data pre-training module comprises:

[0131] A cluster analysis submodule, used to perform cluster analysis based on the operation data and the environmental data using an improved DBSCAN algorithm and remove noise data to obtain a clustering result;

[0132] The normalization processing submodule is used to perform normalization processing based on the clustering results using the interval value method to obtain normalized data to construct a data set.

[0133] The cluster analysis submodule is specifically used for:

[0134] Based on the operation data and the environment data, a neighborhood radius and a minimum number of points are calculated using a genetic algorithm;

[0135] Based on the operation data and the environment data, any data object is selected to determine whether it has been classified into a certain cluster or is determined to be noise data. If so, the data object is reselected; otherwise, the number of data points within the neighborhood radius of the data object as the center point is counted;

[0136] A determination is made based on whether the number of data points within the neighborhood radius of the data object as the center point is less than the minimum number of points. If so, the data object is marked as a boundary point or a center point; otherwise, a new cluster is established with the data object as the core point;

[0137] Based on the unlabeled data objects within the neighborhood radius of the data object as the center point, the number of data points within the neighborhood radius is selected to determine whether it is less than the minimum number of points. If so, the unlabeled data objects are reselected; otherwise, the data objects within the neighborhood radius that are not included in other clusters are added to the new cluster to obtain the clustering result.

[0138] The operating data and environmental data of the distribution transformer in the data acquisition module include: key parameter data and insulation state data in the operating conditions of the dry-type distribution transformer; load information of the oil-immersed distribution transformer, gas dissolved concentration data in the transformer oil, ambient temperature, humidity monitoring quantity and corresponding operating status of the equipment.

[0139] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0140] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0141] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0142] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0143] The above are merely embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are included in the scope of the claims of the present invention to be approved.

Claims

1. A distribution transformer equipment status evaluation method based on machine learning, characterized in that: include: Obtain operating and environmental data of distribution transformers; Based on the operation data and the environmental data, cluster analysis is performed using an improved DBSCAN algorithm and normalization is performed using an interval value method to obtain a data set; Calculating based on the data set combined with a pre-built state evaluation model to obtain the operating state of the distribution transformer; The state evaluation model is constructed by using the data set as input and utilizing the improved CatBoost algorithm to evaluate the operating state of the distribution transformer.

2. The method according to claim 1, characterized in that: The method also includes a construction process of a state evaluation model, wherein the construction process includes: Taking the data set as input and the operating status of the distribution transformer as output; Generate multiple feature rankings based on the data set for learning, and calculate the conversion value of the classification feature; Based on the conversion values ​​of the classification features, a combination of classification features is established using a greedy strategy and a tree structure is selected as a weak learning classification tree; An ordered Boosting algorithm is used to calculate the gradient, and the gradient is used to train the weak learning classification tree; Among them, the operating status of the distribution transformer includes: healthy operation, low temperature overheating, medium temperature overheating, high temperature overheating, partial discharge, low energy discharge and high energy discharge.

3. The method according to claim 1, characterized in that: The improved DBSCAN algorithm is used to perform cluster analysis based on the operation data and the environmental data and the interval value method is used to perform normalization processing to obtain a data set, including: Based on the operation data and the environmental data, cluster analysis is performed using an improved DBSCAN algorithm and noise data is eliminated to obtain a clustering result; Based on the clustering results, normalization processing is performed using an interval value method to obtain normalized data to construct a data set.

4. The method according to claim 3, characterized in that: The clustering analysis is performed based on the operation data and the environmental data using the improved DBSCAN algorithm and noise data is eliminated to obtain a clustering result, including: Based on the operation data and the environment data, a neighborhood radius and a minimum number of points are calculated using a genetic algorithm; Based on the operation data and the environment data, any data object is selected to determine whether it has been classified into a certain cluster or is determined to be noise data. If so, the data object is reselected; otherwise, the number of data points within the neighborhood radius of the data object as the center point is counted; A determination is made based on whether the number of data points within the neighborhood radius of the data object as the center point is less than the minimum number of points. If so, the data object is marked as a boundary point or a center point; otherwise, a new cluster is established with the data object as the core point; Based on the unlabeled data objects within the neighborhood radius of the data object as the center point, the number of data points within the neighborhood radius is selected to determine whether it is less than the minimum number of points. If so, the unlabeled data objects are reselected; otherwise, the data objects within the neighborhood radius that are not included in other clusters are added to the new cluster to obtain the clustering result.

5. The method according to claim 1, characterized in that: The operating data and environmental data of the distribution transformer include: key parameter data and insulation state data in the operating conditions of the dry-type distribution transformer; load information of the oil-immersed distribution transformer, gas dissolved concentration data in the transformer oil, ambient temperature, humidity monitoring data and corresponding operating status of the equipment.

6. A distribution transformer equipment status evaluation system based on machine learning, characterized in that: include: A data acquisition module, used to acquire operating data and environmental data of the distribution transformer; A data pre-training module, used to perform cluster analysis based on the operation data and the environmental data using an improved DBSCAN algorithm and perform normalization processing using an interval value method to obtain a data set; A model solving module, used to calculate based on the data set in combination with a pre-built state evaluation model to obtain the operating state of the distribution transformer; The state evaluation model is constructed by using the data set as input and utilizing the improved CatBoost algorithm to evaluate the operating state of the distribution transformer.

7. The system according to claim 6, characterized in that: It also includes a model building module, which is specifically used to: Taking the data set as input and the operating status of the distribution transformer as output; Generate multiple feature rankings based on the data set for learning, and calculate the conversion value of the classification feature; Based on the conversion values ​​of the classification features, a combination of classification features is established using a greedy strategy and a tree structure is selected as a weak learning classification tree; An ordered Boosting algorithm is used to calculate the gradient, and the gradient is used to train the weak learning classification tree; Among them, the operating status of the distribution transformer includes: healthy operation, low temperature overheating, medium temperature overheating, high temperature overheating, partial discharge, low energy discharge and high energy discharge.

8. The system according to claim 6, characterized in that: The data pre-training module comprises: A cluster analysis submodule, used to perform cluster analysis based on the operation data and the environmental data using an improved DBSCAN algorithm and remove noise data to obtain a clustering result; The normalization processing submodule is used to perform normalization processing based on the clustering results using the interval value method to obtain normalized data to construct a data set.

9. The system according to claim 8, characterized in that: The cluster analysis submodule is specifically used for: Based on the operation data and the environment data, a neighborhood radius and a minimum number of points are calculated using a genetic algorithm; Based on the operation data and the environment data, any data object is selected to determine whether it has been classified into a certain cluster or is determined to be noise data. If so, the data object is reselected; otherwise, the number of data points within the neighborhood radius of the data object as the center point is counted; A determination is made based on whether the number of data points within the neighborhood radius of the data object as the center point is less than the minimum number of points. If so, the data object is marked as a boundary point or a center point; otherwise, a new cluster is established with the data object as the core point; Based on the unlabeled data objects within the neighborhood radius of the data object as the center point, the number of data points within the neighborhood radius is selected to determine whether it is less than the minimum number of points. If so, the unlabeled data objects are reselected; otherwise, the data objects within the neighborhood radius that are not included in other clusters are added to the new cluster to obtain the clustering result.

10. The system according to claim 6, characterized in that: The operating data and environmental data of the distribution transformer in the data acquisition module include: key parameter data and insulation state data in the operating conditions of the dry-type distribution transformer; load information of the oil-immersed distribution transformer, gas dissolved concentration data in the transformer oil, ambient temperature, humidity monitoring quantity and corresponding operating status of the equipment.

Citation Information

Cited By

  • Power distribution network equipment state sensing and full life cycle health assessment method

    CN120632427A

  • A method for state perception and life cycle health assessment of distribution network equipment

    CN120632427B