A method and device for detecting insulation status of a high-voltage composite lightning arrester

By using a clustering algorithm to filter out noise and atypical data points of high-voltage composite lightning arresters and construct a high-quality training data set, the accuracy problem of insulation status detection of high-voltage composite lightning arresters was solved, and the accuracy and reliability of detection were improved.

CN120256986BActive Publication Date: 2025-09-05BEIJING JINGUAN INTELLIGENT ELECTRICAL TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510732577.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-05
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately identify the insulation status of high-voltage composite lightning arresters, which may lead to serious accidents such as breakdown and explosion, affecting the reliability of the power grid.

Method used

By clustering historical electrical data based on a target clustering algorithm, noise and atypical insulation data points are identified and filtered out, and a high-quality training dataset is constructed for training the target classification model to improve the accuracy of insulation status detection.

Benefits of technology

The accuracy and reliability of insulation status detection of high-voltage composite lightning arresters are improved, false alarms and missed alarms are reduced, and the safe operation of the power grid is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256986B_ABST
    Figure CN120256986B_ABST
Patent Text Reader

Abstract

The present application provides a method and device for detecting the insulation status of a high-voltage composite lightning arrester, which belongs to the technical field of lightning arrester detection. The method includes: obtaining target data, which is the current electrical data of the high-voltage composite lightning arrester to be detected; classifying and processing the target data based on a target classification model, wherein the first data set is obtained in the following manner: clustering the data in the second data set based on a target clustering algorithm to obtain a clustering result, and determining atypical insulation data points from each boundary point based on the cluster radius of the cluster where each boundary point is located and the distance between each boundary point and the cluster center of its cluster in the clustering result; screening the second data set based on noise insulation data points and atypical insulation data points, and using the screened data set as the first data set. The present application can improve the data quality of the training set of the lightning arrester insulation detection model, thereby improving the accuracy of lightning arrester insulation detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of lightning arrester detection, and more specifically, relates to a method and device for detecting the insulation status of a high-voltage composite lightning arrester. Background Art

[0002] High-voltage composite arresters are critical devices in power systems that protect equipment from lightning and switching overvoltages. Their insulation performance is directly related to the safe operation of the power grid. In practical applications, failure to promptly detect insulation defects in high-voltage composite arresters can lead to serious accidents such as arrester breakdown and explosion, threatening power grid reliability.

[0003] With the development of smart grid and condition-based maintenance technologies, the power industry has an increasing demand for online monitoring and precise diagnosis of lightning arresters. Therefore, an accurate and reliable method for detecting the insulation condition of lightning arresters is needed. Summary of the Invention

[0004] The purpose of this application is to provide a method and device for detecting the insulation status of a high-voltage composite lightning arrester, so as to improve the data quality of the training set of the lightning arrester insulation status detection model, thereby improving the accuracy of the lightning arrester insulation status detection.

[0005] A first aspect of an embodiment of the present application provides a method for detecting the insulation state of a high-voltage composite lightning arrester, comprising:

[0006] Acquire target data, which is current electrical data of the high-voltage composite lightning arrester to be tested;

[0007] Classify the target data based on the target classification model to obtain a classification result, which includes: the current insulation state information of the high-voltage composite lightning arrester to be detected;

[0008] The target classification model is trained based on the first data set;

[0009] The first data set is obtained in the following way:

[0010] Clustering the data in the second data set based on the target clustering algorithm to obtain a clustering result, the second data set including historical electrical data corresponding to a plurality of high-voltage composite lightning arresters and their corresponding insulation state information, the clustering result including: noise-type insulation data points;

[0011] Determine atypical insulation data points from each boundary point based on the cluster radius of the cluster in which each boundary point is located and the distance between each boundary point and the cluster center of the cluster in which it is located in the clustering result; the data corresponding to the atypical insulation data points are data in the second data set that do not conform to the preset fault mode;

[0012] The second data set is screened based on noise-type insulation data points and atypical-type insulation data points, and the screened data set is used as the first data set.

[0013] A second aspect of the embodiments of the present application provides a device for detecting the insulation state of a high-voltage composite lightning arrester, comprising:

[0014] A data acquisition module is used to acquire target data, which is the current electrical data of the high-voltage composite lightning arrester to be tested;

[0015] The insulation detection module is used to classify the target data based on the target classification model to obtain the classification results, which include: the current insulation status information of the high-voltage composite lightning arrester to be detected;

[0016] The target classification model is trained based on the first data set;

[0017] The first data set is obtained in the following way:

[0018] Clustering the data in the second data set based on the target clustering algorithm to obtain a clustering result, the second data set including historical electrical data corresponding to a plurality of high-voltage composite lightning arresters and their corresponding insulation state information, the clustering result including: noise-type insulation data points;

[0019] Determine atypical insulation data points from each boundary point based on the cluster radius of the cluster in which each boundary point is located and the distance between each boundary point and the cluster center of the cluster in which it is located in the clustering result; the data corresponding to the atypical insulation data points are data in the second data set that do not conform to the preset fault mode;

[0020] The second data set is screened based on noise-type insulation data points and atypical-type insulation data points, and the screened data set is used as the first data set.

[0021] A third aspect of an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above-mentioned method for detecting the insulation status of a high-voltage composite lightning arrester are implemented.

[0022] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned method for detecting the insulation status of a high-voltage composite lightning arrester are implemented.

[0023] The beneficial effects of the method and device for detecting the insulation status of a high-voltage composite lightning arrester provided by the embodiments of the present application are:

[0024] The present application clusters the data in the second data set based on the target clustering algorithm, and can accurately identify the noise-like insulation data points in the clustering results, and screen the second data set based on these noise-like insulation data points, thereby obtaining a first data set that does not contain noise-like insulation data points. The data for training the target classification model is made purer, and the probability of the model learning real and effective features can be improved, thereby improving the accuracy of subsequent detection of the current insulation state of the high-voltage composite lightning arrester. The present application is also based on the cluster radius of the cluster where each boundary point is located in the cluster result and the distance between each boundary point and the cluster center of the cluster where it is located, and can accurately determine the atypical insulation data points from each boundary point, and further screen the second data set based on these atypical insulation data points, so that the target classification model will not be disturbed by the atypical fault data during the training process, and can learn a more accurate and complete fault feature pattern, thereby improving the judgment accuracy of the model in practical applications, reducing the judgment bias, false alarms and omissions in practical applications, and further improving the accuracy and reliability of lightning arrester insulation state detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0026] Figure 1 A schematic flow chart of a method for detecting the insulation status of a high-voltage composite lightning arrester provided in one embodiment of the present application;

[0027] Figure 2 A flowchart of a method for determining a first data set provided in one embodiment of the present application;

[0028] Figure 3 A structural block diagram of a high-voltage composite lightning arrester insulation state detection device provided in one embodiment of the present application;

[0029] Figure 4 A schematic block diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0030] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0031] In order to make the purpose, technical solutions and advantages of this application clearer, specific embodiments will be described below with reference to the accompanying drawings.

[0032] Please refer to Figure 1 , Figure 1 This is a flow chart of a method for detecting the insulation status of a high-voltage composite lightning arrester provided in one embodiment of the present application. The method can be performed by an electronic device and may include:

[0033] S101: Acquire target data, where the target data is current electrical data of the high-voltage composite lightning arrester to be tested.

[0034] In this embodiment, a high-voltage composite lightning arrester is a device used to protect electrical equipment in power systems from overvoltage damage. It is composed of a composite of multiple insulating materials and is capable of operating in high-voltage environments. Electrical data can include voltage, current, insulation resistance, partial discharge, and other data, and can reflect the insulation status of the lightning arrester.

[0035] S102: Classify the target data based on the target classification model to obtain a classification result, which includes: current insulation state information of the high-voltage composite lightning arrester to be detected. The target classification model is obtained by training based on the first data set.

[0036] In this embodiment, the target classification model can be a random forest model, a support vector machine model, or an artificial neural network model. Classification processing refers to using the algorithm and rules of the target classification model to analyze and judge the input target data and classify it into different categories. In this embodiment, the categories are different insulation status categories of high-voltage composite lightning arresters. The classification result is the output result obtained after the target classification model classifies the target data, that is, the insulation status category to which the target data belongs, such as normal, minor fault, major fault, etc. In addition to the insulation status information, it can also include the confidence level of the insulation status information, that is, the degree of certainty of the insulation status information obtained by the target classification model. The insulation status can also be normal, internal moisture, or insulation aging. Specifically, the insulation status detection result is limited by the training data of the target classification model and whether its input data is related to the output data. In this embodiment, the target data obtained are voltage, current, insulation resistance, partial discharge, etc., so the corresponding insulation status detection result can be normal, minor fault, major fault, etc. If further determination of whether the insulation status is normal and the subdivision of the fault under abnormal conditions is required, more relevant electrical data or environmental data needs to be collected, which will not be repeated in this embodiment.

[0037] Furthermore, it can be seen from the above embodiment that the target classification model is obtained by training based on the first data set. In this embodiment, the first data set is a data set obtained by screening the historical electrical data corresponding to multiple high-voltage composite lightning arresters. It should be noted that the models or electrical parameters of the multiple high-voltage composite lightning arresters should be as similar as possible to the high-voltage composite lightning arrester to be tested to ensure the adaptability of the model to the lightning arrester to be tested. Specifically, refer to Figure 2 . Figure 2 A flowchart of a method for determining a first data set provided in one embodiment of the present application is provided; the first data set is obtained through Sa, Sb, and Sc, wherein:

[0038] Sa: Clustering is performed on the data in the second data set based on a target clustering algorithm to obtain clustering results. The second data set includes historical electrical data corresponding to multiple high-voltage composite lightning arresters and their corresponding insulation status information. The clustering results include: noise-type insulation data points.

[0039] In this embodiment, a target clustering algorithm is an algorithm for grouping feature objects into different classes or clusters. It is applied in this scenario to analyze and process the electrical data of a high-voltage composite lightning arrester. The goal is to divide the data into different categories based on similarity and to further filter or determine the data based on the clustering results. In this embodiment, the target clustering algorithm is preferably the DBSCAN density clustering algorithm because its clustering results contain core points, boundary points, and noise points, which facilitates the subsequent determination of the data set. In this embodiment, a core point can be understood as a data point in a relatively dense area within a cluster, which is the cluster center of the cluster or a data point that is relatively close to the cluster center. Taking the electrical data of a high-voltage composite lightning arrester as an example, assume that in the two-dimensional voltage-current data space, there are multiple data points within a region (representing the electrical state of the lightning arrester at different times) with relatively similar voltage and current values. These data points are clustered together to form a cluster. The data points at the center of this cluster or close to the center are core points. The electrical states of the lightning arrester represented by these core points have a high degree of similarity and are typical representatives of this type of electrical state.

[0040] In this embodiment, boundary points can be understood as points that are farther from the cluster center and located at the edge of the cluster. Taking the aforementioned voltage-current data space as an example, at the edge of a cluster, some data points, while somewhat related to other points within the cluster, are farther from the cluster center than the core points. The electrical state of the lightning arrester represented by the boundary point also belongs to the category represented by the cluster, but is relatively on the edge of that category. In other words, it may not belong to the category at all, but only appears to belong to the category in the data. Therefore, further judgment is required.

[0041] In this embodiment, noise points can be understood as data points that do not belong to any cluster, which are usually outliers or error data generated during the data collection process. In the electrical data collection of high-voltage composite lightning arresters, occasional failures of measuring instruments, external interference, etc. may cause some data points that are extremely different from other normal data. These data points cannot form meaningful clusters with other data points and are regarded as noise points. For example, the current value measured at a certain moment obviously deviates from the normal operating range and has no similarity or correlation with the data at other moments. Such data points may be determined as noise points.

[0042] In this embodiment, the second data set contains a collection of historical electrical data corresponding to a plurality of high-voltage composite lightning arresters, covering electrical information at different times and under different operating conditions, such as data on relevant parameters such as voltage, current, insulation resistance, and partial discharge. Clustering processing refers to the use of a target clustering algorithm to calculate and analyze the data in the second data set, and to divide the data into different clusters or categories based on metrics such as the distance between the data. During the clustering process, the algorithm will automatically identify the distribution characteristics of the data, classify similar data points into the same cluster, and divide dissimilar data points into different clusters, thereby forming a clustering result. For example, for the electrical parameters of the lightning arrester, clustering processing can divide data points with similar electrical parameter characteristics into the same cluster, while data points with larger differences will be divided into different clusters, thereby obtaining a clustering result.

[0043] It should be noted that different target clustering algorithms have different corresponding hyperparameters. For example, in the DBSCAN algorithm of this embodiment, the required hyperparameters are the radius of the domain and the minimum number of data points in the neighborhood of a core point. When the target clustering algorithm is the K-means algorithm, the required hyperparameters are the number of clusters and the maximum number of iterations. In this embodiment, the hyperparameters of the target clustering algorithm can be set based on the reference values ​​of the model or determined through multiple experiments.

[0044] In this embodiment, the clustering results contain core points, boundary points, and noise points. Core points can be understood as points at the cluster center or data points close to the cluster center. Boundary points can be understood as points far from the cluster center, that is, points close to the edge of the cluster. Noise points do not belong to any cluster and are noise during data collection. In this application, noise points in the clustering results are defined as noise-insulated data points.

[0045] Sb: Determine atypical insulation data points from each boundary point based on the cluster radius of the cluster where each boundary point is located and the distance between each boundary point and the cluster center of the cluster where it is located in the clustering result; the data corresponding to the atypical insulation data point is the data in the second data set that does not conform to the preset fault mode.

[0046] In this embodiment, it is considered that there are probably some points of low quality in the arrester electrical data collected historically, that is, atypical insulation data points, that is, the data corresponding to the data point may be classified as either an insulation fault or an insulation fault, or it can also be understood that the data corresponding to the atypical insulation data point is the characteristic data in the second data set that does not conform to the preset fault mode. The preset fault mode is set based on the prior knowledge and experience of the possible fault types of the high-voltage composite arrester and the corresponding electrical data characteristics. For example, it is known that when an arrester has an insulation resistance drop fault, its electrical data will show a specific change pattern, which is a preset fault mode. The data corresponding to the atypical insulation data point does not conform to the characteristic data of these preset fault modes, representing special and uncommon situations, or abnormal data points that have not yet been clearly identified as a known fault mode.

[0047] In this embodiment, the cluster radius summarizes the distribution of data points within a cluster and represents the average deviation of most data points from the cluster center. If a data point is close to the cluster center, it means that it has similar characteristics to most data points in the cluster because it is within the main range of the cluster. Conversely, if the distance is farther, it may have different characteristics.

[0048] Therefore, for the lightning arrester insulation status detection scenario, it can be understood that for boundary points, they are inherently at the edge of the cluster and have certain special characteristics. When the distance from a boundary point to the cluster center is much greater than the cluster radius, it indicates that the boundary point is not only at the edge of the cluster, but also has significantly different characteristics from other data points in the cluster. Because in a normal cluster, the distance from most boundary points to the cluster center should fluctuate around the cluster radius, if a boundary point significantly exceeds this range, it is an atypical data point. In this scenario, it is represented by an atypical insulation data point for the lightning arrester. It can be understood that when a lightning arrester data point is far away from the cluster center of its cluster and the cluster is small, it indicates that this data point is atypical fault data, that is, an atypical insulation data point. Because it deviates from the majority of data points, it can be filtered out in subsequent processing.

[0049] Therefore, atypical insulation data points can be screened out through the above judgment method based on cluster radius and the distance from the boundary point to the cluster center.

[0050] Sc: The second data set is screened based on noise-related insulation data points and atypical insulation data points, and the first data set is determined based on the screened data sets.

[0051] In this embodiment, noise-related insulation data points and atypical insulation data points can be removed from the second data set to obtain a filtered data set. Subsequently, this data set can be directly determined as the first data set, or the data in the data set can be further filtered or expanded to ultimately obtain the first data set. It can be understood that the first data set is a purer data set that better represents normal or common insulation states and known fault modes. This screening can reduce the interference of abnormal features on subsequent analysis and model training, thereby improving the accuracy and reliability of insulation state assessment and fault diagnosis of high-voltage composite arresters.

[0052] It can be concluded from the above that the embodiment of the present application can accurately identify the noise-type insulation data points in the clustering results by clustering the data in the second data set based on the target clustering algorithm, and filter the second data set based on these noise-type insulation data points, thereby obtaining a first data set that does not contain noise-type insulation data points. The data for training the target classification model is made purer, and the probability of the model learning real and effective features can be improved, thereby improving the accuracy of subsequent insulation state detection. The embodiment of the present application is also based on the cluster radius of the cluster where each boundary point is located in the cluster result and the distance between each boundary point and the cluster center of the cluster where it is located, and can accurately determine the atypical insulation data points from each boundary point, and further filter the second data set based on these atypical insulation data points, so that the target classification model will not be disturbed by the atypical fault data during the training process, and can learn more accurate and complete fault feature patterns, thereby improving the judgment accuracy of the model in practical applications, reducing the judgment bias, false alarms and omissions in practical applications, and improving the accuracy and reliability of lightning arrester insulation state detection.

[0053] In one embodiment of the present application, determining atypical insulating data points from each boundary point based on the cluster radius of the cluster in which each boundary point is located and the distance between each boundary point and the cluster center of the cluster in which it is located in the clustering result includes:

[0054] For each boundary point, based on the cluster radius of the cluster where the boundary point is located and the distance of the boundary point, suspected atypical insulation data points are determined from each boundary point; the distance of any boundary point is the distance between the boundary point and the cluster center of the cluster where it is located; the probability that the suspected atypical insulation data point belongs to the atypical insulation data point is greater than a preset threshold;

[0055] In response to the number of data points in the cluster where the suspected atypical insulation data point is located being greater than a preset number, the suspected atypical insulation data point is determined to be an atypical insulation data point.

[0056] In this embodiment, a suspected atypical insulation data point refers to a point with a probability greater than a preset threshold that it may be an atypical insulation data point. This means that a suspected atypical insulation data point may be either an atypical insulation data point or a typical insulation data point. The data corresponding to the typical insulation data point is characteristic data in the target data that conforms to a preset fault mode. This is because the aforementioned steps only screen each boundary point based on the cluster radius of the cluster in which the boundary point is located and the distance corresponding to the boundary point (i.e., the distance between the boundary point and the cluster center of the cluster in which it is located). This may lead to a situation where certain fault data is classified as an atypical insulation data point in this application due to its low frequency of occurrence, but in actual classification, this fault data is meaningful, i.e., belongs to the typical insulation data point. Therefore, this application considers first classifying it as a suspected atypical insulation data point, and then determining whether it is an atypical insulation data point based on the number of data points in the cluster. In this embodiment, the preset threshold can be set based on experience.

[0057] In this embodiment, when the number of data points in the cluster where the suspected atypical insulation data point is located is greater than the preset number, it means that this data point does not belong to the aforementioned situation where it is classified as an atypical insulation data point due to its low frequency of occurrence, so it can be classified as an atypical insulation data point to exclude it.

[0058] On the contrary, in response to the number of data points in the cluster where the suspected atypical insulation data point is located being less than or equal to the preset number, the suspected atypical insulation data point is determined as a data point in the first data set.

[0059] From the above, it can be concluded that the present application effectively reduces the risk of misjudgment of atypical insulation data points by first screening out suspected atypical insulation data points based on the cluster radius of the cluster where the boundary point is located and the distance between the boundary point and the cluster center, and then performing a secondary judgment based on the number of data points in the cluster. Since some fault data may be mistakenly classified as atypical insulation data points during the initial screening due to their low frequency of occurrence, these data may be of great significance in actual classification. By considering the factor of the number of data points in the cluster, this type of meaningful fault data is avoided from being mistakenly excluded from the training data, the accuracy of data screening is improved, and it is ensured that the first data set can more comprehensively cover various valuable fault characteristics.

[0060] In one embodiment of the present application, determining suspected atypical insulation data points from each boundary point based on the cluster radius of the cluster where the boundary point is located and the distance of the boundary point includes:

[0061] In response to the cluster radius of the cluster where the boundary point is located being smaller than a preset radius, and the distance between the boundary point and the cluster center of the cluster where the boundary point is located being greater than a first distance, determining the boundary point as a suspected atypical insulation data point;

[0062] In response to the cluster radius of the cluster where the boundary point is located being greater than or equal to the preset radius, and the distance between the boundary point and the cluster center of the cluster where the boundary point is located being greater than a second distance, determining the boundary point as a suspected atypical insulation data point;

[0063] The first distance is smaller than the second distance.

[0064] In this embodiment, when the cluster radius of the cluster containing a boundary point is smaller than a preset radius, and the distance between the boundary point and the cluster center of the cluster is greater than a first distance, the boundary point is screened as a suspected atypical insulation data point. In other words, when the cluster radius is small, it indicates that the data points within the cluster are relatively concentrated. In this case, if the distance between a boundary point and the cluster center is greater than the first distance, it indicates that the boundary point has significantly deviated from the normal data distribution range within the cluster and is likely to have characteristics different from other data points in the cluster. Therefore, it is considered a suspected atypical insulation data point.

[0065] In this embodiment, when the cluster radius of the cluster containing the boundary point is greater than or equal to the preset radius, and the distance between the boundary point and the cluster center of the cluster in which it resides is greater than a second distance, the boundary point is screened as a suspected atypical insulation data point. In other words, when the cluster radius is large, the data points within the cluster are relatively dispersed, and a larger distance threshold (i.e., the second distance) is required to determine whether the boundary point deviates from the normal range. If the distance between the boundary point and the cluster center is greater than the second distance, it indicates that even in a cluster with relatively dispersed data, the boundary point still deviates from the normal data distribution, and therefore it is also screened as a suspected atypical insulation data point. The preset radius, first distance, and second distance can be determined based on experience or multiple experiments.

[0066] From the above, it can be concluded that the present application combines the two key factors of the cluster radius of the cluster where the boundary point is located and the distance between the boundary point and the cluster center, and sets different distance thresholds (first distance and second distance) for different cluster radii to screen suspected atypical insulation data points, thereby achieving a measurement of the degree of deviation of data point features, improving the accuracy of screening suspected atypical insulation data points, and avoiding misjudgment and missed judgment.

[0067] In one embodiment of the present application, screening the second data set based on noise-related insulation data points and atypical insulation data points, and determining the first data set based on the screened data set includes:

[0068] The second data set is screened based on the noise-like insulation data points and the atypical insulation data points in the clustering results, and the screened data set is used as the third data set;

[0069] Expanding the data in the third data set that meets a preset condition to obtain an expanded data set; the preset condition includes: the number of data points in the cluster where the data is located is less than or equal to a preset number;

[0070] The expanded dataset and the third dataset are merged to obtain the first dataset.

[0071] In this embodiment, considering that after the noise-type insulation data points and the atypical insulation data points are removed from the second data set, there is a high probability that there will be a situation in the remaining clusters where the amount of data before the removal is sufficient, but the amount of data in the clusters after the removal is insufficient, the embodiment of the present application can also expand the data that meets the preset conditions to ensure that the amount of data in each training cluster can meet the training requirements.

[0072] The expansion method can be based on generator generation or data augmentation technology, such as performing slight transformations or interpolations on the data to generate new data that is similar to but not identical to the original data, thereby obtaining an expanded data set, and finally merging the expanded data set and the third data set to obtain the first data set. The generator shown in the embodiment of the present application is obtained after pre-training the generative adversarial model.

[0073] From the above, it can be concluded that the present application ensures that each cluster has a sufficient amount of data for model training by identifying the data in which the number of data points in the cluster is less than or equal to the preset number and expanding it, which helps the model learn more comprehensive and stable features, avoids underfitting of the model due to insufficient data, and thus improves the accuracy and generalization ability of the target classification model. The present application expands and merges the specific data to obtain the first data set, which not only removes interference factors such as noise-type insulation data points and atypical insulation data points, but also ensures the quantity and quality of each cluster data, providing a better training data set for the target classification model, enabling the model to more accurately learn the characteristics and laws of the insulation state of the lightning arrester, and then more accurately judge the insulation state of the lightning arrester in actual detection, thereby improving the accuracy and reliability of the lightning arrester insulation state detection.

[0074] In one embodiment of the present application, if there are multiple expanded data sets, the method may further include:

[0075] For each expanded data set, determining a generation effect corresponding to the expanded data set based on a first similarity between each data in the expanded data set and the original expanded data corresponding to the expanded data set, so as to obtain a generation effect corresponding to each expanded data set; the original expanded data is the data corresponding to the expanded data set in the third data set that meets a preset condition; the first similarity is in direct proportion to the generation effect;

[0076] Determine the corresponding attention weights based on the generation effects of each expanded dataset;

[0077] The target classification model is trained based on the first dataset in the following way:

[0078] Determining the attention weights corresponding to the respective data in the first data set based on the attention weights corresponding to the respective expanded data sets and the preset attention weights corresponding to the third data set;

[0079] The initial classification model is trained based on the first data set and the attention weights corresponding to each data in the first data set to obtain a target classification model.

[0080] In this embodiment, it is also considered that the quality of the expanded data is likely to be uneven, and the data itself is generated. Therefore, when training the initial classification model, different attention weights should be set for each expanded data set to achieve better training results.

[0081] In this embodiment, the original expanded data corresponding to each data in the expanded data set and the expanded feature can be converted into the form of a data vector. The process of converting data into a data vector can be to directly form a vector by combining the various feature values ​​of each data point. Because the data in this application is numerical data, for example, a data point has three features, namely voltage value, current value and resistance value, then these three values ​​can be combined into a three-dimensional vector, such as (V, I, R).

[0082] In this embodiment, the cosine similarity or Euclidean distance between the two can be calculated as the similarity of the single data, and the average of the similarities of the data in the expanded dataset can be used as the first similarity of the expanded dataset. The first similarity is directly proportional to the generation effect. The specific proportional relationship can be determined based on multiple experiments or the first similarity can be directly determined as the generation effect.

[0083] In this embodiment, different augmented data sets have different generation effects and different importance to model training. In order to give full play to the role of high-quality augmented data in the model training process and avoid the interference of low-quality augmented data, it is necessary to assign attention weights to each augmented data set. Augmented data sets with good generation effects will be assigned higher attention weights, indicating that these data will be referenced more during model training; augmented data sets with poor generation effects will be assigned lower attention weights to reduce their impact on model training. This allocation method can dynamically adjust the importance of augmented data in model training according to its quality, so that the model can be more inclined to learn more valuable data, thereby improving training efficiency and effectiveness.

[0084] It should be noted that the preset attention weight of the third dataset is generally set to a value greater than the corresponding attention weight of each expanded dataset, because the expanded dataset itself is generated data, and its authenticity and reliability are lower than the real data.

[0085] During the training process, the attention mechanism needs to be set up first. Secondly, the initial classification model will learn the data in the first dataset based on these attention weights. Since the first dataset is obtained by merging the various expanded datasets and the third dataset, the attention weights corresponding to each data in the first dataset can be determined based on the attention weights corresponding to each expanded dataset and the preset attention weights corresponding to the third dataset. Specifically, the attention weights corresponding to each expanded dataset can be directly assigned to the data in the corresponding expanded dataset in the first dataset, and the same applies to the third dataset. Data with high attention weights will have a greater impact in the process of updating model parameters, and data with low attention weights will have a relatively small impact. In this way, the information of the original data and the expanded data is integrated, while taking into account the quality differences of different datasets, so that the trained target classification model can more accurately classify the insulation status of the high-voltage composite lightning arrester, thereby improving the generalization ability and classification accuracy of the model.

[0086] From the above, it can be concluded that the embodiment of the present application determines the generation effect corresponding to the expanded data set by calculating the first similarity between each data in the expanded data set and the original expanded data, thereby realizing the evaluation of the quality of the expanded data, and allocating attention weights to the expanded data set based on the generation effect, so that high-quality expanded data can be fully utilized in model training, while low-quality expanded data is appropriately weakened, avoiding the interference of low-quality data on model training and improving training efficiency and effect.

[0087] In one embodiment of the present application, the target classification model is obtained by training an initial classification model based on the first data set, and the initial classification model is a random forest model;

[0088] The training process of the initial classification model includes:

[0089] extracting data features from a first data set;

[0090] In response to the number of relevant feature pairs in the first data set being greater than a preset number, the depth of the decision tree in the random forest model is increased based on the first step length; wherein a relevant feature pair is a feature pair consisting of two features whose correlation is greater than the second similarity.

[0091] In this embodiment, considering that the depth of the decision tree in the random forest model determines the complexity of the model, and in the insulation state detection scenario of the lightning arrester, when the insulation state is abnormal, its voltage, current, insulation resistance, or partial discharge value generally changes synchronously, that is, there are many complex correlations between the features, so the depth of the decision tree can be determined based on the number of related feature pairs. The preset number, the first step length, and the second similarity can be determined based on experience or based on multiple experiments. The depth of the decision tree can be set based on the reference value of the model. When the aforementioned conditions are met, the reference value can be increased based on the preset first step length.

[0092] When the number of related data pairs in the target data exceeds the preset number, it indicates that there are more complex associations and patterns between the data. Increasing the depth of the decision tree can give the random forest model more opportunities to learn these complex patterns, thereby improving the model's ability to fit the data and classification accuracy. This application refers to increasing the depth of the decision tree in the random forest model based on the first step length as a one-time step, not continuously increasing the depth of the decision tree based on the first step length.

[0093] In one embodiment of the present application, the training process of the initial classification model further includes:

[0094] Improve the node splitting threshold of the decision tree in the random forest model based on the second step size;

[0095] Controls splitting of decision trees in random forest models based on an adjusted node splitting threshold.

[0096] In this embodiment, the node splitting threshold is a parameter that controls the splitting and growth of the decision tree, and the node splitting threshold is a condition for determining whether to continue splitting the node. The reference value set in the random forest model takes into account the noise points and outliers in the data, and prevents the decision tree from learning the noise and outliers as meaningful information. Therefore, the setting is relatively conservative. However, in this application, since the noise points and outliers have been removed, the node splitting threshold can be appropriately increased, and the second step length can be determined based on multiple experiments. During the training process of the random forest model, whether each node of the decision tree is split is determined based on the adjusted node splitting threshold. For each internal node, the impact of the splitting of different features on the evaluation index is calculated. Only when the improvement of the evaluation index brought about by the splitting of a certain feature exceeds the adjusted node splitting threshold, the feature will be selected for node splitting, thereby constructing a more reasonable decision tree structure. In this way, the growth of the decision tree can be controlled to avoid the decision tree being too complex or simple, thereby improving the accuracy and reliability of the random forest model in detecting the insulation status of the high-voltage composite lightning arrester.

[0097] From the above, we can conclude that this application avoids the problem of the model being too simple to capture complex feature relationships or too complex to overfit by dynamically adapting the model complexity according to data features, ensuring good performance of the model under different data conditions. This application can better control the splitting of the decision tree in the random forest model by increasing the node splitting threshold, making the model more focused on meaningful features and patterns, reducing model sensitivity, and enhancing the model's generalization ability.

[0098] Corresponding to the above embodiment, a method for detecting the insulation state of a high-voltage composite lightning arrester is provided. Figure 3 This is a structural block diagram of a high-voltage composite lightning arrester insulation state detection device provided by an embodiment of the present application. For ease of explanation, only the parts related to the embodiment of the present application are shown. Figure 3 The high-voltage composite lightning arrester insulation state detection device 20 includes: a data acquisition module 21 and an insulation detection module 22.

[0099] The data acquisition module 21 is used to acquire target data, which is the current electrical data of the high-voltage composite lightning arrester to be tested;

[0100] The insulation detection module 22 is used to classify the target data based on the target classification model to obtain a classification result, which includes: the current insulation status information of the high-voltage composite lightning arrester to be detected;

[0101] The target classification model is trained based on the first data set;

[0102] The first data set is obtained in the following way:

[0103] Clustering the data in the second data set based on the target clustering algorithm to obtain a clustering result, the second data set including historical electrical data corresponding to a plurality of high-voltage composite lightning arresters and their corresponding insulation state information, the clustering result including: noise-type insulation data points;

[0104] Determine atypical insulation data points from each boundary point based on the cluster radius of the cluster in which each boundary point is located and the distance between each boundary point and the cluster center of the cluster in which it is located in the clustering result; the data corresponding to the atypical insulation data points are data in the second data set that do not conform to the preset fault mode;

[0105] The second data set is screened based on noise-type insulation data points and atypical-type insulation data points, and the first data set is determined based on the screened data set.

[0106] In one embodiment of the present application, determining atypical insulating data points from each boundary point based on the cluster radius of the cluster in which each boundary point is located and the distance between each boundary point and the cluster center of the cluster in which it is located in the clustering result includes:

[0107] For each boundary point, based on the cluster radius of the cluster where the boundary point is located and the distance of the boundary point, suspected atypical insulation data points are determined from each boundary point; the distance of any boundary point is the distance between the boundary point and the cluster center of the cluster where it is located; the probability that the suspected atypical insulation data point belongs to the atypical insulation data point is greater than a preset threshold;

[0108] In response to the number of data points in the cluster where the suspected atypical insulation data point is located being greater than a preset number, the suspected atypical insulation data point is determined to be an atypical insulation data point.

[0109] In one embodiment of the present application, determining suspected atypical insulation data points from each boundary point based on the cluster radius of the cluster where the boundary point is located and the distance of the boundary point includes:

[0110] In response to the cluster radius of the cluster where the boundary point is located being smaller than a preset radius, and the distance between the boundary point and the cluster center of the cluster where the boundary point is located being greater than a first distance, determining the boundary point as a suspected atypical insulation data point;

[0111] In response to the cluster radius of the cluster where the boundary point is located being greater than or equal to the preset radius, and the distance between the boundary point and the cluster center of the cluster where the boundary point is located being greater than a second distance, determining the boundary point as a suspected atypical insulation data point;

[0112] The first distance is smaller than the second distance.

[0113] In one embodiment of the present application, screening the second data set based on noise-related insulation data points and atypical insulation data points, and determining the first data set based on the screened data set includes:

[0114] The second data set is screened based on the noise-like insulation data points and the atypical insulation data points in the clustering results, and the screened data set is used as the third data set;

[0115] Expanding the data in the third data set that meets a preset condition to obtain an expanded data set; the preset condition includes: the number of data points in the cluster where the data is located is less than or equal to a preset number;

[0116] The expanded dataset and the third dataset are merged to obtain the first dataset.

[0117] In one embodiment of the present application, a high-voltage composite lightning arrester insulation state detection device 20 further includes: an attention weight module, configured to determine, for each expanded data set, a generation effect corresponding to the expanded data set based on a first similarity between each data in the expanded data set and the original expanded data corresponding to the expanded data set, so as to obtain the generation effect corresponding to each expanded data set; the original expanded data is the data corresponding to the expanded data set that meets a preset condition in a third data set; the first similarity is in direct proportion to the generation effect;

[0118] Determine the corresponding attention weights based on the generation effects of each expanded dataset;

[0119] The target classification model is trained based on the first dataset in the following way:

[0120] Determining the attention weights corresponding to the respective data in the first data set based on the attention weights corresponding to the respective expanded data sets and the preset attention weights corresponding to the third data set;

[0121] The initial classification model is trained based on the first data set and the attention weights corresponding to each data in the first data set to obtain a target classification model.

[0122] In one embodiment of the present application, the target classification model is obtained by training an initial classification model based on the first data set, and the initial classification model is a random forest model;

[0123] A high-voltage composite lightning arrester insulation state detection device 20 further includes: an initial classification model training module for extracting data features from the first data set;

[0124] In response to the number of relevant feature pairs in the first data set being greater than a preset number, the depth of the decision tree in the random forest model is increased based on the first step length; wherein a relevant feature pair is a feature pair consisting of two features whose correlation is greater than the second similarity.

[0125] In one embodiment of the present application, the target classification model is a target random forest model, and the target random forest model is obtained by training the random forest model based on the first data set;

[0126] The insulation detection module 22 is specifically configured to input target data into a target random forest model; each decision tree in the target random forest model votes on the target data to obtain multiple voting results;

[0127] The result with the highest number of votes in the voting results is taken as the classification result.

[0128] See also Figure 4 , Figure 4This is a schematic block diagram of an electronic device provided in one embodiment of the present application. Figure 4 The electronic device 300 in the embodiment shown may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memory 304 is used to store computer programs, which include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. The processor 301 is configured to call the program instructions to execute the functions of the modules / units in the above-mentioned device embodiments, such as Figure 3 The functions of the data acquisition module 21 and the insulation detection module 22 are shown.

[0129] It should be understood that in the embodiment of the present application, the processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0130] The input device 302 may include a touchpad, a fingerprint collection sensor (for collecting user fingerprint information and fingerprint direction information), a microphone, etc. The output device 303 may include a display (LCD, etc.), a speaker, etc.

[0131] The memory 304 may include a read-only memory and a random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store a preset number.

[0132] In a specific implementation, the processor 301, input device 302, and output device 303 described in the embodiment of the present application can execute the implementation method described in an embodiment of a high-voltage composite lightning arrester insulation status detection method provided in the embodiment of the present application, and can also execute the implementation method of the electronic device described in the embodiment of the present application, which will not be repeated here.

[0133] In another embodiment of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, all or part of the process of the method in the above embodiment is implemented. The computer program can also be used to instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of each of the above method embodiments are implemented. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium.

[0134] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the aforementioned embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, the computer-readable storage medium can include both an internal storage unit of the electronic device and an external storage device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or is about to be output.

[0135] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0136] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the electronic devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0137] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces or units, or can be an electrical, mechanical or other form of connection.

[0138] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present application.

[0139] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0140] The above are only specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for detecting the insulation status of a high-voltage composite lightning arrester, characterized in that: include: Acquiring target data, wherein the target data is current electrical data of the high-voltage composite lightning arrester to be detected; Classify the target data based on the target classification model to obtain a classification result, wherein the classification result includes: current insulation state information of the high-voltage composite lightning arrester to be detected; Wherein, the target classification model is obtained by training based on the first data set; The first data set is obtained by: Clustering the data in the second data set based on a target clustering algorithm to obtain a clustering result, wherein the second data set includes historical electrical data corresponding to a plurality of high-voltage composite lightning arresters and insulation state information corresponding to each of the plurality of high-voltage composite lightning arresters, and the clustering result includes: noise-type insulation data points; Determining atypical insulation data points from each boundary point based on the cluster radius of the cluster in which each boundary point is located and the distance between each boundary point and the cluster center of the cluster in which it is located in the clustering result; the data corresponding to the atypical insulation data point is the data in the second data set that does not conform to the preset fault mode; screening the second data set based on the noise-type insulation data points and the atypical-type insulation data points, and determining the first data set based on the screened data set; The step of determining atypical insulation data points from each boundary point based on the cluster radius of the cluster in which each boundary point is located and the distance between each boundary point and the cluster center of the cluster in which each boundary point is located comprises: For each boundary point, based on the cluster radius of the cluster to which the boundary point belongs and the distance of the boundary point, suspected atypical insulation data points are determined from each boundary point; the distance of any boundary point is the distance between the boundary point and the cluster center of the cluster to which it belongs; the probability that the suspected atypical insulation data point belongs to the atypical insulation data point is greater than a preset threshold; In response to the number of data points in the cluster where the suspected atypical insulation data point is located being greater than a preset number, determining the suspected atypical insulation data point as the atypical insulation data point; The step of determining suspected atypical insulation data points from each boundary point based on the cluster radius of the cluster where the boundary point is located and the distance to the boundary point includes: In response to the cluster radius of the cluster where the boundary point is located being smaller than a preset radius, and the distance between the boundary point and the cluster center of the cluster where the boundary point is located being greater than a first distance, determining the boundary point as a suspected atypical insulation data point; In response to the cluster radius of the cluster where the boundary point is located being greater than or equal to the preset radius, and the distance between the boundary point and the cluster center of the cluster where the boundary point is located being greater than a second distance, determining the boundary point as a suspected atypical insulation data point; The first distance is smaller than the second distance.

2. A method for detecting the insulation state of a high-voltage composite lightning arrester according to claim 1, characterized in that: The step of screening the second data set based on the noise-related insulation data points and the atypical insulation data points, and determining the first data set based on the screened data set, includes: screening the second data set based on the noise-like insulation data points and the atypical insulation data points in the clustering results, and using the screened data set as the third data set; Expanding the data in the third data set that meets a preset condition to obtain an expanded data set; the preset condition includes: the number of data points in the cluster where the data is located is less than or equal to the preset number; The expanded data set and the third data set are merged to obtain the first data set.

3. A method for detecting the insulation state of a high-voltage composite lightning arrester according to claim 2, characterized in that: If there are multiple expanded data sets, the method further includes: For each expanded data set, determining a generation effect corresponding to the expanded data set based on a first similarity between each data in the expanded data set and the original expanded data corresponding to the expanded data set, so as to obtain a generation effect corresponding to each expanded data set; the original expanded data is the data corresponding to the expanded data set in the third data set that meets the preset condition; and the first similarity is in direct proportion to the generation effect; Determine the corresponding attention weights based on the generation effects of the respective expanded data sets; The target classification model is trained based on the first data set in the following manner: Determining the attention weights corresponding to the respective data in the first data set based on the attention weights corresponding to the respective expanded data sets and the preset attention weights corresponding to the third data set; The initial classification model is trained based on the first data set and the attention weights corresponding to each data in the first data set to obtain the target classification model.

4. A method for detecting the insulation state of a high-voltage composite lightning arrester according to claim 1, characterized in that: The target classification model is obtained by training an initial classification model based on the first data set, and the initial classification model is a random forest model; The training process of the initial classification model includes: extracting data features from the first data set; In response to the number of relevant feature pairs in the first data set being greater than a preset number, the depth of the decision tree in the random forest model is increased based on the first step length; wherein a relevant feature pair is a feature pair consisting of two features whose correlation is greater than a second similarity.

5. A method for detecting the insulation status of a high-voltage composite lightning arrester according to claim 4, characterized in that: The target classification model is a target random forest model, and the target random forest model is obtained by training the random forest model based on the first data set; The target data is classified based on the target classification model to obtain a classification result, including: Inputting the target data into the target random forest model; each decision tree in the target random forest model votes on the target data to obtain multiple voting results; The result with the highest number of votes in the voting results is taken as the classification result.

6. A high-voltage composite lightning arrester insulation state detection device, characterized in that: include: A data acquisition module is used to acquire target data, wherein the target data is the current electrical data of the high-voltage composite lightning arrester to be tested; The insulation detection module is used to classify the target data based on the target classification model to obtain a classification result, wherein the classification result includes: current insulation status information of the high-voltage composite lightning arrester to be detected; Wherein, the target classification model is obtained by training based on the first data set; The first data set is obtained by: Clustering the data in the second data set based on a target clustering algorithm to obtain a clustering result, wherein the second data set includes historical electrical data corresponding to a plurality of high-voltage composite lightning arresters and insulation state information corresponding to each of the plurality of high-voltage composite lightning arresters, and the clustering result includes: noise-type insulation data points; Determining atypical insulation data points from each boundary point based on the cluster radius of the cluster in which each boundary point is located and the distance between each boundary point and the cluster center of the cluster in which it is located in the clustering result; the data corresponding to the atypical insulation data point is the data in the second data set that does not conform to the preset fault mode; screening the second data set based on the noise-type insulation data points and the atypical-type insulation data points, and determining the first data set based on the screened data set; The step of determining atypical insulation data points from each boundary point based on the cluster radius of the cluster in which each boundary point is located and the distance between each boundary point and the cluster center of the cluster in which each boundary point is located comprises: For each boundary point, based on the cluster radius of the cluster to which the boundary point belongs and the distance of the boundary point, suspected atypical insulation data points are determined from each boundary point; the distance of any boundary point is the distance between the boundary point and the cluster center of the cluster to which it belongs; the probability that the suspected atypical insulation data point belongs to the atypical insulation data point is greater than a preset threshold; In response to the number of data points in the cluster where the suspected atypical insulation data point is located being greater than a preset number, determining the suspected atypical insulation data point as the atypical insulation data point; The step of determining suspected atypical insulation data points from each boundary point based on the cluster radius of the cluster where the boundary point is located and the distance to the boundary point includes: In response to the cluster radius of the cluster where the boundary point is located being smaller than a preset radius, and the distance between the boundary point and the cluster center of the cluster where the boundary point is located being greater than a first distance, determining the boundary point as a suspected atypical insulation data point; In response to the cluster radius of the cluster where the boundary point is located being greater than or equal to the preset radius, and the distance between the boundary point and the cluster center of the cluster where the boundary point is located being greater than a second distance, determining the boundary point as a suspected atypical insulation data point; The first distance is smaller than the second distance.

7. A high-voltage composite lightning arrester insulation state detection device according to claim 6, characterized in that: The target classification model is obtained by training an initial classification model based on the first data set, and the initial classification model is a random forest model; A high-voltage composite lightning arrester insulation state detection device further includes: an initial classification model training module for extracting data features from the first data set; In response to the number of relevant feature pairs in the first data set being greater than a preset number, the depth of the decision tree in the random forest model is increased based on the first step length; wherein a relevant feature pair is a feature pair consisting of two features whose correlation is greater than a second similarity.

8. A high-voltage composite lightning arrester insulation state detection device according to claim 7, characterized in that: The target classification model is a target random forest model, and the target random forest model is obtained by training the random forest model based on the first data set; The insulation detection module is specifically configured to input the target data into the target random forest model; each decision tree in the target random forest model votes on the target data to obtain multiple voting results; The result with the highest number of votes in the voting results is taken as the classification result.

Citation Information

Patent Citations

  • A feature selection method for clustering algorithm based on density clustering

    CN109543775A

  • Abnormal point proportion optimization method and device based on spectral clustering and computer equipment

    CN109871886A