Big Data-Based Method for Analysis and Verification of Electrical Parameters of Transmission Lines
By preprocessing and clustering the electrical and environmental parameter data of transmission lines, and calculating the differentiation and differentiation characterization coefficients, the problem of inaccurate differentiation of transmission line status is solved, and more efficient fault identification and handling are achieved.
Patent Information
- Application Number
- CN202510730002.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-06-03
AI Technical Summary
In existing technologies, the clustering analysis of electrical and environmental parameter data of transmission lines is not accurate enough, which makes it impossible to accurately distinguish between normal and abnormal states, thus affecting the accuracy of subsequent analysis and verification.
By collecting electrical and environmental parameter data of transmission lines, performing preprocessing, and then conducting cluster analysis, the difference coefficient and difference characterization coefficient are calculated. Based on these coefficients, the consistency and difference of the clusters are determined, and the clusters are prioritized for analysis, with a focus on high-risk areas.
It improves the accuracy of cluster analysis of electrical and environmental parameter data of transmission lines, enabling more accurate identification of abnormal states, optimization of resource allocation, and improvement of fault diagnosis and handling efficiency.
Smart Images

Figure CN120596870B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method for analyzing and verifying electrical parameters of power transmission lines based on big data. Background Technology
[0002] In recent years, the development of big data technology has provided new ideas and methods for the analysis and verification of electrical parameters of transmission lines. By collecting electrical and environmental parameter data of transmission lines and processing and analyzing this data using big data analytics, a more comprehensive and accurate understanding of the operating status of transmission lines can be achieved. However, in practical applications, some problems still need to be solved, such as how to extract effective features from massive amounts of data to improve the accuracy of cluster analysis. Therefore, this study proposes a big data-based method for the analysis and verification of electrical parameters of transmission lines. Through preprocessing, cluster analysis, and analysis of electrical and environmental parameter data, this method improves the accuracy and reliability of the analysis and verification of electrical parameters of transmission lines. This is of great significance for ensuring the safe operation of transmission lines and improving the stability and efficiency of the power system.
[0003] For example, Chinese Patent Application Publication No. CN118731582A discloses a method, device, and transmission line monitoring system for locating transmission line faults. The method includes: acquiring vibration signals from the transmission line; determining the vibration amplitude and frequency of the transmission line based on the vibration signals; determining that the transmission line vibration is abnormal when the vibration amplitude and / or frequency are both outside a predetermined range; acquiring electrical transmission parameter signals of the transmission line with abnormal vibration; determining whether there are electrical parameter anomalies in the transmission line based on the electrical transmission parameter signals, including harmonic anomalies, transient overcurrent anomalies, and connection hardware faults; calculating the correlation metric value between the abnormal data corresponding to the vibration anomaly and the abnormal data corresponding to the electrical parameter anomaly; and determining the fault location range based on the acquisition locations of the vibration signal and the electrical transmission parameter signal when the correlation metric value is greater than a predetermined threshold. This method solves the problem of low accuracy in locating transmission line faults in the prior art.
[0004] However, existing technologies have the problem that the clustering analysis of electrical and environmental parameter data of transmission lines is not accurate enough, which makes it impossible to accurately distinguish between the normal and abnormal states of transmission lines, thus causing deviations in subsequent analysis and verification. Summary of the Invention
[0005] To address this issue, the present invention provides a big data-based method for analyzing and verifying electrical parameters of transmission lines. This method overcomes the problem in existing technologies where the clustering analysis of electrical and environmental parameter data of transmission lines is not accurate enough, leading to an inability to accurately distinguish between normal and abnormal states of transmission lines, which in turn causes deviations in subsequent analysis and verification.
[0006] To achieve the above objectives, this invention provides a method for analyzing and verifying electrical parameters of transmission lines based on big data, comprising:
[0007] Step S1: Collect electrical parameter data and environmental parameter data of the transmission line, and preprocess the electrical parameter data and environmental parameter data to obtain standard data;
[0008] Step S2 involves performing cluster analysis on the standard data to divide it into several clusters. This includes extracting electrical parameter features and environmental parameter features from the standard data, comparing the differences in features of each data point, and combining the data points according to the differences to obtain clusters. The clusters must meet predetermined division criteria.
[0009] Step S3: Calculate the difference coefficient based on the difference in characteristic parameters between data points within the cluster, and determine whether the cluster meets the consistency standard based on the difference coefficient.
[0010] Step S4: Select the analysis and processing method for each cluster, including: if the cluster meets the consistency criteria, identify the characteristics of the data points within the cluster, calculate the difference characterization coefficient based on the characteristics to determine whether the cluster meets the difference criteria, prioritize the clusters that meet the difference criteria, select the transmission lines corresponding to any data point with the largest difference in several clusters from high to low priority for key analysis and verification, and mark them; if the cluster does not meet the difference criteria, conduct a comprehensive analysis and verification of the transmission lines corresponding to all data points within the cluster.
[0011] Step S5: Classify the results of key analysis and verification and comprehensive analysis and verification respectively, establish an abnormal problem list and a normal status list, compare the problem characteristics in the consistent standard clusters and inconsistent standard clusters, determine the accuracy of the previous cluster analysis, and redetermine the cluster analysis weight coefficients if the accuracy is less than the preset accuracy.
[0012] Furthermore, the preprocessing of the electrical parameter data and environmental parameter data includes removing outliers from the data using the Raida criterion, filling missing data using linear interpolation, and mapping the data to the [0,1] interval using the max-min normalization method.
[0013] Further, in step S2, the predetermined partitioning criteria include the average Euclidean distance between data points within a cluster being less than a first preset threshold and the Euclidean distance between the centers of any two clusters being greater than a second preset threshold.
[0014] Further, in step S3, calculating the difference coefficient based on the difference in characteristic parameters between data points within a cluster includes:
[0015] Calculate the Euclidean distance between any two data points within a cluster and form a distance matrix;
[0016] Calculate the average of the distance matrix and use it as the average difference between data points within the cluster;
[0017] Calculate the Euclidean distance from each data point within a cluster to the cluster center, and form a center distance vector;
[0018] Calculate the standard deviation of the center distance vector and use it as an indicator of the dispersion of data points around the cluster center;
[0019] The difference coefficient is obtained by weighting and summing the average difference and the dispersion index.
[0020] Further, in step S3, determining whether the cluster meets the consistency criterion based on the difference coefficient includes:
[0021] If the difference coefficient is less than or equal to the preset difference coefficient, the cluster is determined to meet the consistency standard;
[0022] If the difference coefficient is greater than the preset difference coefficient, it is determined that the cluster does not meet the consistency standard.
[0023] Furthermore, the preset differentiation coefficient is determined based on the average value of the differentiation coefficients of each cluster that has completed cluster analysis for the same transmission line.
[0024] Further, in step S4, calculating the difference characterization coefficient based on the features includes:
[0025] Calculate the difference between the key feature parameters of each data point and the corresponding feature parameters of the cluster center to obtain the difference vector;
[0026] The difference vector is normalized, and the norm of the normalized difference vector is calculated, which is used as a measure of the degree of difference of the data point relative to the cluster center.
[0027] Calculate the average value of the measure of difference among all data points within a cluster, and use it as the difference characterization coefficient of the cluster.
[0028] Further, in step S4, determining that a cluster meets the difference criteria includes that the difference characterization coefficient of the cluster is greater than or equal to a preset difference characterization coefficient threshold.
[0029] Further, in step S4, prioritizing the clusters that meet the difference criteria includes:
[0030] Calculate a priority score for each cluster, the priority score being obtained by weighted summation based on the following factors;
[0031] The magnitude of the difference characterization coefficient;
[0032] The degree of anomaly in the electrical parameters of data points within a cluster;
[0033] The degree of influence of environmental parameters on data points within a cluster;
[0034] Clusters are sorted in descending order based on priority scores, with clusters having higher priority scores appearing first.
[0035] Furthermore, in step S5, the accuracy of the preliminary clustering analysis is determined based on the proportion of data points correctly classified into the normal state list in the consistent standard cluster to the total number of data points in that cluster, the proportion of data points incorrectly classified into the abnormal problem list in the consistent standard cluster to the total number of data points in that cluster, the proportion of data points correctly classified into the abnormal problem list in the inconsistent standard cluster to the total number of data points in that cluster, and the proportion of data points incorrectly classified into the normal state list in the inconsistent standard cluster to the total number of data points in that cluster.
[0036] Compared with the prior art, the beneficial effects of the present invention are that, by statistically analyzing the data characteristics of electrical and environmental parameters and setting reasonable first and second preset thresholds, the present invention can make the data points within the clusters more similar in terms of comprehensive characteristics and the differences between different clusters more obvious. This helps to accurately classify the transmission line data into appropriate clusters, improve the accuracy and rationality of clustering, and thus obtain more valuable clustering results that better reflect the operating status of the transmission lines.
[0037] Furthermore, this invention calculates the differentiation coefficient by combining two dimensions: average difference and dispersion index. This allows for a comprehensive and integrated assessment of the consistency of data points within a cluster. The average difference reflects the overall degree of difference between data points within a cluster, while the dispersion index reflects the density of the distribution of data points around the cluster center. The differentiation coefficient obtained by weighted summation of the two avoids the limitations of single-index evaluation and more accurately reflects the quality of the cluster. This method improves the accuracy of cluster analysis of electrical and environmental parameter data of transmission lines, enabling accurate differentiation between normal and abnormal states of transmission lines and thus selecting key analysis areas.
[0038] Furthermore, this invention calculates the difference vector between each data point and the cluster center, and performs normalization and norm calculation to accurately quantify the degree of difference of each data point relative to the cluster center. This method not only considers the absolute differences of key characteristic parameters (such as voltage, current, and power), but also eliminates the influence of different parameter dimensions through normalization, making the difference measurement more objective and comparable. Taking the average value of the difference measurement of all data points within the cluster as the difference characterization coefficient can comprehensively reflect the data distribution characteristics of the entire cluster. The higher the difference characterization coefficient, the greater the degree of deviation of the data points within the cluster from the center, and the stronger the data dispersion; conversely, it indicates that the data points are relatively concentrated and the clustering effect is better. In the analysis and verification of transmission lines, clusters that meet the difference criteria often indicate potential problems or anomalies. By prioritizing and analyzing these clusters, resource allocation can be effectively optimized, concentrating limited human and material resources on the transmission lines that require the most attention, and improving the efficiency of fault diagnosis and handling.
[0039] Furthermore, this invention achieves a multi-dimensional comprehensive evaluation of clusters by weighted summation of three factors: the difference characterization coefficient, the degree of electrical parameter anomaly, and the degree of environmental parameter influence. This method not only considers the degree of difference between data points and cluster centers but also combines the two key factors of electrical parameter anomaly and environmental parameter influence, making the evaluation results more comprehensive and objective, and able to more accurately reflect the potential risks of transmission lines. By calculating priority scores and sorting them in descending order, high-risk clusters that require special attention can be accurately identified. The degree of electrical parameter anomaly directly reflects whether the current operating status of the transmission line deviates from the normal range, while the degree of environmental parameter influence considers the potential threats of the external environment to the line. The difference characterization coefficient reveals the stability of the clusters from the perspective of data distribution. The combination of these three factors helps to identify potential fault points in advance. Through the above method, the transmission lines with the highest degree of anomaly can be selected from the key analysis clusters for subsequent analysis, improving the accuracy of cluster analysis of electrical and environmental parameter data of transmission lines, as well as the accuracy and efficiency of subsequent electrical parameter analysis of transmission lines. Attached Figure Description
[0040] Figure 1 This is a flowchart illustrating the workflow of the big data-based method for analyzing and verifying electrical parameters of power transmission lines according to an embodiment of the present invention.
[0041] Figure 2 This is a flowchart illustrating the process of calculating the difference coefficient in the big data-based power transmission line electrical parameter analysis and verification method according to an embodiment of the present invention.
[0042] Figure 3 This is a flowchart illustrating the process of calculating the difference characterization coefficient in the big data-based power transmission line electrical parameter analysis and verification method of this invention.
[0043] Figure 4 This is a flowchart illustrating the process of determining whether a cluster meets the consistency standard in the big data-based power transmission line electrical parameter analysis and verification method according to an embodiment of the present invention. Detailed Implementation
[0044] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0045] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0046] Please see Figures 1-4 As shown, Figure 1 This is a flowchart illustrating the workflow of the big data-based method for analyzing and verifying electrical parameters of power transmission lines according to an embodiment of the present invention. Figure 2 This is a flowchart illustrating the process of calculating the difference coefficient in the big data-based power transmission line electrical parameter analysis and verification method according to an embodiment of the present invention. Figure 3 This is a flowchart illustrating the process of calculating the difference characterization coefficient in the big data-based power transmission line electrical parameter analysis and verification method of this invention. Figure 4 This is a flowchart illustrating the process of determining whether a cluster meets the consistency standard in the big data-based power transmission line electrical parameter analysis and verification method according to an embodiment of the present invention.
[0047] The present invention provides a method for analyzing and verifying electrical parameters of transmission lines based on big data, comprising:
[0048] Step S1: Collect electrical parameter data and environmental parameter data of the transmission line, and preprocess the electrical parameter data and environmental parameter data to obtain standard data;
[0049] Step S2 involves performing cluster analysis on the standard data to divide it into several clusters. This includes extracting electrical parameter features and environmental parameter features from the standard data, comparing the differences in features of each data point, and combining the data points according to the differences to obtain clusters. The clusters must meet predetermined division criteria.
[0050] Step S3: Calculate the difference coefficient based on the difference in characteristic parameters between data points within the cluster, and determine whether the cluster meets the consistency standard based on the difference coefficient.
[0051] Step S4: Select the analysis and processing method for each cluster, including: if the cluster meets the consistency criteria, identify the characteristics of the data points within the cluster, calculate the difference characterization coefficient based on the characteristics to determine whether the cluster meets the difference criteria, prioritize the clusters that meet the difference criteria, select the transmission lines corresponding to any data point with the largest difference in several clusters from high to low priority for key analysis and verification, and mark them; if the cluster does not meet the difference criteria, conduct a comprehensive analysis and verification of the transmission lines corresponding to all data points within the cluster.
[0052] Step S5: Classify the results of key analysis and verification and comprehensive analysis and verification respectively, establish an abnormal problem list and a normal status list, compare the problem characteristics in the consistent standard clusters and inconsistent standard clusters, determine the accuracy of the previous cluster analysis, and redetermine the cluster analysis weight coefficients if the accuracy is less than the preset accuracy.
[0053] The electrical parameter data in this embodiment of the invention includes, but is not limited to, “voltage, current and power (active power and reactive power) of transmission lines”, and the environmental parameter data includes, but is not limited to, “temperature, humidity and air pressure”. The preprocessing of the electrical parameter data and environmental parameter data includes removing outliers from the data using the Laida criterion, filling missing data using linear interpolation, and mapping the data to the [0,1] interval using the max-min normalization method.
[0054] In this embodiment of the invention, the key analysis and verification includes a detailed inspection of the electrical parameters of the transmission lines corresponding to the largest difference data points in the clusters that meet the difference criteria, such as checking the stability and accuracy of parameters like voltage, current, and power; assessing the equipment condition of the transmission lines, including changes in electrical parameters such as resistance, inductance, and capacitance, as well as issues like equipment aging and damage; analyzing the impact of environmental parameters on the transmission lines, such as the influence of environmental factors like temperature, humidity, and wind speed on electrical parameters like line insulation performance and conductor resistance; and checking the operational records of the transmission lines, including historical fault records and maintenance records. This involves understanding the operational status and potential problems of the transmission line; comprehensive analysis and verification including a full inspection of the electrical parameters of all data points within clusters that do not meet the difference criteria, covering parameters such as voltage, current, power, resistance, inductance, and capacitance; a comprehensive inspection of the transmission line equipment; an assessment of the transmission line's operating environment, including the impact of long-term changes in environmental parameters on the line, and interference from the surrounding environment; and in-depth analysis of the transmission line's operational data, including comparisons of operational data from different time periods and comparisons with operational data from other similar transmission lines, to identify potential problems and patterns.
[0055] Specifically, in step S2, the predetermined partitioning criteria include the average Euclidean distance between data points within a cluster being less than a first preset threshold and the Euclidean distance between the centers of any two clusters being greater than a second preset threshold.
[0056] In this embodiment of the invention, the first preset threshold and the second preset threshold can be determined by the following method: First, statistical analysis is performed on the collected electrical and environmental parameter data of the transmission line to understand the distribution range, dispersion, and other characteristics of the data. For example, the range and fluctuation of parameters such as voltage, current, power, temperature, humidity, and air pressure are analyzed. The purpose of clustering is to divide the data into clusters with similar characteristics to better analyze the operating status of the transmission line. Therefore, the threshold setting should ensure that the divided clusters have significant distinguishability in terms of electrical and environmental parameters, while ensuring that the data within each cluster is consistent. Consistency; through multiple experiments, try different threshold combinations, observe the quality of clustering results, and use some evaluation indicators, such as cluster compactness (the density of data points within a cluster) and separation (the degree of separation between different clusters), to determine the optimal threshold range; for example, suppose that after analyzing a large amount of data, the following data characteristics are obtained: among electrical parameters, the voltage range is [100, 220] kV, the current range is [50, 500] A, the active power range is [10, 1000] MW, and the reactive power range is [-500, 500] Mvar; among environmental parameters, the temperature range is [ The temperature range is -20 to 40 degrees Celsius, the humidity range is 20% to 80%, and the air pressure range is 90 to 110 kPa. After multiple experiments, it was found that when the first preset threshold is set to 0.5 and the second preset threshold is set to 1.5, a relatively ideal clustering result can be obtained. At this time, the average Euclidean distance of the data points within the cluster is less than 0.5, which means that the data points within the cluster are relatively similar in terms of the comprehensive characteristics of electrical and environmental parameters. For example, in a cluster, the differences in voltage, current, power, temperature, humidity, air pressure, and other parameters of each data point are small, indicating that the transmission lines corresponding to these data points are operating in a relatively stable environment. Under similar conditions, the Euclidean distance between the centers of any two clusters is greater than 1.5, ensuring that there are obvious differences between different clusters. For example, one cluster may represent a power transmission line operating under high temperature and high humidity and high load, while another cluster may represent a power transmission line operating under low temperature and low humidity and low load. By setting such a threshold, power transmission lines in different operating states can be effectively distinguished. The first preset threshold is preferably 0.5, and the second preset threshold is preferably 1.5. However, the above values are not limited to these. Those skilled in the art can also adjust the values of the corresponding preset thresholds according to actual needs.
[0057] This invention, through statistical analysis of the data characteristics of electrical and environmental parameters and by setting reasonable first and second preset thresholds, enables data points within clusters to be more similar in terms of comprehensive characteristics and makes the differences between different clusters more obvious. This helps to accurately classify transmission line data into appropriate clusters, improve the accuracy and rationality of clustering, and thus obtain more valuable clustering results that better reflect the operating status of transmission lines.
[0058] Specifically, in step S3, the step of calculating the difference coefficient based on the difference in feature parameters between data points within a cluster includes:
[0059] Step S3301: Calculate the Euclidean distance between any two data points within a cluster to form a distance matrix;
[0060] Step S3302: Calculate the average value of the distance matrix and use it as the average difference of data points within the cluster;
[0061] Step S3303: Calculate the Euclidean distance from each data point within a cluster to the cluster center, forming a center distance vector;
[0062] Step S3304: Calculate the standard deviation of the center distance vector and use it as an indicator of the dispersion of data points around the cluster center;
[0063] Step S3305: The average difference and the dispersion index are weighted and summed to obtain the difference coefficient.
[0064] In this embodiment of the invention, it is assumed that there is a cluster containing 5 data points, each data point having 3 features (taking voltage, current, and power from electrical parameters as an example). Assume there is a cluster C, where the data points are: D1 = [100, 200, 300] (representing voltage 100, current 200, and power 300, and so on); D2 = [120, 220, 320]; D3 = [90, 180, 280]; D4 = [110, 210, 310]; D5 = [105, 205, 305]. The Euclidean distance between any two data points within the cluster is calculated using the Euclidean distance formula. A distance matrix is formed, and the other elements are calculated similarly to obtain the distance matrix. The sum of all elements in the distance matrix is 348.08, and the number of elements in the distance matrix is 25. Therefore, the average variance is 13.9232. First, the cluster center D0 = [105, 205, 305] is calculated, and then the distance from other data points to the cluster center is calculated to obtain the center distance vector V = [8.66, 15.81, 22.36, 14.14, 0]. The average value of the center distance vector is 12.194, and the standard deviation is 7.49. Assuming that the weight of the average variance α = 0.4 and the weight of the dispersion index β = 0.6, the variance coefficient is 10.06328.
[0065] In this embodiment of the invention, the weights of the average variance and the dispersion index can be determined by the following method: Multiple cluster analyses are performed using historical data, and different weight combinations are tried. The degree of matching between the clustering results and the actual situation is observed, and the weight combination that best matches the actual situation is selected. For example, the clustering results can be compared with known transmission line fault records or actual operating conditions, and indicators such as accuracy and recall can be calculated. The weights are determined by optimizing these indicators. For instance, historical data of 100 clusters are collected, each cluster having a corresponding actual operating status label (normal or abnormal). Initial weights α = 0.5 and β = 0.5 are set, and cluster analysis is performed on these 100 clusters to calculate the variance coefficient. The cluster status is determined based on the difference coefficient, and then compared with the actual labels, resulting in an accuracy of 70%. Next, the weights are adjusted to α = 0.6 and β = 0.4, and the analysis is repeated, increasing the accuracy to 75%. Further adjustments to α = 0.45 and β = 0.55 result in an accuracy of 72%. After multiple trials, it was found that the accuracy is highest when α = 0.6 and β = 0.4. Therefore, in this scenario, the weight of the average difference is determined to be 0.6, and the weight of the dispersion index is determined to be 0.4. The preferred values for the average difference and the dispersion index are 0.4, but these values are not limited to these, and those skilled in the art can adjust them according to actual needs.
[0066] Specifically, in step S3, when determining whether the cluster meets the consistency criteria, the cluster meets the consistency criteria based on the comparison result between the difference coefficient and the preset difference coefficient.
[0067] When the difference coefficient is less than or equal to the preset difference coefficient, the cluster is determined to meet the consistency standard;
[0068] If the difference coefficient is greater than the preset difference coefficient, it is determined that the cluster does not meet the consistency standard.
[0069] In this embodiment of the invention, the preset difference coefficient is the average of the difference coefficients of each cluster that has completed cluster analysis for the same transmission line. However, the value is not limited to this, and those skilled in the art can adjust the value according to actual needs.
[0070] This invention calculates the differentiation coefficient by combining two dimensions: average difference and dispersion index. This allows for a comprehensive and integrated assessment of the consistency of data points within a cluster. The average difference reflects the overall degree of difference between data points within a cluster, while the dispersion index reflects the density of the distribution of data points around the cluster center. The differentiation coefficient obtained by weighted summation of the two avoids the limitations of single-index evaluation and more accurately reflects the quality of the cluster. This method improves the accuracy of cluster analysis of electrical and environmental parameter data of transmission lines, enabling accurate differentiation between normal and abnormal states of transmission lines and thus selecting key analysis areas.
[0071] Specifically, in step S4, the step of calculating the difference characterization coefficient based on the features includes:
[0072] Step S4401: Calculate the difference between the key feature parameters of each data point and the corresponding feature parameters of the cluster center to obtain the difference vector;
[0073] Step S4402: Normalize the difference vector and calculate the norm of the normalized difference vector, using it as a measure of the degree of difference between the data point and the cluster center.
[0074] Step S4403: Calculate the average value of the difference measure of all data points within the cluster and use it as the difference characterization coefficient of the cluster.
[0075] The key feature parameters of each data point in this embodiment of the invention include, but are not limited to, "voltage, current and power"; assuming we have a cluster C containing 4 data points, each data point has 3 key feature parameters (here we still take voltage, current and power in electrical parameters as an example), the specific data are as follows: D1 = [110, 220, 330]; D2 = [120, 230, 340]; D3 = [100, 210, 320]; D4 = [130, 240, 350];
[0076] Calculate the cluster center D0, and average the feature parameters to obtain D0 = [115, 225, 335]. Calculate the difference between the key feature parameters of each data point and the corresponding feature parameters of the cluster center to obtain the difference vector:
[0077] The difference vector between D1 and D0 is V1 = [110-115, 220-225, 330-335] = [-5, -5, -5];
[0078] The difference vector between D2 and D0 is V2 = [120-115, 230-225, 340-335] = [5, 5, 5].
[0079] The difference vector between D3 and D0 is V3 = [100-115, 210-225, 320-335] = [-15, -15, -15];
[0080] The difference vector between D4 and D0 is V4 = [130-115, 240-225, 350-335] = [15, 15, 15]. This difference vector is normalized, and its norm is calculated. This norm is used as a measure of the difference between the data point and the cluster center. For V1, its norm... The normalized vector is The difference measure of this data point is 1. Similarly, the difference measure of V2 is 1, the difference measure of V3 is 1, and the difference measure of V4 is 1. Calculate the average difference measure of all data points in the cluster and use it as the difference characterization coefficient of the cluster. Difference characterization coefficient = (1+1+1+1) / 4 = 1.
[0081] Specifically, in step S4, when determining whether a cluster meets the difference criteria, the cluster meets the difference criteria based on the comparison result between the difference characterization coefficient of the cluster and the preset difference characterization coefficient threshold.
[0082] When the difference characterization coefficient of a cluster is greater than or equal to a preset difference characterization coefficient threshold, the cluster is determined to meet the difference criteria.
[0083] When the difference characterization coefficient of a cluster is less than the preset difference characterization coefficient threshold, the cluster is determined to not meet the difference standard.
[0084] In this embodiment of the invention, the preset difference characterization coefficient threshold is the average value of the difference characterization coefficients of each cluster of several identical transmission lines under different operating conditions. However, the above value is not limited to this, and those skilled in the art can adjust the value according to actual needs.
[0085] This invention calculates the difference vector between each data point and the cluster center, and then performs normalization and norm calculations to accurately quantify the degree of difference of each data point relative to the cluster center. This method not only considers the absolute differences of key characteristic parameters (such as voltage, current, and power), but also eliminates the influence of different parameter dimensions through normalization, making the difference measurement more objective and comparable. The average value of the difference measurement of all data points within the cluster is taken as the difference characterization coefficient, which can comprehensively reflect the data distribution characteristics of the entire cluster. The higher the difference characterization coefficient, the greater the degree of deviation of the data points within the cluster from the center, and the stronger the data dispersion; conversely, it indicates that the data points are relatively concentrated and the clustering effect is better. In the analysis and verification of transmission lines, clusters that meet the difference criteria often indicate potential problems or anomalies. By prioritizing and analyzing these clusters, resource allocation can be effectively optimized, concentrating limited human and material resources on the transmission lines that require the most attention, and improving the efficiency of fault diagnosis and handling.
[0086] Specifically, in step S4, prioritizing clusters that meet the difference criteria includes:
[0087] Calculate a priority score for each cluster, the priority score being obtained by weighted summation based on the following factors;
[0088] The magnitude of the difference characterization coefficient;
[0089] The degree of anomaly in the electrical parameters of data points within a cluster;
[0090] The degree of influence of environmental parameters on data points within a cluster;
[0091] Clusters are sorted in descending order based on priority scores, with clusters having higher priority scores appearing first.
[0092] In this embodiment of the invention, the degree of electrical parameter anomaly of data points within a cluster can be determined by comparing them with standard values. For example, for various electrical parameters of a transmission line, such as voltage, current, and power, there are corresponding rated values or normal operating ranges. The electrical parameters of each data point within the cluster are compared with these standard values. For example, for a transmission line with a rated voltage of 220kV, if the voltage value of a certain data point deviates significantly from 220kV, exceeding the normal fluctuation range (e.g., ±10%), it indicates that the voltage parameter of that data point is abnormal. By calculating the degree of deviation of the electrical parameter of each data point from the standard value and comprehensively evaluating the deviation of each electrical parameter, the degree of electrical parameter anomaly of that data point is obtained. For the entire cluster, the average or weighted average of the degree of electrical parameter anomaly of all data points can be calculated to represent the degree of electrical parameter anomaly of the data points within the cluster. The degree of influence of environmental parameters on the data points within the cluster can be determined by establishing an influence model of environmental parameters (such as temperature, humidity, and wind speed) on electrical parameters. For example, an increase in temperature will lead to an increase in conductor resistance. Larger humidity levels affect parameters such as current and power; increased humidity may reduce the insulation performance of insulators, leading to electrical faults. Based on these influence models, the degree of influence of environmental parameters on electrical parameters for each data point within a cluster is analyzed. For each data point, the influence index of environmental parameters on electrical parameters is calculated based on its environmental parameter value. For a cluster, the average or weighted average of the influence index of environmental parameters for all data points can be calculated to represent the degree of influence of environmental parameters on data points within the cluster. For example, taking the influence of temperature on conductor resistance as an example, the relationship between conductor resistance R and temperature T is known to be R = R0(1 + α(T-T0)) (R0 is the resistance at reference temperature T0, and α is the temperature coefficient of resistance). The temperature of a certain data point is Ti. The influence of temperature change on resistance is calculated according to this formula, which in turn affects parameters such as current and power, thus obtaining the degree of influence of environmental parameters (temperature) on electrical parameters for that data point. Taking into account the influence of other environmental parameters, the degree of influence of environmental parameters on each data point is obtained, thus obtaining the degree of influence of environmental parameters on data points within the cluster.
[0093] In this embodiment of the invention, the weights of the difference characterization coefficient, the weights of the electrical parameter anomalies of data points within clusters, and the weights of the environmental parameter influence of data points within clusters can be determined by the following methods: collecting a large amount of historical data of transmission lines, combining it with actual fault records, analyzing the correlation between different factors and fault occurrence, and inferring the weights by statistically analyzing the contribution of each factor in fault cases. For example, statistically analyzing the proportion of faults directly caused by electrical parameter anomalies, the proportion of faults caused by environmental factors, and the correspondence between the difference characterization coefficient anomalies and faults in historical faults, thereby determining the weights. For example, analyzing the transmission line data of a certain region over the past 3 years, a total of 200 fault events were recorded. Statistics show that 120 failures were directly caused by abnormal electrical parameters, 50 failures were caused by long-term environmental parameter effects leading to equipment performance degradation, and 30 failures had early warnings from the difference characterization coefficient before obvious electrical abnormalities appeared. Based on this, the weights of each factor were calculated as follows: the weight of the degree of electrical parameter abnormality β = 0.6; the weight of the degree of environmental parameter influence γ = 0.25; and the weight of the magnitude of the difference characterization coefficient α = 0.15. However, the above values are not limited to these, and those skilled in the art can adjust these values according to actual needs.
[0094] In this embodiment of the invention, a local power company collected electrical parameter data (voltage, current, power, etc.) and environmental parameter data (temperature, humidity, and wind speed) of transmission lines over a period of time. After data preprocessing and cluster analysis, multiple clusters meeting the difference criteria were obtained. For each cluster, the difference characterization coefficient was calculated according to steps S4401-S4403. For example, the difference characterization coefficient of cluster A is 0.8, and the difference characterization coefficient of cluster B is 0.6. For each data point within a cluster, its electrical parameters are compared with standard values, and the degree of deviation is calculated. For example, if there are 5 data points in cluster A, the degree of deviation of the voltage, current, and power of each data point from the standard values is calculated. Assume that the voltage deviation of data point 1 is 15%, the current deviation is 10%, and the power deviation is 8%; the voltage deviation of data point 2 is 10%, the current deviation is 8%, and the power deviation is 6%, and so on. To calculate the degree of electrical parameter anomaly for each data point, we can use a comprehensive approach, such as simply adding them together and taking the average. For example, the degree of electrical parameter anomaly for data point 1 is (15% + 10% + 8%) / 3 = 11%. Calculate the average degree of electrical parameter anomaly for all data points within a cluster. If the degrees of electrical parameter anomaly for the 5 data points in cluster A are 11%, 9%, 10%, 8%, and 12%, then the degree of electrical parameter anomaly for cluster A is (11% + 9% + 10% + 8%) / 3 = 11%. (10%) + 12%) / 5 = 10%; Establish a model of the influence of environmental parameters on electrical parameters, such as the influence of temperature on conductor resistance (R = R0[1 + α × (T - T0)]) and the influence of humidity on insulator insulation performance; For each data point in a cluster, calculate the influence index of environmental parameters on electrical parameters based on its environmental parameter values. For example, if the temperature of data point 1 is 30℃, the reference temperature T0 is 25℃, the temperature coefficient of resistance α is 0.004, and R0 is 10Ω, then the influence of temperature on resistance is R = 10 × [1 + 0.004 × (30 - 25)] = 10.2Ω, which in turn affects parameters such as current and power, thus obtaining the degree of influence of the environmental parameter (temperature) on electrical parameters at this data point.Considering the combined impact of multiple environmental parameters, the degree of influence of each environmental parameter at each data point is calculated. The average value of the environmental parameter influence index for all data points within a cluster is calculated. The degree of influence of the environmental parameters for the five data points in cluster A are 10%, 8%, 9%, 7%, and 11%, respectively. Therefore, the degree of influence of the environmental parameters in cluster A is (10% + 8% + 9% + 7% + 11%) / 5 = 9%. Based on the above statistical analysis, the weight of the difference characterization coefficient is determined to be α = 0.15, and the weight of the degree of electrical parameter anomaly of the data points within the cluster is determined to be β = 0.6. The influence weight of environmental parameters on data points within a cluster is γ = 0.25; the priority score of each cluster is calculated, and the priority score of cluster A is 0.15 × 0.8 + 0.6 × 10% + 0.25 × 9% = 0.12 + 0.06 + 0.0225 = 0.2025; the priority score of cluster B can be calculated in the same way; the clusters are sorted in descending order according to the priority score, and the clusters with higher priority scores are placed first. For example, if the priority score of cluster A is higher than that of cluster B, then cluster A is placed before cluster B.
[0095] This invention achieves a multi-dimensional comprehensive evaluation of clusters by weighted summation of three factors: the difference characterization coefficient, the degree of electrical parameter anomaly, and the degree of environmental parameter influence. This method not only considers the degree of difference between data points and cluster centers but also combines the two key factors of electrical parameter anomaly and environmental parameter influence, making the evaluation results more comprehensive and objective, and able to more accurately reflect the potential risks of transmission lines. By calculating priority scores and sorting them in descending order, high-risk clusters that require special attention can be accurately identified. The degree of electrical parameter anomaly directly reflects whether the current operating status of the transmission line deviates from the normal range, while the degree of environmental parameter influence considers the potential threats of the external environment to the line. The difference characterization coefficient reveals the stability of the clusters from the perspective of data distribution. The combination of these three factors helps to identify potential fault points in advance. Using the above method, the transmission lines with the highest degree of anomaly can be selected from the key clusters for subsequent analysis, improving the accuracy of cluster analysis of electrical and environmental parameter data of transmission lines, as well as the accuracy and efficiency of subsequent electrical parameter analysis of transmission lines.
[0096] Specifically, in step S5, the accuracy of the preliminary clustering analysis is determined based on the proportion of data points correctly classified into the normal state list in the consistent standard clusters to the total number of data points in the clusters, the proportion of data points incorrectly classified into the abnormal problem list in the consistent standard clusters to the total number of data points in the clusters, the proportion of data points correctly classified into the abnormal problem list in the inconsistent standard clusters to the total number of data points in the clusters, and the proportion of data points incorrectly classified into the normal state list in the inconsistent standard clusters to the total number of data points in the clusters.
[0097] In this embodiment of the invention, the proportion of data points correctly classified into the normal state list in the consistency standard cluster to the total number of data points in the cluster is calculated and denoted as P1.
[0098] Calculate the proportion of data points in the consistent standard clusters that were misclassified into the list of anomalous issues, out of the total number of data points in that cluster, denoted as P2; calculate the proportion of data points in the inconsistent standard clusters that were correctly classified into the list of anomalous issues, out of the total number of data points in that cluster, denoted as P3; calculate the proportion of data points in the inconsistent standard clusters that were misclassified into the list of normal states, out of the total number of data points in that cluster, denoted as P4; then, calculate the accuracy K of the cluster analysis using the following formula:
[0099]
[0100] Where n1 is the total number of data points in the consistent standard cluster, and n2 is the total number of data points in the inconsistent standard cluster.
[0101] In this embodiment of the invention, the redetering of cluster analysis weight coefficients includes redetering the weights of the average difference, the dispersion index, the difference characterization coefficient, the electrical parameter anomaly of data points within the cluster, and the environmental parameter influence of data points within the cluster according to the weight coefficient determination method in this invention. The preset duration can be set to ten times the data acquisition interval of the transmission line, and the preset accuracy is the accurate average value after cluster analysis in the same transmission line. However, the above values are not limited to these, and those skilled in the art can adjust the values according to actual needs.
[0102] This invention determines the accuracy of cluster analysis by calculating the correct and incorrect classification ratios of data points in consistent and inconsistent standard clusters. This allows for a comprehensive and objective evaluation of the accuracy of clustering results. The calculation of different ratios considers the classification of normal and abnormal situations, making the accuracy calculation more reflective of the cluster analysis's ability to distinguish between normal and abnormal data. This helps identify problems in the cluster analysis and allows for targeted improvements. The method for re-determining cluster analysis weight coefficients when the accuracy is lower than the preset accuracy involves adjusting the weights based on the influence of different factors (average variance, dispersion index, variance characterization coefficient, degree of electrical parameter anomaly, and degree of environmental parameter influence) on the clustering results. This makes the cluster analysis more aligned with actual needs. Different transmission lines may have different operating characteristics; by adjusting the weight coefficients, the influence of key factors on the clustering results can be highlighted, improving the adaptability and accuracy of cluster analysis for specific transmission lines. This method improves the accuracy of cluster analysis of electrical and environmental parameter data of transmission lines, as well as the precision and efficiency of subsequent electrical parameter analysis of transmission lines.
[0103] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A method for analyzing and checking electric parameters of a power transmission line based on big data, characterized in that, The method comprises the following steps: S1, collecting electrical parameter data and environmental parameter data of a power transmission line, and preprocessing the electrical parameter data and the environmental parameter data to obtain standard data; S2, performing cluster analysis on the standard data to divide a plurality of cluster clusters, including extracting electrical parameter features and environmental parameter features in the standard data, comparing differences of features of data points, and combining the data points to obtain the cluster clusters according to the differences, wherein the cluster clusters need to meet a predetermined division standard; S3, calculating a differentiation coefficient according to a difference amount of feature parameters between the data points in the cluster clusters, and determining whether the cluster clusters meet a consistency standard according to the differentiation coefficient; S4, selecting an analysis and processing mode for each cluster cluster, including, if the cluster cluster meets the consistency standard, identifying features of the data points in the cluster cluster, calculating a difference representation coefficient according to the features to determine whether the cluster cluster meets a difference standard, and performing priority sorting on the cluster clusters meeting the difference standard, and selecting any maximum difference data point in the cluster clusters according to the priority from high to low to perform key analysis and verification on the power transmission line corresponding to the maximum difference data point, and marking; if the cluster cluster does not meet the difference standard, performing comprehensive analysis and verification on the power transmission lines corresponding to all data points in the cluster cluster; S5, classifying results of the key analysis and verification and the comprehensive analysis and verification respectively, establishing an abnormal problem list and a normal state list, comparing problem features in the cluster clusters meeting the consistency standard and the cluster clusters not meeting the consistency standard, determining accuracy of the early cluster analysis, and re-determining a cluster analysis weight coefficient under the condition that the accuracy is less than a preset accuracy; The preprocessing of the electrical parameter data and the environmental parameter data includes removing outliers in the data by using the Radau criterion, filling missing data by using the linear interpolation method, and mapping the data to the [0, 1] interval by using the maximum-minimum normalization method; In the step S2, the predetermined division standard includes that an average Euclidean distance of the data points in the cluster cluster is less than a first preset threshold value, and a Euclidean distance between centers of any two cluster clusters is greater than a second preset threshold value; In the step S3, the calculation of the differentiation coefficient according to the difference amount of the feature parameters between the data points in the cluster clusters includes: calculating Euclidean distances between any two data points in the cluster cluster to form a distance matrix; calculating an average value of the distance matrix as an average difference amount of the data points in the cluster cluster; calculating Euclidean distances from each data point in the cluster cluster to the cluster center to form a center distance vector; calculating a standard deviation of the center distance vector as a dispersion degree index of the data points around the cluster center; weighting and summing the average difference amount and the dispersion degree index to obtain the differentiation coefficient.
2. The big data based transmission line electrical parameter analysis verification method according to claim 1, characterized in that, In the step S3, the determination of whether the cluster cluster meets the consistency standard according to the differentiation coefficient includes: if the differentiation coefficient is less than or equal to a preset differentiation coefficient, it is determined that the cluster cluster meets the consistency standard; if the differentiation coefficient is greater than the preset differentiation coefficient, it is determined that the cluster cluster does not meet the consistency standard.
3. The big data based transmission line electrical parameter analysis verification method according to claim 2, characterized in that, The preset differentiation coefficient is determined according to an average value of differentiation coefficients of the cluster clusters of the same power transmission line that have completed cluster analysis.
4. The big data based transmission line electrical parameter analysis verification method according to claim 3, characterized in that, In the step S4, the calculating the difference representation coefficient according to the feature includes: calculating the difference between the key feature parameter of each data point and the corresponding feature parameter of the cluster center, to obtain a difference vector; normalizing the difference vector, and calculating the norm of the normalized difference vector as a measure of the difference degree of the data point relative to the cluster center; calculating the average of the difference degree measures of all data points in the cluster, and taking it as the difference representation coefficient of the cluster.
5. The big data based transmission line electrical parameter analysis verification method according to claim 4, wherein, In the step S4, the determining that the cluster meets the difference standard includes that the difference representation coefficient of the cluster is greater than or equal to a preset difference representation coefficient threshold.
6. The big data based transmission line electrical parameter analysis verification method according to claim 5, wherein, In the step S4, the priority sorting of the clusters meeting the difference standard includes: calculating the priority score of each cluster, and the priority score is obtained by weighted summation according to the following factors: the size of the difference representation coefficient; the abnormality degree of the electrical parameters of the data points in the cluster; the influence degree of the environmental parameters of the data points in the cluster; descendingly sorting the clusters according to the priority score, and the cluster with a higher priority score is arranged in front.
7. The big data based transmission line electrical parameter analysis verification method according to claim 6, wherein, In the step S5, the accuracy of the preliminary clustering analysis is determined according to the proportion of the data points correctly classified into the normal state list in the total data points in the consistency standard cluster, the proportion of the data points incorrectly classified into the abnormal problem list in the total data points in the consistency standard cluster, the proportion of the data points correctly classified into the abnormal problem list in the total data points in the inconsistency standard cluster, and the proportion of the data points incorrectly classified into the normal state list in the total data points in the inconsistency standard cluster.
Citation Information
Patent Citations
Power transmission line fault positioning method and device and power transmission line monitoring system
CN118731582A
Method and system for monitoring abnormity of power supply line of new energy photovoltaic power station
CN118783649A
Online intelligent monitoring method and system for power transmission line
CN120030372A