Identification method of key factors of distribution network operation based on PMU measurement data

Through Euclidean distance similarity clustering and multivariate linear regression analysis, the problem of failure to identify the operation mode and key factors of the distribution network in the existing technology is solved, and the quantitative evaluation of the distribution network operation status and the identification of key factors are achieved, guiding future operation control.

CN116738268BActive Publication Date: 2025-09-30STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310546134.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-16
Publication Date
2025-09-30
Estimated Expiration
2043-05-16

AI Technical Summary

Technical Problem

Existing technologies fail to identify the operating modes and key factors at the system level of the distribution network, resulting in the utilization of PMU data remaining in the ex post stage and unable to guide future operations.

Method used

The PMU data is preprocessed by Euclidean distance similarity clustering and adaptive improved K-means clustering. The operation mode of the distribution network and typical power sources are defined. The regression coefficient matrix is ​​analyzed using multiple linear regression, and key indicators are quantified to identify key factors.

Benefits of technology

It is possible to extract typical operating modes and power sources of the distribution network from PMU data, conduct critical quantitative evaluation and identification, focus on monitoring and preventive control of key factors, and assist in abnormal event analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116738268B_ABST
    Figure CN116738268B_ABST
Patent Text Reader

Abstract

The present invention provides a method for identifying key factors in distribution network operation based on PMU measurement data, which relates to the field of distribution networks. The method comprises the following steps: pre-processing PMU measurement data based on the characteristics of Euclidean distance similarity clustering; adaptively improving traditional K-means clustering, iteratively clustering the pre-processed PMU data to obtain high-quality clustering results; defining distribution network operation modes and typical power sources based on the clustering results and their physical meanings; analyzing the influence of typical power sources on typical distribution network operation modes based on multivariate linear regression to obtain a regression coefficient matrix that characterizes the strength of the influence; and defining typical operation mode criticality coefficients and typical power source criticality coefficients based on the regression coefficient matrix as key indicators to identify key factors affecting distribution network operation. The present invention can extract typical distribution network operation modes and power sources from a large and complex amount of distribution network PMU measurement data, and perform quantitative evaluation and identification of key factors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of distribution network PMU data mining and analysis and its application, and specifically to a method for identifying key factors of distribution network operation based on PMU measurement data. Background Art

[0002] With the increasing integration of distributed power sources and random devices such as electric vehicles into distribution networks, the inherent uncertainty of distribution systems has significantly increased, and their operation has become more complex. This poses significant challenges to maintaining observable status, detectable faults, and controllable operation of distribution networks, placing higher demands on the measurement capabilities of distribution systems. Synchronized phasor measurement technology (PMU) significantly improves the real-time, accurate, and synchronized measurement capabilities, providing powerful data support and new decision-making tools for distribution network operation control and energy management.

[0003] Currently, research on the massive amount of operational data provided by PMUs in distribution networks focuses on fault identification and location, topology identification and parameter identification, state estimation and power quality monitoring, harmonic monitoring, and harmonic source location. For example, PMU data is feature mined and event labeled using moving and dynamic time windows. Various machine learning algorithms are applied to the labeled data to obtain trained classifiers / neural networks, which are then used to identify and classify new data. Another example is harmonic state estimation based on hybrid PMU and SCADA measurements, and the harmonic active power at each node is calculated using the harmonic estimation results to identify multiple harmonic sources. Furthermore, a hierarchical and partitioned harmonic responsibility allocation model is developed, leveraging the limited number of PMUs in distribution networks, primarily installed at key nodes.

[0004] However, such studies do not analyze the distribution network from a system-level perspective, fail to identify and judge the operation mode of the distribution network, and fail to identify the key factors that have a significant impact on the operation mode of the distribution network. As a result, the use of PMU data of the distribution network remains in the ex post stage, and it is impossible to use PMU historical data to identify the typical operating status and typical power sources of the distribution network, and analyze the strength of the correlation between them, so as to provide certain guidance for future operations. Summary of the Invention

[0005] In view of the above-mentioned deficiencies in the prior art, the present invention provides a method for identifying key factors of distribution network operation based on PMU measurement data.

[0006] The present invention is achieved through the following technical solutions.

[0007] According to one aspect of the present invention, a method for identifying a distribution network operation mode based on PMU data is provided, comprising:

[0008] S1. Preprocess the PMU measurement data based on the characteristics of Euclidean distance similarity clustering;

[0009] S2, based on the traditional K-means clustering, adaptively improve and iterate the clustering of pre-processed PMU data to obtain high-quality clustering results;

[0010] S3. Based on the high-quality clustering results and their physical meaning, define the distribution network operation mode and typical power sources;

[0011] S4. Analyze the impact of typical power sources on typical distribution network operation modes based on multivariate linear regression, and obtain a regression coefficient matrix that characterizes the strength of the impact;

[0012] S5. Based on the regression coefficient matrix, the critical coefficients of typical operation modes and typical power sources are defined as key indicators to identify the key factors affecting the operation of the distribution network.

[0013] Preferably, the method of pre-processing the PMU measurement data in step S1 based on the characteristics of Euclidean distance similarity clustering is specifically as follows:

[0014] For the distribution network measurement data samples collected by PMU, K-means clustering needs to be performed based on Euclidean distance. The calculation formula of Euclidean distance is as follows:

[0015]

[0016] Where x and y are the measurement data samples of the same physical quantity at two different PMU measurement points; i ,y i are their corresponding i-th elements.

[0017] K-means clustering based on Euclidean distance is used to cluster the time series curves of the measurement data samples based on morphological similarity. To avoid interference caused by differences in the magnitude of measurement values ​​between different measurement points, the clustered data is normalized and preprocessed as follows:

[0018]

[0019] Where, is the normalized data; x i is the data before normalization; x range is the range of the PMU corresponding to the data sample.

[0020] Preferably, the method of performing adaptive improvement on traditional K-means clustering and iterative clustering on the pre-processed PMU data to obtain high-quality clustering results is specifically as follows:

[0021] The number of clusters K of K-means clustering is determined by an adaptive method based on iteration, and the iteration amount is the number of clusters K: based on the clustering quality evaluation index I K , calculate the change rate ΔI of the evaluation index between the adjacent cluster numbers K and K-1 K , set the clustering quality evaluation index change rate threshold ΔI th , if ΔI K ≤ΔI th , then the corresponding K value is the desired number of clusters.

[0022] Clustering quality evaluation index I K As shown below:

[0023]

[0024] Where x is the cluster C k Data samples in c k is the cluster C k The cluster center of .

[0025] Clustering quality evaluation index change rate ΔI K As shown below:

[0026]

[0027] For clustering quality evaluation index change rate threshold ΔI th , the general value range is [0.02,0.1].

[0028] For the clusters obtained by clustering, the cluster center of each cluster is used as the representative of its typical operating characteristics, which is also the final clustering result.

[0029] Preferably, the method of defining the distribution network operation mode and typical power source based on the clustering results and their physical meanings in step S3 is as follows:

[0030] Using the adaptive iterative K-means clustering method described above, PMU measurement data for AC lines, major loads, and switch station PQ flows were clustered. AC lines are the primary structure of the distribution network, so the AC line flow clustering results represent the typical operating mode (flow distribution) of the distribution network. Switch station flows are primarily influenced by renewable energy generation and loads and represent power injection at major nodes. Therefore, the major load and switch station flow clustering results are typical representations of the distribution network's power sources. Based on this analysis, the m-class AC line flow clustering results are defined as the m-class typical operating mode of the distribution network, and the p-class major load and switch station flow clustering results are defined as the p-class typical power sources of the distribution network.

[0031] Preferably, the step S4 is based on a multivariate linear regression method to analyze the influence of typical power sources on the typical operation mode of the distribution network, and obtain a regression coefficient matrix representing the strength of the influence, specifically:

[0032] Take one of the typical distribution network operating modes as the dependent variable and all typical power sources of the distribution network as the independent variables to perform a multiple linear regression analysis. Repeat the above process for all typical distribution network operating modes. The regression equation of the multiple linear regression can be expressed as:

[0033] y=Xβ+ε

[0034] Where y is the dependent variable vector; X is the independent variable matrix, consisting of the row vectors of p data samples (main loads, switch stations) and 1; β is the regression coefficient vector; and ε is the error vector. The specific form is shown below.

[0035]

[0036] Based on the regression equation, the least squares method is used to minimize the error function to obtain the estimated value of the dependent variable. and regression coefficient estimates The error function is as follows:

[0037]

[0038] After performing multiple regression analysis on all typical operation modes of distribution networks, the following regression coefficient matrix is ​​obtained:

[0039]

[0040] Where m is the number of typical operation modes of the distribution network. Considering that the positive and negative signs in the power flow only represent the direction of the power flow, the absolute value of the elements in the matrix |β i,j | represents the degree to which the i-th typical operation mode of the distribution network is affected by the j-th power source, |β i,j |The larger it is, the stronger the impact.

[0041] Preferably, the step S5 defines the typical operation mode critical coefficient and the typical power source critical coefficient based on the regression coefficient matrix as key indicators to identify the key factors affecting the operation of the distribution network, specifically:

[0042] Based on the regression coefficient matrix Define the criticality coefficient γ of typical operation mode i and the typical power source criticality factor δ j :

[0043]

[0044] Criticality coefficient of typical operation mode γ i It describes the influence of the typical power sources of the distribution network on the i-th typical operation mode of the distribution network. The larger the value, the more easily the i-th typical operation mode of the distribution network is affected by various typical power sources of the distribution network, and the stronger the volatility. Therefore, γ i Greater than the threshold γ th The typical operation mode of the distribution network is defined as the key typical operation mode of the distribution network, which needs to be monitored in detail.

[0045] Similarly, the typical power source criticality factor δ j It describes the influence of the jth typical power source of the distribution network on all typical operation modes of the distribution network. The larger the value, the greater the fluctuation of the jth typical power source of the distribution network will cause greater fluctuations in various typical operation modes of the distribution network. j Greater than the threshold δ th The typical power sources of the distribution network are defined as the key power sources of the distribution network and need to be monitored intensively.

[0046] Generally, the judgment threshold γ th ,δ th Set to the mean, that is

[0047] Based on the above technical solution, the present invention has the following beneficial effects: the method for identifying key factors of distribution network operation based on PMU measurement data provided by the present invention performs data mining based on a large amount of complex distribution network PMU measurement data, can extract typical operating modes and power sources of the distribution network, and perform key quantitative evaluation and identification; it can help the distribution network to focus on monitoring and preventive control of key factors according to actual operating conditions and needs, or assist the distribution network in abnormal event analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a schematic flow diagram of the method of the present invention;

[0049] Figure 2 is a multivariate linear regression analysis fitting diagram of the typical operation mode of the distribution network affected by the typical power source in a preferred embodiment of the present invention;

[0050] Figure 3 is a regression coefficient matrix heat map in a preferred embodiment of the present invention;

[0051] Figure 4 This is a key identification heat map of a typical operating mode in a preferred embodiment of the present invention;

[0052] Figure 5 This is a typical power source criticality identification thermal diagram in a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0053] The following is a detailed description of an embodiment of the present invention. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process. It should be noted that those skilled in the art may make various modifications and improvements without departing from the scope of the present invention, and these modifications and improvements fall within the scope of protection of the present invention.

[0054] An embodiment of the present invention provides a method for identifying key factors in distribution network operation based on PMU measurement data, involving distribution network PMU data mining, analysis, and application. The method preprocesses PMU measurement data based on the characteristics of Euclidean distance similarity clustering. Adaptively improve upon traditional K-means clustering, and iteratively cluster the preprocessed PMU data to obtain high-quality clustering results. Based on the clustering results and their physical meanings, the method defines distribution network operation modes and typical power sources. The method analyzes the influence of typical power sources on typical distribution network operation modes based on multivariate linear regression, obtaining a regression coefficient matrix representing the strength of the influence. Based on the regression coefficient matrix, the method defines typical operation mode criticality coefficients and typical power source criticality coefficients as key indicators to identify key factors affecting distribution network operation. The present invention can extract typical distribution network operation modes and power sources from a large and complex amount of distribution network PMU measurement data, and perform quantitative evaluation and identification of their criticality. This method has practical theoretical significance and promotional value for mining PMU measurement data in actual distribution network operation.

[0055] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0056] Please see first Figure 1 , Figure 1 1 is a flow chart of a method for identifying key factors of distribution network operation based on PMU measurement data provided by an embodiment of the present invention. As shown in the figure, the method for identifying key factors of distribution network operation based on PMU measurement data implemented by the present invention includes the following steps:

[0057] Step 1) Preprocess the PMU measurement data based on the characteristics of Euclidean distance similarity clustering.

[0058] As a preferred embodiment, the method of step 1) is:

[0059] For the distribution network measurement data samples collected by PMU, K-means clustering needs to be performed based on Euclidean distance. The calculation formula of Euclidean distance is as follows:

[0060]

[0061] Where x and y are the measurement data samples of the same physical quantity at two different PMU measurement points; i ,y i are their corresponding i-th elements.

[0062] K-means clustering based on Euclidean distance is used to cluster the time series curves of the measurement data samples based on morphological similarity. To avoid interference caused by differences in the magnitude of measurement values ​​between different measurement points, the clustered data is normalized and preprocessed as follows:

[0063]

[0064] Where, is the normalized data; x i is the data before normalization; x range is the range of the PMU corresponding to the data sample.

[0065] In step 2), adaptive improvements are made based on traditional K-means clustering, and clustering iterations are performed on the preprocessed PMU data to obtain high-quality clustering results.

[0066] As a preferred embodiment, the method of step 2) is:

[0067] The number of clusters K of K-means clustering is determined by an adaptive method based on iteration, and the iteration amount is the number of clusters K: based on the clustering quality evaluation index I K , calculate the change rate ΔI of the evaluation index between the adjacent cluster numbers K and K-1 K , set the clustering quality evaluation index change rate threshold ΔI th , if ΔI K ≤ΔI th , then the corresponding K value is the desired number of clusters.

[0068] Clustering quality evaluation index I K As shown below:

[0069]

[0070] Where x is the cluster C k Data samples in c k is the cluster C k The cluster center of .

[0071] Clustering quality evaluation index change rate ΔI K As shown below:

[0072]

[0073] For clustering quality evaluation index change rate threshold ΔI th , the general value range is [0.02,0.1].

[0074] For the clusters obtained by clustering, the cluster center of each cluster is used as the representative of its typical operating characteristics, which is also the final clustering result.

[0075] Step 3), based on the clustering results and their physical meaning, define the distribution network operation mode and typical power sources.

[0076] As a preferred embodiment, the method of step 3) is:

[0077] Using the adaptive iterative K-means clustering method described above, PMU measurement data for AC lines, major loads, and switch station PQ flows were clustered. AC lines are the primary structure of the distribution network, so the AC line flow clustering results represent the typical operating mode (flow distribution) of the distribution network. Switch station flows are primarily influenced by renewable energy generation and loads and represent power injection at major nodes. Therefore, the major load and switch station flow clustering results are typical representations of the distribution network's power sources. Based on this analysis, the m-class AC line flow clustering results are defined as the m-class typical operating mode of the distribution network, and the p-class major load and switch station flow clustering results are defined as the p-class typical power sources of the distribution network.

[0078] Step 4) Based on multiple linear regression, the influence of typical power sources on the typical operation mode of the distribution network is analyzed to obtain a regression coefficient matrix that characterizes the strength of the influence.

[0079] As a preferred embodiment, the method of step 4) is:

[0080] Take one of the typical distribution network operating modes as the dependent variable and all typical power sources of the distribution network as the independent variables to perform a multiple linear regression analysis. Repeat the above process for all typical distribution network operating modes. The regression equation of the multiple linear regression can be expressed as:

[0081] y=Xβ+ε

[0082] Where y is the dependent variable vector; X is the independent variable matrix, consisting of the row vectors of p data samples (main loads, switch stations) and 1; β is the regression coefficient vector; and ε is the error vector. The specific form is shown below.

[0083]

[0084] Based on the regression equation, the least squares method is used to minimize the error function to obtain the estimated value of the dependent variable. and regression coefficient estimates The error function is as follows:

[0085]

[0086] After performing multiple regression analysis on all typical operation modes of distribution networks, the following regression coefficient matrix is ​​obtained:

[0087]

[0088] Where m is the number of typical operation modes of the distribution network. Considering that the positive and negative signs in the power flow only represent the direction of the power flow, the absolute value of the elements in the matrix |β i,j | represents the degree to which the i-th typical operation mode of the distribution network is affected by the j-th power source, |β i,j |The larger it is, the stronger the impact.

[0089] Step 5) Based on the regression coefficient matrix, the critical coefficients of typical operation modes and typical power sources are defined as key indicators to identify key factors affecting the operation of the distribution network.

[0090] As a preferred embodiment, the method of step 5) is:

[0091] Based on the regression coefficient matrix Define the criticality coefficient γ of typical operation mode i and the typical power source criticality factor δ j :

[0092]

[0093] Criticality coefficient of typical operation mode γ i It describes the influence of the typical power sources of the distribution network on the i-th typical operation mode of the distribution network. The larger the value, the more easily the i-th typical operation mode of the distribution network is affected by various typical power sources of the distribution network, and the stronger the volatility. Therefore, γ i Greater than the threshold γ th The typical operation mode of the distribution network is defined as the key typical operation mode of the distribution network, which needs to be monitored in detail.

[0094] Similarly, the typical power source criticality factor δ j It describes the influence of the jth typical power source of the distribution network on all typical operation modes of the distribution network. The larger the value, the greater the fluctuation of the jth typical power source of the distribution network will cause greater fluctuations in various typical operation modes of the distribution network. j Greater than the threshold δ thThe typical power sources of the distribution network are defined as the key power sources of the distribution network and need to be monitored intensively.

[0095] Generally, the judgment threshold γ th ,δ th Set to the mean, that is

[0096]

[0097] Figure 2 This is a multiple linear regression analysis fitting diagram of the typical operation mode of the distribution network in an embodiment of the present invention affected by a typical power source, showing a comparison of the typical operation mode of the distribution network and its multiple linear regression fitting effect; Figure 3 This is a heat map of the regression coefficient matrix in an embodiment of the present invention, which intuitively shows the numerical comparison of each element in the regression coefficient matrix; Figure 4 This is a heat map of criticality identification of typical operating modes in an embodiment of the present invention, showing critical operating modes (value 1) and non-critical operating modes (value 0); Figure 5 This is a typical power source criticality identification heat map in an embodiment of the present invention, showing critical power sources (value is 1) and non-critical power sources (value is 0).

[0098] The method for identifying key factors of distribution network operation based on PMU measurement data provided by the above-mentioned embodiment of the present invention performs data mining based on a large amount of complex distribution network PMU measurement data, and can extract key operating information therefrom, thereby realizing the definition and identification of typical operating states and typical power sources of the distribution network; the analysis of the influence of typical power sources on the typical operating mode of the distribution network based on multivariate linear regression can quantitatively evaluate the criticality of the typical operating states and typical power sources of the distribution network, and can help the distribution network to focus on monitoring and preventive control of key factors according to actual operating conditions and needs, or assist the distribution network in abnormal event analysis.

[0099] Compared with traditional distribution network PMU data mining and analysis, it can analyze from the system level of the distribution network and identify and judge the operation mode of the distribution network; it can identify the key factors that have an important impact on the operation mode of the distribution network, so that the use of PMU data of the distribution network is not just in the ex post stage, but it can use PMU historical data to identify the typical operation status and typical power source of the distribution network, and analyze the strength of the correlation between them, so as to provide certain guidance for future operations.

[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for identifying key factors of distribution network operation based on PMU measurement data, characterized in that: The steps include: S1. Preprocess the PMU measurement data based on the characteristics of Euclidean distance similarity clustering; S2, based on the traditional K-means clustering, adaptively improve and iterate the clustering of pre-processed PMU data to obtain high-quality clustering results; S3. Based on the high-quality clustering results and their physical meaning, define the distribution network operation mode and typical power sources; S4. Analyze the impact of typical power sources on typical distribution network operation modes based on multivariate linear regression, and obtain a regression coefficient matrix that characterizes the strength of the impact; S5. Based on the regression coefficient matrix, define the critical coefficients of typical operation modes and typical power sources as key indicators to identify key factors affecting the operation of the distribution network; The step S3 is specifically as follows: The adaptive improved K-means clustering method is used to cluster the PMU measurement data of AC lines, main loads, and switch station PQ flows, and define The clustering results of AC line power flow are the distribution network Typical operation mode of the class; definition The clustering results of the main loads and switch station flows are the distribution network Class typical power source; The step S4 is specifically as follows: Taking one of the typical operation modes of the distribution network as the dependent variable and all the typical power sources of the distribution network as the independent variables, a multiple linear regression analysis is performed. The above process is repeated for all the typical operation modes of the distribution network. The regression equation of the multiple linear regression is expressed as: Where, is the dependent variable vector; is the independent variable matrix, The row vector of the data samples is composed of 1; is the regression coefficient vector; is the error vector, and its specific form is as follows: Based on the regression equation, the least squares method is used to minimize the error function to obtain the estimated value of the dependent variable. and regression coefficient estimates , the error function is as follows: After performing multiple regression analysis on all typical operation modes of distribution networks, the following regression coefficient matrix is ​​obtained: : Where, is the number of typical operation modes of the distribution network. Considering that the positive and negative signs in the flow only represent the direction of the flow, the absolute value of the elements in the matrix Represents the distribution network The typical operation mode is affected by The degree of influence of various power sources, The larger it is, the stronger the impact; The step S5 is specifically as follows: Based on the regression coefficient matrix , define the criticality coefficient of typical operation mode and typical power source criticality coefficients : Will Greater than threshold The typical operation mode of the distribution network is defined as the key typical operation mode of the distribution network, and it is monitored in key areas; Will Greater than threshold The typical power source of the distribution network is defined as the key power source of the distribution network, and it is monitored in detail; The judgment threshold Set to the mean, that is , .

2. The method for identifying key factors of distribution network operation based on PMU measurement data according to claim 1, characterized in that: The step S1 is specifically as follows: K-means clustering is performed on the distribution network measurement data samples collected by PMU based on the Euclidean distance. The calculation formula of the Euclidean distance is as follows: Where, These are measurement data samples of the same physical quantity at two different PMU measurement points; They correspond to elements; The data to be clustered is normalized and preprocessed as follows: Where, is the normalized data; is the data before normalization; is the range of the PMU corresponding to the data sample.

3. The method for identifying key factors of distribution network operation based on PMU measurement data according to claim 1 or 2, characterized in that: The step S2 is specifically as follows: The number of clusters for K-means clustering , determined by an adaptive method based on iteration, the iteration amount is the number of clusters :Based on clustering quality evaluation index , calculate the number of adjacent clusters and The change rate of evaluation indicators , set the clustering quality evaluation index change rate threshold ,like , then the corresponding The value is the desired number of clusters; Clustering quality evaluation indicators As shown below: Where, It is a cluster Data samples in ; It is a cluster The cluster center of Clustering quality evaluation index change rate As shown below: For the clusters obtained by clustering, the cluster center of each cluster is used as the representative of its typical operating characteristics, which is also the final clustering result.