Topological similarity data distribution verification method for power system

By collecting and analyzing operating data under different topological states in the power system, calculating similarity and screening, combined with the streaming data processing framework, the problems of many topological changes and frequent line switching in the power system are solved, and efficient data analysis and abnormal identification are achieved.

CN120067700APending Publication Date: 2025-05-30CHINA SOUTHERN POWER GRID NEW POWER SYSTEM (BEIJING) RESEARCH INSTITUTE CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510083183.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The topology changes in the power system and frequent line switching have led to large data analysis workloads, and the existing technology is difficult to effectively solve this problem.

Method used

A method for verification of topological similarity data distribution in power system is proposed. By collecting operating data under different topological states, calculating the similarity between topological states, filtering out the historical state closest to the current topological state, and performing fast approximation similarity query and data distribution verification. The real-time monitoring system introduces a streaming data processing framework.

Benefits of technology

It improves the accuracy and efficiency of data analysis, reduces the computational complexity, and can quickly identify abnormal distributions or inconsistent operating states, providing decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067700A_ABST
    Figure CN120067700A_ABST
Patent Text Reader

Abstract

The invention discloses a power system topology similarity data distribution verification method. The method comprises the following steps: collecting operation data of a power system; calculating the similarity between the topological states; screening out a plurality of historical topology states closest to the current topology state; fast approximate similarity query is carried out, and similar topological states are fast matched; the power system operation data distribution under different topological states is verified, the influence of topological changes on the system operation data is analyzed, abnormal distribution or inconsistent operation states are identified, and a streaming data processing framework is introduced. According to the method, different similarity measurement methods are combined, accurate analysis can be carried out for different data features, the accuracy of a data analysis result is ensured, meanwhile, the corresponding speed and efficiency of analysis are improved, the locality sensitive hashing algorithm and the K-nearest neighbor algorithm are introduced, global calculation on a whole historical data set is avoided, and the calculation efficiency is improved. And the calculation complexity is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power systems, and particularly relates to a method for verifying the data distribution of power system topology similarity. Background Art

[0002] A power system is a complex network that includes a large number of generating units, substations, transmission lines, and distribution facilities.

[0003] To ensure the stability and economic operation of the system, the power grid needs to frequently adjust the line configuration and topology. These adjustments may be caused by various factors such as load fluctuations, changes in the status of generating units, maintenance requirements, and sudden failures. Moreover, equipment such as transmission lines and substations in the power grid need to be taken out of service for maintenance or repair due to aging, external environment, or sudden failures. To ensure the continuity and stability of power supply, the power grid must frequently switch lines to bypass faulty lines or arrange standby lines.

[0004] Therefore, the current power system faces problems such as frequent topology changes, frequent line switching, and a large amount of data analysis work for different topological states. For this reason, we propose a method for verifying the data distribution of power system topology similarity to solve the above problems. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for verifying the data distribution of power system topology similarity to solve the problems raised in the above background art.

[0006] To achieve the above purpose, the present invention provides the following technical solution: A method for verifying the data distribution of power system topology similarity, including the following steps:

[0007] Step 1: Collect the operation data of the power system under different topological states to form a multi-dimensional data set;

[0008] Step 2: Calculate the similarity between topological states based on the operation data of the power system under different topological states;

[0009] Step 3: Further screen out multiple historical topological states that are closest to the current topological state according to the calculation results of the similarity to optimize the efficiency of similarity analysis;

[0010] Step 4: Perform a fast approximate similarity query in the topological state data distribution to quickly match similar topological states;

[0011] Step 5: Based on the similarity calculation results, verify the operation data distribution of the power system under different topological states, analyze the impact of topological changes on the system operation data, and identify abnormal distributions or inconsistent operation states;

[0012] Step 6: Introduce a streaming data processing framework through a real-time monitoring system to ensure rapid response to verification and analysis when the data changes in real time and provide decision support.

[0013] Preferably, in Step 1, the collected data also needs to be normalized and denoised.

[0014] Preferably, the specific steps of the data normalization process include:

[0015] Select a normalization method, and the normalization methods include min-max normalization and Z-score standardization. The calculation formula for min-max normalization is:

[0016]

[0017] The calculation formula for Z-score standardization is:

[0018]

[0019] Normalize the operation data in different topological states. According to the selected normalization method, normalize the data in different dimensions such as voltage, current, and frequency.

[0020] Preferably, the specific steps of the data denoising process include:

[0021] Noise identification. The types of noise include extreme outliers and random errors generated during the sampling process;

[0022] Select a denoising method. The denoising methods include the moving average method, median filtering method, and wavelet transform denoising. The median filtering method replaces the values in the sliding window with the median. The wavelet transform denoising performs frequency domain analysis on the data through wavelet decomposition, removes the high-frequency noise part, and retains the low-frequency effective signal. The moving average method averages the data through a sliding window to smooth the fluctuations within a short period. Its calculation formula is:

[0023]

[0024] Preferably, the calculation methods for the similarity between topological states in Step 2 include cosine similarity, Jaccard similarity, and Pearson correlation coefficient. The cosine similarity is used to calculate the similarity of topological structure feature vectors and measure the angle between two topological state vectors. The Jaccard similarity is used to calculate the similarity of line switch states under different topological states and analyze the intersection and union of line states. The Pearson correlation coefficient is used to analyze the linear correlation between the operation data of the power system under different topological states.

[0025] Preferably, the specific steps for the cosine similarity to calculate the similarity of topological structure feature vectors include:

[0026] Encode different topological states of the power system and convert each topological state into a feature vector;

[0027] For two topological state vectors A and B, the cosine similarity calculation formula is:

[0028]

[0029] where A·B represents the dot product of the two vectors, and ║A║ and ║B║ represent the Euclidean norms of the vectors respectively;

[0030] The value of the cosine similarity is between 0 and 1, and the similarity value from 0 to 1 indicates that the topological state structures are gradually similar.

[0031] Preferably, the specific steps for calculating the similarity of the topological structure feature vectors by the Jaccard similarity include:

[0032] Encode the line switch states of the power system;

[0033] For two topological states A and B, the Jaccard similarity calculation formula is:

[0034]

[0035] where A∩B is the intersection of the line switch states, representing the number of lines that are both open, and A∪B is the union of the line switch states, representing the number of lines with any one state being open;

[0036] The value of the Jaccard similarity is between 0 and 1, and the similarity value from 0 to 1 indicates that the topological state structures are gradually similar.

[0037] Preferably, the specific steps for analyzing the linear correlation of the operation data in the topological state by the Pearson correlation coefficient include:

[0038] Collect the power operation data under different topological states to form data vectors corresponding to the topological states;

[0039] For the operation data X and Y of two topological states, the Pearson correlation coefficient calculation formula is:

[0040]

[0041] where, X i and Y i represent the operation data corresponding to the topological states X and Y respectively, and are their means respectively;

[0042] The value of the Pearson correlation coefficient ranges from -1 to 1. A value of 1 indicates a perfect positive correlation between two data sets, -1 indicates a perfect negative correlation between two data sets, and 0 indicates no linear correlation between two data sets.

[0043] Preferably, the specific steps of step three based on the K-nearest neighbor algorithm are as follows:

[0044] According to the calculated similarity, all historical topological states are sorted from high to low according to the similarity, and the top K historical topological states with the highest similarity are selected;

[0045] The selected K historical topological states are used as a reference set for similarity analysis, and further detailed data analysis is carried out to make a more in-depth comparison of these selected K states.

[0046] Preferably, the specific steps of step four based on the locality-sensitive hashing method are as follows:

[0047] Select a suitable family of hash functions according to the similarity metric;

[0048] Use multiple hash functions to generate hash tables and map all historical topological states into hash buckets;

[0049] Hash-map the target topological state to find all candidate states that fall into the same hash bucket;

[0050] Calculate the similarity of the topological states in the candidate set and filter out the closest multiple topological states;

[0051] Output the set of topological states of approximate nearest neighbors and conduct further analysis.

[0052] Preferably, the specific steps of step six are as follows:

[0053] System architecture design, select a suitable streaming processing framework, and design the data flow architecture;

[0054] Data collection, access real-time data sources, and format and standardize the data;

[0055] Data flow transmission, set up a message queue, and set data partitioning and replication;

[0056] Real-time data processing, implement the streaming data processing logic and the similarity calculation module;

[0057] Anomaly detection and identification, implement the anomaly detection algorithm and data distribution verification;

[0058] Real-time monitoring and alarm, set up a monitoring dashboard and an alarm mechanism;

[0059] Data storage and retrospective analysis, data storage scheme and historical data integration;

[0060] System maintenance and optimization, performance monitoring and optimization, and regular update of the algorithm model.

[0061] The technical effects and advantages of the present invention:

[0062] By combining different similarity measurement methods, the present invention can perform precise analysis for different data characteristics, ensure the accuracy of data analysis results, improve the response speed and efficiency of analysis, and at the same time introduce the locality-sensitive hashing algorithm and the K-nearest neighbor algorithm, avoiding global calculation on the entire historical data set and reducing the computational complexity. Brief Description of the Drawings

[0063] Figure 1 It is a schematic diagram of the method flow of the present invention. Detailed Embodiments

[0064] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0065] The present invention provides a Figure 1 power system topology similarity data distribution verification method as shown below, including the following steps:

[0066] Step 1: Collect the operation data of the power system under different topological states to form a multi-dimensional data set. The collected data also needs to be normalized and denoised. Through normalization, data with different dimensions is converted into the same dimension range, avoiding excessive influence of certain features on the analysis results due to different orders of magnitude, making the data more suitable for similarity calculation and model training. By denoising the data, noise data can be eliminated to ensure the accuracy and reliability of the data;

[0067] Step 2: Calculate the similarity between topological states based on the operation data of the power system under different topological states;

[0068] Step 3: Based on the K-nearest neighbor algorithm, further screen out multiple historical topological states closest to the current topological state according to the calculation results of the similarity to reduce the amount of calculation and optimize the similarity analysis efficiency;

[0069] Step 4: Perform a fast approximate similarity query in the topological state data distribution to improve the calculation efficiency of high-dimensional data and quickly match similar topological states;

[0070] Step Five: Based on the similarity calculation results, verify the distribution of power system operation data under different topological states, analyze the impact of topological changes on system operation data, and identify abnormal distributions or inconsistent operation states;

[0071] Step Six: Introduce a streaming data processing framework through a real-time monitoring system to ensure rapid response to verification and analysis when data changes in real time and provide decision support.

[0072] This application can perform precise analysis for different data characteristics by combining different similarity measurement methods, ensuring the accuracy of data analysis results. At the same time, it improves the response speed and efficiency of analysis. By introducing the Locality-Sensitive Hashing algorithm and the K-Nearest Neighbor algorithm simultaneously, it avoids global calculations on the entire historical dataset and reduces the computational complexity.

[0073] Among them, the specific steps of data normalization processing include:

[0074] Select a normalization method. The normalization methods include min-max normalization and Z-score standardization. Min-max normalization maps data to the range [0, 1] and is applicable to situations where all features have similar importance requirements. The calculation formula for min-max normalization is:

[0075]

[0076] Z-score standardization normalizes data into a standard normal distribution with a mean of 0 and a standard deviation of 1, which is applicable to situations where the data has a large fluctuation range or different distributions. The calculation formula for Z-score standardization is:

[0077]

[0078] where μ is the mean and σ is the standard deviation;

[0079] Normalize the operation data in different topological states. According to the selected normalization method, normalize the data in different dimensions such as voltage, current, and frequency to ensure that features in different dimensions have similar scales when performing similarity calculations.

[0080] Among them, the specific steps of data denoising processing include:

[0081] Noise identification. The types of noise include extreme outliers and random errors generated during the sampling process. Extreme outliers can be instantaneous voltage and current spikes;

[0082] Select a denoising method. The denoising methods include the moving average method, the median filtering method, and wavelet transform denoising. The median filtering method replaces the values in the sliding window with the median to reduce the influence of outliers and is especially suitable for dealing with discrete outliers. Wavelet transform denoising performs frequency-domain analysis on the data through wavelet decomposition, removes the high-frequency noise part, and retains the low-frequency effective signal. The moving average method averages the data through a sliding window to smooth the fluctuations within a short period. Its calculation formula is:

[0083]

[0084] According to the type of data noise, select an appropriate denoising method. For example, for sudden spike noise, median filtering can be used; for periodic noise, the moving average or wavelet transform can be used.

[0085] Through the normalization and denoising of the data, it can ensure that the operation data under different topological states of the system has good quality and consistency, thus providing reliable input for subsequent similarity calculation and data analysis.

[0086] Among them, the methods for calculating the similarity between topological states in step two include cosine similarity, Jaccard similarity, and Pearson correlation coefficient. Cosine similarity is used to calculate the similarity of topological structure feature vectors and measure the angle between two topological state vectors. Jaccard similarity is used to calculate the similarity of line switch states under different topological states and analyze the intersection and union of line states. Pearson correlation coefficient is used to analyze the linear correlation between the operation data of the power system under different topological states. By using methods such as cosine similarity, Jaccard similarity, and Pearson correlation coefficient to conduct similarity analysis on different topological states, it can quickly evaluate the influence of different topological structures on the operation data of the power system, reduce the analysis workload, and improve the analysis efficiency.

[0087] Among them, the specific steps for calculating the similarity of topological structure feature vectors by cosine similarity include:

[0088] Encode different topological states of the power system and convert each topological state into a feature vector;

[0089] For two topological state vectors A and B, the formula for their cosine similarity is:

[0090]

[0091] Where A·B represents the dot product of the two vectors, and ║A║ and ║B║ represent the Euclidean norms of the vectors respectively;

[0092] The value of cosine similarity ranges from 0 to 1. The similarity value from 0 to 1 indicates that the topological state structures are gradually becoming more similar. The closer the similarity is to 1, the more similar the structures of the two topological states are, and the smaller the angle between the vectors. The closer the similarity is to 0, the greater the difference in the topological states.

[0093] Among them, the specific steps for calculating the similarity of topological structure feature vectors using Jaccard similarity include:

[0094] Encode the line switch states of the power system;

[0095] For two topological states A and B, the Jaccard similarity calculation formula is:

[0096]

[0097] Among them, A∩B is the intersection of the line switch states, indicating the number of lines that are both on, and A∪B is the union of the line switch states, indicating the number of lines that are on in any one state;

[0098] The value of Jaccard similarity ranges from 0 to 1. The similarity value from 0 to 1 indicates that the topological state structures are gradually becoming more similar. The closer the value is to 1, the more similar the line switch states are under the two topological states, and the larger the intersection of the lines; the closer the value is to 0, the greater the difference in the topological states.

[0099] Among them, the specific steps for analyzing the linear correlation of operation data under topological states using Pearson correlation coefficient include:

[0100] Collect the power operation data under different topological states to form data vectors corresponding to the topological states;

[0101] For the operation data X and Y of two topological states, the Pearson correlation coefficient calculation formula is:

[0102]

[0103] Among them, X i and Y i respectively represent the operation data corresponding to topological states X and Y, and are their respective means;

[0104] The value of the Pearson correlation coefficient ranges from -1 to 1. 1 indicates that the two data sets are completely positively correlated, -1 indicates that the two data sets are completely negatively correlated, and 0 indicates that the two data sets have no linear correlation.

[0105] Among them, the K-nearest neighbor algorithm is used to screen out the K historical states that are closest to the current topological state according to the similarity, which greatly reduces the range of the data set that needs to be analyzed in depth, reduces the computational amount, and optimizes the speed and efficiency of similarity analysis. The specific steps of Step 3 based on the K-nearest neighbor algorithm are as follows:

[0106] According to the calculated similarity, all historical topological states are sorted from high to low according to the similarity, and the top K historical topological states with the highest similarity are selected. The value of K represents the number of historical topological states that need to be screened out as the closest to the current topological state. Usually, K is a small positive integer, and common values are 3, 5, 7, etc., which specifically depend on the actual application scenario and data distribution.

[0107] The K historical topological states screened out are used as a reference set for similarity analysis, and further detailed data analysis is carried out. By only analyzing these K closest states, the computational amount of analysis can be significantly reduced, avoiding comprehensive calculation of the entire historical data set, and making a more in-depth comparison of these selected K states, such as calculating the correlation of their operation data and evaluating the operation effect.

[0108] After sorting the historical topological states, if further analysis of these most similar states is needed, a voting mechanism or weighted average method can be used. For example, in the majority voting method, if most of the K nearest neighbors belong to a certain topological feature category, it can be considered that the current topological state is most similar to this type of feature; in the weighted average method, different weights are given according to the similarity size of each neighbor, and weighted average is used to determine the feature or parameter most similar to the target topological state.

[0109] Among them, locality-sensitive hashing is a method that maps high-dimensional data to a low-dimensional space through a hash function, which is used to efficiently perform approximate nearest neighbor queries. In the distribution of power system topological state data, using locality-sensitive hashing can quickly match similar topological states and improve the efficiency of high-dimensional data processing. The specific steps of Step 4 based on the locality-sensitive hashing method are as follows:

[0110] Select a suitable family of hash functions according to the similarity metric.

[0111] Use multiple hash functions to generate hash tables, and map all historical topological states to hash buckets. To improve the accuracy of queries, multiple hash tables are usually generated, and each hash table uses a different hash function for mapping to ensure that similar topological states have a higher matching probability in the same table. According to the selected family of hash functions, generate hash functions h 1 ,h 2 ,...,h k, where each hash function is used to map the high-dimensional vector to a low-dimensional hash bucket. For example, when using cosine similarity, the hash function can divide the data space through a random hyperplane and project the high-dimensional vector onto a low-dimensional binary representation. Map each topological state feature vector to the corresponding bucket through multiple hash functions to construct one or more hash tables;

[0112] Perform a hash mapping on the target topological state to find all candidate states that fall into the same hash bucket. Map the target topological state (i.e., the topological state currently to be queried) vector through the same hash function as when constructing the hash table to obtain its hash bucket position. In the hash table, retrieve all historical topological states that fall into the same bucket according to the hash value of the target topological state. These states are potential approximate nearest neighbors;

[0113] Calculate the similarity of the topological states in the candidate set and filter out the closest multiple topological states;

[0114] Output the set of approximate nearest neighbor topological states and perform further analysis. Output several historical topological states that are most similar to the current topological state. These states can be used for further analysis, such as the scheduling and prediction of power systems. Since locality-sensitive hashing only outputs a candidate nearest neighbor set and reduces the computational amount through the hash table, subsequent precise calculations are only performed on a smaller amount of data, significantly improving the computational efficiency;

[0115] The performance of locality-sensitive hashing is directly related to the number of hash functions. Increasing the number of hash functions

[0116] `

[0117] It can improve the precision rate, but may also increase the time complexity. By reasonably selecting the value of k, the hash table can not only effectively narrow down the candidate set, but also ensure the accuracy of the approximate result. The increase in the number of hash tables can improve the recall rate, and more hash tables can reduce the possibility of missing approximate nearest neighbors. Locality-sensitive hashing can significantly improve the speed of approximate nearest neighbor queries when dealing with high-dimensional data, and is especially suitable for the complex topological state data distribution in power systems. It reduces the query time by using multiple hash tables, making the calculation more efficient when facing large-scale historical data. In a power system with frequent topological changes, locality-sensitive hashing can be used to quickly find the historical state most similar to the current topological state, facilitating historical data analysis and decision support. In the operation of power systems, locality-sensitive hashing can be used to quickly find historical abnormal states similar to the current operating state to assist in early warning and risk assessment. Through the locality-sensitive hashing algorithm, approximate similarity queries can be quickly performed in a high-dimensional space, significantly improving the matching speed of similar topological states, effectively dealing with the complex high-dimensional data characteristics in power systems, and optimizing the query efficiency of large-scale historical data. By introducing locality-sensitive hashing and the k-nearest neighbor algorithm, global calculations on the entire historical data set are avoided, reducing the computational complexity, which is suitable for scenarios with frequent topological changes and a large amount of operating data in power systems

[0118] In step five, the power system operation data distribution under different topological states is verified. Verification means that when the topological state changes, the consistency and rationality of the system operation data are verified. By comparing the operation data distributions under different topological states, the stability and consistency of the data are ensured. The verification steps are as follows:

[0119] Obtain the operation data of different topological states: Collect the operation data (such as voltage, current, power, load, etc.) under the current topological state and historical similar topological states;

[0120] Compare the data distributions under different topological states: Based on the similarity calculation results, select similar historical topological states and compare the corresponding operation data to verify whether there are significant differences in the data distributions under the same or similar topological structures;

[0121] Set a reasonable threshold range: Based on system design and historical data, set a reasonable range of operation data fluctuations (such as the normal fluctuations of voltage and current within a certain interval);

[0122] Detect data deviation: Compare the distribution of the current operation data with that of the historical operation data to check whether the current data exceeds the preset threshold range to ensure the consistency between the current system state and the historical state under similar conditions.

[0123] Among the analyses of the impact of topological changes on system operation data, the analysis steps are mainly used to evaluate the specific impact of different topological changes on system operation data, identify the trends and characteristics of system operation under different topological states, and the analysis steps are as follows:

[0124] Utilize the similarity calculation results: Based on the similar topological states calculated using cosine similarity, Jaccard similarity, Pearson correlation coefficient, etc., compare the similar historical states with the current state;

[0125] Analyze the changes in operation data under different topological states: By comparing the operation data of topological states with higher similarity, analyze the changes in key parameters (such as voltage, current, load, etc.) of the power system under topological changes;

[0126] Find data trends: By comparing the distribution of operation data under different topological states, identify the key characteristics of the power system (such as voltage fluctuations, load change trends, etc.);

[0127] Local sensitive analysis: Combine the K-nearest neighbor algorithm to extract data features from K similar historical topological states, and analyze the similarities and differences of data under different topological states,

[0128] Multidimensional data analysis: Utilize the distribution characteristics of multidimensional data to analyze the impact of topological changes on operation data in each dimension (such as time, geographical location, load, etc.)

[0129] Identification of influencing factors: By analyzing the operation data in different states before and after topological changes, identify the parameters that have a greater impact on system operation (such as the switches of certain specific lines, the impact of load changes on the overall system);

[0130] Operation reliability analysis: Evaluate whether the reliability of the system changes under different topological states, such as whether the stability and security of the system are affected when certain lines are closed or opened.

[0131] Among the identification of abnormal distributions or inconsistent operation states, for identifying abnormal operation states or inconsistent distribution characteristics caused by topological changes through the analysis of operation data, the identification steps are as follows:

[0132] Abnormal pattern recognition: By comparing with the normal operation data in historical similar topological states, identify the abnormal data patterns in the current topological state. For example, if the distribution of current or voltage deviates from the normal distribution of historical data, there may be abnormalities;

[0133] Set abnormal thresholds: Compare historical data and set the abnormal threshold ranges for each key parameter (such as voltage fluctuations, sudden load increases, etc.). If the current data exceeds this threshold during detection, it is regarded as abnormal;

[0134] Historical data inconsistency analysis: By analyzing the historical data of similar topological states, significant differences in the distribution of current operating data and historical data are identified. If there are significant differences between the current data and the historical data under similar topologies, it may indicate that there are abnormalities or potential faults in the system;

[0135] Using statistical methods for distribution analysis: Utilize statistical methods such as standard deviation, skewness, and kurtosis to analyze the data distribution and find abnormal points that deviate from the normal distribution;

[0136] Generating anomaly warnings: When anomalies or inconsistent distributions are detected in the operating data of the power system, warning messages are generated to indicate possible system risks or potential faults;

[0137] Real-time monitoring and adjustment: Feed the detected abnormal distribution back into the control system to adjust the operating strategy of the power system in real time and prevent potential faults from occurring.

[0138] Among them, the specific steps of step six are as follows:

[0139] System architecture design, select a suitable streaming processing framework, choose a suitable streaming processing framework according to requirements, such as Apache Kafka, Apache Flink, Apache Storm, or Spark Streaming, etc. Design the data flow architecture, determine the data source, data flow direction, processing nodes, and data storage scheme to ensure that the entire system can process and analyze data in real time;

[0140] Data collection, access real-time data sources, configure data collection tools, connect to real-time monitoring devices, sensors, SCADA systems, etc. of the power system to ensure that real-time operating data of the power system can be obtained. Data formatting and standardization, format the collected data to unify the data format and structure for subsequent processing;

[0141] Data flow transmission, build a message queue, use a message queue (such as Kafka) to push real-time data to the processing nodes to ensure efficient data transmission between various processing components. Set data partitioning and replication, configure data partitioning and replication to improve the reliability and fault tolerance of data transmission;

[0142] Real-time data processing, implement streaming data processing logic and similarity calculation module, write data processing logic in the stream processing framework, use streaming processing APIs (such as Flink's DataStream API) for data filtering, aggregation, transformation, and calculation, integrate the previously mentioned similarity calculation methods (cosine similarity, Jaccard similarity, Pearson correlation coefficient, etc.), and calculate and analyze the similarity of data in real time;

[0143] Anomaly detection and recognition, implementing anomaly detection algorithms and data distribution verification, integrating anomaly detection algorithms (such as threshold detection, machine learning-based classifiers) into the real-time processing logic, monitoring anomalies in the data stream in real time, verifying the distribution of power system operation data through real-time analysis, and identifying abnormal distributions or inconsistent operating states;

[0144] Real-time monitoring and alarming, setting up a monitoring dashboard and an alarming mechanism, using data visualization tools (such as Grafana, Kibana, etc.) to build a real-time monitoring dashboard, displaying the key operating parameters and abnormal states of the power system in real time, configuring an alarm system, and automatically triggering an alarm to notify relevant personnel (through emails, text messages, application notifications, etc.) when an abnormal state or inconsistent distribution is detected;

[0145] Data storage and retrospective analysis, data storage solutions and historical data integration, storing the processed real-time data in a database (such as the time series database InfluxDB, the NoSQL database MongoDB, etc.) for subsequent analysis and retrospective, regularly combining the real-time processing results with historical data for comprehensive analysis and modeling to improve the intelligence level of the system;

[0146] System maintenance and optimization, performance monitoring and optimization, and regular update of the algorithm model, monitoring the performance metrics of the stream processing system in real time, such as latency, throughput, etc., regularly optimizing the processing logic and resource configuration to ensure the efficiency and stability of the system, and regularly updating the anomaly detection and analysis algorithms according to the changes in real-time data to improve the accuracy and adaptability of the model.

[0147] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for verifying the distribution of power system topology similarity data, characterized in that: The following steps are involved: Step 1: Collect the operation data of the power system under different topological states to form a multidimensional data set; Step 2: Based on the power system operation data under different topological states, calculate the similarity between the topological states; Step 3: Based on the similarity calculation results, further select multiple historical topological states that are closest to the current topological state to optimize the similarity analysis efficiency; Step 4: Perform a fast approximate similarity query in the topological state data distribution to quickly match similar topological states; Step 5: Based on the similarity calculation results, the distribution of power system operation data under different topological states is verified, the impact of topological changes on system operation data is analyzed, and abnormal distribution or inconsistent operation states are identified; Step 6: Introduce a streaming data processing framework through a real-time monitoring system to ensure rapid response of verification analysis and provide decision support when data changes in real time.

2. A method for verifying the distribution of power system topology similarity data according to claim 1, characterized in that: In the step 1, the collected data needs to be normalized and denoised; The specific steps of the data normalization process include: Select a normalization method, which includes minimum-maximum normalization and Z-score normalization. The calculation formula of the minimum-maximum normalization is: The calculation formula for the Z-score standardization is: The operating data of different topological states are normalized, and the data of different dimensions such as voltage, current and frequency are normalized according to the selected normalization method.

3. A method for verifying the distribution of power system topology similarity data according to claim 2, characterized in that: The specific steps of the data denoising process include: Noise identification, where the types of noise include extreme outliers and random errors generated during the sampling process; Select a denoising method, which includes moving average method, median filtering method and wavelet transform denoising. The median filtering method replaces the value of the sliding window with the median value. The wavelet transform denoising performs frequency domain analysis on the data through wavelet decomposition to remove the high-frequency noise part and retain the low-frequency effective signal. The moving average method averages the data through a sliding window to smooth the fluctuations in a short period of time. The calculation formula is:

4. A method for verifying the distribution of power system topology similarity data according to claim 1, characterized in that: The calculation method for the similarity between topological states in step 2 includes cosine similarity, Jaccard similarity and Pearson correlation coefficient. The cosine similarity is used to calculate the similarity of topological structure feature vectors and measure the angle between two topological state vectors. The Jaccard similarity is used to calculate the similarity of line switch states under different topological states and analyze the intersection and union of line states. The Pearson correlation coefficient is used to analyze the linear correlation between power system operation data under different topological states.

5. A method for verifying the distribution of power system topology similarity data according to claim 4, characterized in that: The specific steps of calculating the similarity of the topological structure feature vectors by the cosine similarity include: Encode different topological states of the power system and convert each topological state into a feature vector; For two topological state vectors A and B, the cosine similarity calculation formula is: Where A·B represents the dot product of two vectors, ║A║ and ║B║ represent the Euclidean norm of the vectors respectively; The value of cosine similarity is between 0 and 1. The similarity value from 0 to 1 indicates that the topological state structure is gradually similar.

6. A method for verifying the distribution of power system topology similarity data according to claim 4, characterized in that: The specific steps of calculating the similarity of the topological structure feature vectors by the Jaccard similarity include: Encode the circuit breaker status of the power system; For two topological states A and B, the Jaccard similarity calculation formula is: Where A∩B is the intersection of the line switch states, indicating the number of lines that are both open, and A∪B is the union of the line switch states, indicating the number of lines that are open in either state; The value of Jaccard similarity is between 0 and 1, and the similarity value from 0 to 1 indicates that the topological state structure is gradually similar.

7. A method for verifying the distribution of power system topology similarity data according to claim 4, characterized in that: The specific steps of analyzing the linear correlation of the running data in the topological state by the Pearson correlation coefficient include: Collect power operation data under different topological states to form data vectors corresponding to the topological states; For the running data X and Y of two topological states, the Pearson correlation coefficient calculation formula is: Among them, X i and Y i Respectively represent the running data corresponding to the topological states X and Y, and are their means respectively; The value of the Pearson correlation coefficient is between -1 and 1, where 1 means that the two data sets are completely positively correlated, -1 means that the two data sets are completely negatively correlated, and 0 means that the two data sets have no linear correlation.

8. A method for verifying the distribution of power system topology similarity data according to claim 1, characterized in that: The specific steps of step 3 based on the K nearest neighbor algorithm are: According to the calculated similarity, all historical topological states are sorted from high to low according to the similarity, and the top K historical topological states with the highest similarity are selected; The K selected historical topological states are used as a reference set for similarity analysis, and further detailed data analysis is performed to make a deeper comparison of these K selected states.

9. A method for verifying the distribution of power system topology similarity data according to claim 1, characterized in that: The specific steps of step 4 based on the local sensitive hashing method are: Choose a suitable family of hash functions based on similarity metrics; Use multiple hash functions to generate a hash table and map all historical topological states into hash buckets; Hash map the target topology state to find all candidate states that fall into the same hash bucket; Calculate the similarity of the topological states of the candidate set and select the closest multiple topological states; Output a collection of topological states of approximate nearest neighbors for further analysis.

10. A method for verifying the distribution of power system topology similarity data according to claim 1, characterized in that: The specific steps of step six are: System architecture design, select appropriate stream processing framework, and design data stream architecture; Data collection, real-time data source access, data formatting and standardization; Data stream transmission, building message queues, setting up data partitioning and replication; Real-time data processing, implementing streaming data processing logic and similarity calculation module; Anomaly detection and identification, implementation of anomaly detection algorithms and data distribution verification; Real-time monitoring and alarm, setting up monitoring dashboard and alarm mechanism; Data storage and retrospective analysis, data storage solutions and historical data integration; System maintenance and optimization, performance monitoring and optimization, and regular updates of algorithm models.

Citation Information

Cited By

  • Electric power topology rapid identification method based on depth map matching

    CN120386985A

  • A fast identification method of power topology based on deep graph matching

    CN120386985B