A statistical method for power quality data integrity
By statistically analyzing and remedying the data integrity of power quality monitoring points and areas, and combining it with a tree-shaped power grid topology, the problems of large data volume and insufficient data integrity in power quality data monitoring have been solved, achieving efficient and safe data supplementation and analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ZHONGNENG NEW ELECTRIC TECHNOLOGY CO LTD
- Filing Date
- 2024-08-22
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies for power quality data monitoring suffer from problems such as insufficient storage space due to large data volume, low query efficiency, difficulty in handling abnormal data, and insufficient data integrity. Furthermore, existing gap filling methods have low accuracy, high computational resource requirements, and security risks.
By statistically analyzing the daily and monthly probabilities of data integrity at monitoring points and monitoring areas, and combining the multi-level regional forms of the tree-shaped power grid topology to conduct regional daily and monthly statistics, and using a weighted summation method to supplement the data, we ensure data integrity and locate and remedy the causes of missing data.
It enables intuitive and detailed analysis of power quality data, improves the representativeness and completeness of the data, reduces the need for network connectivity, and enhances security and the accuracy of data supplementation.
Smart Images

Figure CN119003971B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power quality data acquisition and analysis technology, and specifically to a statistical method for power quality data integrity. Background Technology
[0002] As a system for monitoring power quality data, the power quality monitoring terminal can effectively analyze the power grid topology within a region. During the analysis process, one characteristic is the large amount of data. Large data volumes can lead to a series of problems such as insufficient storage space, low query efficiency, and difficulty in handling abnormal data. Another important issue is the integrity of power quality data.
[0003] Power quality data is the foundation for various subsequent power quality analyses. Incomplete data will lead to incomplete and unreliable power quality analysis results. In practical applications, power quality monitoring systems are prone to data loss due to their wide coverage, complex network structure, and complex data interfaces. Data integrity has become a very important factor affecting the normal operation of power quality monitoring systems.
[0004] Currently, there are two main methods for supplementing missing power quality data: First, using big data analysis to obtain power quality data and fill in missing values; second, directly reading the missing power quality data from the power quality monitoring terminal and performing five complete data supplementation operations.
[0005] However, the measures currently being taken have at least three problems:
[0006] 1) Simply adding missing power quality data five times results in low accuracy and poor performance;
[0007] 2) Big data computing requires excessive computing resources, and it often requires network connectivity, which may raise security and confidentiality issues.
[0008] 3) The power quality data is not intuitive enough, making it difficult to conduct targeted analysis or processing. Summary of the Invention
[0009] To address the three issues mentioned above, the purpose of this invention is to propose a statistical method for power quality data integrity. By statistically analyzing the daily and monthly probabilities of data integrity at monitoring points and monitoring areas, intuitive power quality data can be effectively obtained, facilitating refined analysis down to each monitoring point on each day. Simultaneously, a weighted summation method combined with a multi-level regional structure based on a tree-like power grid topology is used for regional daily and monthly statistics, making the power quality data more representative and complete. Furthermore, corresponding data supplementation is performed to address the causes of missing data integrity, effectively mitigating the adverse effects of data integrity gaps.
[0010] This was achieved through the following technical solutions:
[0011] A statistical method for power quality data integrity is provided for power quality processing at each monitoring point and in each monitoring area of the power grid. Each monitoring area includes multiple monitoring points. The method includes the following steps:
[0012] S1. Perform daily and monthly statistics for monitoring points: Read the source data of each monitoring point in the power grid from the power quality acquisition terminal, and based on each source data, perform the following calculations for each monitoring point according to the formula. Daily statistics on data integrity are performed to obtain the daily probability of data integrity for each monitoring point. Here, ID represents the daily probability of data integrity for the monitoring point, VD represents the actual amount of data transmitted by the monitoring point on that day, and OD represents the amount of data that the monitoring point should transmit on that day, calculated over a statistical time interval of 1 minute or 3 minutes. When the daily statistics for monitoring points meet the statistical requirements for one month, the formula is used... Monthly statistics on data integrity are performed for each monitoring point to obtain the monthly probability of data integrity for each monitoring point; where IM represents the monthly probability of data integrity for the monitoring point, m represents the number of days, and ID... m This represents the ID for day m, where M represents the total number of days in a month;
[0013] S2. Perform daily and monthly regional statistics using a tree-like power grid topology: Set up multiple hierarchical regions in the form of a tree topology, with each hierarchical region forming a parent-child relationship from the highest to the lowest level. In any lowest-level region, use a weighted average method with the corresponding ID as the first data item to calculate the daily probability of regional data integrity for the lowest-level region. Simultaneously, use a weighted average method with the corresponding IM as the second data item to calculate the monthly probability of regional data integrity for the lowest-level region. Each monitoring region in step S1 is listed as a lowest-level region and calculated sequentially according to the parent-child relationship. The sequential calculation includes: the daily probability of data integrity for each parent region is obtained by weighted averaging the daily probabilities of data integrity for the corresponding child regions, and the monthly probability of data integrity for each parent region is obtained by weighted averaging the monthly probabilities of data integrity for the corresponding child regions. Each hierarchical region is set with a corresponding data integrity threshold. Based on the above threshold, the daily and monthly probabilities of regional data integrity for each hierarchical region are judged to determine each hierarchical region that needs data supplementation.
[0014] By statistically analyzing the daily and monthly probabilities of data integrity at monitoring points and in monitoring areas, intuitive data on power quality can be effectively obtained, facilitating refined analysis down to each monitoring point on each day. Furthermore, by employing a weighted summation method combined with a multi-level regional structure based on a tree-like power grid topology, regional daily and monthly statistics are performed, making the power quality data more representative and complete. In addition, the calculation process does not require a network connection, effectively enhancing security and confidentiality.
[0015] Preferably, the multiple hierarchical regions, from the highest to the lowest level, include Level 1, Level 2, Level 3, and Level 4 regions. The daily and monthly probabilities of data integrity for each monitoring region are calculated starting from Level 4. Calculating data integrity layer by layer for different regions provides crucial information for locating missing data, data replenishment, and equipment failure.
[0016] Preferably, in step S2, the thresholds set for each hierarchical region can be the same or different. When each threshold is different, the thresholds are decreased in order from the lowest level region to the highest level region. Setting the thresholds to be the same facilitates processing, while setting different thresholds allows for targeted processing for different stages and needs.
[0017] Preferably, in step S2, after identifying each tiered region that requires supplementary data recruitment, the reason for the need for supplementary data recruitment in each tiered region is recorded; wherein, the reasons include at least cached data loss, network congestion delay, and equipment failure. Recording the reasons for supplementary data recruitment facilitates subsequent processing by staff, thereby mitigating the adverse effects caused by data loss.
[0018] Preferably, the reason for data supplementation is that when cached data is missing, the location of the missing cached data is determined, and the corresponding monitoring point is located based on the location of the missing cached data. The five sets of data before the monitoring point was missing and the five sets of data that were normally submitted after the missing data are weighted and averaged to obtain the missing data. The missing data is then resubmitted for statistics. By weighting and averaging the data before and after the monitoring point is missing, the obtained data is very close to the data at the actual missing location. Resubmitting this data effectively improves the accuracy of data supplementation.
[0019] Preferably, when data loss is caused by network congestion, the transmission bandwidth is increased. Alternatively, after determining the location of the missing data due to network congestion, the corresponding monitoring point is located based on the location of the missing data. The five sets of data before the missing data point and the five sets of data that were normally uploaded after the missing data point are weighted and averaged to obtain the missing data. The missing data is then re-uploaded for statistical analysis. Supplementing the missing data can reduce errors caused by network congestion.
[0020] The beneficial effects of this invention compared to the prior art are:
[0021] The technical solution of this invention effectively obtains intuitive data on power quality by statistically analyzing the daily and monthly probabilities of data integrity at monitoring points and monitoring areas. This also facilitates refined analysis, allowing for precision down to each monitoring point on each day. Furthermore, by using a weighted summation method combined with a multi-level regional structure based on a tree-like power grid topology for daily and monthly regional statistics, the power quality data becomes more representative and complete. In addition, corresponding data supplementation is performed to address the causes of missing data integrity, effectively mitigating the adverse effects of such data loss. Attached Figure Description
[0022] Figure 1 A flowchart for a statistical method for power quality data integrity;
[0023] Figure 2 This diagram illustrates the results of power quality statistics performed by users in Jiangsu Province, based on statistical methods for ensuring the integrity of power quality data. Detailed Implementation
[0024] The technical solutions of the present invention will now be described in detail with reference to the accompanying drawings. Example
[0025] like Figure 1 The diagram shows a flowchart of a statistical method for power quality data integrity. This method is used to analyze and process power quality at each monitoring point and in each monitoring area of the power grid. Each monitoring area includes multiple monitoring points. At each monitoring point, relevant power quality data can be obtained using devices such as power quality acquisition terminals. The power quality acquisition terminal can be the PQS2000 power quality fusion terminal, which is a power quality analyzer that can effectively acquire power quality data at the monitoring points.
[0026] The method specifically includes the following steps:
[0027] S1. Conduct daily and monthly statistics for monitoring points:
[0028] The system reads source data from each monitoring point in the power grid using a power quality acquisition terminal. This source data includes parameters such as current, voltage, amplitude-frequency characteristics, harmonics, three-phase imbalance, and power quality events. Based on each source data point, the system applies a formula to each monitoring point. Daily statistics of data integrity can obtain the daily probability of data integrity for each monitoring point; where ID represents the daily probability of data integrity of the monitoring point data, VD represents the actual number of data sent by the monitoring points put into operation on the current day, and "sent" means transmitted, OD represents the number of data that should be sent when calculating according to the statistical time interval for the monitoring points put into operation on the current day, and the statistical time interval is 1 minute or 3 minutes. The statistical time interval is set to 1 minute or 3 minutes to provide fine data sampling and facilitate improving the accuracy of analysis. In actual calculation, the daily probability of data integrity for each monitoring point of the previous day can be statistically calculated before 1 o'clock every day.
[0029] When the daily statistics of the monitoring points meet the statistical quantity for one month, at this time, the monthly probability of data integrity for each monitoring point can be obtained. According to the formula Monthly statistics of data integrity are carried out for each monitoring point to obtain the monthly probability of data integrity for each monitoring point; where IM represents the monthly probability of data integrity of the monitoring point, m represents the number of days, ID m represents the ID of the mth day, and M represents the total number of days in each month. In actual calculation, the monthly probability of data integrity for each monitoring point of the previous month can be statistically calculated on the 1st day of each month.
[0030] S2. Conduct regional daily statistics and regional monthly statistics in the form of a tree-shaped power grid topology structure:
[0031] Set multiple hierarchical regions in the form of a tree topology. Between each hierarchical region, a parent-child relationship is formed in sequence from the highest-level region to the lowest-level region. Taking four hierarchical regions as an example, it includes a first-level region, a second-level region, a third-level region, and a fourth-level region. The first-level region is the parent of the second-level region, the second-level region is the parent of the third-level region, and the third-level region is the parent of the fourth-level region. Calculate the daily probability and monthly probability of regional data integrity starting from the fourth-level region for each monitoring region. This tree topology structure facilitates calculating the data integrity of different regions level by level, and further provides a very important basis for data missing location, data supplement, and equipment fault location, quickly finding the area where errors occur, and then conducting targeted analysis. It should be noted that the multiple hierarchical regions are at least three hierarchical regions, which can be four, or more, and the specific number is determined by the actual situation.
[0032] For the above four hierarchical regions, there may be n first-level regions, m second-level regions, i third-level regions, and j fourth-level regions. n, m, i, and j are all positive integers and n < m < i < j. For example: the first-level region represents the provincial region, with a total of 10. Each first-level region includes at least one second-level region; the second-level region represents the municipal region, with a total of 20. Each second-level region includes at least one third-level region; the third-level region represents the town level, with a total of 40. Each third-level region includes at least one fourth-level region; the fourth-level region represents the village level, with a total of 80.
[0033] In this embodiment, in any lowest-level region, a weighted average method is used, with the corresponding ID as the first data item, to calculate the daily probability of regional data integrity for that region. Simultaneously, a weighted average method is used, with the corresponding IM as the second data item, to calculate the monthly probability of regional data integrity for that region. Based on the ID of each monitoring point in step S1, a weighted average method is applied according to the formula... Calculate the daily probability of complete regional data for each monitoring area, where KD represents the daily probability of complete regional data for the monitoring area, and d represents the number of monitoring points in the monitoring area.
[0034] Taking the four graded regions mentioned above as an example, each monitoring region in step S1 is listed as a fourth-level region and calculated sequentially according to the parent-child relationship. The sequential calculation includes: the daily probability of data integrity of each parent region is obtained by weighted averaging of the daily probability of data integrity of the corresponding child region, and the monthly probability of data integrity of each parent region is obtained by weighted averaging of the monthly probability of data integrity of the corresponding child region.
[0035] Each tiered region has a corresponding data integrity threshold. Based on these thresholds, the daily and monthly probabilities of data integrity in each tiered region are used to determine which tiered regions require supplementary data collection. For example, a uniform threshold of 60% can be set. If the integrity level is not lower than 60%, the data integrity is within the actual requirement range; if it is lower than 60%, supplementary data collection is required. Alternatively, when setting different thresholds, each threshold can be decreased sequentially from the lowest to the highest tiered region.
[0036] It should be noted that when data integrity falls below the corresponding threshold and is considered unqualified, data resubmission is a necessary measure. Data integrity refers to the accurate, complete, and consistent state of data throughout its entire lifecycle. This is a key factor in ensuring data quality and reliability, and is crucial for actual users and manufacturers because it directly impacts decision-making, business operations, and customer trust. In steps S1 and S2, when calculating the daily, monthly, and regional daily and monthly probabilities of data integrity for monitoring points and regions, if data is found to be below the threshold, immediate data resubmission is required, re-uploading the erroneous data from the corresponding monitoring point or region.
[0037] In this embodiment, in step S2, after identifying each graded region that requires data supplementation, the reason for the need for data supplementation in each graded region is recorded. Here, since the power quality acquisition terminal used generally comes with its own system logs and monitoring data, the reason for the need for data supplementation can be determined based on this information, which is also the reason when the probability of insufficient data integrity is low. These reasons may include missing cached data, network congestion and delays, electromagnetic interference, and equipment failure.
[0038] When the cause of data loss is missing cached data, determine the location of the missing cached data, locate the corresponding monitoring point based on the location of the missing cached data, and calculate the weighted average of the 5 sets of data before the missing data and the 5 sets of data that were normally uploaded after the missing data to obtain the missing data. Then, re-upload the missing data for statistical analysis. If the cause is network congestion and delay, the same methods used for missing cached data can be employed, or other methods can be used: slowing down the time between each data reception or improving the bandwidth of the original transmission link to ensure complete data transmission. If the cause is electromagnetic interference, standard electromagnetic interference protection coatings can be added to locations requiring electromagnetic interference protection, depending on the actual equipment requirements. If the cause is equipment failure, simply notify maintenance personnel for inspection and repair.
[0039] like Figure 2 The image shows the results of power quality statistics in Jiangsu Province, based on a statistical method for assessing power quality data integrity. Jiangsu's power quality integrity score is 79.37%, derived from a weighted average of data from 13 cities in Jiangsu. Due to the large amount of data, only partial data from Nantong and Kunshan are shown here as representatives. The statistics date is September 13, 2023, at 2:10 PM, representing the daily probability of power quality integrity for the previous day. The threshold for the daily probability of power quality integrity is 60%; a score above 60% is considered acceptable, while a score below 60% is considered unacceptable.
[0040] Jiangsu Province is designated as a Level 1 region, Nantong and Kunshan as Level 2 regions, and Nantong Maoyou Timber, Nantong Yingnei IoT Technology, and Nantong Xiangyu Marine Equipment Co., Ltd. (all within Nantong) and Kunshan Semiconductor Co., Ltd. (within Kunshan) are all classified as Level 3 regions. Nantong Maoyou Timber monitors multiple points across four Level 4 regions: load output NPQS1, power supply line #2, power supply line #1, and load input NPQS2. Nantong Yingnei IoT Technology monitors multiple points within the main power supply cabinet (a Level 4 region). Nantong Xiangyu Marine Equipment Co., Ltd. monitors three Level 4 regions: integrated transformer #2 distribution line, integrated transformer #1 distribution line, and No. 1 shipyard transformer line. Kunshan Semiconductor Co., Ltd. monitors based solely on the main power supply line (a Level 4 region).
[0041] In the third-level area of Kunshan Semiconductor Co., Ltd., power quality monitoring was conducted at multiple monitoring points along the incoming line. The network parameters were 00-B7-8D-00-B9-58, the power grid provider was NARI Group Corporation, and the monitoring terminal equipment used at the monitoring points was npqs_571w. Communication was normal. The npqs_571w acquired source data from each monitoring point and calculated the daily probability of power quality integrity for each monitoring point, then performed a weighted average. The daily probability of power quality integrity for the incoming line on September 12, 2023, was found to be 99.15%. Therefore, the daily probability of power quality integrity for Kunshan Semiconductor Co., Ltd., and the daily probability of power quality integrity for Kunshan City were also found to be 99.15%.
[0042] Then, following the same method, the daily probabilities of power quality integrity for Nantong Maoyou Timber, Nantong Yingnei Internet of Things Technology, and Nantong Xiangyu Marine Equipment Co., Ltd. were calculated to be 36.96%, 0.00%, and 98.66%, respectively. After weighted summation, the daily probability of power quality integrity for Nantong City was obtained as 45.21%.
[0043] Next, the daily probability of power quality integrity for the other 11 cities was calculated using the same method. The daily probability of power quality integrity for the 13 cities was then weighted and averaged to obtain the daily probability of power quality integrity for Jiangsu as 79.37%.
[0044] In summary, this invention effectively obtains intuitive power quality data by statistically analyzing the daily and monthly probabilities of data integrity at monitoring points and monitoring areas. This also facilitates refined analysis, allowing for precision down to each monitoring point on each day. Furthermore, by employing a weighted summation method combined with a multi-level regional structure based on a tree-like power grid topology for daily and monthly regional statistics, the power quality data becomes more representative and complete. In addition, by addressing the causes of missing data, corresponding data supplementation effectively mitigates the adverse effects of data integrity deficiencies, demonstrating significant advancements in this field.
[0045] The above embodiments are merely illustrative of the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solutions based on the technical concept proposed in this invention shall fall within the scope of protection of this invention.
Claims
1. A statistical method for power quality data integrity for power quality processing of each monitoring point and each monitoring zone in a power grid, each monitoring zone comprising a plurality of monitoring points, characterized in that, The method includes the following steps: S1. Conduct daily and monthly statistics for monitoring points: The power quality acquisition terminal reads each source data of each monitoring point in the power grid, and based on each source data, the data integrity daily statistics of each monitoring point are carried out according to the formula to obtain the data integrity daily probability of each monitoring point; wherein ID represents the data integrity daily probability of the monitoring point, VD represents the actual data quantity of the monitoring point sent on the day, and OD represents the data quantity of the monitoring point that should be sent in the statistical time interval, and the statistical time interval is 1 minute or 3 minutes. When the daily statistics of monitoring points meet the statistical requirements for one month, the formula is used. Monthly statistics on data integrity are performed for each monitoring point to obtain the monthly probability of data integrity for each monitoring point; where IM represents the monthly probability of data integrity for the monitoring point, m represents the number of days, and ID... m This represents the ID for day m, where M represents the total number of days in a month; S2. Performing daily and monthly regional statistics using a tree-shaped power grid topology: Multiple hierarchical regions are set up in the form of a tree topology, and each hierarchical region forms a parent-child relationship from the highest level region to the lowest level region. In any lowest level region, the daily probability of regional data integrity of the lowest level region is calculated by using the corresponding ID as the first data item using the weighted average method. At the same time, the monthly probability of regional data integrity of the lowest level region is calculated by using the corresponding IM as the second data item using the weighted average method. In step S1, each monitoring area is listed as the lowest level area and calculated sequentially according to the parent-child relationship. The sequential calculation includes: the daily probability of data integrity of each parent area is obtained by weighted average of the daily probability of data integrity of the corresponding child area, and the monthly probability of data integrity of each parent area is obtained by weighted average of the monthly probability of data integrity of the corresponding child area. Each graded region has a corresponding data integrity threshold. Based on the threshold, the daily and monthly probabilities of regional data integrity in each graded region are used to determine the graded regions that need to be supplemented with data.
2. The statistical method for power quality data integrity according to claim 1, characterized in that, Multiple hierarchical regions, from the highest to the lowest level, include Level 1, Level 2, Level 3, and Level 4 regions. The daily and monthly probabilities of regional data integrity are calculated for each monitoring region, starting from Level 4.
3. The statistical method for power quality data integrity according to claim 1, characterized in that, In step S2, the thresholds set for each graded region may be the same or different. When each threshold is set differently, each threshold is decreased in order from the lowest graded region to the highest graded region.
4. The statistical method for power quality data integrity according to claim 1, characterized in that, In step S2, after identifying each graded region that needs to supplement data recruitment, the reason for supplementing data recruitment in each graded region is recorded; wherein, the reason includes at least cached data loss, network congestion delay and equipment failure.
5. The statistical method for power quality data integrity according to claim 4, characterized in that, The reason for data replenishment is that when cached data is missing, the location of the missing cached data is determined, the corresponding monitoring point is located based on the location of the missing cached data, and the five sets of data before the monitoring point was missing and the five sets of data that were normally submitted after the missing data were weighted and averaged to obtain the missing data. The missing data is then resubmitted for statistics.
6. The statistical method for power quality data integrity according to claim 4, characterized in that, The reason for data supplementation is that when there is network congestion and delay, the bandwidth during transmission is increased, or after determining the location of the data loss caused by network congestion and delay, the corresponding monitoring point is located based on the location of the data loss. The five sets of data before the monitoring point was lost and the five sets of data that were normally sent after the loss are weighted and averaged to obtain the missing data. The missing data is then resent for statistics.