A Method and System for Power Data Quality Governance Based on Multi-Source Data Fusion and Rule Engine

CN122571398APending Publication Date: 2026-08-14国家电网有限公司客户服务中心
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

这种方式难以应对复杂业务场景下的数据质量问题

Benefits of technology

首先,本发明技术方案实现了数据质量的量化评估,打破了传统技术中质量判定标准不统一的技术偏见,提供了客观的量化依据;其次,权重动态可调机制适配了不同业务场景的差异化需求,提升了模型的场景适配度与鲁棒性;最后,以量化评分为基础的分级治理策略,实现了数据流转的平滑切换与自动恢复,保障了系统运行的稳定高效。。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122571398A_ABST
    Figure CN122571398A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for power data quality governance based on multi-source data fusion and a rule engine, belonging to the field of power data governance technology. The method first collects multi-source power data in real time from equipment, business, external environment, and historical data, and extracts features; then, it performs spatiotemporal alignment and correlation mapping of the data, and identifies anomalies through a multi-dimensional model using statistics, trend consistency verification, and multi-source cross-validation; next, relying on a rule engine, it divides the data into five quality levels (Q0 to Q4) based on a set of rules such as completeness and accuracy, using weighted quantitative scoring; finally, it triggers multi-level progressive governance strategies according to the level, and ensures stable data flow through smooth switching and automatic recovery mechanisms. This invention achieves real-time monitoring and intelligent governance of power data quality, improving the level of intelligent governance and business support capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power data governance technology, specifically to a power data quality governance method and system based on multi-source data fusion and rule engine. Background Technology

[0002] With the deepening of smart grid construction, the amount of data generated in power system operation, equipment monitoring, and user electricity consumption is growing exponentially. This data comes from diverse sources, has complex structures, and is updated frequently, making it a crucial foundation for grid dispatching, equipment maintenance, load forecasting, and other operations.

[0003] However, current power data quality management largely relies on offline cleaning and periodic verification, typically only conducting quality checks after data is stored or after business anomalies are exposed. This reactive approach fails to promptly detect and prevent contaminated data from flowing into downstream systems, leading to biased business decisions and incurring high repair costs. Furthermore, most systems use fixed, single rules to assess data quality. This approach struggles to address data quality issues in complex business scenarios. For example, while a voltage value at a given moment might be within a reasonable range, a significant deviation from historical trends cannot be identified by a single rule. Additionally, current technologies lack multi-source data fusion analysis: they fail to correlate and verify data from different systems. For instance, a sudden load increase on a line, without supporting meteorological or user behavior data, makes it difficult to determine whether it's a genuine load change or a data acquisition anomaly. Moreover, current quality governance strategies often involve fixed processing triggered by fixed rules, lacking differentiated responses to different risk levels, making it difficult to balance data availability with governance effectiveness.

[0004] Therefore, it is necessary to improve existing technologies to enhance the intelligence level of data governance and its business support capabilities. Summary of the Invention

[0005] The purpose of this invention is to provide a power data quality governance method and system based on multi-source data fusion and rule engine. By constructing a closed-loop governance system of perception-fusion-evaluation-decision-governance, it realizes real-time monitoring and intelligent governance of power data quality.

[0006] To achieve the objectives of this invention, the technical solution provided by this invention is as follows: First aspect This invention provides a power data quality governance method based on multi-source data fusion and a rule engine. Includes the following steps: Step S1: Real-time acquisition and feature extraction of multi-source power data; The multi-source power data includes equipment-side data, business-side data, environmental and external data, as well as metadata and historical data; Step S2: Perform spatiotemporal alignment and correlation mapping on multi-source power data, and identify data anomalies through a multi-dimensional anomaly detection model; Step S3: Using the rule engine, the data quality is graded and evaluated according to the preset rule set, and the quality level is divided. Step S4: Trigger corresponding governance strategies based on different quality levels, and ensure stable data flow through smooth switching and recovery mechanisms.

[0007] Second aspect This invention provides a power data quality governance system based on multi-source data fusion and a rule engine, used to execute the power data quality governance method based on multi-source data fusion and a rule engine, comprising the following: The acquisition and extraction unit is used for real-time acquisition and feature extraction of multi-source power data; The multi-source power data includes equipment-side data, business-side data, environmental and external data, as well as metadata and historical data; An anomaly identification unit is used to perform spatiotemporal alignment and correlation mapping of multi-source power data, and to identify data anomalies through a multi-dimensional anomaly detection model. The grading unit is used to evaluate data quality by classifying it according to a set of preset rules, using a rule engine; The strategy unit is used to trigger corresponding governance strategies based on different quality levels, and ensures stable data flow through smooth switching and recovery mechanisms.

[0008] Compared with the prior art, the beneficial effects of the present invention are as follows: First, the technical solution of this invention enables quantitative assessment of data quality, breaking the technical bias of inconsistent quality judgment standards in traditional technologies and providing objective quantitative evidence. Second, the dynamically adjustable weight mechanism adapts to the differentiated needs of different business scenarios, improving the model's scenario adaptability and robustness. Finally, the hierarchical governance strategy based on quantitative scoring enables smooth switching and automatic recovery of data flow, ensuring the stable and efficient operation of the system. Attached Figure Description

[0009] Figure 1 This is a schematic diagram of the power data quality governance method based on multi-source data fusion and rule engine provided in an embodiment of the present invention. Detailed Implementation

[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0011] It should be noted that the acquisition of data and collection of information in this application are legal, compliant, or obtained with the consent of the subject of the data collection.

[0012] The core idea of ​​this invention is to identify abnormal data patterns through multi-source data fusion, dynamically evaluate the data quality level using a rule engine, and trigger corresponding governance strategies based on the level.

[0013] like Figure 1 As shown, this embodiment provides a power data quality governance method based on multi-source data fusion and rule engine, including the following steps: Step S1: Real-time acquisition and feature extraction of multi-source power data; The multi-source power data expands upon the existing data acquisition platform by adding feature dimensions strongly correlated with data quality. These include equipment-side data, business-side data, environmental and external data, and metadata and historical data. Equipment-side data includes real-time measurement data such as voltage, current, power, and frequency uploaded by smart meters, FTUs, and RTUs. Business-side data includes user profiles, electricity usage categories, load characteristics, and power outage records from marketing systems, dispatch systems, and electricity consumption collection systems. Environmental and external data includes meteorological data, holiday information, and geographic information to help determine whether load anomalies are caused by external factors. Metadata and historical data include data source information, acquisition time, transmission link status, and historical quality tags to assess data reliability and trend anomalies.

[0014] Step S2: Perform spatiotemporal alignment and correlation mapping on multi-source power data, and identify data anomalies through a multi-dimensional anomaly detection model; This step is the core of data quality governance. The system uses multi-source data fusion technology to identify the logical consistency and business rationality among the data.

[0015] The spatiotemporal alignment and association mapping is used to align data from different sources according to timestamps, device IDs, and geographical locations to build a unified spatiotemporal data view.

[0016] The multi-dimensional anomaly detection model includes a statistical model, a trend consistency verification model, and a multi-source cross-validation model. The statistical model uses box plots and sliding window mean drift detection to identify outliers. The trend consistency verification model uses historical data to train a regression model and compares the deviation between actual and predicted values ​​to determine anomalies; the anomaly determination formula for the trend consistency verification is: like If so, it is considered abnormal. ; in, This is the sensitivity coefficient. The standard deviation of historical residuals. α As a trend factor, For the rate of change, The measurement value of the device at time t. For predicted values, It's a hyperparameter. This is the traversal index.

[0017] The multi-source cross-validation model determines abnormal data collection or actual business changes by comparing and correlating data from different systems, devices, and environments. For example, if a line experiences a sudden increase in load, and there are no changes on adjacent lines or supporting weather or user behavior data, it is determined to be a data quality issue.

[0018] The consistency index formula for the multi-source cross-validation is: ; like If the data is below the threshold, it is considered inconsistent. in, , These are measurements taken from the same device on different systems at the same time.

[0019] Step S3: Using the rule engine, the data quality is graded and evaluated according to the preset rule set, and the quality level is divided. The anomaly detection results are input into the rule engine, which then performs a graded evaluation of the data quality based on a preset rule set. The rule engine supports dynamic configuration and expansion. The preset rule set includes integrity rules, accuracy rules, consistency rules, timeliness rules, and logical rules; Integrity rules: whether key fields are missing, and whether data points are continuous; Accuracy rules: Whether the value exceeds a reasonable range or contradicts the data from relevant equipment; Consistency rule: Whether the data of the same device is consistent across different systems; Timeliness rules: Has the data expired or been uploaded? Is the delay too large? Logical rules: such as whether it is reasonable for the current to be non-zero when the voltage is zero; Based on the triggering of the rules, the data quality is divided into five levels, as follows: The quality grades are divided into five levels: Q0 Excellent, Q1 Usable, Q2 Questionable, Q3 Low Quality, and Q4 Invalid.

[0020] Q0: High-quality: Data is complete, accurate, consistent, and timely, and can be directly used for core business.

[0021] Q1: Available: There are minor anomalies, but they do not affect the macro trend analysis.

[0022] Q2: Questionable: The data shows obvious anomalies and requires manual review or supplementary verification.

[0023] Q3: Low quality: The data is unreliable and needs to be cleaned or removed, but can still be retained for traceability.

[0024] Q4: Invalid: The data is completely unusable and needs to be discarded or re-collected.

[0025] In this step, a weighted summation model is used to quantitatively assess data quality. The overall data quality score is defined as the weighted sum of the scores for each dimension, as shown in the following formula: ; Q is the comprehensive data quality score, with a value range of [0, 100], which is used to quantify the overall quality level of power data. The higher the score, the better the data quality. It is mapped to five quality levels from Q0 to Q4 through a preset threshold. , , , These are the weighting coefficients for the dimensions of completeness, accuracy, consistency, and timeliness, respectively, all ranging from [0,1], and satisfying the following conditions: + + =1 is a dynamically adjustable parameter that can be flexibly adjusted according to the quality control priority of the business scenario.

[0026] Step S4: Trigger corresponding governance strategies based on different quality levels, and ensure stable data flow through smooth switching and recovery mechanisms.

[0027] The governance strategy is a multi-level, gradual governance strategy, including the following: Q0 response (no action): Data is directly entered into the database for use by business systems.

[0028] Q1 Response (Marking Hint): Add a quality label to the data field to indicate to downstream systems that there is a minor anomaly in the data, but no action is required.

[0029] Q2 Response (Completed and Corrected): Interpolation completion: Missing points are completed using linear interpolation or historical values.

[0030] Model correction: Correcting outliers using data from nearby devices or predictive models.

[0031] Manual review pending: Abnormal data is pushed to the data governance platform for manual review.

[0032] Q3 Response (Blocking and Isolation): Data isolation: Low-quality data is stored in an isolated area and is not used for business calculations, but only for source tracing and analysis.

[0033] Trigger an alarm: Notify maintenance personnel to investigate the data source or data acquisition device.

[0034] Q4 Response (Forced Cleaning): Data discarding: Invalid data is discarded directly to avoid polluting the data lake.

[0035] Source-side re-sampling: Triggers data source retransmission or device self-test process. All governance policies are smoothly switched through a buffer, and the governance policies are automatically lifted after data quality is restored.

[0036] Test Example 1 Scenario 1: A municipal power grid dispatch center monitors the load data of a 10kV line in real time.

[0037] (1) Data acquisition and feature extraction: The system collected the current load value of the line as 5.2MW, the historical average value for the same period was 3.8MW, the load of adjacent lines did not change significantly, and the meteorological data showed no special weather.

[0038] (2) Multi-source fusion anomaly detection: The statistical model detected that this value exceeded the historical average for the same period by 3 times the standard deviation and marked it as an outlier.

[0039] The trend consistency model predicted a value of 3.9MW, with a deviation of 33%, exceeding the threshold.

[0040] Multi-source cross-validation: The marketing system showed no large-scale user startups or shutdowns during this period, and the weather was normal, so it was initially determined to be a data anomaly.

[0041] (3) Rule engine evaluation: Triggered an accuracy anomaly rule.

[0042] The logical consistency exception rule was triggered (sudden load increase without reasonable explanation).

[0043] Overall rating: Q3 (low quality).

[0044] (4) Implementation of governance strategies: The system stores this data point in isolation and does not participate in scheduling calculations.

[0045] An alarm is triggered to notify maintenance personnel to check the data acquisition terminal on this line.

[0046] At the same time, the system automatically calls up data from nearby lines and historical models to generate a correction value for reference.

[0047] (5) Recovery mechanism: If the maintenance personnel find that the communication module of the data acquisition terminal is faulty, the data will be restored to normal after the repair, the system will automatically release the isolation and restore the data to the database.

[0048] Test Example 2 A power company's transformer substation line loss management personnel discovered that the daily line loss rate of a certain substation was abnormally high, reaching 15% (normally it is 3-5%). It is necessary to determine whether this is a real line loss problem or a data quality problem.

[0049] (1) Data acquisition and feature extraction: The system synchronously acquires multi-source data for this area: Power supply: Real-time measurement data from the main meter of the transformer area (SCADA system).

[0050] Electricity sales volume: Data from the marketing system, frozen daily data from user electricity meters.

[0051] Transformer area records include the number of connected users, user types, transformer capacity, and line-transformer relationships.

[0052] Environmental and external data: daily temperature, whether it is a holiday, and whether there is any special weather.

[0053] Historical line loss data: the trend of line loss rate over the past week and month.

[0054] Equipment status: Communication status of the main meter in the distribution area, success rate of user meter data collection, and presence of fault alarms.

[0055] (2) Multi-source fusion anomaly detection: Statistical model: A line loss rate of 15% exceeding the historical mean by ±3 standard deviations is marked as an outlier.

[0056] Trend consistency: Based on historical line loss and the current power supply, the predicted line loss should be around 4%, but the deviation is too large.

[0057] Multi-source cross-validation: Upon checking the time alignment between the main meter data of the distribution area and the user meter data, it was found that some user meter data were not uploaded during the daily freeze period.

[0058] Comparison of file information: There have been no new users or file changes in this area recently.

[0059] Environmental factors: The temperature was normal that day, it was not a holiday, and external causes such as sudden load changes were excluded.

[0060] Equipment status: The communication of the main meter in the distribution area is normal, but the success rate of user meter data collection is 92%, with data missing for 5 users.

[0061] (3) Rule engine evaluation: Integrity rule triggered: Some users' electricity meter data is missing.

[0062] Accuracy rule triggered: Line loss rate exceeds reasonable range.

[0063] Consistency rule triggered: The sum of power supply and power sales (including missing estimates) does not match the distribution area model.

[0064] The rules engine's overall assessment indicates that the abnormal line loss is mainly caused by missing data, not a genuine line loss issue, and the data quality level is rated as Q2 (questionable).

[0065] (4) Implementation of governance strategies: Completeness and Correction: The system automatically retrieves the historical electricity consumption data of missing users for the same period (average electricity consumption on the same day of the previous week) for interpolation and completes the data, and recalculates the electricity sales volume.

[0066] The line loss rate in the background area has been corrected to 4.2%, returning it to the normal range.

[0067] The following is a note indicating that the original data field has a quality label indicating a problem due to missing data collection. The corrected value will be used in the line loss statistics report, and an explanation will be provided.

[0068] Trigger an alarm: Notify maintenance personnel to handle the 5 meters that failed to collect data and record the fault type.

[0069] (5) Recovery Mechanism: After the maintenance personnel repair the meter data collection, subsequent data uploads are normal, the system automatically removes the completion strategy, and the data quality level is restored to Q0. At the same time, the event record is stored in the knowledge base for optimization of the missing value completion model.

[0070] Furthermore, this embodiment provides a power data quality governance system based on multi-source data fusion and a rule engine, used to execute the power data quality governance method based on multi-source data fusion and a rule engine, characterized by including the following: The acquisition and extraction unit is used for real-time acquisition and feature extraction of multi-source power data; The multi-source power data includes equipment-side data, business-side data, environmental and external data, as well as metadata and historical data; An anomaly identification unit is used to perform spatiotemporal alignment and correlation mapping of multi-source power data, and to identify data anomalies through a multi-dimensional anomaly detection model. The grading unit is used to evaluate data quality by classifying it according to a set of preset rules, using a rule engine; The strategy unit is used to trigger corresponding governance strategies based on different quality levels, and ensures stable data flow through smooth switching and recovery mechanisms.

[0071] The device-side data includes real-time measurement data of voltage, current, power, and frequency, as well as rate of change and fluctuation characteristics, uploaded by smart meters, FTUs, and RTUs; the business-side data includes user profiles, electricity usage categories, load characteristics, and power outage records from the marketing system, dispatch system, and electricity consumption data collection system; the environmental and external data includes meteorological data, holiday information, and geographic information; and the metadata and historical data include data source information, collection time, transmission link status, and historical quality tags.

[0072] Finally, it should be noted that the above embodiments are merely illustrative and explanatory of the present invention, and are not intended to limit the present invention to the scope of the described embodiments. Furthermore, those skilled in the art will understand that the present invention is not limited to the above embodiments, and many more variations and modifications can be made based on the teachings of the present invention, all of which fall within the scope of protection claimed by the present invention.

Claims

1. A power data quality governance method based on multi-source data fusion and rule engine. Its features are, Includes the following steps: Step S1: Real-time acquisition and feature extraction of multi-source power data; The multi-source power data includes equipment-side data, business-side data, environmental and external data, as well as metadata and historical data; Step S2: Perform spatiotemporal alignment and correlation mapping on multi-source power data, and identify data anomalies through a multi-dimensional anomaly detection model; Step S3: Using the rule engine, the data quality is graded and evaluated according to the preset rule set, and the quality level is divided. Step S4: Trigger corresponding governance strategies based on different quality levels, and ensure stable data flow through smooth switching and recovery mechanisms.

2. The power data quality governance method based on multi-source data fusion and rule engine according to claim 1, characterized in that, In step S1, the device-side data includes real-time measurement data of voltage, current, power, and frequency uploaded by smart meters, FTUs, and RTUs; the business-side data includes user profiles, electricity consumption categories, load characteristics, and power outage records from the marketing system, dispatch system, and electricity consumption data collection system; the environmental and external data includes meteorological data, holiday information, and geographic information; and the metadata and historical data include data source information, collection time, transmission link status, and historical quality tags.

3. The power data quality governance method based on multi-source data fusion and rule engine according to claim 2, characterized in that, In step S2, the multi-dimensional anomaly detection model includes a statistical model, a trend consistency verification model, and a multi-source cross-validation model. The statistical model uses box plots and sliding window mean drift detection to identify outliers. The trend consistency verification model uses historical data to train a regression model and compares the deviation between actual and predicted values ​​to determine anomalies. The multi-source cross-validation model determines data collection anomalies or real business changes by comparing the correlation of data from different systems, devices, and environments.

4. The power data quality governance method based on multi-source data fusion and rule engine according to claim 3, characterized in that, The anomaly detection formula for the trend consistency check is as follows: like If so, it is considered abnormal. ; in, This is the sensitivity coefficient. The standard deviation of historical residuals. α As a trend factor, For the rate of change, The measurement value of the device at time t. For predicted values, It's a hyperparameter. This is the traversal index.

5. The power data quality governance method based on multi-source data fusion and rule engine according to claim 4, characterized in that, The consistency index formula for the multi-source cross-validation is: ; like If the data is below the threshold, it is considered inconsistent. in, , These are measurements taken from the same device on different systems at the same time.

6. The power data quality governance method based on multi-source data fusion and rule engine according to claim 5, characterized in that, In step S3, the preset rule set includes integrity rules, accuracy rules, consistency rules, timeliness rules, and logical rules; The quality grades are divided into five levels: Q0 Excellent, Q1 Usable, Q2 Questionable, Q3 Low Quality, and Q4 Invalid.

7. The power data quality governance method based on multi-source data fusion and rule engine according to claim 6, characterized in that, In step S3, a weighted summation model is used to quantitatively assess data quality. The overall data quality score is defined as the weighted sum of the scores for each dimension, as shown in the following formula: ; Q is the comprehensive data quality score, with a value range of [0,100], which is used to quantify the overall quality level of power data. The higher the score, the better the data quality. It is mapped to five quality levels from Q0 to Q4 through a preset threshold. , , , These are the weighting coefficients for the dimensions of completeness, accuracy, consistency, and timeliness, respectively, all ranging from [0,1], and satisfying the following conditions: + + =1 is a dynamically adjustable parameter that can be flexibly adjusted according to the quality control priority of the business scenario.

8. The power data quality governance method based on multi-source data fusion and rule engine according to claim 7, characterized in that, In step S4, the governance strategy is a multi-level progressive governance strategy, including the following: Q0 level: data is directly entered into the database; Q1 level: quality label is added as a prompt, no processing is performed; Q2 level: interpolation completion, model correction, and manual review are performed; Q3 level: data is isolated and stored, and operation and maintenance alarms are triggered; Q4 level: data is discarded, and source end re-collection is triggered. All governance policies are switched smoothly through a buffer, and the governance policies are automatically lifted once data quality is restored.

9. A power data quality governance system based on multi-source data fusion and rule engine, used to execute the power data quality governance method based on multi-source data fusion and rule engine as described in any one of claims 1-8, characterized in that, Including the following: The acquisition and extraction unit is used for real-time acquisition and feature extraction of multi-source power data; The multi-source power data includes equipment-side data, business-side data, environmental and external data, as well as metadata and historical data; An anomaly identification unit is used to perform spatiotemporal alignment and correlation mapping of multi-source power data, and to identify data anomalies through a multi-dimensional anomaly detection model. The grading unit is used to evaluate data quality and classify it into quality levels based on a preset set of rules using a rule engine. The strategy unit is used to trigger corresponding governance strategies based on different quality levels, and ensures stable data flow through smooth switching and recovery mechanisms.

10. A power data quality governance system based on multi-source data fusion and rule engine according to claim 9, characterized in that, The device-side data includes real-time measurement data of voltage, current, power, and frequency uploaded by smart meters, FTUs, and RTUs; the business-side data includes user profiles, electricity consumption categories, load characteristics, and power outage records from the marketing system, dispatch system, and electricity consumption data collection system; the environmental and external data includes meteorological data, holiday information, and geographic information; and the metadata and historical data include data source information, collection time, transmission link status, and historical quality tags.