Multi-source data quality evaluation method and related device for power supply reliability analysis
Patent Information
- Application Number
- CN202611043042.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-25
AI Technical Summary
然而,这类评价体系主要面向通用数据库或综合信息系统场景,其设计目标与供电可靠性分析的实际应用场景存在明显的不匹配,难以刻画数据质量问题对系统平均停电持续时间指数(SAIDI)、系统平均停电次数指数(SAIFI)等可靠性指标计算结果的传导影响,使得现有数据质量评价方法无法针对性地体现供电可靠性分析对数据质量的特殊要求,从而造成了目前供电可靠性分析准确度不稳定的技术问题
本申请的方案通过构建可靠性影响敏感度评分和过程可追溯性评分等供电可靠性分析特有维度的评估指标,直接量化数据质量偏差对供电可靠性指标计算结果的传导影响,有效解决了通用数据质量评价体系与业务场景脱节的问题,具有能够针对性评估多源数据质量对供电可靠性分析结果的影响程度,提升数据质量评估与供电可靠性分析业务场景的适配性,从而保障分析结果可靠性和一致性的优点。
Smart Images

Figure CN122820002A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power system supply reliability assessment technology, and in particular to a multi-source data quality assessment method and related apparatus for power supply reliability analysis. Background Technology
[0002] With the continuous construction of distribution automation, smart meters, feeder terminals and online monitoring devices, the data sources on which power supply reliability analysis depends are constantly increasing, and the amount and types of data are growing rapidly. This has led to a shift in power supply reliability assessment from the traditional assessment model that relies on structured ledgers and manual power outage records to a comprehensive assessment model based on multi-source data.
[0003] In the aforementioned comprehensive evaluation model, ensuring the quality of multi-source data is fundamental to the analysis. Current power supply reliability analysis typically uses general dimensions such as accuracy, completeness, and consistency to evaluate data quality. However, these evaluation systems are primarily geared towards general databases or integrated information systems. Their design goals are significantly mismatched with the actual application scenarios of power supply reliability analysis. They struggle to characterize the impact of data quality issues on the calculation results of reliability indicators such as the System Average Outage Duration Index (SAIDI) and the System Average Outage Frequency Index (SAIFI). This makes existing data quality evaluation methods unable to specifically reflect the unique data quality requirements of power supply reliability analysis, resulting in the technical problem of unstable accuracy in current power supply reliability analysis. Summary of the Invention
[0004] This application provides a multi-source data quality assessment method and related equipment for power supply reliability analysis, which is used to specifically assess the impact of multi-source data quality on power supply reliability analysis results, improve the adaptability of data quality assessment to power supply reliability analysis business scenarios, and thus ensure the reliability and consistency of analysis results.
[0005] To achieve the above-mentioned objectives, the first aspect of this application provides a multi-source data quality assessment method for power supply reliability analysis, comprising: Acquire multi-source data to be evaluated, wherein the multi-source data includes: equipment ledger data, operation measurement and status data, fault and power outage event data, and user and load data; The multi-source data is preprocessed, wherein the preprocessing includes: data cleaning, data alignment and standardization, wherein the data alignment includes: time alignment and object mapping; Based on the preprocessed multi-source data, and combined with the preset data quality assessment index calculation logic, the data quality index score of the multi-source data is calculated, wherein the data quality index score includes: reliability impact sensitivity score and process traceability score; The multi-source data quality assessment result is determined by weighting the scores of each data quality indicator and combining them with the preset data quality assessment threshold.
[0006] Preferably, the calculation method for the reliability-affected sensitivity score includes: Based on any one or more types of raw data from the multi-source data, a preset perturbation term is applied to the raw data to obtain perturbed data; Based on the disturbance data and the corresponding original data, and combined with the power supply reliability index calculation logic, the corresponding first reliability index vector and second reliability index vector are calculated, wherein the first reliability index vector is the power supply reliability index vector corresponding to the original data, and the second reliability index vector is the power supply reliability index vector corresponding to the disturbance data. Calculate the relative deviation between the first reliability index vector and the second reliability index vector to determine the reliability impact sensitivity score corresponding to the multi-source data.
[0007] Preferably, the calculation method for the process traceability score includes: Based on the fault and power outage event data in the multi-source data, an event stage set corresponding to each event record is generated according to the event identifier, timestamp, and event stage label in the fault and power outage event data; The completeness of the event record is determined by the ratio of the number of recorded stages in the event stage set corresponding to the event record to the preset number of complete stages. Verify the time order of each record stage in the event stage set, determine the number of compliant stages in the event stage set that satisfy the time order constraint, and then determine the logical consistency degree corresponding to the event record based on the ratio of the number of compliant stages to the number of record stages. The process traceability score is obtained by weighting the completeness of the stage and the consistency of logic.
[0008] Preferably, the data quality index scoring further includes: data availability score and structural consistency score.
[0009] Preferably, the calculation method for the data availability score includes: The data fields of the multi-source data are validated to determine the number of valid fields, and then the field availability is determined based on the ratio of the number of valid fields to the total number of fields. Based on the completeness of each data field, the number of valid objects is counted, and then the availability of objects is determined by the ratio of the number of valid objects to the preset theoretical number of objects. A data availability score is obtained by weighting the availability of the field with the availability of the object.
[0010] Preferably, the structural consistency score is calculated using the following method: According to the preset data consistency rules, the consistency of the multi-source data is checked, and the number of violations and the number of conformities of each data record in the multi-source data are counted. The basic consistency is determined by the complement of the ratio of the number of rule violations to the number of rule compliances; Object matching is performed among various types of data in the multi-source data. The number of valid objects that are successfully matched across types is counted. The consistency of topology and mapping is determined based on the ratio of the number of valid objects to the total number of objects. The structural consistency score is obtained by weighting the basic consistency and the topology and mapping consistency.
[0011] Preferably, before determining the multi-source data quality assessment result based on the weighted sum of the scores of each data quality indicator and a preset data quality assessment threshold, the process further includes: Based on the data quality index scores corresponding to each evaluation dimension, multiple evaluation dimension perturbation datasets are obtained respectively. The evaluation dimension perturbation dataset is formed by applying a perturbation term to one of the evaluation dimensions based on a preset original dataset, and the evaluation dimensions targeted by each evaluation dimension perturbation dataset do not overlap. Based on the disturbance datasets for each evaluation dimension and the original dataset, and following the power supply reliability index calculation logic, the disturbance reliability index results corresponding to the disturbance datasets for each evaluation dimension and the benchmark reliability index results corresponding to the original dataset are calculated respectively. Based on the relative deviations between the results of each disturbance reliability index and the results of the benchmark reliability index, the influence weights corresponding to each evaluation dimension are determined.
[0012] A second aspect of this application provides a multi-source data quality assessment device for power supply reliability analysis, comprising: A multi-source data acquisition unit is used to acquire multi-source data to be evaluated, wherein the multi-source data includes: equipment ledger data, operation measurement and status data, fault and power outage event data, and user and load data; A data preprocessing unit is used to preprocess the multi-source data, wherein the preprocessing includes: data cleaning, data alignment and standardization, wherein the data alignment includes: time alignment and object mapping; The data quality index calculation unit is used to calculate the data quality index score of the multi-source data based on the preprocessed multi-source data and in combination with the preset data quality assessment index calculation logic. The data quality index score includes: reliability impact sensitivity score and process traceability score. The data quality assessment unit is used to determine the multi-source data quality assessment result based on the weighted sum of the scores of each data quality indicator and a preset data quality assessment threshold.
[0013] A third aspect of this application provides a multi-source data quality assessment terminal for power supply reliability analysis, comprising: a memory and a processor; The memory is used to store program code, which corresponds to the multi-source data quality assessment method for power supply reliability analysis as provided in the first aspect of this application. The processor is used to read and execute the program code to implement the multi-source data quality assessment method for power supply reliability analysis.
[0014] The fourth aspect of this application provides a computer-readable storage medium storing program code that is read and executed by a processor to implement the multi-source data quality assessment method for power supply reliability analysis as provided in the first aspect of this application.
[0015] As can be seen from the above technical solutions, this application has the following advantages: The proposed solution directly quantifies the transmission impact of data quality deviations on the calculation results of power supply reliability indicators by constructing evaluation indicators specific to power supply reliability analysis, such as reliability impact sensitivity scores and process traceability scores. This effectively solves the problem of the disconnect between general data quality evaluation systems and business scenarios. It has the advantages of being able to specifically assess the impact of multi-source data quality on power supply reliability analysis results, improving the adaptability of data quality assessment to power supply reliability analysis business scenarios, and thus ensuring the reliability and consistency of analysis results. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating an embodiment of a multi-source data quality assessment method for power supply reliability analysis provided in this application.
[0018] Figure 2 This is a schematic diagram of the architecture of an embodiment of a multi-source data quality assessment device for power supply reliability analysis provided in this application.
[0019] Figure 3This is a schematic diagram of the architecture of a multi-source data quality assessment terminal embodiment for power supply reliability analysis provided in this application. Detailed Implementation
[0020] In traditional power supply reliability analysis, the data quality evaluation system mainly uses general dimensions such as data accuracy, completeness, and consistency for measurement. However, the design goals of this system are significantly mismatched with the actual application scenarios of power supply reliability analysis. This makes it impossible to effectively characterize the transmission impact of data quality problems on the calculation results of reliability indicators such as the System Average Outage Duration Index (SAIDI) and the System Average Outage Frequency Index (SAIFI). As a result, the accuracy of power supply reliability analysis results fluctuates, which in turn affects the reliability of the analysis work.
[0021] For example, in a typical urban power distribution network reliability assessment scenario, multi-source data comprises equipment ledger data collected by the distribution automation system, operational measurement and status data recorded by smart meters, fault and outage event data reported by feeder terminals, and user and load data generated by user load monitoring devices. When event stage labels are missing or timestamps are inaccurate in the fault and outage event data, the general evaluation system can only identify it as a completeness or time consistency problem, but cannot reflect the specific impact mechanism of this problem on the SAIDI calculation logic. Furthermore, data association errors caused by inconsistent equipment identification during object mapping lead to deviations between the outage range statistics and the actual physical topology, resulting in unreliability of reliability index calculation results and ultimately affecting the basis for operation and maintenance strategy formulation. If the above problems are not addressed, data quality evaluation will fail to reflect the special requirements of power supply reliability analysis for data quality, resulting in a lack of targeted guidance for data governance work. This leads to a continuous decrease in the credibility of power supply reliability analysis results, which may cause deviations in power grid operation decisions and thus adversely affect the stability of power supply services.
[0022] In view of this, embodiments of this application provide a multi-source data quality assessment method and related equipment for power supply reliability analysis, which is used to specifically assess the impact of multi-source data quality on the power supply reliability analysis results, improve the adaptability of data quality assessment to power supply reliability analysis business scenarios, and thus ensure the reliability and consistency of analysis results.
[0023] To make the inventive objectives, features, and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0024] First, a detailed description of an embodiment of a multi-source data quality assessment method for power supply reliability analysis provided in this application is as follows: Please see Figure 1 The multi-source data quality assessment method for power supply reliability analysis proposed in this application includes the following steps: Step 101: Obtain the multi-source data to be evaluated; Step 102: Preprocess the multi-source data; Step 103: Based on the preprocessed multi-source data and combined with the preset data quality assessment index calculation logic, calculate the data quality index score of the multi-source data; The data quality index scores include: reliability impact sensitivity score and process traceability score; Step 104: Determine the multi-source data quality assessment result based on the weighted sum of the scores of each data quality indicator and the preset data quality assessment threshold.
[0025] For ease of understanding, the following explains some key terms in this embodiment: Multi-source data refers to the collection of data from different systems, formats, and time dimensions involved in power supply reliability analysis. Specifically, it can include equipment ledger data, used to record the static attributes and topological relationships of power equipment; operational measurement and status data, used to reflect real-time power consumption, voltage, current, and equipment status information of the power system; fault and outage event data, used to record event information such as the type of fault, outage range, and duration of power system outages; and user and load data, used to describe user electricity consumption behavior and load characteristics. These data collectively constitute the foundational information for power supply reliability analysis.
[0026] Data preprocessing is the process of processing and transforming raw, multi-source data to improve its quality and usability. This preprocessing typically includes data cleaning, data alignment, and standardization. Data cleaning aims to identify and correct errors, missing values, or outliers in the data. Data alignment addresses inconsistencies in time and objects between different data sources. Time alignment ensures that different time series data remain synchronized over time, while object mapping ensures that data describing the same physical entity from different data sources can be correctly correlated. Standardization aims to eliminate differences in units and numerical ranges between different data sources or different metrics, making them comparable.
[0027] The data quality assessment indicator calculation logic refers to the set of rules and algorithms used to quantify the quality of multi-source data. This logic calculates scores for various data quality indicators based on pre-defined assessment dimensions after preprocessing the data.
[0028] Data quality metric scores are numerical representations that quantify data quality, reflecting data performance across specific dimensions. In this embodiment, the scores include a reliability impact sensitivity score and a process traceability score. The reliability impact sensitivity score measures the degree to which a particular type of data quality deviation affects the calculation results of power supply reliability metrics. The process traceability score assesses the extent to which data records support the integrity and logical consistency of business processes.
[0029] Data quality assessment thresholds are preset standard values used to determine whether data quality meets requirements. By comparing the calculated data quality index scores with these thresholds, the overall quality level of multi-source data can be determined.
[0030] The multi-source data quality assessment result is a final judgment obtained based on the scores of various data quality indicators and their weighted sum, combined with a preset assessment threshold. It is used to indicate whether the current multi-source data is suitable for power supply reliability analysis.
[0031] More specifically, this embodiment provides a multi-source data quality assessment method for power supply reliability analysis, the main technical features of which are described in detail below: First, this method involves acquiring multi-source data to be evaluated. This multi-source data is the basis for power supply reliability analysis and comes from diverse sources and is complex in type.
[0032] Secondly, the acquired multi-source data undergoes preprocessing. This preprocessing is a crucial step in improving data quality. Specifically, data cleaning can be achieved through manual verification, where data analysts check each data record to identify and correct obvious errors or outliers. For example, missing equipment model information in the equipment ledger can be manually added. Regarding data alignment, time alignment can be achieved through simple date and timestamp matching, such as converting all data timestamps to Coordinated Universal Time (UTC) format. Object mapping can be accomplished by establishing a static lookup table that predefines the correspondence between identifiers of the same physical entities in different data sources; for example, manually associating transformer numbers from different systems. Standardization can be achieved using linear scaling, uniformly mapping all numerical data to a fixed interval of [0, 1] to eliminate dimensional differences.
[0033] Furthermore, based on the preprocessed multi-source data and combined with the pre-defined data quality assessment index calculation logic, a data quality index score for the multi-source data is calculated. This calculation logic consists of a predefined series of rules and formulas. For example, the reliability impact sensitivity score can be determined simply through expert experience, i.e., domain experts subjectively score the importance of reliability analysis based on different data types. The process traceability score can be calculated based on the completeness of fields recorded in the fault and power outage event data. For example, the fill rate of key fields (such as event start time, end time, and power outage range) in each event record is statistically analyzed and used as the scoring basis. These scores aim to quantify the quality level of the data from different dimensions.
[0034] Finally, based on the weighted sum of the scores for each data quality indicator and a preset data quality assessment threshold, the multi-source data quality assessment result is determined. This weighted sum can be calculated by assigning a fixed weight value to each data quality indicator score; these weight values can be preset by the system administrator based on experience. For example, the reliability impact sensitivity score might be assigned a weight of 0.6, while the process traceability score might be assigned a weight of 0.4. These weighted scores are then summed. This sum is compared to a preset data quality assessment threshold. For example, if the total score is higher than a preset value (e.g., 70 points), the assessment result is judged as "good data quality"; otherwise, it is judged as "poor data quality." This assessment result provides a data quality reference for subsequent power supply reliability analysis.
[0035] Furthermore, the above technical solution will be explained in more detail below through a more specific example: Suppose a power company needs to conduct an annual analysis of the power supply reliability of its distribution network, but its data comes from multiple independent systems, including an equipment management system, a SCADA system, a power outage management system, and a customer information system. To ensure the accuracy of the analysis results, a quality assessment of this multi-source data is required.
[0036] First, the multi-source data acquisition unit acquires the multi-source data to be evaluated. Specifically, equipment ledger data can be exported as a CSV file from the equipment management system's database; operational measurement and status data can be obtained from the SCADA system's historical data library via an API interface; fault and power outage event data can be manually exported as an Excel spreadsheet from the power outage management system; and user and load data can be downloaded in batches periodically from the customer information system. This data is then aggregated into a data lake, awaiting further processing.
[0037] Next, the data preprocessing unit preprocesses this aggregated multi-source data. During the data cleaning phase, the system identifies and marks records in the equipment ledger that lack key parameters (such as rated voltage) and prompts operators to manually enter them. Abnormal peaks or troughs in the operational measurement data are identified using a simple threshold filtering method and marked as suspicious data. In the data alignment phase, time alignment ensures that the timestamps of all operational measurement data and fault event data are unified to second-level precision; for example, millisecond-level timestamps in the SCADA system are truncated to second-level precision. Object mapping uses a pre-configured mapping table to associate the device IDs in the power outage management system with the device IDs in the equipment management system, ensuring that records of the same device in different data sources can be correctly identified. Standardization processing converts load data of different units (such as kilowatts and megawatts) into standard units and performs linear normalization to ensure consistent numerical ranges.
[0038] Subsequently, the data quality index calculation unit calculates data quality index scores based on this preprocessed multi-source data and in conjunction with preset data quality assessment index calculation logic. For example, for the reliability impact sensitivity score, a simple rule can be set: if the missing rate of key equipment (such as transformers and circuit breakers) in the equipment ledger data exceeds 5%, 20 points are deducted directly from this score; if the missing rate of a certain type of measuring point (such as line current) in the operational measurement data exceeds 10%, 15 points are deducted. For the process traceability score, it can be checked whether each event record in the fault and power outage event data contains the three key fields of "fault occurrence time," "fault recovery time," and "power outage range." If an event record lacks one of these fields, its traceability score will be reduced accordingly.
[0039] Finally, the data quality assessment unit determines the multi-source data quality assessment result based on the weighted sum of the scores for each data quality indicator, combined with a preset data quality assessment threshold. Assuming the weight of the reliability impact sensitivity score is 0.6 and the weight of the process traceability score is 0.4, if the calculated reliability impact sensitivity score is 80 and the process traceability score is 70, then the weighted sum will be 80. 0.6 + 70 0.4 = 48 + 28 = 76 points. If the preset evaluation threshold is 75 points, then the final evaluation result will be judged as "data quality qualified". This result shows that the current multi-source data quality is sufficient to support relatively accurate power supply reliability analysis, thus effectively solving the technical problem of unstable analysis accuracy caused by data quality issues.
[0040] Based on the above examples, the multi-source data quality assessment method for power supply reliability analysis proposed in this embodiment demonstrates significant progress in its technical concept. In existing technologies, data quality assessment often employs general dimensions, such as focusing only on data completeness or consistency, without deeply considering the actual impact of data quality issues on the calculation results of power supply reliability indicators. For example, in traditional methods, a missing equipment ledger field might be simply marked as "incomplete," but its specific impact on SAIDI or SAIFI cannot be quantified.
[0041] This embodiment directly integrates data quality assessment with the core requirements of power supply reliability analysis by introducing reliability impact sensitivity scoring and process traceability scoring. Reliability impact sensitivity scoring quantifies the transmission impact of different data quality issues on the calculation results of power supply reliability indicators, making the assessment results more targeted. For example, in the above example, the direct deduction mechanism for the missing data rate in equipment ledgers can intuitively reflect the potential negative impact of this missing data on reliability analysis, which contrasts sharply with the general assessment method that simply marks data as "incomplete."
[0042] Furthermore, process traceability scoring focuses on assessing how data records support the integrity and logical consistency of business processes, which is crucial for the accurate analysis of fault and power outage event data. In the example above, assessing traceability by checking the completeness of key fields in fault event records ensures that the quality of event data meets the needs of subsequent fault analysis and cause tracing, thus providing a more reliable data foundation for power supply reliability analysis.
[0043] In some of the above embodiments, this application proposes a multi-source data quality assessment method for power supply reliability analysis. This method can acquire and preprocess multi-source data and calculate data quality index scores, including a reliability impact sensitivity score. Based on this, this application further proposes a calculation method for the reliability impact sensitivity score.
[0044] This approach first selects one or more types of raw data from the multi-source data for analysis. The multi-source data refers to a collection of data from various sources and of various types used for power supply reliability analysis, such as equipment ledger data, operational measurement and status data, fault and outage event data, and user and load data. Selecting "any one or more types of raw data" for evaluation makes the evaluation process flexible, allowing for sensitivity analysis of specific data types that have a significant impact on power supply reliability, based on actual needs and concerns.
[0045] Subsequently, a preset perturbation term is applied to the selected original data to obtain perturbed data. Applying a preset perturbation term means introducing artificially set deviations or errors that simulate data quality problems into the original data. These perturbation terms can be random errors, systematic biases, missing data, outliers, etc. For example, a certain percentage of value can be added or removed from a field in the original data, some data records can be randomly deleted, or some data can be replaced with erroneous values. Perturbed data is the data after the perturbation term has been applied; it simulates the state of the original data when certain quality problems exist.
[0046] Next, based on the disturbance data and the corresponding original data, and combined with the power supply reliability index calculation logic, the corresponding first reliability index vector and second reliability index vector are calculated.
[0047] The power supply reliability index calculation logic refers to the mathematical models and algorithms used to evaluate the power supply reliability of a power system. It represents standard methods or models for assessing power supply reliability, such as calculating the system's average outage time, average outage frequency, and average outage time per user. These can be reliability assessment models defined based on IEEE standards or State Grid enterprise standards, such as Monte Carlo simulation or analytical methods, or customized reliability assessment algorithms developed based on the specific characteristics and operational data of a regional power grid. The first reliability index vector refers to the set of reliability assessment results obtained using unperturbed raw data through the power supply reliability index calculation logic, representing the power supply reliability level under ideal data quality. The second reliability index vector refers to the set of reliability assessment results obtained using perturbed data through the same power supply reliability index calculation logic, representing the power supply reliability level when data quality issues exist.
[0048] Next, the relative deviation between the first reliability index vector and the second reliability index vector is calculated to determine the reliability impact sensitivity score corresponding to the multi-source data. The relative deviation is a quantitative representation of the difference between the first and second reliability index vectors, usually presented as a percentage or ratio. For example, it could be the ratio of the absolute value of the difference between corresponding elements of the two vectors to the absolute value of corresponding elements of the original vector. The reliability impact sensitivity score is a numerical value calculated based on the relative deviation, used to quantify the degree of impact of multi-source data quality problems on the power supply reliability analysis results. The higher the score, the greater the impact of data quality problems on the reliability analysis results.
[0049] Furthermore, when the raw data contains multiple data categories, the final reliability impact sensitivity score is based on the average of the relative deviations corresponding to each category. When perturbation analysis is performed on multiple categories of raw data, each category will generate a corresponding relative deviation. By calculating the arithmetic mean of the relative deviations corresponding to each category of raw data, a comprehensive score reflecting the overall data quality's impact on reliability can be obtained. Alternatively, other statistical methods, such as weighted averages, can be used to assign different weights based on the importance of different data types.
[0050] For example, let the power supply reliability index vector be: ; in, Can represent , , Indicators such as average power outage duration; For the first By applying a perturbation to the data quality dimension, we obtain the perturbed metrics:
[0051] Then the first The degree of influence of class data on reliability indicators is defined as follows: ; in, It is the first The reliability of the data affects sensitivity. It is the first Each reliability metric weight, It is a very small positive number that prevents the denominator from being zero; The overall sensitivity can be obtained by averaging the values of all evaluated objects j.
[0052] Finally, this overall sensitivity is used as the sensitivity score for the reliability impact of the multi-source data.
[0053] Through the above technical solution, this application provides an effective method for quantifying the sensitivity scoring of multi-source data reliability impact. This method, by simulating the impact of data quality issues on power supply reliability analysis results, can accurately identify which data types or data quality issues are more sensitive to the final power supply reliability assessment results. This makes data quality assessment no longer isolated, but closely integrated with core business objectives (power supply reliability), thus providing clear priorities and directions for data governance and quality improvement work. For example, when a certain type of data is found to be highly sensitive to reliability indicators, resources can be prioritized to improve the quality of that type of data, thereby achieving greater accuracy improvement in reliability analysis at a lower cost. This quantitative assessment method significantly enhances the scientific rigor and guidance of data quality assessment, avoids blind investment and resource waste, and ensures the accuracy and effectiveness of power supply reliability analysis.
[0054] In some of the above embodiments, this application proposes a multi-source data quality assessment method for power supply reliability analysis, including process traceability scoring. Based on this, this application further proposes a calculation method for the process traceability score, including: based on fault and power outage event data in the multi-source data, generating an event stage set corresponding to each event record according to the event identifier, timestamp, and event stage label in the fault and power outage event data; determining the stage completeness corresponding to the event record based on the ratio of the number of recorded stages in the event stage set corresponding to the event record to a preset number of complete stages; verifying the time order of each recorded stage in the event stage set, determining the number of compliant stages in the event stage set that satisfy the time order constraint, and then determining the logical consistency corresponding to the event record based on the ratio of the number of compliant stages to the number of recorded stages; and obtaining the process traceability score based on the weighted sum of the stage completeness and the logical consistency.
[0055] This calculation method first generates a set of event stages for each event record based on fault and power outage event data from multiple sources, according to event identifiers, timestamps, and event stage labels. Fault and power outage event data refers to data recording power system faults and power outages, typically including detailed information such as the event's occurrence time, restoration time, involved equipment, fault cause, and handling process. The event identifier is a unique code used to identify each fault or power outage event; it can be a unique serial number automatically generated by the system or a composite identifier combining equipment ID and timestamp. The timestamp is the precise time information recording the occurrence of each stage in the event data, such as fault occurrence time, isolation operation time, and power restoration time; its format can be uniform UTC time or local time. Event stage labels are descriptive markers for different states or actions during event handling, such as "fault occurrence," "protection action," "isolation," "emergency repair," and "power restoration." These labels help discretize the continuous event process into identifiable stages. An event record corresponds to an event stage set, which is an ordered set that organizes all related timestamps and event stage labels under the same event identifier. For example, it can be a list or array, where each element represents an event stage and contains its timestamp and label.
[0056] Next, the completeness of the event record's stages is determined by comparing the number of recorded stages in the event stage set corresponding to the event record with the preset number of complete stages. This step quantifies the completeness of the stage information contained in a single event record. The number of recorded stages refers to the actual number of event stages extracted from the event stage set corresponding to the event record. The preset number of complete stages is a number of stages that should be included in an ideal or minimum event handling process, predefined according to industry standards, operating procedures, or expert experience. For example, for a typical line fault, four complete stages might be preset: "fault occurrence," "isolation," "repair," and "restoration." By calculating the ratio of the number of recorded stages to the preset number of complete stages, a value between 0 and 1 can be obtained; the higher the value, the more complete the stage information of the event record.
[0057] Subsequently, the chronological order of each recorded stage in the event stage set is verified to determine the number of compliant stages that satisfy the chronological order constraint. Then, based on the ratio of the number of compliant stages to the number of recorded stages, the logical consistency of the event record is determined. This step is used to evaluate whether the chronological order of each stage in the event record conforms to logical and practical operational specifications. Verifying the chronological order of each recorded stage in the event stage set means checking whether the timestamps of adjacent stages in the event stage set are in an increasing relationship; that is, the timestamp of the later stage must be later than or equal to the timestamp of the earlier stage. For example, the timestamp of the isolation operation cannot be earlier than the timestamp of the fault occurrence. The number of compliant stages that satisfy the chronological order constraint refers to the number of stages in the event stage set that are deemed to conform to the logical order after passing the chronological order verification. If the timestamp of a certain stage violates the order, that stage and its subsequent affected stages may not be counted as compliant stages. By calculating the ratio of the number of compliant stages to the number of recorded stages, a value between 0 and 1 can be obtained. The higher the value, the more consistent the logical order of the event records.
[0058] Finally, the process traceability score is obtained by weighting the completeness of the stage and the consistency of logic. This step combines the two dimensions of completeness of stage and consistency of logic to form a comprehensive process traceability assessment index. The weighted sum involves multiplying each of the completeness of stage and consistency of logic by preset weighting coefficients, and then adding the results. These weighting coefficients can be adjusted according to actual business needs and the degree of importance placed on completeness and logic. For example, logical consistency can be considered more important than stage completeness, thus giving it a higher weight. The final process traceability score is a comprehensive numerical value that can intuitively reflect the quality level of event data in terms of record completeness and temporal logic.
[0059] For example, for a power outage event, define a set of critical stages: ; It includes stages such as fault occurrence, fault reporting, fault location, fault isolation, fault repair, and power restoration. Number of stages; Let the first The number of recorded stages for each event is: The stage completeness is: ; Further considering the logical rationality of the stages, let the number of stages satisfying the time sequence constraint be . Then the logical consistency is: ; Therefore, process traceability is defined as: ; in, These are the weighting coefficients.
[0060] For example, as a specific implementation, suppose there is a fault and power outage event data record with the event identifier "E20231026001". The raw data of this event includes the following timestamp and event stage label: - "Fault Occurred", timestamp: 2023-10-26 10:00:00; - “Protective Action”, timestamp: 2023-10-26 10:00:05; - "Isolation Operation", timestamp: 2023-10-26 10:05:00; - "Power restored", timestamp: 2023-10-26 11:00:00; Based on this information, a set of event stages corresponding to the event record can be generated. Assume the preset total number of stages is 5, including "Fault Occurrence," "Protection Action," "Isolation Operation," "On-site Repair," and "Power Restoration." In this example, the actual number of recorded stages is 4 (missing the "On-site Repair" stage). Therefore, the stage completeness of this event record can be calculated as 4 / 5 = 0.8. Next, the time sequence of each recorded stage in the event stage set is verified: - “Fault Occurred” (10:00:00) → “Protection Action” (10:00:05): The sequence is correct.
[0061] - “Protective Action” (10:00:05) → “Isolation Operation” (10:05:00): The sequence is correct.
[0062] - “Isolation Operation” (10:05:00) → “Power Restoration” (11:00:00): The sequence is correct.
[0063] All four stages of the records satisfy the chronological order constraint, therefore the number of compliant stages is 4. The logical consistency of this event record can be calculated as 4 / 4 = 1.0. Finally, assuming the weight of stage completeness is 0.6 and the weight of logical consistency is 0.4, the process traceability score of this event can be calculated as (0.8). 0.6) + (1.0 0.4) = 0.48 + 0.4 = 0.88. Through the above calculation, a comprehensive process traceability score for this specific event record can be obtained, thus providing a quantitative basis for power supply reliability analysis.
[0064] Through the above technical solutions, this application provides a detailed and quantitative method for calculating process traceability scores. This method can comprehensively assess the completeness and temporal logic of fault and power outage event data, avoiding biases in power supply reliability analysis caused by incomplete event records or disordered chronological order. By combining stage completeness and logical consistency, the data quality assessment results are more accurate and reliable, thereby significantly improving the accuracy and credibility of power supply reliability analysis and providing stronger data support for power system operation decisions.
[0065] In some of the above embodiments, this application proposes a method for evaluating the quality of multi-source data by using reliability impact sensitivity scores and process traceability scores. Based on this, this application further proposes that the data quality index scores also include data availability scores and structural consistency scores.
[0066] Data availability scoring measures the integrity and accessibility of a dataset. It refers to the extent to which data is accessible and usable by users or systems at a specific point in time, as well as the integrity of the data records themselves. This score can be used to measure whether the data meets the minimum data requirements for power supply reliability analysis tasks. One implementation is to determine this by the ratio of valid records to the total number of records in the dataset, or by checking the fill rate of key fields. Another implementation is to conduct a comprehensive evaluation based on indicators such as the online time of the data storage system, data backup and recovery capabilities, to ensure that data can be accessed and used promptly when needed. Structural consistency scoring assesses the uniformity and standardization of data in terms of format, type, range, and correlation between different data sources. It refers to whether the data follows predefined data models, schemas, or business rules, and whether there are logical conflicts or mismatches between different data tables or datasets. This score can reflect structural quality issues. One implementation is to define a series of data validation rules and check whether data records conform to these rules, such as whether the data type is correct, whether the numerical range is reasonable, and whether related fields match. Another approach is to establish a data dictionary and metadata management mechanism to ensure that all data is stored and managed in accordance with unified standards, and to identify and correct structural inconsistencies through regular data audits.
[0067] The further proposed solution in this embodiment introduces data availability and structural consistency scores on top of the existing reliability impact sensitivity score and process traceability score, thereby constructing a more comprehensive and detailed multi-source data quality assessment system. The introduction of the data availability score ensures that the assessment process fully considers the completeness and accessibility of the data, guaranteeing that the data used for power supply reliability analysis is sufficient and usable, avoiding analytical biases caused by missing data. The introduction of the structural consistency score focuses on the standardization of the internal logic and external relationships of the data, ensuring the structural uniformity of data from different sources and of different types. This is crucial for integrating multi-source data to perform complex power supply reliability model calculations, effectively avoiding calculation errors or misjudgments caused by inconsistent data structures. By comprehensively considering these four types of scores and determining the final assessment result based on their weighted sum, this application can more accurately identify data quality problems affecting power supply reliability analysis, providing more guiding evidence for subsequent data governance and optimization, thereby improving the overall accuracy and reliability of power supply reliability analysis.
[0068] Based on the above embodiments, this application further proposes a method for calculating data availability scores, including: The data fields of the multi-source data are validated to determine the number of valid fields. Then, the field availability is determined based on the ratio of the number of valid fields to the total number of fields. The number of valid objects is counted based on the object completeness of each data field. Then, the object availability is determined based on the ratio of the number of valid objects to the preset theoretical number of objects. The data availability score is obtained by weighted sum of the field availability and the object availability.
[0069] The process of validating data fields in the multi-source data to determine the number of valid fields aims to assess the integrity of each data record in the multi-source data. Data field validation involves checking whether each field in a data record contains valid values or conforms to preset data format, type, and range requirements. For example, for equipment ledger data, fields such as "Equipment Number," "Commissioning Date," and "Rated Voltage" can be validated to ensure they are not empty, are numeric, or are within a reasonable range. The number of valid fields refers to the total number of fields in all data records that meet the integrity or compliance requirements after validation. This validation can be automated using preset data dictionaries, data patterns, or business rules. Alternatively, a series of data integrity rules, such as non-empty constraints, uniqueness constraints, and referential integrity constraints, can be defined to scan each field in the multi-source data and count the number of fields that conform to these rules as the number of valid fields.
[0070] The ratio of valid fields to total fields determines field availability, a step used to quantify the overall completeness of data fields. Field availability is a metric that measures the field fill rate and effectiveness in a dataset, calculated by dividing the number of valid fields by the total number of fields. The total number of fields refers to the total number of all fields in all data records, including valid and invalid fields. A higher ratio indicates better data field completeness and higher data availability. Alternatively, different weights can be assigned to fields of varying importance, and then the weighted ratio of valid fields to weighted total fields can be calculated for a more refined reflection of field availability.
[0071] Based on the completeness of each data field, the number of valid objects is counted. This step aims to assess the completeness of each independent data entity in the dataset. A data object refers to a complete record representing a specific entity (such as a device, a user, or a fault event) in multi-source data. Object completeness refers to whether all fields contained in a data object are valid or meet completeness requirements. For example, for a device object in equipment ledger data, if its key fields such as "device number," "device type," and "installation location" are all valid, the object is considered complete. The number of valid objects refers to the total number of data objects in the entire dataset that meet the preset completeness criteria. As an alternative implementation, a threshold can be defined, for example, a data object is counted as a valid object only when more than 80% of its key fields are valid.
[0072] The ratio of the number of valid objects to the preset theoretical number of objects determines object availability. This step quantifies the degree of matching between the actual valid data objects in the dataset and the data objects that should theoretically exist. Object availability is an indicator that measures data coverage and accuracy, calculated by dividing the number of valid objects by the preset theoretical number of objects. The preset theoretical number of objects refers to the total number of data objects that should theoretically exist based on business needs, system design, or external reference data. The higher the ratio, the more comprehensive the data coverage of actual objects, and the higher the data availability. As another implementation method, the preset theoretical number of objects can be dynamically determined based on historical data trends, planned data, or comparison with authoritative external data sources.
[0073] A data usability score is obtained by weighting the field availability and object availability. This step combines the two dimensions of field availability and object availability to form a unified data usability score. Weighted summation refers to linearly combining field availability and object availability according to preset weights. By using weighted summation, the influence of different dimensions on the data usability score can be flexibly adjusted according to actual business needs and data characteristics, resulting in a more comprehensive usability assessment result that better reflects the actual application scenario. Alternatively, a non-linear combination method can be used, such as through fuzzy logic or machine learning models, to comprehensively judge and output the data usability score based on field availability, object availability, and other relevant factors.
[0074] For example, suppose the set of key fields required for a certain analysis task is: ; For the The number of valid fields in each data record is [number]. Then the field availability is defined as: ; Furthermore, considering the integrity of the object layer, let... For the theoretical number of objects, If the number of valid objects is given, then the object availability is: ; Comprehensive analysis defines usability as: ; in, These are the weighting coefficients.
[0075] Further, a concrete example will be used to illustrate this. Suppose we need to evaluate the data availability of equipment ledger data for a power supply company. First, the multi-source data acquisition unit acquires the equipment ledger data, which includes fields such as equipment number, equipment type, commissioning date, rated voltage, and installation location. The data quality index calculation unit will first validate these data fields. For example, it checks whether the "equipment number" field is empty and correctly formatted, whether the "commissioning date" is in a valid date format, and whether the "rated voltage" is a numerical value within a reasonable range. Through validation, the number of valid fields that meet the requirements is counted across all records. Assuming there are a total of 10,000 fields, and 9,500 are valid, then the field availability is 0.95. Next, the data quality index calculation unit will count the number of valid objects based on the completeness of each equipment record. For example, if the three key fields of "equipment number," "equipment type," and "installation location" of a equipment record are all valid, then the equipment object is considered complete. Assuming the power supply area should theoretically have 1000 devices, and the number of valid devices counted is 980, then the object availability is 0.98. Finally, the data quality indicator calculation unit can calculate a data availability score of 0.6 based on preset weights (e.g., field availability weight 0.6, object availability weight 0.4). 0.95 + 0.4 0.98 = 0.57 + 0.392 = 0.962.
[0076] The above technical solution enables a comprehensive and accurate assessment of data availability from multiple sources. This solution considers not only the completeness of data fields, ensuring the micro-quality of data records, but also the coverage of data objects, guaranteeing the macro-representation of the data to actual business entities. This layered and comprehensive assessment approach allows data availability scores to more precisely reflect the actual value of data in power supply reliability analysis, effectively avoiding analytical biases caused by missing data fields or incomplete data objects. This provides stronger and more reliable data support for power supply reliability analysis, improving the accuracy and credibility of the assessment results.
[0077] In some of the above embodiments, this application proposes a data quality index scoring method for the multi-source data quality assessment method for power supply reliability analysis. This method may further include a structural consistency score to measure the consistency of multi-source data across devices, users, events, and topology associations. Based on this, this application further proposes a calculation method for the structural consistency score, including: According to preset data consistency rules, consistency verification is performed on the multi-source data, and the number of rule violations and rule compliance for each data record in the multi-source data is counted. Basic consistency is determined based on the complement of the ratio of the number of rule violations to the number of rule compliance. Object matching is performed between various types of data in the multi-source data, and the number of valid objects that successfully match across types is counted. Then, the topology and mapping consistency is determined based on the ratio of the number of valid objects to the total number of objects. The structural consistency score is obtained by weighted sum of the basic consistency and the topology and mapping consistency.
[0078] The preset data consistency rules are a set of specifications used to define the internal structure and logical relationships of multi-source data, aiming to ensure data integrity, accuracy, and consistency. For example, preset rules may require that equipment numbers in equipment ledger data be unique and conform to a specific encoding format, or that measured values in operational measurement and status data must be within a reasonable physical range. In another implementation, preset rules may require that the start time of a power outage in fault and power outage event data cannot be later than the end time of the power outage, or that the user ID in user and load data must be associated with a connection point ID in the equipment ledger data. Consistency verification of the multi-source data refers to checking each data record in the multi-source data according to the preset data consistency rules to determine whether it conforms to these rules. For example, all equipment records in the equipment ledger data can be traversed to check whether each equipment number is unique and verify whether its format conforms to a preset regular expression. In another implementation, all measurement records in the operational measurement and status data can be checked to determine whether their measured values exceed preset upper and lower thresholds or whether null values exist. The statistical analysis of the number of rule violations and the number of rule compliance for each data record in the multi-source data refers to the quantitative statistical analysis of the verification results for each data record after consistency verification. The number of rule violations refers to the number of times a data record does not meet a preset rule, while the number of rule compliance refers to the number of times a data record meets the preset rule. For example, for a device record, if its device number is not unique and the device type field is empty, then the number of rule violations is 2. In another implementation, if a fault event record meets all rules regarding time sequence and event type encoding, then its number of rule compliance is equal to the total number of preset rules. Based on the complement of the ratio of the number of rule violations to the number of rule compliance, the basic consistency is determined as an indicator of the internal structural regularity of the multi-source data. Its calculation method is: 1 - (number of rule violations / (number of rule violations + number of rule compliance)). For example, if a data record has 2 rule violations and 8 rule compliances, then its basic consistency is 1 - (2 / (2+8)) = 0.8. Object matching across different types of data in the multi-source data refers to identifying and associating records representing the same physical or logical entity in different data types. For example, matching the transformer ID in the equipment ledger data with the transformer measurement point ID in the operation measurement and status data can confirm that they point to the same physical transformer. In another implementation, user address information in user and load data can be matched with geographic coordinates in Geographic Information System (GIS) data to establish an association between users and the power grid topology. The number of valid objects successfully matched across different data types refers to the number of unique entities that have successfully established an association between different data types.For example, if there are 100 transformer records in the equipment ledger, and 95 measurement point records in the operational measurement data can be successfully matched with 95 of these 100 transformers, then the number of valid objects is 95. Based on the ratio of the number of valid objects to the total number of objects, topology and mapping consistency is determined as an indicator of the completeness and accuracy of object associations between different data sources. The calculation method is: number of valid objects / total number of objects. For example, if a total of 100 devices need to be matched between different data sources, and 90 are successfully matched, then the topology and mapping consistency is 0.9. The structural consistency score, obtained by weighting the basic consistency and the topology and mapping consistency, is the final indicator comprehensively reflecting the internal normalization of the data and its cross-data source association. By assigning different weights to basic consistency and topology and mapping consistency and performing a weighted sum, a comprehensive score can be obtained.
[0079] For example, let the first The total number of consistency rules that records need to satisfy is The number of rule violations is Then, basic consistency is defined as: ; Further considering topology and mapping consistency, let the number of objects successfully matched across systems be... The total number of objects is Then the consistency between topology and mapping can be expressed as: ; The definition of overall structural consistency is: ; in, These are the weighting coefficients; As a specific implementation method, suppose a reliability analysis of the power supply system in a certain area is required, involving equipment ledger data, operational measurement and status data, and fault and power outage event data. To calculate the structural consistency score, predefined data consistency rules must first be defined. For example, it can be stipulated that the "Equipment Number" field in the equipment ledger data must be a unique string with a length of 8 characters; the "Equipment Type" field must be a value from a predefined list (such as "Transformer" or "Circuit Breaker"); and the "Commissioning Date" cannot be later than the current date. For operational measurement and status data, it can be stipulated that the "Measurement Value" field must be a numeric type between 0 and 1000; and the "Measurement Time" field must be a valid timestamp. For fault and power outage event data, it can be stipulated that the "Event ID" must be unique, and the "Power Outage Start Time" must be earlier than the "Power Outage End Time." During consistency verification, the system will iterate through each equipment record in the equipment ledger. If a record has a duplicate equipment number or its equipment type is not in the predefined list, the number of rule violations for that record will increase. Simultaneously, the number of records that satisfy the rules will also be counted. For example, if a device record has 5 rules, 3 of which are met and 2 are violated, then its basic consistency is 1 - (2 / 5) = 0.6. After statistically analyzing all data records, the basic consistency of the entire dataset can be obtained. Next, object matching is performed in the multi-source data. For example, the device ledger data records detailed information about all transformers, while the operation measurement and status data records the real-time voltage, current, and other measured values of each transformer. By matching the "Transformer ID" in the device ledger with the "Device ID to which the measurement point belongs" in the operation measurement data, it can be identified which transformers have corresponding measurement data and which do not. Assuming there are 100 transformers in the device ledger, and 90 of them have corresponding measurement point records in the operation measurement data, then the number of valid objects is 90, and the total number of objects is 100. At this point, the topology and mapping consistency is 90 / 100 = 0.9. Finally, the calculated basic consistency (e.g., 0.8) and topology and mapping consistency (e.g., 0.9) are weighted and summed. If the weight for basic consistency is set to 0.7, and the weight for topology and mapping consistency is set to 0.3, then the final structural consistency score is 0.7. 0.8 + 0.3 0.9 = 0.56 + 0.27 = 0.83. This score will be used as one of the indicators for evaluating the quality of multi-source data, to comprehensively judge the structural quality of the data.
[0080] This embodiment introduces a structural consistency score to comprehensively evaluate the structural quality of multi-source data, addressing issues of insufficient internal data standardization and lack of cross-data source correlation. First, a detailed internal verification of the multi-source data is performed using pre-defined data consistency rules. These rules cover the format, range, uniqueness of data fields, and logical relationships between records, ensuring the intrinsic quality of each data record. During verification, the system accurately counts the number of rules violated and followed for each data record, thus calculating basic consistency. Basic consistency reflects the degree of standardization within the data set; a higher value indicates a more rigorous internal data structure. Second, to address the issue of inaccurate object mapping between different data sources, this solution performs object matching between various types of data in the multi-source data. For example, by using common identifiers or attributes, equipment in the equipment ledger data is associated with measurement points in the operational measurement data. The ratio of the number of successfully matched valid objects to the total number of objects is used to determine topology and mapping consistency. Topology and mapping consistency quantifies the completeness and accuracy of data association between different data sources, which is crucial for building a unified power grid model and conducting cross-domain analysis. Finally, the basic consistency and topology and mapping consistency are weighted and summed to obtain the structural consistency score. This comprehensive scoring mechanism not only considers the standardization of the data itself but also the accuracy of the correlation between data, thus providing a comprehensive and objective structural quality assessment result. In this way, this scheme can identify deep-seated data structure problems that affect the accuracy of power supply reliability analysis, providing a clear direction for subsequent data repair and optimization, and ensuring the structural integrity and logical coherence of multi-source data during reliability analysis.
[0081] In some embodiments described above in this application, a multi-source data quality assessment result is determined based on a weighted sum of scores for each data quality indicator, combined with a preset data quality assessment threshold. Building upon this, this application further proposes the following steps before determining the multi-source data quality assessment result based on a weighted sum of scores for each data quality indicator, combined with a preset data quality assessment threshold: Based on the data quality index scores corresponding to each evaluation dimension, multiple evaluation dimension perturbation datasets are obtained. Each evaluation dimension perturbation dataset is formed by applying a perturbation term to one of the evaluation dimensions based on a preset original dataset, and the evaluation dimensions targeted by each evaluation dimension perturbation dataset do not overlap. Based on each evaluation dimension perturbation dataset and the original dataset, according to the power supply reliability index calculation logic, the perturbation reliability index results corresponding to each evaluation dimension perturbation dataset and the baseline reliability index results corresponding to the original dataset are calculated. Based on the relative deviation between each perturbation reliability index result and the baseline reliability index result, the influence weight corresponding to each evaluation dimension is determined.
[0082] The data quality index scores, corresponding to various evaluation dimensions, refer to treating different aspects of data quality assessment, such as reliability impact sensitivity, process traceability, data availability, and structural consistency, as independent evaluation perspectives. These dimensions each measure different aspects of data quality, and each dimension may have varying degrees of impact on the final power supply reliability analysis results. Obtaining multiple evaluation dimension perturbation datasets is to systematically simulate the impact of changes in different data quality dimensions on the power supply reliability analysis results. This can be generated by introducing random noise, missing values, error values, or making specific proportions of modifications to a pre-set original dataset for a particular evaluation dimension. For example, data missing values can be simulated to affect the data availability dimension, or data errors can be simulated to affect the structural consistency dimension. The pre-set original dataset is the foundational data for perturbation analysis. It is typically pre-processed and considered high-quality multi-source data. It can be a set of actual operating data that has undergone rigorous screening and verification over a historical period, or a benchmark dataset generated through simulation or modeling based on actual system characteristics and operating patterns. This approach, which involves applying a perturbation term to only one evaluation dimension, emphasizes the singularity of each perturbation operation. That is, only one evaluation dimension is perturbed at a time, isolating and analyzing the independent impact of that dimension. For example, when generating a perturbed dataset for the "data availability" dimension, a certain proportion of records in the original dataset can be randomly deleted without altering the characteristics of other dimensions. Each evaluation dimension's perturbation dataset targets a non-overlapping evaluation dimension, ensuring that each perturbed dataset independently reflects the impact of changes in a single evaluation dimension. This avoids confusion between perturbations of different dimensions, thus enabling accurate assessment of the independent contribution of each dimension. This can be achieved by defining a strict set of perturbation rules to ensure that applying a perturbation to one dimension does not unintentionally affect the characteristics of other dimensions.
[0083] The power supply reliability index calculation logic refers to a standard method or model used to evaluate the reliability of a power system. For example, it can calculate indicators such as the system's average outage time, average outage frequency, and average outage time per user. This can be a reliability assessment model based on IEEE standards or State Grid enterprise standards, such as Monte Carlo simulation or analytical methods, or a customized reliability assessment algorithm developed based on the characteristics and operational data of a specific regional power grid. The disturbance reliability index result refers to the reliability assessment result obtained through the power supply reliability index calculation logic using a dataset disturbed by a specific evaluation dimension. For example, when data availability decreases, the calculated system average outage time may increase; this increased value is the disturbance reliability index result. The benchmark reliability index result refers to the reliability assessment result obtained through the power supply reliability index calculation logic using a preset original dataset, serving as a benchmark for comparison. It represents the system's performance when data quality is good. The relative deviation is used to quantify the impact of disturbances on reliability indicators. It is typically expressed as the ratio of the difference between the disturbance result and the baseline result to the baseline result. This can be obtained by calculating the absolute value of (disturbance reliability indicator result - baseline reliability indicator result) / baseline reliability indicator result, or by using statistical methods such as mean squared error or percentage error. Determining the influence weight for each evaluation dimension means that the larger the relative deviation, the greater the impact of changes in the data quality of that evaluation dimension on the power supply reliability analysis results; therefore, its weight should be higher. The relative deviation can be used directly as a weight, or it can be normalized before being used as a weight. For example, the relative deviations of all dimensions can be summed, and then the relative deviation of each dimension can be divided by the sum to obtain the normalized weight.
[0084] The following is a concrete example to illustrate this. Suppose that when conducting a quality assessment of multi-source data for a regional power grid, it is necessary to determine the weights of four evaluation dimensions: "Reliability Impact Sensitivity Score," "Process Traceability Score," "Data Availability Score," and "Structural Consistency Score." First, a set of historical operational data that has undergone rigorous cleaning and verification and is considered to be of high quality can be selected as the pre-set raw dataset. Next, perturbation datasets are generated for each evaluation dimension: For the "Reliability Impact Sensitivity" dimension, a small perturbation is applied to the equipment parameters in a randomly selected portion of the equipment ledger data in the original dataset, forming the first evaluation dimension perturbation dataset; for the "Process Traceability" dimension, timestamps or event stage labels are randomly deleted from some event records in the fault and power outage event data of the original dataset to simulate a decrease in data traceability, forming the second evaluation dimension perturbation dataset; for the "Data Availability" dimension, sensor measurements are randomly deleted from the operational measurement and status data of the original dataset to simulate data loss, forming the third evaluation dimension perturbation dataset; for the "Structural Consistency" dimension, equipment connection relationships inconsistent with those in the geographic information system data are intentionally introduced into the equipment ledger data of the original dataset to simulate structural inconsistencies, forming the fourth evaluation dimension perturbation dataset. When generating these perturbation datasets, it is ensured that each perturbation applies only to its corresponding evaluation dimension, and that perturbation operations between different dimensions do not affect each other. Then, the original dataset and the four disturbance datasets are input into a pre-defined power supply reliability index calculation model to calculate the baseline reliability index result corresponding to the original dataset, as well as the disturbance reliability index result corresponding to each of the four disturbance datasets. For example, the SAIDI of the system average outage time for the original dataset might be 0.5 hours / year, while the SAIDI for the "Reliability Impact Sensitivity" disturbance dataset is 0.52 hours / year, the "Process Traceability" disturbance dataset is 0.58 hours / year, the "Data Availability" disturbance dataset is 0.65 hours / year, and the "Structural Consistency" disturbance dataset is 0.70 hours / year. Finally, relative deviations are calculated based on these results; for example, the relative deviation for "Reliability Impact Sensitivity" is 0.04, for "Process Traceability" it is 0.16, for "Data Availability" it is 0.30, and for "Structural Consistency" it is 0.40. Normalizing these relative deviations yields the impact weights for each evaluation dimension.
[0085] Through the above technical solution, this application overcomes the problems of traditional data quality assessment methods, such as the strong subjectivity of weight settings and the inability to accurately reflect the impact of data quality on business operations. By systematically applying perturbations to each evaluation dimension and analyzing their impact on power supply reliability indicators, the weight of each data quality indicator can be objectively and quantitatively determined. This business-impact-based weight allocation mechanism enables the final multi-source data quality assessment results to more accurately reflect the actual impact of data quality on power supply reliability analysis. This provides power system managers with more instructive insights into data quality status, helping to prioritize the improvement of data quality issues that have the greatest impact on power supply reliability, optimize resource allocation, and improve the overall level of power supply reliability.
[0086] The above is a detailed description of an embodiment of a multi-source data quality assessment method for power supply reliability analysis provided in this application. The following is a detailed description of related embodiments of a multi-source data quality assessment device, terminal, and storage medium for power supply reliability analysis provided in this application.
[0087] Please see Figure 2 This application provides a multi-source data quality assessment device for power supply reliability analysis, comprising: The multi-source data acquisition unit 201 is used to acquire multi-source data to be evaluated, wherein the multi-source data includes: equipment ledger data, operation measurement and status data, fault and power outage event data, and user and load data; The data preprocessing unit 202 is used to preprocess the multi-source data, wherein the preprocessing includes: data cleaning, data alignment and standardization, wherein the data alignment includes: time alignment and object mapping; The data quality index calculation unit 203 is used to calculate the data quality index score of the multi-source data based on the preprocessed multi-source data and in combination with the preset data quality assessment index calculation logic. The data quality index score includes: reliability impact sensitivity score and process traceability score. The data quality assessment unit 204 is used to determine the multi-source data quality assessment result based on the weighted sum of the scores of each data quality indicator and in combination with the preset data quality assessment threshold.
[0088] This embodiment incorporates two data quality assessment indicators specifically designed for power supply reliability analysis into the evaluation system. This application directly links data quality issues with the calculation results of reliability indicators, thereby matching the data quality evaluation method with the actual application scenarios of power supply reliability analysis. Specifically, the reliability impact sensitivity score quantifies the relative deviation of the original data disturbance on the reliability indicator vector, accurately reflecting the degree of impact of data quality issues on the analysis results; the process traceability score verifies the stage completeness and logical consistency of the event stage set, ensuring that fault and power outage event data meet the traceability requirements of business processes. Based on the above technical concept, the technical problem of unstable accuracy in power supply reliability analysis is effectively solved, improving the reliability and accuracy of power supply reliability analysis, thus providing more accurate data quality assurance for power system operation and management.
[0089] This application further proposes an embodiment of a multi-source data quality assessment terminal for power supply reliability analysis. For example... Figure 3 As shown, the terminal includes a memory 33 and a processor 31, which are connected via a communication bus 34. The memory 33 stores program code corresponding to the aforementioned multi-source data quality assessment method for power supply reliability analysis; the processor 31 reads and executes the program code to implement the aforementioned multi-source data quality assessment method for power supply reliability analysis.
[0090] Specifically, the multi-source data quality assessment terminal for power supply reliability analysis is a specially designed computing device whose core function is to provide a stable and efficient operating environment for the aforementioned data quality assessment methods. This terminal can be a standalone hardware device, such as an industrial control computer, a server, or an embedded system, or it can be a virtualized instance deployed on a cloud computing platform. Its main function is to receive, process, and output multi-source data quality assessment results related to power supply reliability analysis.
[0091] The memory is a hardware component in the terminal used to store digital information and instructions. It is responsible not only for long-term storage of the program code implementing the data quality assessment method, but also for temporary storage of intermediate data, evaluation results, and other necessary data generated during data processing. The memory can be volatile memory, such as random access memory (RAM), used to provide high-speed data access during program execution; or it can be non-volatile memory, such as solid-state drives (SSDs), hard disk drives (HDDs), or flash memory, used for persistent storage of the operating system, applications, and configuration data.
[0092] The processor is the central processing unit of the terminal, responsible for interpreting and executing instructions in memory and performing various computational tasks. It is the "brain" of the entire data quality assessment process, driving every step of data acquisition, preprocessing, indicator calculation, and result determination. The processor can be a general-purpose central processing unit (CPU), such as a multi-core processor, to support parallel computing and improve processing efficiency; or it can be a graphics processing unit (GPU) or a field-programmable gate array (FPGA) to provide superior performance in scenarios requiring extensive parallel computing or specific hardware acceleration.
[0093] The program code stored in the memory is a concrete implementation of the multi-source data quality assessment method for power supply reliability analysis described above. This program code transforms the various steps and logical relationships defined in the method into a computer-executable sequence of instructions. For example, the program code might include data cleaning algorithms, data alignment logic, a formula for calculating the reliability impact sensitivity score, rules for determining process traceability scores, and logic for comparing the weighted sum of the final assessment results with a threshold. This correspondence ensures that the terminal can accurately reproduce and execute the assessment method during execution.
[0094] The processor transforms abstract evaluation methods into actual computational operations by reading and executing program code from memory. Following the preset instruction sequence in the program code, the processor retrieves data from memory, performs data processing, calculations, and logical judgments, and writes the results back to memory or outputs them to other devices. This process enables data quality evaluation methods to operate automatically and efficiently without human intervention, thus providing continuous and accurate data quality support for power supply reliability analysis.
[0095] The aforementioned technical solution integrates a multi-source data quality assessment method for power supply reliability analysis into a specific terminal device, enabling the method to operate automatically and continuously without manual intervention. This significantly improves the efficiency and real-time performance of data quality assessment, ensuring that the fundamental data upon which power supply reliability analysis relies remains at an acceptable quality level. The introduction of this terminal provides robust physical and software support for data governance and reliability management in power systems, allowing data quality issues to be detected and corrected promptly. This effectively avoids deviations in power supply reliability analysis caused by data quality defects, thereby enhancing the stability and security of power system operation.
[0096] Furthermore, this application also proposes a computer-readable storage medium storing program code, which is used to be read and executed by a processor to implement the multi-source data quality assessment method for power supply reliability analysis.
[0097] Specifically, the computer-readable storage medium is a physical entity capable of storing data in digital form and accessible by a computer system or processing device. Its function is to provide persistent storage space for program code, ensuring the integrity of program logic even after a system power outage or restart. As possible implementations, this medium can be non-volatile memory, such as a hard disk drive, solid-state drive, flash memory, read-only memory (ROM), or optical disc; or it can be volatile memory, such as random access memory (RAM), temporarily storing code and data during program execution. The presence of program code indicates that the computer-readable storage medium internally stores a set of instructions for implementing specific functions. This program code is a concrete manifestation of the multi-source data quality assessment method for power supply reliability analysis, transforming abstract algorithms and logic into a sequence of instructions that a computer can recognize and execute. The program code can exist in various forms, for example, it can be compiled machine code or an executable file, source code or bytecode of an interpreted language (such as Python, Java, etc.), or stored as a script file. The program code is used by the processor to read and execute, describing the dynamic operation process of the program code. The processor is the core computing unit of a computer system, responsible for fetching instructions from the storage medium and decoding, processing, and controlling them according to their logical order. Through the processor's reading and execution, the program code stored on the medium is activated, thereby driving the computer system to complete predetermined tasks. Specifically, the processor accesses the storage medium through hardware interfaces such as the memory controller, loads program instructions and related data into its internal registers or cache, and then processes these instructions through its execution unit to perform operations such as data processing, logical judgment, and result output. The multi-source data quality assessment method for power supply reliability analysis clearly defines the final functional goal of the program code: by executing this code, it completely implements all the steps and logic of the aforementioned multi-source data quality assessment method for power supply reliability analysis. This means that the program code will cover all aspects from multi-source data acquisition, data preprocessing, calculation of data quality indicators (such as reliability impact sensitivity score, process traceability score, data availability score, and structural consistency score), to the determination of the weights of evaluation dimensions, and the final output of the multi-source data quality assessment results.
[0098] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the terminals, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0099] In the several embodiments provided in this application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units through some interfaces, and may be electrical, mechanical, or other forms.
[0100] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0101] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0102] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0103] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0104] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0105] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A multi-source data quality assessment method for power supply reliability analysis, characterized in that, include: Acquire multi-source data to be evaluated, wherein the multi-source data includes: equipment ledger data, operation measurement and status data, fault and power outage event data, and user and load data; The multi-source data is preprocessed, wherein the preprocessing includes: data cleaning, data alignment and standardization, wherein the data alignment includes: time alignment and object mapping; Based on the preprocessed multi-source data, and combined with the preset data quality assessment index calculation logic, the data quality index score of the multi-source data is calculated, wherein the data quality index score includes: reliability impact sensitivity score and process traceability score; The multi-source data quality assessment result is determined by weighting the scores of each data quality indicator and combining them with the preset data quality assessment threshold.
2. The multi-source data quality assessment method for power supply reliability analysis according to claim 1, characterized in that, The calculation method for the reliability impact sensitivity score includes: Based on any one or more types of raw data from the multi-source data, a preset perturbation term is applied to the raw data to obtain perturbed data; Based on the disturbance data and the corresponding original data, and combined with the power supply reliability index calculation logic, the corresponding first reliability index vector and second reliability index vector are calculated, wherein the first reliability index vector is the power supply reliability index vector corresponding to the original data, and the second reliability index vector is the power supply reliability index vector corresponding to the disturbance data. Calculate the relative deviation between the first reliability index vector and the second reliability index vector to determine the reliability impact sensitivity score corresponding to the multi-source data.
3. The multi-source data quality assessment method for power supply reliability analysis according to claim 1, characterized in that, The calculation method for the process traceability score includes: Based on the fault and power outage event data in the multi-source data, an event stage set corresponding to each event record is generated according to the event identifier, timestamp, and event stage label in the fault and power outage event data; The completeness of the event record is determined by the ratio of the number of recorded stages in the event stage set corresponding to the event record to the preset number of complete stages. Verify the time order of each record stage in the event stage set, determine the number of compliant stages in the event stage set that satisfy the time order constraint, and then determine the logical consistency degree corresponding to the event record based on the ratio of the number of compliant stages to the number of record stages. The process traceability score is obtained by weighting the completeness of the stage and the consistency of logic.
4. The multi-source data quality assessment method for power supply reliability analysis according to claim 1, characterized in that, The data quality index scoring also includes: data availability score and structural consistency score.
5. The multi-source data quality assessment method for power supply reliability analysis according to claim 4, characterized in that, The calculation method for the data availability score includes: The data fields of the multi-source data are validated to determine the number of valid fields, and then the field availability is determined based on the ratio of the number of valid fields to the total number of fields. Based on the completeness of each data field, the number of valid objects is counted, and then the availability of objects is determined by the ratio of the number of valid objects to the preset theoretical number of objects. A data availability score is obtained by weighting the availability of the field with the availability of the object.
6. The multi-source data quality assessment method for power supply reliability analysis according to claim 4, characterized in that, The calculation method for the structural consistency score includes: According to the preset data consistency rules, the consistency of the multi-source data is checked, and the number of violations and the number of conformities of each data record in the multi-source data are counted. The basic consistency is determined by the complement of the ratio of the number of rule violations to the number of rule compliances; Object matching is performed among various types of data in the multi-source data. The number of valid objects that are successfully matched across types is counted. The consistency of topology and mapping is determined based on the ratio of the number of valid objects to the total number of objects. The structural consistency score is obtained by weighting the basic consistency and the topology and mapping consistency.
7. The multi-source data quality assessment method for power supply reliability analysis according to claim 1, characterized in that, Before determining the multi-source data quality assessment results, based on the weighted sum of scores from various data quality indicators and a preset data quality assessment threshold, the following steps are also included: Based on the data quality index scores corresponding to each evaluation dimension, multiple evaluation dimension perturbation datasets are obtained respectively. The evaluation dimension perturbation dataset is formed by applying a perturbation term to one of the evaluation dimensions based on a preset original dataset, and the evaluation dimensions targeted by each evaluation dimension perturbation dataset do not overlap. Based on the disturbance datasets for each evaluation dimension and the original dataset, and following the power supply reliability index calculation logic, the disturbance reliability index results corresponding to the disturbance datasets for each evaluation dimension and the benchmark reliability index results corresponding to the original dataset are calculated respectively. Based on the relative deviations between the results of each disturbance reliability index and the results of the benchmark reliability index, the influence weights corresponding to each evaluation dimension are determined.
8. A multi-source data quality assessment device for power supply reliability analysis, characterized in that, include: A multi-source data acquisition unit is used to acquire multi-source data to be evaluated, wherein the multi-source data includes: equipment ledger data, operation measurement and status data, fault and power outage event data, and user and load data; A data preprocessing unit is used to preprocess the multi-source data, wherein the preprocessing includes: data cleaning, data alignment and standardization, wherein the data alignment includes: time alignment and object mapping; The data quality index calculation unit is used to calculate the data quality index score of the multi-source data based on the preprocessed multi-source data and in combination with the preset data quality assessment index calculation logic. The data quality index score includes: reliability impact sensitivity score and process traceability score. The data quality assessment unit is used to determine the multi-source data quality assessment result based on the weighted sum of the scores of each data quality indicator and a preset data quality assessment threshold.
9. A multi-source data quality assessment terminal for power supply reliability analysis, characterized in that, include: Memory and processor; The memory is used to store program code, which corresponds to the multi-source data quality assessment method for power supply reliability analysis as described in any one of claims 1 to 7; The processor is used to read and execute the program code to implement the multi-source data quality assessment method for power supply reliability analysis.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code that is read and executed by a processor to implement the multi-source data quality assessment method for power supply reliability analysis as described in any one of claims 1 to 7.