Dynamic cleaning and quality evaluation system for power multi-source heterogeneous data

By generating a dynamic cleaning and quality assessment system that integrates adversarial networks with power expertise, the problems of insufficient adaptability and collaborative processing of multi-source heterogeneous data in power data processing are solved, achieving efficient and comprehensive data quality management and improving system reliability.

CN120804083APending Publication Date: 2025-10-17GUANGDONG POWER GRID CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511142000.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing power data processing technologies lack adaptability, have separation between evaluation and cleaning, lack deep integration of professional knowledge, and have limited collaborative processing capabilities for multi-source heterogeneous data, making it difficult to effectively integrate and utilize data information from different sources.

Method used

By deeply integrating generative adversarial network (GAN) technology with professional knowledge in the power field, a dynamic cleaning and quality assessment system is constructed, including data access, metadata management, generative adversarial network, rule engine and quality assessment module, to achieve adaptive cleaning rule generation and optimization, and combine multi-dimensional evaluation standards to conduct data quality assessment.

Benefits of technology

It realizes the integrated processing of data quality assessment and repair, improves data processing efficiency, enhances adaptive capabilities, comprehensively covers all aspects of power data quality, and reduces equipment misoperation rate and system failure rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804083A_ABST
    Figure CN120804083A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power data processing, in particular to a dynamic cleaning and quality evaluation system for power multi-source heterogeneous data, which comprises a data access module for acquiring and preprocessing the power multi-source heterogeneous data. The metadata management module extracts metadata features and constructs a data quality evaluation index system, the generative adversarial network module comprises a generator and a discriminator, the generator repairs problem data, the discriminator evaluates the quality of the repaired data and feeds back the quality to the generator, and the rule engine module generates an adaptive cleaning rule based on the metadata features. The system is applied to the repairing process of the generator, the quality evaluation module evaluates the quality of the repaired data based on a multi-dimensional standard and generates an evaluation report, the data storage module stores the repaired data and the evaluation report, the system realizes integrated processing of data quality evaluation and repairing, the data processing efficiency is improved, and the data quality evaluation and repairing efficiency is improved. Compared with a traditional method, the efficiency is improved by 3-5 times, and an effective solution is provided for power multi-source heterogeneous data processing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of electric power data processing, specifically to a dynamic cleaning and quality evaluation system for electric power multi-source heterogeneous data. BACKGROUND

[0002] With the continuous development of smart grid construction, the data generated by the power system presents the characteristics of multi-source, heterogeneity and dynamic change. These data come from SCADA systems, equipment account systems, historical databases and various types of offline quasi-real-time data sources. The data format is diverse and the quality is uneven. The existing electric power data processing technology mainly has the following problems: Firstly, traditional data quality evaluation methods often rely on static rules and manual experience, lack of self-adaptive ability, and are difficult to cope with the variability of electric power data. Secondly, data evaluation and cleaning are usually two separate processes, resulting in low processing efficiency and poor cleaning effect. Thirdly, the existing technology lacks deep integration of professional knowledge in the field of electric power, and cannot fully understand and process the special properties of electric power data. Finally, the traditional method has limited ability to cooperatively process multi-source heterogeneous data, making it difficult to effectively integrate and utilize data information from different sources.

[0003] With the development of artificial intelligence technology, deep learning methods have shown great potential in the field of data processing, but how to effectively apply these technologies to electric power data quality management is still a problem to be solved. SUMMARY

[0004] The purpose of the present application is to provide a dynamic cleaning and quality evaluation system for electric power multi-source heterogeneous data, aiming to solve the problems of insufficient self-adaptive ability, separation of evaluation and cleaning, lack of deep integration of professional knowledge, and limited ability to cooperatively process multi-source data in the existing technology.

[0005] The present application proposes a dynamic cleaning and quality evaluation system for electric power multi-source heterogeneous data, comprising: a data access module for collecting electric power multi-source heterogeneous data and pre-processing the electric power multi-source heterogeneous data; a metadata management module connected to the data access module for extracting metadata features from the electric power multi-source heterogeneous data and constructing an electric power data quality evaluation index system; a generative adversarial network module connected to the metadata management module, including a generator and a discriminator, wherein: the generator is used to receive electric power multi-source heterogeneous data containing problems and output repaired electric power multi-source heterogeneous data; the discriminator is used to evaluate the quality of the repaired electric power multi-source heterogeneous data and feed back the evaluation results to the generator; a rule engine module, connected with the metadata management module and the generative adversarial network module, configured to generate adaptive cleaning rules based on the metadata features and apply the adaptive cleaning rules to the repair process of the generator; a quality assessment module, connected with the discriminator, configured to perform quality assessment on the repaired power multi-source heterogeneous data based on multi-dimensional assessment criteria and generate a quality assessment report; a data storage module, connected with the generator and the quality assessment module, configured to store the repaired power multi-source heterogeneous data and the quality assessment report.

[0006] Preferably, the metadata management module comprises: a metadata extraction unit, configured to extract metadata features of a basic layer, a relationship layer, a business layer and a quality layer from the power multi-source heterogeneous data; an index system construction unit, configured to construct a quality assessment index system of seven dimensions of integrity, accuracy, timeliness, consistency, reasonableness, standardization and uniqueness based on the metadata features; a metadata constraint generation unit, configured to generate metadata constraint conditions based on the quality assessment index system; a metadata update unit, configured to dynamically update the metadata features and the metadata constraint conditions according to the evaluation results of the quality assessment module.

[0007] Preferably, in the generative adversarial network module: the generator comprises a feature extraction layer, an anomaly detection layer, a data repair layer and a quality enhancement layer, and is configured to perform hierarchical processing on the power multi-source heterogeneous data containing problems; the discriminator comprises a true-false discrimination branch and a quality assessment branch, wherein the true-false discrimination branch is configured to judge data authenticity, and the quality assessment branch is configured to perform multi-dimensional quality assessment; the generator and the discriminator promote each other through adversarial training, and jointly improve the data repair and quality assessment capabilities.

[0008] Preferably, the rule engine module comprises: an association rule mining unit, configured to mine multi-dimensional association rules such as attribute similarity, value dependence, co-occurrence pattern and time sequence association from the power multi-source heterogeneous data; a cleaning rule generation unit, configured to generate adaptive cleaning rules based on the multi-dimensional association rules and the metadata constraint conditions; a rule optimization unit, configured to dynamically optimize the adaptive cleaning rules according to the evaluation results of the quality assessment module; a rule execution unit, configured to apply the adaptive cleaning rules to the repair process of the generator.

[0009] As preferred, the quality assessment module comprises: an integrity assessment unit for assessing the field filling rate and data coverage degree of the repaired power multi-source heterogeneous data; an accuracy assessment unit for assessing the compliance of the repaired power multi-source heterogeneous data with actual values; a timeliness assessment unit for assessing the timeliness of the repaired power multi-source heterogeneous data; a consistency assessment unit for assessing the consistency degree of the repaired power multi-source heterogeneous data between different data sources; a rationality assessment unit for assessing whether the repaired power multi-source heterogeneous data conforms to business logic and value range; a specification assessment unit for assessing whether the repaired power multi-source heterogeneous data conforms to standard specifications; a uniqueness assessment unit for assessing whether the repaired power multi-source heterogeneous data has duplicate records.

[0010] As preferred, the generative adversarial network module further comprises: a training data management unit for constructing a training data set with quality labels; a target function management unit for managing generator target functions, discriminator target functions, and quality assessment target functions; a parameter optimization unit for optimizing network parameters of the generator and the discriminator according to the target functions; a convergence control unit for stopping the training process when the target functions reach a preset convergence condition.

[0011] As preferred, the data access module comprises: a multi-source data collection unit for collecting power multi-source heterogeneous data from SCADA systems, device account systems, historical databases, and offline quasi-real-time data sources; a data standardization unit for performing format unification and encoding conversion on the power multi-source heterogeneous data; a preliminary cleaning unit for performing outlier detection and missing value processing on the power multi-source heterogeneous data; a data classification unit for classifying and storing the power multi-source heterogeneous data according to data types and sources.

[0012] As preferred, it further comprises: a feedback optimization module connected with the quality assessment module and the generative adversarial network module, for collecting evaluation results of the quality assessment module and user feedback, and optimizing parameters of the generative adversarial network module; The feedback optimization module comprises: A performance monitoring unit for monitoring the running state and performance indicators of each component of the system; A problem identification unit for identifying problems and deficiencies in the system based on the monitoring results of the performance monitoring unit; An optimization strategy generation unit for generating optimization strategies for the problems and deficiencies; A parameter adjustment unit for adjusting system parameters according to the optimization strategies to improve system performance.

[0013] As a preferred embodiment, the quality assessment module and the generator form an evaluation-repair linkage mechanism, comprising: A quality problem classification system for classifying quality problems by dimension, severity, repair difficulty and priority; A repair strategy selection mechanism for selecting the most suitable repair strategy according to the characteristics of the quality problem; A context-aware repair mechanism for making repair decisions considering the time, space, business and system context of the data; A repair effect verification mechanism for verifying the degree of quality improvement of the repaired data.

[0014] As a preferred embodiment, it further comprises: An application service module connected to the data storage module for providing data quality service interfaces for external systems; The application service module comprises: A quality report service unit for generating and publishing data quality assessment reports; A data repair service unit for receiving data repair requests from external systems and returning repair results; An API interface unit for providing standardized service interfaces to support external system calls for data quality services; A visualization unit for displaying data quality status and repair effects in chart form.

[0015] The present application integrates the generative adversarial network (GAN) technology with professional knowledge in the power field, and constructs an innovative data quality management framework. The system has the following beneficial effects: 1. The system realizes integrated processing of data quality assessment and repair, improves data processing efficiency, and improves efficiency by 3-5 times compared with traditional methods; 2. The system innovatively designs a seven-dimensional power data quality assessment framework, which comprehensively covers all aspects of power data quality; 3. Through the application of GAN network, the system has strong adaptive ability and can handle more than 95% of abnormal situations; 4. The metadata constraint-driven adaptive cleaning rule engine realizes the automatic generation and optimization of cleaning rules; 5. A deep cycle optimization learning mechanism is designed to enable the system to continuously optimize and evolve; 6. The power data quality is greatly improved, indirectly improving the reliability of the power system operation, and reducing the equipment misoperation rate and system failure rate. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 The overall structure diagram of the power multi-source heterogeneous data dynamic cleaning and quality evaluation system of the present application; Figure 2 The structure diagram of the metadata management module of the present application; Figure 3 The structure diagram of the generator and discriminator training process of the present application; Figure 4 The structure diagram of the rule engine module of the present application; Figure 5 The structure diagram of the quality evaluation module of the present application; Figure 6 The data flow process diagram of the present application; Figure 7 The structure diagram of the generator and discriminator training process of the present application; Figure 8 The system implementation application effect comparison diagram of the present application; DETAILED DESCRIPTION

[0017] Please refer to the accompanying Figures 1-8 , the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0018] Referring to Figure 1 , the power multi-source heterogeneous data dynamic cleaning and quality evaluation system provided by the present application mainly includes a data access module 1, a metadata management module 2, a generative adversarial network module 3, a rule engine module 4, a quality evaluation module 5, a data storage module 6, a feedback optimization module 7 and an application service module 8. The modules are connected with each other through data flow and control flow to form a complete data quality management system.

[0019] In a preferred embodiment of the present application, the data access module 1 is used to collect power multi-source heterogeneous data and pre-process the power multi-source heterogeneous data. The sources of power multi-source heterogeneous data are diverse, including but not limited to SCADA system data, equipment account system data, historical database data and offline quasi-real-time data, etc. The data access module 1 unifies the formats of these data and performs preliminary cleaning to lay a foundation for subsequent processing.

[0020] The metadata management module 2 is connected with the data access module 1, used for extracting metadata features from the power multi-source heterogeneous data, and constructing an index system of power data quality evaluation. Metadata is data describing data, containing information such as data structure, relationship, business rules, etc., and is an important basis for data quality evaluation.

[0021] The generative adversarial network module 3 is connected with the metadata management module 2, including a generator 31 and a discriminator 32. The generator 31 receives power multi-source heterogeneous data containing problems and outputs repaired power multi-source heterogeneous data; the discriminator 32 evaluates the quality of the repaired power multi-source heterogeneous data and feeds back the evaluation results to the generator 31. Through this adversarial learning mechanism, the system can continuously improve the ability of data repair and quality evaluation.

[0022] The rule engine module 4 is connected with the metadata management module 2 and the generative adversarial network module 3, used for generating adaptive cleaning rules based on metadata features, and applying the adaptive cleaning rules to the repair process of the generator 31. The rule engine module 4 enables the system to combine power field professional knowledge for data processing, improving the accuracy and professionalism of processing.

[0023] The quality evaluation module 5 is connected with the discriminator 32, used for quality evaluation of the repaired power multi-source heterogeneous data based on multi-dimensional evaluation standards, and generating a quality evaluation report. The evaluation results are not only used to guide data repair, but also provide data quality reference for data users.

[0024] The data storage module 6 is connected with the generator 31 and the quality evaluation module 5, used for storing the repaired power multi-source heterogeneous data and the quality evaluation report. These data and reports can be used for subsequent analysis and use.

[0025] In addition, the present application also includes a feedback optimization module 7 and an application service module 8, respectively used for system optimization and external service.

[0026] Reference Figure 2 The metadata management module 2 includes a metadata extraction unit 21, an index system construction unit 22, a metadata constraint generation unit 23 and a metadata update unit 24.

[0027] The metadata extraction unit 21 is used for extracting metadata features of the basic layer, the relationship layer, the business layer and the quality layer from the power multi-source heterogeneous data. The basic layer metadata includes basic attributes such as data type, length, value range; the relationship layer metadata includes inter-field association, foreign key relationship, combination constraint, etc.; the business layer metadata includes business rules, business processes, state conversion, etc.; and the quality layer metadata includes quality standards, historical quality performance, quality expectations, etc.

[0028] Preferably, the metadata extraction process employs multiple technical means, including database structure analysis, data mining, business system interface invocation, etc. For example, for SCADA system data, the basic layer metadata can be extracted by analyzing the database table structure; the relationship layer metadata can be formed by discovering the correlation between fields through correlation analysis algorithms; the business layer metadata can be constructed by calling the business system API to obtain business rules; and the quality layer metadata can be formed by analyzing the historical data quality evaluation results.

[0029] The index system construction unit 22 is configured to construct a quality evaluation index system of seven dimensions, including integrity, accuracy, timeliness, consistency, rationality, standardization, and uniqueness, based on the metadata characteristics. These seven dimensions comprehensively cover all aspects of power data quality, forming a complete evaluation framework.

[0030] The metadata constraint generation unit 23 is configured to generate metadata constraint conditions based on the quality evaluation index system. These constraint conditions are important basis for data cleaning and quality evaluation. For example, for substation equipment operation state data, the following constraint conditions can be generated: the equipment operation state value must be in the range of {0, 1, 2, 3}, where 0 represents shutdown, 1 represents operation, 2 represents maintenance, and 3 represents failure; the equipment operation state change must follow specific state transition rules, such as not being able to change from operation state to failure state without intermediate transition state, etc.

[0031] The metadata update unit 24 is configured to dynamically update the metadata characteristics and metadata constraint conditions according to the evaluation results of the quality evaluation module 5. This dynamic updating mechanism enables the system to continuously adapt to changes in data characteristics, maintaining long-term effectiveness.

[0032] Referring to Figure 3 The generator 31 in the generative adversarial network module 3 includes a feature extraction layer 311, an anomaly detection layer 312, a data repair layer 313, and a quality enhancement layer 314, which are used to perform hierarchical processing on power multi-source heterogeneous data containing problems.

[0033] The feature extraction layer 311 adopts a multi-layer neural network structure to extract key features of the data. For power data, feature extraction is particularly important because different types of power data have different feature patterns. For example, for substation equipment operation data, important features include equipment type, operation state, operation parameters, etc.; for line load data, important features include time distribution, load change pattern, etc.

[0034] The anomaly detection layer 312 uses various anomaly detection algorithms to identify abnormal points in the data based on the extracted features. In one embodiment of the present application, the anomaly detection employs a hybrid method based on statistics and machine learning, including the Z-score method, the Local Outlier Factor (LOF) method, and the Isolation Forest algorithm, etc. Different detection algorithms are selected for different types of anomalies. For example, for outliers, the Z-score method can be used; for density anomalies, the LOF method can be used; for complex pattern anomalies, the Isolation Forest algorithm can be used.

[0035] The data repair layer 313 adopts different repair strategies for different types of anomalies according to the anomaly detection results. For example, for missing values, a context-based interpolation method can be used; for outliers, a pattern-based replacement method can be used; for inconsistent values, a rule-based adjustment method can be used. In actual applications, the selection of repair strategies also takes into account the business characteristics and usage scenarios of the data.

[0036] The quality enhancement layer 314 further optimizes the repaired data to improve the overall data quality. Quality enhancement employs various technical means, including data smoothing, noise filtering, consistency checking, etc. For example, for time series data, a sliding window average method can be used for smoothing; for multi-source data, consistency rules can be used for checking and adjusting.

[0037] The discriminator 32 includes a true-false discrimination branch 321 and a quality evaluation branch 322. The true-false discrimination branch 321 is used to determine whether the data is real data, which is the basic function of the GAN network; the quality evaluation branch 322 is used to evaluate the data quality in multiple dimensions, which is one of the innovations of the present application.

[0038] During the training process, the generator 31 and the discriminator 32 promote each other through adversarial training, and jointly improve the data repair and quality evaluation capabilities. Specifically, the discriminator 32 continuously improves the ability to distinguish between real and fake data and evaluate data quality, while the generator 31 continuously improves the ability to generate high-quality repair data. This adversarial mechanism enables the system to learn and improve autonomously without explicit programming.

[0039] The objective function used in the training process includes the generator loss function, the discriminator loss function, and the quality evaluation loss function. These three loss functions together constitute the optimization objective of the system, and by balancing the relationship between the three, the optimal performance of the system is achieved.

[0040] The generator loss function can be represented as: , wherein, represents the repair output of the generator for the input data Z, represents the score of the discriminator on the repaired data, IE represents the expected value operator for calculating the average value of a random variable, represents the input data distribution, describing the probability distribution characteristics of the input data. This loss function measures the authenticity of the data generated by the generator, and the smaller the value, the closer the generated data is to the real data.

[0041] The discriminator loss function can be represented as: , wherein, represents the real data sample, represents the real data distribution, describing the probability distribution characteristics of the real data. This loss function measures the ability of the discriminator to distinguish between true and false data, and the smaller the value, the stronger the discrimination ability.

[0042] The quality evaluation loss function can be represented as: , wherein, represents the score of the repaired data on the i-th quality dimension, represents the target score of the i-th dimension, represents the weight of the i-th dimension, and the value range is , and . This loss function measures the quality level of the repaired data, and the smaller the value, the closer the quality is to the target requirement.

[0043] The comprehensive objective function is: , wherein, and are balance factors for adjusting the relative importance of each loss function. In practical applications, these parameters can be adjusted according to the specific data characteristics and business requirements. For example, for key business data with high quality requirements, the value of can be increased to strengthen the impact of quality evaluation; for complex data distribution scenarios, the value of can be increased to strengthen the training of the discriminator.

[0044] In addition, the generative adversarial network module 3 also includes a training data management unit 33, an objective function management unit 34, a parameter optimization unit 35, and a convergence control unit 36.

[0045] The training data management unit 33 is configured to build a training data set with quality labels. The training data includes positive samples (high-quality data) and negative samples (problem data), each of which has a quality label of seven dimensions. For example, for a substation equipment operation data sample, its quality label can be: integrity 0.95, accuracy 0.92, timeliness 0.98, consistency 0.90, rationality 0.93, standardization 0.97, and uniqueness 1.0.

[0046] The objective function management unit 34 is configured to manage the generator objective function, the discriminator objective function, and the quality evaluation objective function. In different training stages, the weights of the objective functions can be dynamically adjusted to achieve optimal training effect.

[0047] The parameter optimization unit 35 is configured to optimize the network parameters of the generator 31 and the discriminator 32 according to the objective functions. The optimization adopts a gradient descent type algorithm, such as the Adam optimizer, and the learning rate is initially set to 0.0002 and a dynamic adjustment strategy is adopted.

[0048] The convergence control unit 36 is configured to stop the training process when the objective function reaches a preset convergence condition. The convergence condition includes loss function value stability, evaluation index reaching the standard, and repair effect meeting the requirements, etc. For example, when the loss function changes less than 0.001 for 10 consecutive training rounds, and the quality scores of the seven dimensions are all above 0.9, the training is considered to be converged.

[0049] Reference Figure 4 The rule engine module 4 includes an association rule mining unit 41, a cleaning rule generation unit 42, a rule optimization unit 43, and a rule execution unit 44.

[0050] The association rule mining unit 41 is configured to mine multi-dimensional association rules such as attribute similarity, value dependence, co-occurrence pattern, and time sequence association from power multi-source heterogeneous data. Association rule mining is the basis of the rule engine, which provides the association knowledge between data.

[0051] In an embodiment of the present application, the attribute similarity analysis is based on field name, data type, and value characteristics for similarity calculation. For example, for two fields A and B, the similarity can be represented as: , Wherein, represents the name similarity, which measures the similarity of the field name, represents the type similarity, which measures the similarity of the field data type, represents the value similarity, which measures the similarity of the field value characteristics, , and are weight coefficients, satisfying In power data processing, it is common to set , , to highlight the importance of value characteristics.

[0052] Value dependency analysis is used to discover functional and conditional dependencies between fields. For example, in substation equipment data, there is a conditional dependency between equipment operating status and equipment maintenance records: when the equipment maintenance record shows that the equipment is under maintenance, the equipment operating status should be in the maintenance state.

[0053] Co-occurrence pattern analysis is used to identify field co-occurrence patterns and rules in data. For example, certain equipment parameters usually change simultaneously or maintain a certain relationship. In power systems, load increase is usually accompanied by voltage drop and line loss increase, and this co-occurrence pattern can be used for data verification and repair.

[0054] Time series correlation analysis is used to discover correlation patterns in time series data. For example, electricity consumption load data usually presents obvious intra-day, intra-week and seasonal patterns, which can be used for anomaly detection and data repair.

[0055] The cleaning rule generation unit 42 is used to generate adaptive cleaning rules based on multi-dimensional association rules and metadata constraint conditions. Cleaning rules include two parts: verification rules and repair rules. Verification rules are used to identify data problems, and repair rules are used to solve data problems.

[0056] In the rule generation process, the system first identifies the type of quality problem in the data, then selects the appropriate rule template according to the problem type, optimizes the rule parameters according to the data characteristics, and finally combines multiple basic rules into a composite rule. For example, for the missing value problem in power equipment operating data, the following rules can be generated: Verification rule: IF equipment operating status IS NULL THEN mark as missing value problem Repair rule: IF equipment operating status IS NULL AND last time status = running AND next time status = running AND latest maintenance record time > current time THEN set status to running The rule optimization unit 43 is used to dynamically optimize adaptive cleaning rules according to the evaluation results of the quality evaluation module 5. The optimization process includes rule parameter adjustment, rule structure optimization and rule set update, etc. For example, if a repair rule leads to a decrease in data consistency, the system will automatically adjust the parameters or structure of the rule, or even remove the rule.

[0057] The rule execution unit 44 is configured to apply the adaptive cleaning rules to the repair process of the generator 31. The rule execution adopts the reasoning engine technology, which supports efficient execution of complex rules. During the execution process, the system records the rule application and the effect, which provides the basis for subsequent optimization.

[0058] With reference to Figure 5 The quality assessment module 5 includes an integrity assessment unit 51, an accuracy assessment unit 52, a timeliness assessment unit 53, a consistency assessment unit 54, a rationality assessment unit 55, a normative assessment unit 56, and a uniqueness assessment unit 57.

[0059] The integrity assessment unit 51 is configured to assess the field filling rate and data coverage degree of the repaired power multi-source heterogeneous data. The integrity assessment focuses on whether the data is missing, which is the basic guarantee of data quality.

[0060] In an embodiment of the present application, the integrity score can be represented as: , Wherein, represents the number of filled fields, indicating the number of non-empty fields in the data record, represents the total number of fields, indicating the number of all fields in the data record, represents the number of covered data items, indicating the number of actually collected data items, represents the number of expected covered data items, indicating the number of data items that should be collected according to business requirements, and are weight coefficients, satisfying In practical applications, usually , to highlight the importance of data coverage.

[0061] The accuracy assessment unit 52 is configured to assess the conformity of the repaired power multi-source heterogeneous data with the actual value. The accuracy assessment focuses on the authenticity and correctness of the data, which is the core indicator of data quality.

[0062] The accuracy score can be represented as: , Wherein, represents the actual value of the data item, represents the reference value or standard value, n represents the number of data items, and R represents the data value range, i.e. the maximum possible variation range of the data. For data that cannot obtain reference values, indirect assessment methods based on rules and patterns can be used.

[0063] The timeliness evaluation unit 53 is configured to evaluate the timeliness of the repaired power multi-source heterogeneous data. The timeliness evaluation focuses on the update frequency and delay of the data, and is particularly important for scenarios with high real-time requirements.

[0064] The timeliness score can be represented as: , wherein, denotes the current time, denotes the data update time, denotes the timeliness decay coefficient, which is set according to business requirements. For example, for substation real-time monitoring data, the timeliness decay coefficient can be set as , which means that the timeliness score decreases by about 63% for every 10-minute delay.

[0065] The consistency evaluation unit 54 is configured to evaluate the consistency of the repaired power multi-source heterogeneous data between different data sources. The consistency evaluation focuses on the internal consistency of the data and the consistency with other data sources.

[0066] The consistency score can be represented as: , wherein, denotes the number of consistent data items, which refers to the number of data items meeting the consistency rules, denotes the total number of data items, which refers to the number of all data items that need to be checked for consistency. The consistency judgment is based on pre-defined consistency rules, such as inter-field relationships, data source mapping, etc.

[0067] The rationality evaluation unit 55 is configured to evaluate whether the repaired power multi-source heterogeneous data meets the business logic and value range. The rationality evaluation focuses on the business rationality and value rationality of the data.

[0068] The rationality score can be represented as: ,

[0069] wherein, denotes the number of rational data items, which refers to the number of data items meeting the business rules and value range, denotes the total number of data items, which refers to the number of all data items that need to be checked for rationality. The rationality judgment is based on business rules and data distribution characteristics. For example, the load of a substation should not exceed 120% of its design capacity, and exceeding this value will be judged as irrational.

[0070] The specification evaluation unit 56 is configured to evaluate whether the repaired power multi-source heterogeneous data meets the standard specifications. The specification evaluation focuses on whether the data format, encoding, unit, etc. meet the predetermined standards.

[0071] The specification score can be represented as: , wherein, represents the number of data items meeting the standard, indicating the number of data items meeting the data standard specification, represents the total number of data items, indicating the number of all data items that need to be checked for specification. The specification judgment is based on data standard specifications, such as field naming rules, data format standards, etc.

[0072] The uniqueness evaluation unit 57 is used to evaluate whether the repaired power multi-source heterogeneous data has duplicate records. The uniqueness evaluation focuses on the redundancy of the data.

[0073] The uniqueness score can be represented as: , wherein, represents the number of duplicate data items, indicating the number of data items identified as duplicates, represents the total number of data items, indicating the number of all data items that need to be checked for uniqueness. The uniqueness judgment is based on uniqueness constraints, such as primary key uniqueness, business key uniqueness, etc.

[0074] The quality evaluation module 5 calculates the comprehensive quality score based on the scores of the above seven dimensions: , wherein, represents the score of the i-th dimension, with a value range of , represents the weight of the i-th dimension, satisfying The weight setting should be adjusted according to the specific application scenario and business requirements. For example, for real-time monitoring data, the timeliness weight can be set higher; for historical statistical data, the accuracy and completeness weights can be set higher.

[0075] The data access module 1 includes a multi-source data collection unit 11, a data standardization unit 12, a preliminary cleaning unit 13, and a data classification unit 14.

[0076] The multi-source data collection unit 11 is used to collect power multi-source heterogeneous data from SCADA systems, equipment account systems, historical databases, and offline quasi-real-time data sources. Data collection supports multiple ways, including database connection, API call, file import, etc.

[0077] In an embodiment of the present application, the SCADA system data is collected in real time, with a collection frequency of 1 second; the equipment account system data is collected in a change triggered manner, triggering collection when the equipment information changes; the historical database data is imported in batches, imported daily at a fixed time; the offline quasi-real-time data is polled in a timing polling manner, polled once every minute.

[0078] The data standardization unit 12 is used to unify the format and convert the encoding of the multi-source heterogeneous power data. Standardization is the basis of data integration, ensuring that data from different sources can be processed uniformly.

[0079] Standardization includes standardizing field names, converting data types, and unifying coding methods. For example, the fields representing device status in different systems are uniformly named "deviceStatus," and the status codes are standardized into numeric codes (0 - out of service, 1 - operating, 2 - maintenance, 3 - fault).

[0080] The preliminary cleaning unit 13 is used to detect outliers and process missing values ​​in the multi-source heterogeneous power data. Preliminary cleaning is the first step to improve data quality and address obvious data issues.

[0081] Outlier detection uses a combination of statistical methods and rule-based methods. For example, for numerical data, you can use Detect outliers based on the principle; for categorical data, a predefined list of valid values ​​can be used to detect outliers.

[0082] Missing values ​​can be handled using a variety of strategies, including deletion, filling, and ignoring. The specific strategy chosen depends on the data characteristics and business needs. For example, missing values ​​in key business fields can be filled using model predictions based on historical data; missing values ​​in non-critical fields can be filled using simple means or modes.

[0083] The data classification unit 14 is used to classify and store the multi-source heterogeneous power data according to data type and source. Classified storage facilitates subsequent targeted processing.

[0084] Classification dimensions include data type (e.g., time series data, event data, status data, etc.) and data source (e.g., SCADA system, equipment inventory system, etc.). For example, time series data from a SCADA system and status data from an equipment inventory system are stored separately, facilitating the application of different processing strategies to each.

[0085] The feedback optimization module 7 is connected to the quality assessment module 5 and the generative adversarial network module 3, and is used to collect the evaluation results and user feedback of the quality assessment module 5 and optimize the parameters of the generative adversarial network module 3.

[0086] The feedback optimization module 7 includes a performance monitoring unit 71 , a problem identification unit 72 , an optimization strategy generation unit 73 and a parameter adjustment unit 74 .

[0087] The performance monitoring unit 71 is configured to monitor the running state and performance indicators of each component of the system. The monitoring indicators include processing efficiency, resource usage, quality evaluation results, etc. For example, the repair accuracy of the generator 31, the discrimination accuracy of the discriminator 32, the processing time of the system as a whole, etc.

[0088] The problem identification unit 72 is configured to identify problems and deficiencies in the system based on the monitoring results of the performance monitoring unit 71. The problem identification adopts a method combining threshold judgment and trend analysis. For example, when the repair accuracy is lower than 85% or decreases by more than 5% for 3 consecutive days, it is determined that there is a problem of insufficient repair capability.

[0089] The optimization strategy generation unit 73 is configured to generate optimization strategies for problems and deficiencies. The optimization strategies are generated based on a preset strategy library and historical experience. For example, for the problem of insufficient repair accuracy, strategies such as increasing training data, adjusting network structure, or optimizing loss function can be generated.

[0090] The parameter adjustment unit 74 is configured to adjust system parameters according to the optimization strategies to improve system performance. Parameter adjustment supports both automatic adjustment and manual intervention. For example, automatically adjust parameters such as network learning rate and batch size; set specific loss function weights or network structure parameters by experts.

[0091] An evaluation-repair linkage mechanism is formed between the quality evaluation module 5 and the generator 31, including a quality problem classification system 91, a repair strategy selection mechanism 92, a context-aware repair mechanism 93, and a repair effect verification mechanism 94.

[0092] The quality problem classification system 91 is configured to classify quality problems according to dimensions, severity, repair difficulty, and priority. The classification results guide the selection and priority order of repair strategies.

[0093] In an embodiment of the present application, the quality problems are classified according to dimensions, severity, repair difficulty, and priority. The dimensions include integrity problems, accuracy problems, timeliness problems, consistency problems, reasonableness problems, specification problems, and uniqueness problems; the severity includes three levels of severe, medium, and slight; the repair difficulty includes three levels of high, medium, and low; and the priority includes four levels of urgent, high, medium, and low.

[0094] For example, the absence of key equipment state in the substation equipment state data can be classified as an integrity problem, a severe level, a medium repair difficulty, and an urgent priority. Such classification helps the system to judge the urgency and difficulty of repair and reasonably allocate resources.

[0095] The repair strategy selection mechanism 92 is configured to select the most suitable repair strategy according to the characteristics of the quality problem. The selection of the repair strategy considers factors such as problem type, data characteristics, and context information.

[0096] The repair strategy library contains various predefined strategies, such as rule-based repair, model-based repair, historical data-based repair, etc. For example, for missing values in time series data, a time series model-based repair strategy can be selected; for incorrect values in classification data, a rule-based repair strategy can be selected.

[0097] The context-aware repair mechanism 93 is used to make repair decisions considering the time, space, business, and system context of the data. Context information provides an important reference for repair, improving the accuracy of repair.

[0098] The temporal context includes the timestamp, historical trends, etc. of the data; the spatial context includes the spatial distribution, adjacent area situation, etc. of the data; the business context includes the business environment, business process, etc. in which the data is located; and the system context includes the role of the data in the system, the relationship with other data, etc.

[0099] For example, when repairing the status data of a substation equipment, the system considers the historical state change of the equipment (temporal context), the status of associated equipment (spatial context), the current business operation (business context), and the system operation mode (system context), and comprehensively judges the most reasonable repair value.

[0100] The repair effect verification mechanism 94 is used to verify the degree of quality improvement of the repaired data. The verification result is used to evaluate the effectiveness of the repair strategy and guide subsequent optimization.

[0101] The verification methods include quality score comparison, business rule verification, and expert review, etc. For example, compare the quality score changes before and after repair; check whether the repaired data meets all business rules; invite experts to sample and review the repair effect.

[0102] The application service module 8 is connected with the data storage module 6, and is used to provide a data quality service interface for external systems.

[0103] The application service module 8 includes a quality report service unit 81, a data repair service unit 82, an API interface unit 83, and a visualization unit 84.

[0104] The quality report service unit 81 is used to generate and publish data quality evaluation reports. The report content includes quality score, problem distribution, repair suggestion, etc., providing data quality reference for data users.

[0105] In an embodiment of the present application, the quality report includes overall quality score, sub-item score of seven dimensions, problem distribution statistics, quality trend analysis, and improvement suggestions, etc. The report supports multiple formats such as PDF, HTML, and Excel, etc., meeting the needs of different scenarios.

[0106] The data repair service unit 82 is used to receive data repair requests from external systems and return repair results. This service enables the data repair capabilities of the system to be called by external systems, expanding the application scope of the system.

[0107] The service interface supports both batch repair and real-time repair modes. Batch repair is suitable for repairing large-scale historical data, with longer processing time but high throughput; real-time repair is suitable for repairing a small amount of real-time data, with short response time but small throughput.

[0108] The API interface unit 83 is used to provide standardized service interfaces to support external systems calling data quality services. The interface uses the RESTful style and supports HTTP / HTTPS protocols, making it easy to integrate and call.

[0109] The main interfaces include data quality assessment interface, data repair interface, rule management interface, and report generation interface, etc. Each interface has detailed parameter descriptions and calling examples to facilitate developers.

[0110] The visualization unit 84 is used to display data quality status and repair effects in chart form. Visualization is intuitive and clear, making it easy for users to understand and analyze data quality.

[0111] Visualization content includes quality score dashboard, problem distribution pie chart, quality trend line chart, and repair effect comparison chart, etc. Interactive operations such as data filtering, view switching, and drilling down are supported to improve user experience.

[0112] Referring to Figure 6 The workflow of the system mainly includes data access, metadata management, rule generation, data evaluation and repair, result storage, and service provision steps.

[0113] First, the data access module 1 collects power multi-source heterogeneous data from multiple data sources and performs preprocessing. The preprocessed data is fed into the metadata management module 2 for metadata extraction and index system construction, and the generator 31 for quality evaluation and repair.

[0114] The metadata features extracted by the metadata management module 2 and the index system constructed are passed to the rule engine module 4 for generating adaptive cleaning rules. These rules are applied to the repair process of the generator 31 to improve the accuracy and professionalism of the repair.

[0115] The generator 31 repairs the data containing problems and outputs the repaired data. The discriminator 32 evaluates the repaired data, and the evaluation results are fed back to the generator 31 for optimizing the repair process, and are also passed to the quality evaluation module 5 for multi-dimensional quality evaluation.

[0116] The quality assessment module 5 conducts a comprehensive assessment of the repaired data based on seven-dimensional assessment criteria, generating a quality assessment report. The assessment results are stored in the data storage module 6 and fed back to the feedback optimization module 7 for system optimization.

[0117] The data storage module 6 stores the repaired data and quality assessment reports, providing data support for subsequent use. The application service module 8 provides various service interfaces for external systems based on stored data and reports, expanding the application value of the system.

[0118] The entire workflow forms a closed-loop system, continuously improving the system's data processing capabilities and quality management levels through continuous data flow and feedback optimization.

[0119] Referring to Figure 8 The system has achieved remarkable results in practical applications. Compared with traditional methods, the system has significantly improved in data quality assessment accuracy, data repair efficiency, system adaptability, and operation and maintenance costs.

[0120] In terms of data quality assessment accuracy, the system achieves more than 95%, while traditional methods only achieve 70-75%. In particular, when dealing with complex abnormal situations, the system's advantage is more obvious.

[0121] In terms of data repair efficiency, the system's processing speed is 3-5 times that of traditional methods, greatly improving data processing efficiency. For example, processing 10GB of power data, traditional methods take about 2 hours, while the system only takes about 30 minutes.

[0122] In terms of system adaptability, the system can handle more than 95% of abnormal situations, while traditional methods usually only cover 60-70%. This adaptability enables the system to handle various complex and changing data situations.

[0123] In terms of operation and maintenance costs, the system reduces manual intervention through automated data assessment and repair, reducing operation and maintenance costs by more than 50%. At the same time, high-quality data also reduces downstream system abnormalities and failures, indirectly reducing overall operation and maintenance costs.

[0124] In addition, the application of the system has also improved the reliability of the power system operation. After the system is deployed, the device misoperation rate is reduced by 60%, and the system operation failure rate is reduced by 40%, providing a strong guarantee for the safe and stable operation of the power system.

[0125] In summary, the dynamic cleaning and quality assessment system for power multi-source heterogeneous data provided by the present application solves multiple key problems in the field of power data quality management through innovative technical solutions, and has important application value and promotion prospects.

[0126] The above description is only the preferred embodiment of the present application, and is not intended to limit the present application. The present application can have various changes and modifications for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. Dynamic cleaning and quality assessment system for multi-source heterogeneous power data, characterized by: include: A data access module is used to collect and pre-process the heterogeneous data from multiple sources of power; A metadata management module, connected to the data access module, for extracting metadata features from the electric power multi-source heterogeneous data and constructing an electric power data quality assessment indicator system; A generative adversarial network module is connected to the metadata management module and includes a generator and a discriminator, wherein: The generator is used to receive the electric power multi-source heterogeneous data containing problems and output the repaired electric power multi-source heterogeneous data; The discriminator is used to perform quality assessment on the repaired electric power multi-source heterogeneous data and feed back the assessment result to the generator; a rule engine module, connected to the metadata management module and the generative adversarial network module, for generating adaptive cleaning rules based on the metadata features and applying the adaptive cleaning rules to the repair process of the generator; a quality assessment module, connected to the discriminator, for performing a quality assessment on the repaired power multi-source heterogeneous data based on a multi-dimensional assessment standard and generating a quality assessment report; A data storage module is connected to the generator and the quality assessment module, and is used to store the repaired power multi-source heterogeneous data and the quality assessment report.

2. The dynamic cleaning and quality assessment system for electric power multi-source heterogeneous data according to claim 1 is characterized in that: The metadata management module includes: A metadata extraction unit, configured to extract metadata features of a base layer, a relationship layer, a business layer, and a quality layer from the multi-source heterogeneous power data; An indicator system construction unit, used to construct a quality assessment indicator system in seven dimensions: completeness, accuracy, timeliness, consistency, rationality, standardization, and uniqueness based on the metadata characteristics; A metadata constraint generating unit, configured to generate metadata constraint conditions based on the quality assessment indicator system; The metadata updating unit is configured to dynamically update the metadata features and the metadata constraints according to the evaluation result of the quality evaluation module.

3. The dynamic cleaning and quality assessment system for multi-source heterogeneous power data according to claim 1 is characterized in that: In the generative adversarial network module: The generator includes a feature extraction layer, an anomaly detection layer, a data repair layer and a quality enhancement layer, and is used to perform hierarchical processing on the multi-source heterogeneous power data containing problems; The discriminator includes a true-false discrimination branch and a quality assessment branch, wherein the true-false discrimination branch is used to judge the authenticity of the data, and the quality assessment branch is used for multi-dimensional quality assessment; The generator and the discriminator promote each other through adversarial training, and jointly improve data restoration and quality assessment capabilities.

4. The dynamic cleaning and quality assessment system for electric power multi-source heterogeneous data according to claim 1 is characterized in that: The rule engine module includes: An association rule mining unit, configured to mine multi-dimensional association rules such as attribute similarity, value dependency, co-occurrence pattern, and time series association from the multi-source heterogeneous power data; A cleaning rule generating unit, configured to generate an adaptive cleaning rule based on the multidimensional association rule and the metadata constraint condition; a rule optimization unit, configured to dynamically optimize the adaptive cleaning rules according to the evaluation result of the quality evaluation module; A rule execution unit is used to apply the adaptive cleaning rule to the repair process of the generator.

5. The dynamic cleaning and quality assessment system for electric power multi-source heterogeneous data according to claim 1 is characterized in that: The quality assessment module includes: an integrity assessment unit, configured to assess the field fill rate and data coverage of the repaired electric power multi-source heterogeneous data; an accuracy evaluation unit, configured to evaluate the conformity of the repaired multi-source heterogeneous power data with the actual value; A timeliness evaluation unit, configured to evaluate the timeliness of the repaired power multi-source heterogeneous data; A consistency evaluation unit, configured to evaluate the consistency of the repaired electric power multi-source heterogeneous data between different data sources; A rationality evaluation unit, configured to evaluate whether the repaired electric power multi-source heterogeneous data complies with business logic and value range; a normative evaluation unit, configured to evaluate whether the repaired electric power multi-source heterogeneous data complies with standard specifications; The uniqueness evaluation unit is used to evaluate whether there are duplicate records in the repaired power multi-source heterogeneous data.

6. The dynamic cleaning and quality assessment system for electric power multi-source heterogeneous data according to claim 3 is characterized in that: The generative adversarial network module also includes: Training data management unit, used to build training datasets with quality labels; An objective function management unit, used to manage the generator objective function, the discriminator objective function, and the quality assessment objective function; A parameter optimization unit, configured to optimize the network parameters of the generator and the discriminator according to the objective function; The convergence control unit is used to stop the training process when the objective function reaches the preset convergence condition.

7. The dynamic cleaning and quality assessment system for electric power multi-source heterogeneous data according to claim 1 is characterized in that: The data access module includes: Multi-source data acquisition unit, used to collect multi-source heterogeneous power data from SCADA systems, equipment inventory systems, historical databases, and offline quasi-real-time data sources; A data standardization unit, configured to perform format unification and code conversion on the multi-source heterogeneous power data; a preliminary cleaning unit, configured to perform outlier detection and missing value processing on the multi-source heterogeneous power data; The data classification unit is used to classify and store the multi-source heterogeneous power data according to data type and source.

8. The dynamic cleaning and quality assessment system for electric power multi-source heterogeneous data according to claim 1 is characterized in that: Also includes: A feedback optimization module, connected to the quality assessment module and the generative adversarial network module, for collecting the evaluation results and user feedback of the quality assessment module and optimizing the parameters of the generative adversarial network module; The feedback optimization module includes: Performance monitoring unit, used to monitor the operating status and performance indicators of each system component; a problem identification unit, configured to identify problems and deficiencies in the system based on the monitoring results of the performance monitoring unit; An optimization strategy generating unit, configured to generate an optimization strategy for the aforementioned problems and deficiencies; The parameter adjustment unit is used to adjust system parameters according to the optimization strategy to improve system performance.

9. The dynamic cleaning and quality assessment system for electric power multi-source heterogeneous data according to claim 1 is characterized in that: The quality assessment module and the generator form an assessment-repair linkage mechanism, including: A quality problem classification system, which is used to classify quality problems by dimension, severity, difficulty of repair, and priority; Repair strategy selection mechanism, used to select the most suitable repair strategy based on the characteristics of the quality problem; Context-aware repair mechanisms that consider the temporal, spatial, business, and system context of data to make repair decisions; The repair effect verification mechanism is used to verify the quality improvement of the repaired data.

10. The dynamic cleaning and quality assessment system for electric power multi-source heterogeneous data according to claim 1 is characterized in that: Also includes: An application service module, connected to the data storage module, for providing a data quality service interface for an external system; The application service module includes: Quality reporting service unit, used to generate and publish data quality assessment reports; The data repair service unit is used to receive data repair requests from external systems and return repair results; API interface unit, used to provide standardized service interfaces and support external systems to call data quality services; The visualization unit is used to display the data quality status and repair effect in the form of charts.

Citation Information

Cited By

  • Method and device for automatically normalizing road network data

    CN121563325A