Power enterprise production data quality analysis and optimization method and system
By preprocessing, quantitative evaluation and repairing abnormal data of power enterprises, the production data quality problems of power enterprises are solved, and data quality improvement and production process optimization are achieved.
Patent Information
- Application Number
- CN202411907312.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-06
AI Technical Summary
Production data quality problems in power enterprises are common, resulting in distortion of information in production links, affecting decision-making efficiency and production optimization. The existing technology lacks comprehensive and systematic evaluation and optimization methods, making it difficult to accurately locate and quantify data quality problems.
Provides a method for the quality analysis and optimization of production data of power enterprises, including data preprocessing, quantitative evaluation, abnormal data identification and repair, quantitative evaluation through the accuracy, completeness, timeliness and consistency of data, identify and repair missing, errors and redundant data, and display evaluation scores and optimization results through visual tools.
It effectively improves data quality, improves the efficiency of production processes, reduces the probability of failures, enhances the convenience and visibility of data management, and helps enterprises discover potential system or operational problems and make improvements.
Smart Images

Figure CN119938656A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data quality management, and in particular to a method and system for analyzing and optimizing production data quality of an electric power enterprise. Background Art
[0002] In the daily production management of power companies, a large amount of data related to equipment operation, production planning, maintenance records, quality inspection, etc. is involved. These data are crucial to the company's operational efficiency, energy utilization, equipment maintenance strategy and production safety. However, the production data of power companies usually comes from multiple systems. Due to different collection sources, complex data structure and possible failures during transmission, data quality problems are relatively common. Low data quality will lead to information distortion in the production process, which in turn affects decision-making efficiency and production optimization. The current data quality optimization lacks a comprehensive and systematic evaluation and optimization method, which makes it difficult to accurately locate and quantify data quality problems, affecting data management and problem tracing. Summary of the invention
[0003] In view of the deficiencies of the prior art, the present invention provides a method for analyzing and optimizing production data quality of an electric power enterprise.
[0004] The present invention provides a method for analyzing and optimizing the quality of production data of an electric power enterprise, comprising the following steps:
[0005] S1: Obtaining production data of each production link of the power enterprise and preprocessing the production data;
[0006] S2: Quantitatively evaluate the accuracy, completeness, timeliness and consistency of the production data to obtain a comprehensive evaluation score;
[0007] S3: within a preset time range, if the comprehensive evaluation score of the production data is less than a preset threshold, the production data is identified, abnormal data types in the production data are identified, including missing data, erroneous data and redundant data, and the positions of the abnormal data types are marked, and the factors causing the abnormal data are determined according to the positions of the abnormal data;
[0008] S4: Optimizing and repairing the abnormal data type;
[0009] S5: The optimized production data is stored in the database, and the comprehensive evaluation score of the production data and the optimization results of the production data are displayed through visualization tools.
[0010] Preferably, in step S1, the production data is preprocessed, specifically, a data cleaning algorithm is used to identify and remove noise data; according to preset business data rules, data irrelevant to the production link is eliminated, and data related to the production link is retained.
[0011] Preferably, in the step S2, specifically, according to the preset accuracy evaluation standard value, the accuracy score is obtained by calculating the deviation value between the production data and the preset accuracy evaluation standard value; the completeness score is obtained by counting the proportion of missing values in the production data; the timeliness score is obtained by calculating the interval between the production data generation time and the current time; the consistency score is obtained by detecting the logical consistency of the data in different links; the accuracy score, completeness score, timeliness score and consistency score are weightedly summed to obtain the comprehensive evaluation score of the production data.
[0012] Preferably, in step S3, abnormal data types in the production data are identified, including missing data, erroneous data and redundant data. Specifically, within a preset time range, if the comprehensive evaluation score of the production data is less than a preset threshold, the abnormal data is identified; missing data in the production data is identified through rule matching, including records with empty or unfilled data fields; erroneous data in the production data is identified through data verification rules, including data with values exceeding a preset range, data that does not conform to the format, or data with logical conflicts; redundant data in the production data is identified through repeatability detection, including duplicate records or redundant fields; an abnormal data report is generated, which includes the type, location, factors causing the abnormal data, and quantity statistics of the abnormal data.
[0013] Preferably, in step S4, specifically, if the identified abnormal data type is missing data, an interpolation algorithm is used to repair it and a substitute value is generated to fill the missing data; if the identified abnormal data type is erroneous data, rule verification is used to repair it; if the identified abnormal data type is redundant data, data fusion is used for optimization; a data optimization report is generated based on the optimization and repair results of the abnormal data, and the report includes a comparison of data before and after optimization.
[0014] Preferably, in step S5, specifically, the optimized production data is stored in a database; the comprehensive evaluation score of the production data is displayed through a bar chart; the production data optimization result is displayed through a comparison chart before and after the abnormal data is repaired, and an interactive interface is provided to allow the user to filter and view specific optimization results according to the time range, production link and abnormal data type.
[0015] The present invention provides a power enterprise production data quality analysis and optimization system, which is characterized by comprising the following:
[0016] Data processing module: used to obtain production data of each production link of the power enterprise and pre-process the production data;
[0017] Data quality assessment module: used to connect with the data processing module to quantitatively assess the accuracy, completeness, timeliness and consistency of the production data to obtain a comprehensive assessment score;
[0018] Data type identification module: used to connect with the data quality assessment module, and within a preset time range, if the comprehensive assessment score of the production data is less than a preset threshold, identify the production data, identify the abnormal data type in the production data, including missing data, erroneous data and redundant data, mark the location of the abnormal data type, and determine the factors causing the abnormal data according to the location of the abnormal data;
[0019] Abnormal data repair and optimization module: used to connect with the data type identification module to optimize and repair the abnormal data type;
[0020] Visualization module: used to connect with the data processing module, the data type identification module, the abnormal data repair and optimization module and the data type identification module, store the optimized production data in the database, and display the comprehensive evaluation score of the production data and the production data optimization results through visualization tools.
[0021] The present invention discloses a method for analyzing and optimizing the quality of production data of an electric power enterprise, which has the following beneficial effects: quantitative evaluation is performed through the four dimensions of data accuracy, completeness, timeliness and consistency, and abnormal data such as missing, erroneous and redundant data are identified through preset thresholds, thereby effectively improving the quality of data. By analyzing the type, location and generation factors of abnormal data, including equipment failure, sensor problems, manual input errors, etc., it provides good important information for enterprises to solve problems, effectively improves the efficiency of production processes, and reduces the probability of failures; users can filter time ranges, production links and abnormal data types through an interactive interface, dynamically view the latest evaluation and optimization results, and further enhance the convenience and visibility of data management. It is applicable to all aspects of production data of electric power enterprises, including data analysis and optimization such as equipment operation, production planning, quality inspection, etc., and has strong versatility. Through data preprocessing, abnormal identification and repair, the method can effectively serve the data quality management and production process optimization of enterprises, and improve the overall production management level; through accurate data quality evaluation and problem repair, staff can make decisions based on reliable production data, and improve the scientificity and accuracy of decisions. At the same time, the analysis report of abnormal data helps enterprises discover potential system or operation problems, so as to make targeted improvements. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0023] Figure 1 A method flow chart of a method for analyzing and optimizing production data quality of a power enterprise provided by the present invention; DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0025] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0026] refer to Figure 1 The present invention provides a method for analyzing and optimizing the production data quality of an electric power enterprise, which is characterized by comprising the following steps:
[0027] S1: Obtain production data from all production links of power enterprises and pre-process the production data; production data includes equipment operation data such as current, voltage, power, temperature and other sensor data, production plan data such as production work orders, output plans, etc., and quality inspection data such as product qualification rate and number of defective products; S2: Quantitatively evaluate the accuracy, completeness, timeliness and consistency of production data to obtain a comprehensive evaluation score;
[0028] S3: Within a preset time range, if the comprehensive evaluation score of the production data is less than a preset threshold, the production data is identified to identify the abnormal data types in the production data, including missing data, erroneous data, and redundant data, and the location of the abnormal data type is marked, and the factors causing the abnormal data are determined according to the location of the abnormal data;
[0029] S4: Optimize and repair abnormal data types;
[0030] S5: The optimized production data is stored in the database, and the comprehensive evaluation score of the production data and the optimization results of the production data are displayed through visualization tools.
[0031] The present invention provides a method for analyzing and optimizing the quality of production data of power enterprises, which can quantitatively evaluate the accuracy, completeness, timeliness and consistency of data, and identify abnormal data such as missing, erroneous and redundant data through preset thresholds, thereby effectively improving the quality of data. By analyzing the type, location and generation factors of abnormal data, the efficiency of the production process is effectively improved and the probability of failure is reduced; users can filter the time range, production links and abnormal data types through an interactive interface, dynamically view the latest evaluation and optimization results, and further enhance the convenience and visibility of data management.
[0032] In a preferred embodiment, the production data is preprocessed in step S1, specifically, a data cleaning algorithm is used to identify and remove noise data; according to preset business data rules, data irrelevant to the production link is eliminated, and data related to the production link is retained.
[0033] In a preferred embodiment, in step S2, specifically, according to the preset accuracy evaluation standard value, the accuracy score is obtained by calculating the deviation value between the production data and the preset accuracy evaluation standard value; the integrity score is obtained by counting the proportion of missing values in the production data; the timeliness score is obtained by calculating the interval between the production data generation time and the current time; the consistency score is obtained by detecting the logical consistency of the data in different links; the accuracy score, integrity score, timeliness score and consistency score are weighted and summed to obtain the comprehensive evaluation score of the production data. Among them, the weight of each indicator is set according to the specific needs and priority of the production data of the power enterprise.
[0034] Specifically: The formula for calculating accuracy is:
[0035]
[0036] X i is the value of the i-th data;
[0037] X ref To assess standard values for preset accuracy;
[0038] △ max is the maximum allowable deviation;
[0039] n: Total amount of data.
[0040] The formula for calculating the complete score is as follows:
[0041]
[0042] N total is the total amount of data;
[0043] N miss is the number of missing data.
[0044] The formula for calculating timeliness is:
[0045]
[0046] Tcur is the current time;
[0047] Tgen is the data generation time;
[0048] Tthr is the maximum allowed time interval.
[0049] The formula for calculating the consistency score is as follows:
[0050]
[0051] Nincon is the number of inconsistent data items;
[0052] Ntotal is the total amount of data. The preset business logic rules can be set according to business needs, such as field rules: the order quantity cannot be negative; time rules: start time < end time, data generation time must comply with sequential logic; upstream and downstream data rules: input data and output data in the production system must comply with conservation relationships or constraints, etc.
[0053] In a preferred embodiment, in step S3, specifically: within a preset time range, if the comprehensive evaluation score of the production data is less than a preset threshold, the abnormal data identification process is started; through rule matching, missing data in the production data is identified, including records with empty values or unfilled data fields, and the specific location of the missing data is marked;
[0054] Specifically, it traverses each field of the data set to check whether there are empty values, NULL, or illegal characters such as NA and -; check the required fields and determine whether the fields are missing based on the preset business rules.
[0055] Through data verification rules, erroneous data in production data can be identified, including values outside the reasonable range, data that does not conform to the format, or data with logical conflicts, and the specific location of the erroneous data can be marked;
[0056] Specifically, set the field value range, and mark the data beyond the field value range as erroneous data; check whether the field value meets the preset format requirements, such as date, time, number and other formats; detect whether there are conflicts or inconsistencies between data based on business logic, such as production end time < production start time, mark it as erroneous data.
[0057] Through duplication detection, redundant data in production data can be identified, including duplicate records or redundant fields, and the specific location of redundant data can be marked; redundant field information in the data table can be checked to ensure that there is no redundant duplication of field content. For example, if the device name and device alias contain duplicate content, they are marked as redundant fields.
[0058] Generate an abnormal data report, which includes the type, location and quantity statistics of abnormal data.
[0059] Specifically, add a unique timestamp to the collected production data, reconstruct the time sequence, identify the data generation time, transmission delay, and possible time interval anomalies. By comparing the time distribution pattern of the data, find abnormal nodes with large deviations from the expected time. Label the production data with data source labels such as device ID, collection point code, operator number, etc., and bind the data to its source device or link. Through label matching and reverse tracing, track the specific source location of the abnormal data, such as equipment, sensors, operating system, or manual input link.
[0060] Determine the factors causing abnormal data based on the location of abnormal data.
[0061] Specifically, the upstream and downstream links where the abnormal data is located are analyzed by using the data flow diagram and the production process relationship diagram, and the equipment, personnel or operation links that are highly associated with the abnormal data are identified. If the data of a certain device is abnormal, check its upstream input data and downstream impact range to identify the associated nodes. Based on the timestamp of the abnormal data, the data change trend before and after the abnormal data occurs is analyzed by the time window method to find the triggering event or abnormal fluctuation node of the abnormal data. For example, if the operation data of a certain device suddenly drops, combined with the device log or load data in the time period, it is found that the anomaly is caused by equipment overload. Introduce Granger causal analysis or causal inference model to judge the causal relationship of abnormal data through data characteristics and upstream and downstream impact relationships. If the quality inspection data is abnormal, analyze the relationship with the equipment operation data and infer that the anomaly may be caused by equipment failure. Based on the results of causal analysis, output optimization suggestions: for hardware problems, it is recommended to maintain the equipment or replace the sensor; for data collection problems, optimize the data collection frequency or transmission protocol; for operation problems, it is recommended to strengthen personnel training or update the operation process. In a preferred embodiment, step S4 is specifically as follows:
[0062] If missing data is identified, an interpolation algorithm is used to repair it and generate a replacement value to fill the missing data; specifically, linear interpolation, spline interpolation, mean substitution and other methods are selected for repair according to the data characteristics; if erroneous data is identified, rule verification is used to repair it;
[0063] Specifically, for out-of-range values, format errors and logical conflicts, format correction and logic adjustment are used to repair them. For example, upper and lower limit constraints are applied to out-of-range values, format errors are handled using regular expressions, and logical conflicts are repaired through rule correction. If redundant data is identified, data fusion is used for optimization; specifically, if it is a complete duplicate detection, all duplicate records are deleted and only one is retained. If partial duplication is detected, duplicate records are detected by matching key fields such as order number and equipment number, and conflicting fields are merged. If it is a redundant field, delete the field that has no practical meaning, or merge the fields with duplicate content. According to the generated data optimization report, the report includes a comparison of data before and after optimization.
[0064] In a preferred embodiment, step S5 specifically includes: storing the optimized production data in a database; displaying the comprehensive evaluation score of the production data through a bar graph; visually displaying the accuracy, completeness, timeliness and consistency scores of the data quality; displaying the production data optimization results through a comparison chart before and after the abnormal data is repaired, so that users can intuitively understand the data optimization effect; providing an interactive interface to allow users to filter and view specific optimization results according to time range, production link and abnormal data type; integrating visualization tools with the database to realize real-time updating and dynamic display of data, ensuring that users can obtain the latest data quality evaluation and optimization results in a timely manner.
[0065] The present invention provides a power enterprise production data quality analysis and optimization system, which is characterized by comprising the following:
[0066] Data processing module: used to obtain production data of each production link of the power enterprise and pre-process the production data;
[0067] Data quality assessment module: used to connect with the data processing module to quantitatively assess the accuracy, completeness, timeliness and consistency of production data and obtain a comprehensive assessment score;
[0068] Data type identification module: used to connect with the data quality assessment module. Within the preset time range, if the comprehensive assessment score of the production data is less than the preset threshold, the production data is identified to identify the abnormal data types in the production data, including missing data, erroneous data and redundant data, and the location of the abnormal data type is marked. The factors causing the abnormal data are determined according to the location of the abnormal data.
[0069] Abnormal data repair and optimization module: used to connect with the data type identification module to optimize and repair abnormal data types;
[0070] Visualization module: used to connect with the data processing module, data type identification module, abnormal data repair and optimization module and data type identification module, store the optimized production data in the database, and display the comprehensive evaluation score of the production data and the optimization results of the production data through visualization tools.
[0071] The present invention discloses a method and system for analyzing and optimizing the quality of production data of an electric power enterprise. The method and system are used to quantitatively evaluate the data in four dimensions: accuracy, completeness, timeliness and consistency. The method identifies abnormal data such as missing, erroneous and redundant data through preset thresholds, thereby effectively improving the quality of the data. By analyzing the type, location and generation factors of abnormal data, combined with a data flow diagram and a time window method, the source of data anomalies can be quickly located, including equipment failure, sensor problems, manual input errors, etc., so as to provide specific optimization suggestions for enterprises, such as equipment maintenance, data acquisition frequency adjustment or personnel training. This process effectively improves the efficiency of the production process and reduces the probability of failure. Users can filter the time range, production links and abnormal data types through an interactive interface, dynamically view the latest evaluation and optimization results, and further enhance the convenience and visibility of data management. The method is applicable to all aspects of production data of electric power enterprises, including data analysis and optimization such as equipment operation, production planning, and quality inspection, and has strong versatility. Through data preprocessing, abnormal identification and repair, the method can effectively serve the data quality management and production process optimization of enterprises, and improve the overall production management level. Through accurate data quality evaluation and problem repair, staff can make decisions based on reliable production data, and improve the scientificity and accuracy of decision-making. At the same time, analysis reports on abnormal data help companies identify potential system or operational problems and make targeted improvements.
Claims
1. A method for analyzing and optimizing production data quality of an electric power enterprise, characterized in that: The steps include: S1: Obtaining production data of each production link of the power enterprise and preprocessing the production data; S2: Quantitatively evaluate the accuracy, completeness, timeliness and consistency of the production data to obtain a comprehensive evaluation score; S3: within a preset time range, if the comprehensive evaluation score of the production data is less than a preset threshold, the production data is identified, abnormal data types in the production data are identified, including missing data, erroneous data and redundant data, and the positions of the abnormal data types are marked, and the factors causing the abnormal data are determined according to the positions of the abnormal data; S4: Optimizing and repairing the abnormal data type; S5: The optimized production data is stored in the database, and the comprehensive evaluation score of the production data and the optimization results of the production data are displayed through visualization tools.
2. A method for analyzing and optimizing production data quality of a power enterprise according to claim 1, characterized in that: In the step S1, the production data is preprocessed, specifically, a data cleaning algorithm is used to identify and remove noise data; according to preset business data rules, data irrelevant to the production link is eliminated, and data related to the production link is retained.
3. A method for analyzing and optimizing production data quality of a power enterprise according to claim 1, characterized in that: In step S2, specifically, according to the preset accuracy evaluation standard value, the accuracy score is obtained by calculating the deviation value between the production data and the preset accuracy evaluation standard value; and the completeness score is obtained by counting the proportion of missing values in the production data; The timeliness score is obtained by calculating the interval between the production data generation time and the current time; the consistency score is obtained by detecting the logical consistency of the data in different links; the accuracy score, completeness score, timeliness score and consistency score are weighted and summed to obtain the comprehensive evaluation score of the production data.
4. A method for analyzing and optimizing production data quality of a power enterprise according to claim 1, characterized in that: In the step S3, the abnormal data types in the production data are identified, including missing data, erroneous data and redundant data. Specifically, within a preset time range, if the comprehensive evaluation score of the production data is less than a preset threshold, the abnormal data is identified; through rule matching, the missing data in the production data is identified, including records with empty values or unfilled data fields; through data verification rules, the erroneous data in the production data is identified, including values outside the preset range, data that do not conform to the format or data with logical conflicts; through repeatability detection, the redundant data in the production data is identified, including duplicate records or redundant fields; and an abnormal data report is generated, which includes the type, location, factors causing the abnormal data and quantity statistics of the abnormal data.
5. The method for analyzing and optimizing production data quality of a power enterprise according to claim 1, characterized in that: In step S4, specifically, if the identified abnormal data type is missing data, an interpolation algorithm is used to repair it and a substitute value is generated to fill the missing data; if the identified abnormal data type is erroneous data, rule verification is used to repair it; if the identified abnormal data type is redundant data, data fusion is used for optimization; Generate a data optimization report based on the optimization and repair results of abnormal data, which includes a comparison of data before and after optimization.
6. A method for analyzing and optimizing production data quality of a power enterprise according to claim 1, characterized in that: In the step S5, specifically, the optimized production data is stored in a database; the comprehensive evaluation score of the production data is displayed through a bar chart; the optimization result of the production data is displayed through a comparison chart before and after the abnormal data is repaired, and an interactive interface is provided to allow the user to filter and view specific optimization results according to the time range, production link and abnormal data type.
7. A power enterprise production data quality analysis and optimization system, characterized in that: These include: Data processing module: used to obtain production data of each production link of the power enterprise and pre-process the production data; Data quality assessment module: used to connect with the data processing module to quantitatively assess the accuracy, completeness, timeliness and consistency of the production data to obtain a comprehensive assessment score; Data type identification module: used to connect with the data quality assessment module, and within a preset time range, if the comprehensive assessment score of the production data is less than a preset threshold, identify the production data, identify the abnormal data type in the production data, including missing data, erroneous data and redundant data, mark the location of the abnormal data type, and determine the factors causing the abnormal data according to the location of the abnormal data; Abnormal data repair and optimization module: used to connect with the data type identification module to optimize and repair the abnormal data type; Visualization module: used to connect with the data processing module, the data type identification module, the abnormal data repair and optimization module and the data type identification module, store the optimized production data in the database, and display the comprehensive evaluation score of the production data and the production data optimization results through visualization tools.
Citation Information
Cited By
Multi-tenant master data quality collaborative management method for smart power grid
CN120596476A
Data standardization and quality detection method for dynamic environment monitoring of cloud side end
CN121691506A