Method and system for multi-source data governance for electricity-energy-carbon-pollution and device
By collecting and cleaning multi-source time-series data, adding multi-level tags, and constructing a multi-level directory structure, the problems of insufficient coverage and low accuracy of pollution supervision data were solved, achieving data comprehensiveness and accuracy, and providing solid data support for pollution supervision.
Patent Information
- Application Number
- CN202411367891.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-29
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-09-29
AI Technical Summary
Existing pollution monitoring data lacks multi-source verification, resulting in insufficient data coverage and low accuracy of monitoring data. Current technologies only collect and process data from real-time electricity consumption and related industry production processes, without considering pollution sources in real life and the actual pollution situation in cities.
Collect multi-source time-series data of objects at different levels, perform multi-dimensional data characteristic analysis, including anomaly characteristics, missing characteristics, and spatiotemporal characteristics, clean the data, add multi-level tags, and store the data based on a multi-level directory structure to ensure the comprehensiveness and accuracy of the data.
Through unified and standardized processing and analysis of multi-source data, the correctness and accuracy of the data were achieved, providing data support and laying the foundation for subsequent data analysis for pollution supervision.
Smart Images

Figure CN119127863B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data governance technology, specifically to a multi-source data governance method, system, and equipment for electricity, energy, carbon, and pollution. Background Technology
[0002] In the course of modern industrial and social development, energy consumption, production activities, and pollution emissions have become focal points of global concern. With increasing environmental awareness, governments and businesses worldwide are actively seeking effective methods to monitor and manage energy consumption and pollutant emissions. Therefore, obtaining comprehensive and accurate environmental pollution-related data is of paramount importance.
[0003] Existing pollution control methods largely rely on direct monitoring of relevant production data. For example, Chinese patent application CN115238823A discloses a standardized governance method for time-series data in the energy industry, implemented through a data acquisition platform, a data storage management platform, and a data analysis and governance platform. Specifically, the data acquisition platform collects time-series data on power generation (wind, solar, hydro, thermal, etc.) and energy production (coal, oil, natural gas, etc.) via an acquisition interface program. The storage management platform primarily includes time-series database storage and structured data storage modules related to the time-series data. The data analysis and governance platform includes an algorithm model management and modeling operation module for data quality auditing, and a data governance module including data encoding, data asset cataloging, and data services. While this solution offers strong real-time performance, it only collects and processes data from real-time electricity consumption and related industry production processes, neglecting actual pollution sources and the actual pollution situation in cities, resulting in insufficient coverage of regulatory data. Furthermore, existing pollution monitoring data lacks multi-source verification, potentially leading to erroneous data and affecting the accuracy of monitoring data. Summary of the Invention
[0004] To overcome the problems of insufficient data coverage and low accuracy of monitoring data due to the lack of multi-source data verification in the aforementioned pollution monitoring data, this invention provides a multi-source data governance method, system, and equipment for electricity-energy-carbon-pollution.
[0005] On the one hand, this invention provides a multi-source data governance method for electricity, energy, carbon, and pollution, including:
[0006] Collect multi-source time-series data of objects at different levels, including power data, energy data, carbon emission data, and pollutant data;
[0007] Multi-dimensional data characteristic analysis is performed on the multi-source time series data to obtain the abnormal characteristics, missing characteristics, and spatiotemporal characteristics of the multi-source time series data; based on the abnormal characteristics, missing characteristics, and spatiotemporal characteristics of the multi-source time series data, data cleaning is performed on the multi-source time series data to obtain cleaned multi-source data;
[0008] The object features, business features, and time features of the cleaned multi-source data are extracted respectively, and multi-level labels are added to the cleaned multi-source data based on the object features, business features, and time features;
[0009] Based on the multi-level tags, a corresponding multi-level directory structure is constructed. Based on the multi-level directory structure and the multi-level tags of the cleaned multi-source data, the cleaned multi-source data is stored to achieve the governance of the multi-source time-series data.
[0010] Optionally, the collection of multi-source time-series data of objects at different levels includes:
[0011] Collect electricity consumption statistics, energy statistics, carbon emission statistics, and pollutant statistics within each administrative region;
[0012] Collect actual electricity consumption data, actual production capacity data, carbon emission monitoring data, and pollutant monitoring data from different enterprises in various industries.
[0013] Optionally, the step of data cleaning the multi-source time series data based on its anomaly characteristics, missing characteristics, and spatiotemporal characteristics to obtain cleaned multi-source data includes:
[0014] Based on the abnormal characteristics of the multi-source time series data, outlier values of the multi-source time series data are removed to obtain data after outlier removal.
[0015] Based on the missing characteristics of the multi-source time series data, missing values of the data after removing anomalies are filled to obtain the missing-filled data.
[0016] Based on the spatiotemporal characteristics of the multi-source time-series data, the time and space of the missing data are aligned to obtain the cleaned multi-source data.
[0017] Optionally, the missing value imputation of the data after removing anomalies based on the missing characteristics of the multi-source time series data to obtain the missing-impacted data includes:
[0018] For data with short-term missing characteristics, linear interpolation or the mean of the same period is used to fill in the missing values.
[0019] For data with intermediate missing characteristics, missing values are filled using neural network prediction or extreme gradient boosting algorithm.
[0020] For data with long-term missing characteristics, missing values are filled using a combination of time series decomposition and extreme gradient boosting algorithm.
[0021] Optionally, the missing value imputation based on time series decomposition combined with the limiting gradient boosting algorithm includes:
[0022] Based on the time series decomposition method, feature analysis of the data to be processed is performed to obtain the long-term trend features and short-term periodic features of the data to be processed.
[0023] Based on the long-term trend characteristics, short-term periodic characteristics, and extreme gradient boosting algorithm of the data to be processed, the missing values of the data to be processed are predicted, and the missing values of the data to be processed are filled in using the prediction results.
[0024] The data to be processed is data in the multi-source time series data that has a long-term missing characteristic.
[0025] Optionally, before removing outliers from the multi-source time series data based on its abnormal characteristics to obtain data after anomaly removal, the method further includes:
[0026] The multi-source time-series data is subjected to data deduplication processing;
[0027] After imputing missing values in the data following the removal of anomalies based on the missing characteristics of the multi-source time series data, and obtaining the imputed data, the method further includes:
[0028] The missing data is then standardized in terms of units and format.
[0029] Optionally, after aligning the missing data in time and space based on the spatiotemporal characteristics of the multi-source time-series data to obtain cleaned multi-source data, the method further includes:
[0030] The cleaned multi-source data is then validated based on data from different sources.
[0031] Optionally, the multi-level tagging includes first-level tags, second-level tags, and third-level tags. Adding multi-level tags to the cleaned multi-source data based on the object characteristics, business characteristics, and time characteristics includes:
[0032] Based on the object characteristics, multi-level object labels are added to the cleaned multi-source data as the first-level labels;
[0033] Based on the aforementioned business characteristics, business category tags are added as secondary tags to the cleaned multi-source data;
[0034] Based on the time characteristics, time granularity labels are added to the cleaned multi-source data as the third-level labels;
[0035] The multi-level object labels include regional labels, industry labels, and enterprise labels, and the business category labels include electricity category labels, energy category labels, carbon emission category labels, and pollutant category labels.
[0036] Optionally, the step of constructing a corresponding multi-level directory structure based on the multi-level tags, and storing the cleaned multi-source data based on the multi-level directory structure and the multi-level tags of the cleaned multi-source data, includes:
[0037] Construct a multi-level directory structure that corresponds one-to-one with the tag level of the multi-level tags;
[0038] Based on the multi-level directory structure, a storage directory for the cleaned multi-source data is constructed, and a target table data model is formed based on the storage directory for data storage.
[0039] On the other hand, the present invention also provides a multi-source data governance system for electricity-energy-carbon-pollution, the system comprising:
[0040] The data collection module is used to collect multi-source time-series data of objects at different levels, including power data, energy data, carbon emission data, and pollutant data.
[0041] The data cleaning module is used to perform multi-dimensional data characteristic analysis on the multi-source time series data to obtain the abnormal characteristics, missing characteristics, and spatiotemporal characteristics of the multi-source time series data; and to perform data cleaning on the multi-source time series data based on the abnormal characteristics, missing characteristics, and spatiotemporal characteristics of the multi-source time series data to obtain cleaned multi-source data.
[0042] The tag adding module is used to extract the object features, business features, and time features of the cleaned multi-source data respectively, and add multi-level tags to the cleaned multi-source data based on the object features, business features, and time features;
[0043] The data storage module is used to construct a corresponding multi-level directory structure based on the multi-level tags, and to store the cleaned multi-source data based on the multi-level directory structure and the multi-level tags of the cleaned multi-source data, so as to realize the governance of the multi-source time-series data.
[0044] On the other hand, the present invention also provides an electronic device, comprising: at least one processor and a memory; the memory and the processor are connected via a bus;
[0045] The memory is used to store one or more programs;
[0046] When the one or more programs are executed by the at least one processor, the multi-source data governance method for electricity-energy-carbon-pollution described in any one of the above embodiments is implemented.
[0047] On the other hand, the present invention also provides a readable storage medium having an executable program stored thereon, wherein when the executable program is executed, it implements the multi-source data governance method for electricity-energy-carbon-pollution described in any one of the above.
[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0049] This invention provides a multi-source data governance method, system, and device for electricity, energy, carbon, and pollution. On one hand, it collects multi-source time-series data from different levels of objects, including electricity data, energy data, carbon emission data, and pollutant data, to obtain diverse types of data related to electricity, energy, carbon, and pollution, ensuring the breadth and comprehensiveness of the collected data. Through multi-dimensional data characteristic analysis of the multi-source time-series data, its anomaly characteristics, missing characteristics, and spatiotemporal characteristics are obtained. Based on these anomaly characteristics, missing characteristics, and spatiotemporal characteristics, the multi-source time-series data is cleaned to obtain cleaned multi-source data, enabling mutual verification between the multiple sources and ensuring the correctness and accuracy of the cleaned data. On the other hand, it adds multi-level tags and constructs a multi-level directory structure based on object characteristics, business characteristics, and time characteristics, and stores data based on this multi-level directory structure. This facilitates data reading and writing processes and data traceability during subsequent data processing, providing data support for subsequent data analysis in pollution supervision. Attached Figure Description
[0050] Figure 1 This is a flowchart illustrating a multi-source data governance method for electricity, energy, carbon, and pollution according to the present invention.
[0051] Figure 2 This is a schematic diagram of a multi-level labeling system as an example of the present invention;
[0052] Figure 3 This is a schematic diagram of the initial data directory as an example of the present invention;
[0053] Figure 4 This is a schematic diagram of a target table data model based on multi-level tag naming, which is an example of the present invention.
[0054] Figure 5 This is a schematic diagram of a target table data model with added multi-level tags, as an example of the present invention.
[0055] Figure 6This invention provides an example of original electricity consumption data and an electricity consumption-time curve simulating various consecutive missing values.
[0056] Figure 7 This is a diagram showing the fill-in result of an example of the limiting gradient boosting algorithm of the present invention;
[0057] Figure 8 This is a diagram illustrating the filling effect of a time series decomposition method combined with the limiting gradient boosting algorithm, as an example of the present invention.
[0058] Figure 9 A heatmap showing the accuracy of different filling methods for different consecutive missing days, as an example of the present invention;
[0059] Figure 10 This is a schematic diagram of the electronic device of the present invention. Detailed Implementation
[0060] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0061] Example 1:
[0062] This invention provides a multi-source data governance method for electricity, energy, carbon, and pollution, as illustrated in the diagram below. Figure 1 As shown, it includes:
[0063] Step S110: Collect multi-source time-series data of objects at different levels, including power data, energy data, carbon emission data, and pollutant data;
[0064] Step S120: Perform multi-dimensional data characteristic analysis on the multi-source time series data to obtain the abnormal characteristics, missing characteristics, and spatiotemporal characteristics of the multi-source time series data; perform data cleaning on the multi-source time series data based on the abnormal characteristics, missing characteristics, and spatiotemporal characteristics of the multi-source time series data to obtain cleaned multi-source data;
[0065] Step S130: Extract the object features, business features, and time features of the cleaned multi-source data respectively, and add multi-level labels to the cleaned multi-source data based on the object features, business features, and time features;
[0066] Step S140: Construct a corresponding multi-level directory structure based on the multi-level tags, and store the cleaned multi-source data based on the multi-level directory structure and the multi-level tags of the cleaned multi-source data to achieve the governance of the multi-source time-series data.
[0067] In this example implementation, objects at different levels include administrative regions, industries, and enterprises. Multi-source time-series data includes electricity data, energy data, carbon emission data, and pollutant data. That is, multi-source data can include electricity data, energy data, carbon emission data, and pollutant data from different enterprises within various industries of each administrative region; for example, electricity consumption data of a certain enterprise, carbon emission data of a certain city, and pollutant monitoring data. Anomaly characteristics can include anomaly location, anomaly distribution, etc.; missing characteristics can include the number and distribution of missing values, etc.; spatiotemporal characteristics can include temporal distribution features and spatial distribution features. Object characteristics can include features such as region, industry, and enterprise; business characteristics can include business type features, such as electricity, energy, carbon emission, and pollutant business characteristics; and time characteristics include features at different time granularities. For example, electricity consumption statistics, energy statistics, carbon emission statistics, and pollutant statistics can be collected from each administrative region; actual electricity consumption data, actual production capacity data, carbon emission monitoring data, and pollutant monitoring data of different enterprises within each industry can be collected to obtain multi-source time-series data for objects at different levels. This example collects multi-source datasets, including electricity, energy consumption, production capacity, carbon emissions, and pollutants, starting from different object levels of the data source. This ensures the comprehensiveness and breadth of coverage of the data, making the foundational data for subsequent pollutant monitoring model construction richer and more diverse. Through multi-dimensional data characteristic analysis of the multi-source time-series data, the abnormal characteristics, missing characteristics, and spatiotemporal characteristics of the multi-source time-series data are obtained. Based on these characteristics, data cleaning is performed to obtain cleaned multi-source data, enabling cross-verification of the multi-source data and ensuring the correctness and accuracy of the cleaned data. This allows for unified and standardized processing and analysis of data from different sources. Furthermore, multi-level tagging and directory structure construction are performed based on object characteristics, business characteristics, and time characteristics. Data storage is based on this multi-level directory structure, facilitating data reading and writing processes and data traceability during subsequent data processing, providing data support for subsequent data analysis in pollution regulation.
[0068] In one example implementation, the data cleaning of the multi-source time series data based on the anomaly characteristics, missing characteristics, and spatiotemporal characteristics of the multi-source time series data in step S120, to obtain cleaned multi-source data, includes:
[0069] Based on the abnormal characteristics of the multi-source time series data, outlier values of the multi-source time series data are removed to obtain data after outlier removal.
[0070] Based on the missing characteristics of the multi-source time series data, missing values of the data after removing anomalies are filled to obtain the missing-filled data.
[0071] Based on the spatiotemporal characteristics of the multi-source time-series data, the time and space of the missing data are aligned to obtain the cleaned multi-source data.
[0072] In this example implementation, data exceeding a preset normal range can be considered as anomalous data. Alternatively, the 3σ rule can be used to identify outliers in multi-source time-series data, and these anomalous data / outliers can be removed or corrected. Missing value imputation methods can be selected based on the missing characteristics. For example, missing value imputation methods can include linear interpolation, Lagrange interpolation, contemporaneous mean substitution, neural network prediction, and extreme gradient boosting (XGBoost) algorithms, etc. Linear interpolation estimates the value of intermediate gaps based on known data points, which is simple, fast, and computationally efficient. Lagrange interpolation accurately passes through given data points in polynomial form, showing good performance in local interpolation scenarios. Contemporaneous mean substitution uses historical data for imputation, offering strong adaptability. Neural network prediction uses artificial neural networks to fill missing data, generating smooth and continuous prediction results by learning data patterns. Extreme gradient boosting, or XGBoost interpolation, can handle nonlinear relationships and multivariate data, while also exhibiting good anti-overfitting capabilities. When handling missing values, the Extreme Gradient Boosting algorithm creates an additional binary tree branch for the missing numerical feature to predict the missing value. The algorithm attempts to fill the missing value with different values and selects the value that maximizes the improvement in model performance as the optimal solution. For example, missing features can include short-term, medium-term, and long-term missing values. This example can filter missing imputation methods based on different missing characteristics of multi-source time series data to obtain the data with the best imputation effect. Furthermore, for unimportant features or data items with minimal impact, the missing value can be ignored, i.e., records containing missing values can be directly deleted.
[0073] For example, the missing value imputation of the data after removing anomalies based on the missing characteristics of the multi-source time series data to obtain the missing-impacted data includes the following process:
[0074] For data with short-term missing characteristics, linear interpolation or the mean of the same period is used to fill in the missing values.
[0075] For data with intermediate missing characteristics, missing values are filled using neural network prediction or extreme gradient boosting algorithm.
[0076] For data with long-term missing characteristics, missing values are filled using a combination of time series decomposition and extreme gradient boosting algorithm.
[0077] In this example implementation, the missing characteristics can be determined based on the number of consecutive missing data points. For example, the missing characteristics can be divided into short-term missing, medium-term missing, and long-term missing based on the number of consecutive missing data points. For instance, the case where the number of consecutive null values is less than or equal to 7 days can be determined as short-term missing; the case where the number of consecutive null values is greater than 7 days but less than or equal to 15 days can be determined as medium-term missing; and the case where the length of consecutive null values is between 15 and 30 days can be determined as long-term missing. For cases exceeding 30 days, the actual situation is checked to determine whether there is a production stoppage or malfunction, etc. For short-term missing values, linear interpolation or the mean substitution method can be used for imputation, which is simple and efficient. For medium-term missing values, neural network prediction or the extreme gradient boosting algorithm can be used for imputation, that is, using a trained neural network model or decision tree model to predict missing values to ensure the accuracy of imputation. For long-term missing values, time series decomposition combined with the extreme gradient boosting algorithm can be used for imputation. The time series decomposition method is used to obtain the long-term trend and short-term periodic characteristics of the data to be processed. These long-term trend and short-term periodic characteristics are used as additional input features of the corresponding model of the extreme gradient boosting algorithm, which increases the mining of long-term and short-term patterns in the data, improves the accuracy of imputation of long-term missing data, and improves data quality.
[0078] For example, the missing value imputation based on time series decomposition combined with the limiting gradient boosting algorithm includes the following process:
[0079] Based on the time series decomposition method, feature analysis of the data to be processed is performed to obtain the long-term trend characteristics and short-term seasonal cycle characteristics of the data to be processed.
[0080] Based on the long-term trend characteristics, short-term seasonal cycle characteristics, and extreme gradient boosting algorithm of the data to be processed, the missing values of the data to be processed are predicted, and the missing values of the data to be processed are filled in using the prediction results.
[0081] In this example implementation, the data to be processed is long-term missing data from multi-source time series data. The data to be processed (i.e., long-term missing time series data) is decomposed into seasonal components, trend components, and residuals using a time series decomposition algorithm, as shown in the following formula:
[0082] y(t)=S(t)+T(t)+R(t)
[0083] Where y(t) represents the data to be processed, T(t) is the trend component, S(t) is the seasonal component, and R(t) is the residual component; the long-term trend features (corresponding to the trend component) and short-term periodic features (corresponding to the seasonal component) are extracted from the data to be processed. For the extreme gradient boosting algorithm, the input features of the model are first determined. For example, the 3-day difference, 7-day difference, 15-day difference, 3-day rolling standard deviation, 7-day rolling standard deviation, 15-day difference, 7-day moving average, and 15-day moving average can be calculated from the data to be processed as the input features of the model corresponding to the extreme gradient boosting algorithm. Among them, the difference refers to the index obtained by calculating the difference between each data point and the data point x1 time units ago, such as x1=3 is the 3-day difference, x1=7 is the 7-day difference, etc.; the rolling standard deviation refers to the index obtained by rolling the calculation of the standard deviation of x2 consecutive data points with a certain step size (e.g., step size of 1) and time window length, such as x2=3 corresponds to the 3-day rolling standard deviation, x2=7 corresponds to the 7-day rolling standard deviation; the moving average refers to the index obtained by calculating the average of the data points over x3 consecutive days, such as x3=3 corresponds to the 3-day moving average, x3=7 corresponds to the 7-day moving average. The aforementioned indices corresponding to differencing, rolling standard deviation, and moving average, along with the data to be processed and the long-term trend and short-term periodic features obtained from the time series decomposition algorithm, are used as input features for the Extreme Gradient Boosting (XGBoost) model to predict and impute missing values in the data. During the XGBoost model training phase, the calculated indices and original samples, along with the long-term trend and short-term periodic features of the samples, are combined to form a training dataset for training the XGBoost model. The XGBoost model is trained using all data samples with the feature set. During training, GridSearchCV is used to select the optimal solution for the model's hyperparameters, and the optimal hyperparameter configuration is used to train the model to predict missing values.
[0084] In one example implementation, before removing outliers from the multi-source time series data based on its anomaly characteristics to obtain data with outliers removed, the method further includes:
[0085] The multi-source time-series data is deduplicated.
[0086] In this example implementation, duplicate data in multi-source time series data can be identified and deleted by means of table comparison or data comparison, thereby achieving data deduplication.
[0087] After imputing missing values in the data following the removal of anomalies based on the missing characteristics of the multi-source time series data, and obtaining the imputed data, the method further includes:
[0088] The missing data is then standardized in terms of units and format.
[0089] In this example implementation, data from different sources may have different units and / or formats, so it is necessary to unify the units and formats to ensure that the units and formats of the same data are consistent; for example, all electricity consumption units are converted to kilowatt-hours, and the date, time, and currency use the same format to facilitate computer processing.
[0090] In some example implementations, after aligning the missing data in time and space based on the spatiotemporal characteristics of the multi-source time series data to obtain cleaned multi-source data, the method further includes:
[0091] The cleaned multi-source data is then validated based on data from different sources.
[0092] In this example implementation, data of the same nature from different sources can be compared and verified to achieve verification between data from different sources; official data can also be used to verify other data. For example, for various time-series pollutant data (such as regional grid data in the air quality monitoring data list), data verification can be performed by combining provincial control monitoring data; for real-time electricity consumption data of enterprises, data verification can be performed by combining the actual production status of enterprises, and data that fails verification can be further checked and corrected.
[0093] In some example implementations, step S130, which involves adding multi-level labels to the cleaned multi-source data based on the object characteristics, the business characteristics, and the time characteristics, includes:
[0094] Based on the object characteristics, multi-level object labels are added to the cleaned multi-source data as the first-level labels;
[0095] Based on the aforementioned business characteristics, business category tags are added as secondary tags to the cleaned multi-source data;
[0096] Based on the aforementioned time characteristics, time granularity labels are added to the cleaned multi-source data as the third-level labels.
[0097] In this example implementation, object characteristics may include regional characteristics, industry characteristics, and enterprise characteristics; business characteristics include various business type characteristics, such as electricity, energy, carbon emissions, and pollutants; time characteristics include time granularity characteristics, such as year, month, day, hour, and minute. Multi-level tags include first-level tags, second-level tags, and third-level tags. Multi-level object tags are added as first-level tags based on object characteristics, including regional tags, industry tags, and enterprise tags; business category tags are added based on business characteristics, including electricity category tags, energy category tags, carbon emission category tags, and pollutant category tags; and time granularity characteristics, such as year, month, day, hour, and minute tags, are added based on time characteristics, thus forming multi-level tags. When constructing multi-level tags, key factors to consider include the statistical subject, administrative region, industry classification, and data time granularity: First, determine the statistical subject, classifying it according to "administrative region, industry, and enterprise," forming the basic level (first-level tag) in the tagging system; based on data business type, divide the data into four categories: "electricity, energy, carbon, and pollution," providing a multi-dimensional perspective for analysis (as second-level tags); according to time granularity, use different time granularities such as "hour, day, month, and year" to meet different analytical needs (as third-level tags). Figure 2 As shown, this is the constructed multi-level tagging system. First-level tags include region, industry, and enterprise; second-level tags include electricity, energy, carbon emissions, and air pollutants. The region-electricity tag includes data on electricity consumption, power generation, green electricity use, and green electricity trading, with each data type categorized by time granularity: year, month, day, and hour. Similarly, the enterprise-electricity tag includes electricity consumption and electricity load data, each with different time granularity tags. Likewise, the data patterns under other tag combinations can be observed. This example uses a multi-level tagging system to accurately and efficiently integrate multi-source data from different levels of objects.
[0098] In some implementations, step S140, which involves constructing a corresponding multi-level directory structure based on the multi-level tags, and storing the cleaned multi-source data based on the multi-level directory structure and the multi-level tags of the cleaned multi-source data, includes:
[0099] Construct a multi-level directory structure that corresponds one-to-one with the tag level of the multi-level tags;
[0100] Based on the multi-level directory structure, a storage directory for the cleaned multi-source data is constructed, and a target table data model is formed based on the storage directory for data storage.
[0101] In this example implementation, a multi-level tagging system is associated with the data storage directory structure, and a one-to-one multi-level directory structure is constructed based on the multi-level tags. Specifically, when constructing the multi-level directory structure, the directory structure and hierarchy are determined according to the hierarchical construction system, dividing the statistical subject, administrative region, and industry into multiple dimensions. This forms a data directory structure with a basic level of "region-industry-enterprise," four intermediate levels of "electricity-energy-carbon-pollution," and internal levels of "hour, day, month, year." Simultaneously, the collected data is archived and organized, and classified and stored according to the multi-level directory structure. Specifically, a storage directory is constructed based on the multi-level directory structure, and a target table data model is formed for data storage. In the target table data model, a detailed description and metadata need to be provided for each data item. For example, the basic information of the collected multi-source data can be organized first, that is, the collected data can be initially archived to ensure basic data retrieval. The preliminary data directory includes file name, indicator, indicator description (field description), unit, precision, source, etc. Figure 3 As shown; after data cleaning and archiving, the basic data information is further archived and organized to form an efficient and scalable data storage and retrieval model; the metadata description in the target table data model includes table code, table name, table description, table source, field name, field type, corresponding standard code, data storage format requirements, primary key, etc., such as Figure 4 As shown, table names can be determined based on multi-level tags. The purpose of metadata is to identify, describe, evaluate, and track resources, enabling effective discovery, understanding, organization, and management of data resources. Finally, the data catalog is integrated with the data tagging system: attaching the data tagging system to the data catalog allows for the construction of a comprehensive, orderly, and easily manageable multi-level catalog structure, supporting complex data analysis needs and ensuring long-term data availability and security. The target table data model with added multi-level tags is as follows: Figure 5 As shown in the example, this example uses multi-level tags to form a target table data model, which enables fast data retrieval within the model while ensuring the accuracy of the retrieved data.
[0102] To achieve effective monitoring and management of energy consumption and pollutant emissions, constructing an efficient and accurate "electricity-pollution" model is crucial for improving the accuracy of environmental supervision. The "electricity-pollution" model aims to integrate multi-source data on electricity consumption, energy consumption, production capacity, carbon emissions, and pollutants to provide a more comprehensive and precise monitoring and management method. This invention acquires multi-source data from multiple levels of objects, covering electricity, energy consumption, production capacity, carbon emissions, and pollutants, ensuring the comprehensiveness and breadth of data coverage, making the foundational data for model construction richer and more diverse. Through data characteristic analysis and data cleaning steps, a data standardization governance system is established, effectively identifying and handling missing values, outliers, and duplicate data, improving data quality and reliability. Data standardization steps ensure consistency in data units and formats, enabling data from different sources to be processed and analyzed in a unified and standardized manner. By hierarchically linking the acquisition of multi-source data from multi-level objects, the analysis of multi-dimensional data characteristics, the addition of multi-level tags, the construction of a multi-level directory structure based on multi-level tags, and the data storage process, a clear hierarchical framework for the data processing process is formed. This ensures that the entire data governance process is orderly and standardized, avoiding conflicts and confusion in processes and dimensions, and ensuring the correspondence between data and its attributes and the accuracy of the data. At the same time, through the fusion and governance of multi-source data using multi-level tags, a high-quality, multi-dimensional data model is constructed, providing solid data support for decision-making in the fields of power, energy, and pollution control.
[0103] Experimental verification
[0104] The effectiveness of the missing value filling method of the present invention will be verified below.
[0105] (1) Taking the electricity meter data of Company A as an example, the 3σ rule is first used to identify outliers in the data, and the data outside the normal distribution are removed or the outliers are adjusted as appropriate. Then, the data after the outlier adjustment is statistically analyzed for consecutive missing values, and the results are shown in Table 1.
[0106] Table 1
[0107]
[0108] Missing values were handled according to the distribution of missing values in Table 1. First, for unimportant features or data items with minimal impact, records containing missing values were directly deleted; then, the remaining missing values were imputed. Taking Company A's electricity consumption data as an example, its original electricity consumption data is as follows: Figure 6 The upper part shows the data, with various missing data scenarios added. Taking the cases of 4, 7, and 20 consecutive missing values as examples, the simulated electricity consumption data with missing values is as follows. Figure 6The lower half of the diagram shows how to impute consecutive missing values of different lengths. While basic linear logical interpolation methods such as linear interpolation and Lagrange interpolation are known to be simple, fast, and efficient, they cannot identify trends in time-series data with medium- to long-term missing values. Therefore, the extreme gradient boosting algorithm is used to impute medium- to long-term missing data. Specifically, features such as 3-day differencing, 7-day differencing, 15-day differencing, 3-day rolling standard deviation, 7-day rolling standard deviation, 15-day differencing, 7-day moving average, and 15-day moving average are calculated from the original time-series data. These calculated features are combined with the original data to form a training set for training the XGBoost model. The XGBoost model is trained using all data samples with the feature set. During training, GridSearchCV is used to select the optimal solution for the model hyperparameters, and the model is trained using the optimal hyperparameter configuration to predict missing values. The final missing value imputation result is shown in the diagram. Figure 7 As shown, from Figure 7 It can be seen that the extreme gradient boosting algorithm can achieve data imputation for medium-length missing data, but its accuracy in imputing long-term continuous missing data is poor. Finally, the feature importance of each feature was determined, with the 7-day rolling standard deviation having the highest importance at 0.46, followed by the 7-day differencing at 0.15. Feature importance measures the contribution of each feature to the model's predictive performance, helping to identify the features most influential on model predictions, thereby optimizing feature selection, model interpretation, and decision-making.
[0109] (2) Missing data in Company A's electricity consumption data were imputed using a combination of time series decomposition and extreme gradient boosting algorithms. The imputed results are as follows: Figure 8 As shown, from Figure 8 As can be seen, the imputation method combining time series decomposition algorithm and extreme gradient boosting algorithm significantly improves the accuracy of imputing long-term missing data. The feature importance of each input feature obtained by the extreme gradient boosting algorithm model is shown in Table 2.
[0110] Table 2
[0111]
[0112] As can be seen from the data in Table 2, the long-term trend feature S(t) and short-term periodic feature T(t) obtained by the time series decomposition method in this invention have high feature importance, indicating that the long-term trend feature and short-term periodic feature as additional features of the extreme gradient boosting algorithm in this invention are important for the accurate prediction of missing values.
[0113] (3) Using the electricity consumption data of Company A as the data sample, and assuming that the number of consecutive missing data is 1 to 30, we will conduct data imputation experiments using linear interpolation, Lagrange interpolation, mean replacement method, neural network interpolation, extreme gradient boosting algorithm and time series splitting method combined with extreme gradient boosting algorithm, respectively. We will calculate the precision of the imputation result of each algorithm compared with the original data. The precision calculation formula is shown below.
[0114]
[0115] Here, TP (True Positives) refers to the number of true positives, i.e., the number of samples correctly fitted to the positive class by the model, and FP (False Positives) refers to the number of false positives, i.e., the number of samples incorrectly fitted to the positive class by the model. By conducting imputation experiments for each method under different consecutive missing values, the precision results of each experiment are finally summarized as follows: Figure 9 The heat map shown, Figure 9 The darker the color block, the better the filling method performs in the case of consecutive null values. Figure 9 The distribution of medium- and dark-colored blocks leads to the following conclusions: short-term missing data can be filled using linear interpolation and the mean of the same period. Medium-term missing data can be filled using neural network imputation or the extreme gradient boosting algorithm, depending on the specific circumstances. Long-term missing data can be filled using a combination of time series decomposition and extreme gradient boosting algorithms. In summary, employing a combination of multiple imputation methods, and selecting the most suitable imputation method based on the continuity length and characteristics of the missing data, can significantly improve the effectiveness of data imputation. This not only improves data integrity but also enhances the reliability and accuracy of data governance.
[0116] Example 2:
[0117] Based on the same inventive concept, this invention also provides a multi-source data governance system for electricity, energy, carbon, and pollution, the system comprising:
[0118] The data collection module is used to collect multi-source time-series data of objects at different levels, including power data, energy data, carbon emission data, and pollutant data.
[0119] The data cleaning module is used to perform multi-dimensional data characteristic analysis on the multi-source time series data to obtain the abnormal characteristics, missing characteristics, and spatiotemporal characteristics of the multi-source time series data; and to perform data cleaning on the multi-source time series data based on the abnormal characteristics, missing characteristics, and spatiotemporal characteristics of the multi-source time series data to obtain cleaned multi-source data.
[0120] The tag adding module is used to extract the object features, business features, and time features of the cleaned multi-source data respectively, and add multi-level tags to the cleaned multi-source data based on the object features, business features, and time features;
[0121] The data storage module is used to construct a corresponding multi-level directory structure based on the multi-level tags, and to store the cleaned multi-source data based on the multi-level directory structure and the multi-level tags of the cleaned multi-source data, so as to realize the governance of the multi-source time-series data.
[0122] In one possible implementation, the data collection module is specifically used for:
[0123] Collect electricity consumption statistics, energy statistics, carbon emission statistics, and pollutant statistics within each administrative region;
[0124] Collect actual electricity consumption data, actual production capacity data, carbon emission monitoring data, and pollutant monitoring data from different enterprises in various industries.
[0125] In one possible implementation, the data cleaning module includes:
[0126] Anomaly handling submodule is used to remove outliers from the multi-source time series data based on the anomaly characteristics of the multi-source time series data, and obtain data after removing anomalies.
[0127] The missing value filling submodule is used to fill in the missing values of the data after removing anomalies based on the missing characteristics of the multi-source time series data, so as to obtain the missing value filled data.
[0128] The spatiotemporal alignment submodule is used to align the missing data in time and space based on the spatiotemporal characteristics of the multi-source time series data, so as to obtain the cleaned multi-source data.
[0129] In one possible implementation, the missing filling submodule includes:
[0130] The short-term imputation subunit is used to impute missing values for data with the missing characteristic of short-term missingness using linear interpolation or the mean replacement method of the same period;
[0131] The intermediate-stage imputation subunit is used to impute missing values for data with intermediate-stage missing characteristics based on neural network prediction or extreme gradient boosting algorithm.
[0132] The long-term imputation subunit is used to impute missing values for data with the missing characteristic of long-term missingness, based on time series decomposition combined with the extreme gradient boosting algorithm.
[0133] In one possible implementation, the long-term filling subunit is specifically used for:
[0134] Based on the time series decomposition method, feature analysis of the data to be processed is performed to obtain the long-term trend features and short-term periodic features of the data to be processed.
[0135] Based on the long-term trend characteristics, short-term periodic characteristics, and extreme gradient boosting algorithm of the data to be processed, the missing values of the data to be processed are predicted, and the missing values of the data to be processed are filled in using the prediction results.
[0136] The data to be processed is data in the multi-source time series data that has a long-term missing characteristic.
[0137] In one possible implementation, the data cleaning module further includes:
[0138] The deduplication submodule is used to perform data deduplication processing on the multi-source time-series data;
[0139] The data cleaning module also includes:
[0140] The standardization submodule is used to standardize the units and format of the data after missing data filling.
[0141] In one possible implementation, the data cleaning module further includes:
[0142] The verification submodule is used to verify the cleaned multi-source data from different sources.
[0143] In one possible implementation, the multi-level tags include first-level tags, second-level tags, and third-level tags, and the tag adding module is specifically used for:
[0144] Based on the object characteristics, multi-level object labels are added to the cleaned multi-source data as the first-level labels;
[0145] Based on the aforementioned business characteristics, business category tags are added as secondary tags to the cleaned multi-source data;
[0146] Based on the time characteristics, time granularity labels are added to the cleaned multi-source data as the third-level labels;
[0147] The multi-level object labels include regional labels, industry labels, and enterprise labels, and the business category labels include electricity category labels, energy category labels, carbon emission category labels, and pollutant category labels.
[0148] In one possible implementation, the data storage module is specifically used for:
[0149] Construct a multi-level directory structure that corresponds one-to-one with the tag level of the multi-level tags;
[0150] Based on the multi-level directory structure, a storage directory for the cleaned multi-source data is constructed, and a target table data model is formed based on the storage directory for data storage.
[0151] Example 3:
[0152] like Figure 10 As shown, the present invention also provides an electronic device, which may be a computer device, a microcontroller device, a smart mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, processor, and transceiver component are connected via a bus; the memory can be used to store executable programs, and an exemplary executable program may include instructions; the processor is used to execute the instructions stored in the memory. The memory can also be used to store data, which can be accessed and / or modified when instructions are executed.
[0153] The processor may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and it is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the storage medium to realize the corresponding method flow or corresponding function, so as to realize the steps of a multi-source data governance method for electricity-energy-carbon-pollution in the above embodiments.
[0154] Example 4
[0155] Based on the same inventive concept, this invention also provides a readable storage medium, specifically an electronic device readable storage medium (Memory). This readable storage medium is a memory device within an electronic device used to store programs and data. It is understood that the storage medium here can include both built-in storage media within the electronic device and extended storage media supported by the electronic device. The storage medium provides storage space, which stores the terminal's operating system. Furthermore, this storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more executable programs (including program code). It should be noted that the storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device. Loading and executing one or more instructions stored in the storage medium by the processor can implement the steps of a multi-source data governance method for electricity-energy-carbon-pollution described in the above embodiments.
[0156] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0157] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0158] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0159] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit its scope of protection. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that after reading the present invention, they can still make various changes, modifications or equivalent substitutions to the specific implementation methods of the application, but these changes, modifications or equivalent substitutions are all within the scope of protection of the claims pending approval.
Claims
1. A multi-source data governance method for electricity, energy, carbon, and pollution, characterized in that, include: Collect multi-source time-series data of objects at different levels, including power data, energy data, carbon emission data, and pollutant data; Multi-dimensional data characteristic analysis is performed on the multi-source time series data to obtain the abnormal characteristics, missing characteristics, and spatiotemporal characteristics of the multi-source time series data; based on the abnormal characteristics, missing characteristics, and spatiotemporal characteristics of the multi-source time series data, data cleaning is performed on the multi-source time series data to obtain cleaned multi-source data; The object features, business features, and time features of the cleaned multi-source data are extracted respectively, and multi-level labels are added to the cleaned multi-source data based on the object features, business features, and time features; Construct a multi-level directory structure that corresponds one-to-one with the tag level of the multi-level tags; Based on the multi-level directory structure, a storage directory for the cleaned multi-source data is constructed, and a target table data model is formed based on the storage directory for data storage, so as to realize the governance of the multi-source time-series data. The process includes, after performing missing data filling and time-space alignment based on the spatiotemporal characteristics of the multi-source time-series data to obtain cleaned multi-source data, the process further includes: verifying the cleaned multi-source data from different sources. The missing data characteristics include short-term missing data, medium-term missing data, and long-term missing data. For data with short-term missing data, linear interpolation or the mean of the same period is used to fill in the missing values. For data with medium-term missing data, neural network prediction or the extreme gradient boosting algorithm is used to fill in the missing values. For data with long-term missing data, time series decomposition is used to perform feature analysis on the data to be processed to obtain the long-term trend features and short-term periodic features of the data to be processed. The input features of the model, as well as the long-term trend features and short-term periodic features of the data to be processed, are input into the extreme gradient boosting algorithm model to predict the missing values of the data to be processed, and the prediction results are used to fill in the missing values of the data to be processed. The data to be processed is the data with long-term missing data in the multi-source time series data, and the input features are the index features calculated based on the data to be processed.
2. The method according to claim 1, characterized in that, The collection of multi-source time-series data of objects at different levels includes: Collect electricity consumption statistics, energy statistics, carbon emission statistics, and pollutant statistics within each administrative region; Collect actual electricity consumption data, actual production capacity data, carbon emission monitoring data, and pollutant monitoring data from different enterprises in various industries.
3. The method according to claim 1 or 2, characterized in that, The data cleaning process, based on the anomaly characteristics, missing characteristics, and spatiotemporal characteristics of the multi-source time-series data, yields cleaned multi-source data, including: Based on the abnormal characteristics of the multi-source time series data, outlier values of the multi-source time series data are removed to obtain data after outlier removal. Based on the missing characteristics of the multi-source time series data, missing values of the data after removing anomalies are filled to obtain the missing-filled data. Based on the spatiotemporal characteristics of the multi-source time-series data, the time and space of the missing data are aligned to obtain the cleaned multi-source data.
4. The method according to claim 3, characterized in that, Before removing outliers from the multi-source time series data based on its abnormal characteristics to obtain the data after anomaly removal, the method further includes: The multi-source time-series data is subjected to data deduplication processing; After imputing missing values in the data following the removal of anomalies based on the missing characteristics of the multi-source time series data, and obtaining the imputed data, the method further includes: The missing data is then standardized in terms of units and format.
5. The method according to claim 1, characterized in that, The multi-level tagging includes first-level tags, second-level tags, and third-level tags. Adding multi-level tags to the cleaned multi-source data based on the object characteristics, business characteristics, and time characteristics includes: Based on the object characteristics, multi-level object labels are added to the cleaned multi-source data as the first-level labels; Based on the aforementioned business characteristics, business category tags are added as secondary tags to the cleaned multi-source data; Based on the time characteristics, time granularity labels are added to the cleaned multi-source data as the third-level labels; The multi-level object labels include regional labels, industry labels, and enterprise labels, and the business category labels include electricity category labels, energy category labels, carbon emission category labels, and pollutant category labels.
6. A multi-source data governance system for electricity, energy, carbon, and pollution, characterized in that, include: The data collection module is used to collect multi-source time-series data of objects at different levels, including power data, energy data, carbon emission data, and pollutant data. The data cleaning module is used to perform multi-dimensional data characteristic analysis on the multi-source time series data to obtain the abnormal characteristics, missing characteristics, and spatiotemporal characteristics of the multi-source time series data; and to perform data cleaning on the multi-source time series data based on the abnormal characteristics, missing characteristics, and spatiotemporal characteristics of the multi-source time series data to obtain cleaned multi-source data. The tag adding module is used to extract the object features, business features, and time features of the cleaned multi-source data respectively, and add multi-level tags to the cleaned multi-source data based on the object features, business features, and time features; The data storage module is used to construct a multi-level directory structure that corresponds one-to-one with the tag level of the multi-level tags; Based on the multi-level directory structure, a storage directory for the cleaned multi-source data is constructed, and a target table data model is formed based on the storage directory for data storage, so as to realize the governance of the multi-source time-series data. The data cleaning module further includes a verification submodule, used to verify the cleaned multi-source data from different sources; the missing characteristics include short-term missing, medium-term missing, and long-term missing; for data with short-term missing characteristics, linear interpolation or the mean of the same period is used to fill in the missing values; for data with medium-term missing characteristics, neural network prediction or the extreme gradient boosting algorithm is used to fill in the missing values; for data with long-term missing characteristics, time series decomposition is used to perform feature analysis on the data to be processed to obtain the long-term trend characteristics and short-term periodic characteristics of the data to be processed; the input features of the model and the long-term trend characteristics and short-term periodic characteristics of the data to be processed are input into the extreme gradient boosting algorithm corresponding model to predict the missing values of the data to be processed, and the prediction results are used to fill in the missing values of the data to be processed; the data to be processed is the data with long-term missing characteristics in the multi-source time series data, and the input features are the index features calculated based on the data to be processed.
Citation Information
Patent Citations
Time series data standardization treatment method for energy industry
CN115238823A
Data governance system and method
CN112699175A
Power equipment state trend prediction method based on multi-factor neural network algorithm
CN116628473A
Thermal power plant intelligent monitoring system data cleaning method, system and program product
CN118519992A