Building power consumption data layered cleaning method, medium and electronic equipment
Through the layered cleaning method, outliers and missing values in building power consumption data are identified and processed, high-quality data processing is achieved and refined energy management is supported.
Patent Information
- Application Number
- CN202510542718.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-28
AI Technical Summary
There are a large number of outliers and missing values in the building power consumption data, which affects the accuracy and reliability of the data, which in turn affects the effectiveness of energy management.
The layered cleaning method is used to identify outliers through clustering algorithms and inter-layer mutual verification is performed to distinguish between true and false outliers. Replace the outliers with null values, and divide the missing values into long-term and local missing values, and fill them with interpolation and support vector regression algorithm respectively.
It improves the accuracy and reliability of building power consumption data, improves the accuracy of outlier recognition rate and missing value filling, and supports refined energy management.
Smart Images

Figure CN120086213A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of energy consumption data processing, and particularly to a method, medium and electronic device for hierarchical cleaning of building electricity consumption data. Background Art
[0002] The energy consumption of public buildings, especially electricity consumption, is an important source of carbon emissions. With the increasing refinement of energy management, accurate analysis of building electricity consumption data has become crucial. However, the actual building electricity consumption data collected often has various problems such as a large number of outliers and missing values, seriously affecting the accuracy and reliability of the data, and further affecting the effect of energy management.
[0003] In response to this, traditional data processing methods usually use simple statistical methods, threshold comparison or empirical judgment to process outliers and missing values, but these methods lack scientificity and systematicness, and it is difficult to effectively identify and process various complex data relationships. Especially for building electricity consumption data with the characteristics of multiple levels and multiple types, traditional processing methods often have the problem of low outlier recognition rate. In particular, building electricity consumption data is affected by multiple factors such as time, weather, and equipment type, which makes the data show complex non-linear relationships, further increasing the difficulty of data processing. Therefore, a more scientific, systematic and effective method for cleaning building electricity consumption data is needed to accurately identify and process outliers and missing values in the data, improve the reliability of the data, and provide support for refined energy management. Summary of the Invention
[0004] Aiming at the problems of poor quality of building sub-metering data, low outlier recognition rate of traditional data cleaning methods and the filling effect of missing values not conforming to data characteristics, the present invention systematically proposes a method, medium and electronic device for hierarchical cleaning of building electricity consumption data from aspects such as data stratification, outlier recognition and cross-layer verification, and classification filling of missing values, meeting the requirements of high-quality data for building energy consumption centralized control platforms and building refined management.
[0005] To achieve the above object, the present invention adopts the following technical solutions: In the first aspect, the present invention provides a method for hierarchical cleaning of building electricity consumption data, including: Obtain building electricity metering data, and stratify the data according to the types of building electricity-consuming equipment, including primary stratification data, secondary stratification data and tertiary stratification data; For each stratified data, use a clustering algorithm to identify outliers; perform cross-layer verification on the outliers in the primary, secondary and tertiary stratified data to delimit pseudo-outliers and true outliers; Keep the original data of the false outliers unchanged and replace the true outliers with null values; for each stratified data, identify the null values after outlier replacement and the null values in the original data as missing values, and classify the missing values into long-term missing values and local missing values; Use the interpolation method to fill the local missing values and use the support vector regression algorithm to predictively fill the long-term missing values to obtain the cleaned electricity consumption data.
[0006] Preferably, the first-level stratified data includes the total building electricity consumption; The second-level stratified data includes lighting and socket electricity consumption, air conditioning system electricity consumption, power system electricity consumption, and special system electricity consumption; The third-level stratified data subordinate to the lighting and socket electricity consumption includes indoor lighting electricity consumption, indoor socket electricity consumption, public area socket electricity consumption, and outdoor landscape lighting electricity consumption; The third-level stratified data subordinate to the indoor socket electricity consumption includes cold and heat source system electricity consumption, air conditioning water system electricity consumption, air conditioning air system electricity consumption, and decentralized air conditioning electricity consumption; The third-level stratified data subordinate to the public area socket electricity consumption includes elevator electricity consumption, water pump electricity consumption, fan electricity consumption, and data center electricity consumption; The third-level stratified data subordinate to the special system electricity consumption includes kitchen and restaurant electricity consumption, electric vehicle charging pile electricity consumption, mechanical garage electricity consumption, production and operation electricity consumption, bathing electricity consumption, and rental area electricity consumption.
[0007] Preferably, for each stratified data, use the clustering algorithm to identify outliers, specifically including: Normalize the first-level, second-level, and third-level stratified data; Use the time feature data and environmental feature data corresponding to the target data as clustering features, and respectively form a first-level stratified clustering feature matrix, a second-level stratified clustering feature matrix, and a third-level stratified clustering feature matrix with the first-level, second-level, and third-level stratified data; each row in the clustering feature matrix contains a sample point, and the sample point has time features and environmental features; Use the clustering algorithm to detect outliers in the first-level, second-level, and third-level stratified clustering feature matrices respectively, and mark the outliers in the first-level, second-level, and third-level stratified data correspondingly.
[0008] Preferably, use the clustering algorithm to detect outliers in the first-level, second-level, and third-level stratified clustering feature matrices respectively, and mark the outliers in the first-level, second-level, and third-level stratified data correspondingly, where the clustering algorithm is the DBSCAN algorithm; specifically including: Obtain the sample points of the first-level, second-level, and third-level stratified clustering feature matrices, calculate the Euclidean distance between the sample points, and form a distance matrix; Based on the distance matrix, if the number of points within the neighborhood radius of the sample point is not less than the minimum number of samples in the neighborhood, then take this sample point as a core point; Recursively find all core points that can be directly or indirectly reached based on the core point, as well as their neighborhood points, to form a cluster; Judge according to the preset indication function. If the sample point does not belong to the cluster, the sample point is an outlier.
[0009] Preferably, the inter-layer cross-check of outliers in the first-level, second-level, and third-level stratified data is performed to delimit pseudo-outliers and true outliers, which specifically includes: If the second-level stratified data and the first-level stratified data at the same moment are both identified as outliers, and the difference between the sum of the second-level stratified data and the first-level stratified data is within the error range of the metering instrument, then the outlier at this moment is determined to be a pseudo-outlier; If the third-level stratified data and the second-level stratified data to which it belongs at the same moment are both identified as outliers, and the difference between the sum of the third-level stratified data and the second-level stratified data to which it belongs is within the error range of the metering instrument, then the outlier at this moment is determined to be a pseudo-outlier; If only one layer of the second-level stratified data and the first-level stratified data at the same moment is identified as an outlier, then the outlier at this moment is determined to be a true outlier; If only one layer of the third-level stratified data and the second-level stratified data to which it belongs at the same moment is identified as an outlier, then the outlier at this moment is determined to be a true outlier; In other cases, it is determined to be a true outlier.
[0010] Preferably, the missing values are divided into long-term missing values and local missing values, specifically: when the electricity consumption data of each layer is empty for N or more consecutive moments, it is determined to be a long-term missing value; when the number of consecutive empty values is less than N, it is determined to be a local missing value; N is a preset positive integer.
[0011] Preferably, the support vector regression algorithm is used to predictively fill the long-term missing values, which specifically includes: Extract the corresponding time features, environmental features, and historical electricity consumption features according to the time point corresponding to the long-term missing value, and preprocess the features; Input the preprocessed features into the pre-trained SVR model to obtain the predicted values of each stratified data.
[0012] In a second aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in a method for hierarchical cleaning of building electricity consumption data described in the first aspect are implemented.
[0013] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in a method for hierarchical cleaning of building electricity consumption data described in the first aspect are implemented.
[0014] In a fourth aspect, the present invention provides a hierarchical cleaning system for building power consumption data, comprising: A data stratification module, configured to obtain building power consumption measurement data and stratify the data according to the types of building electrical equipment, including primary stratification data, secondary stratification data, and tertiary stratification data; An anomaly recognition module, configured to identify outliers for each stratified data using a clustering algorithm; perform inter-layer cross-verification on the outliers in the primary, secondary, and tertiary stratified data to demarcate pseudo-anomalies and true anomalies; A data partitioning module, configured to keep the original data of pseudo-anomalies unchanged and replace true anomalies with null values; for each stratified data, identify both the null values after replacing outliers and the null values in the original data as missing values, and classify the missing values into long-term missing values and local missing values; A data cleaning module, configured to fill local missing values using interpolation and perform predictive filling of long-term missing values using a support vector regression algorithm to obtain cleaned power consumption data.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention stratifies according to equipment types, specifically identifies outliers, and can accurately locate problem data; distinguishes true and pseudo-anomalies through inter-layer cross-verification, avoiding misjudgment caused by single judgment, and improving the accuracy of outlier processing; identifies and classifies outliers and original null values for targeted processing; fills local missing values using interpolation, simply and efficiently restoring the data; then uses a support vector regression algorithm to predictively fill long-term missing values and make reasonable estimates based on data patterns. Overall, the process is perfect, effectively improving the quality of building power consumption data and providing reliable data support for energy consumption analysis, building management, etc.
[0016] (2) In traditional outlier judgment methods, they are often limited to a single dimension or only analyze within the same level of data, making it difficult to comprehensively capture the complex relationships between data. However, the inter-layer cross-verification of the present invention breaks the routine by introducing a cross-level verification mechanism. By comparing the abnormal conditions of data at different levels at the same moment and combining the error range of metering instruments for comprehensive judgment. This cross-level association method fully explores the internal logical connections between data, no longer views each level of data in isolation, can more accurately distinguish true and pseudo-anomalies, effectively avoids misjudgment problems caused by pure algorithm detection or accidental errors, and greatly improves the accuracy and reliability of outlier judgment.
[0017] (3) Traditional data processing methods have a low recognition rate for outliers in building power consumption data and are difficult to effectively handle complex data relationships. The present invention realizes high-precision recognition of outliers by adopting a clustering algorithm and combining an inter-layer cross-check mechanism. The present invention can achieve an outlier recognition rate of more than 93% for building power consumption, far higher than traditional methods, thereby ensuring the accuracy and reliability of the data.
[0018] (4) Regarding the problem of missing values in data, traditional methods often lack scientificity and systematicness. The present invention classifies missing values into long-term missing values and local missing values, and respectively uses a support vector regression algorithm and an interpolation method for predictive filling and direct filling. In particular, for long-term missing values, the predictive filling accuracy of the present invention exceeds 95%, effectively solving the impact of data gaps on energy management.
[0019] The advantages of additional aspects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute a limitation to the present invention.
[0021] Figure 1 The main flowchart of a method for hierarchical cleaning of building power consumption data provided by an embodiment of the present invention; Figure 2 The flowchart of data acquisition and layering provided by an embodiment of the present invention; Figure 3 The flowchart of outlier recognition provided by an embodiment of the present invention; Figure 4 The flowchart of filling missing values provided by an embodiment of the present invention; Figure 5 The effect diagram of filling local missing values provided by an embodiment of the present invention; Figure 6 The effect diagram of predictive filling of long-term missing values provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The present invention will be further described below in conjunction with the drawings and embodiments.
[0023] Embodiment 1 As Figure 1 shown, this embodiment discloses a method for hierarchical cleaning of building power consumption data, including the following steps: S1: Obtain building electricity measurement data, and layer the data according to the types of building electrical equipment, including primary layer data, secondary layer data, and tertiary layer data; S2: For each layer of data, clustering algorithm is used to identify outliers; inter-layer verification is performed on outliers in the primary, secondary and tertiary layered data to define false outliers and true outliers; S3: Keep the original data of false outliers unchanged, and replace the true outliers with null values; for each layer of data, mark the null values after the outliers are replaced and the null values in the original data as missing values, and divide the missing values into long-term missing values and local missing values; S4: Use interpolation method to fill in local missing values, and use support vector regression algorithm to predictively fill in long-term missing values to obtain power consumption data after cleaning.
[0024] Next, combine Figure 1 , a layered cleaning method for building power consumption data disclosed in this embodiment is described in detail.
[0025] In S1, Figure 2 As shown, the building electricity metering data is obtained, and the data is divided into first-level hierarchical data, second-level hierarchical data and third-level hierarchical data according to the type of building electricity equipment. The data are all directly measured data by meters.
[0026] Specifically, the first-level hierarchical data is the total electricity consumption of the building; Secondary hierarchical data include electricity consumption of lighting sockets, electricity consumption of air conditioning systems, electricity consumption of power systems and electricity consumption of special systems; Each type of secondary hierarchical data has subordinate tertiary hierarchical data based on the actual electrical equipment usage of the building.
[0027] As a specific implementation method, the three-level hierarchical data belonging to the lighting socket electricity consumption includes indoor lighting electricity consumption, indoor socket electricity consumption, public area socket electricity consumption and outdoor landscape lighting electricity consumption; The three-level hierarchical data belonging to the indoor socket electricity consumption includes the electricity consumption of the cold and heat source system, the electricity consumption of the air conditioning water system, the electricity consumption of the air conditioning wind system and the electricity consumption of the distributed air conditioning system; The three-level hierarchical data belonging to the power consumption of public area sockets includes elevator power consumption, water pump power consumption, fan power consumption and data room power consumption; The three-level hierarchical data belonging to special system electricity consumption includes kitchen and restaurant electricity consumption, car charging pile electricity consumption, machinery garage electricity consumption, production and operation electricity consumption, bathing electricity consumption and rental area electricity consumption.
[0028] Before identifying outliers in power consumption data, this embodiment first divides the complex building power consumption data into different levels according to equipment type. Data at different levels correspond to different power consumption systems, and targeted anomaly detection and missing value processing can be carried out to avoid confusion, which helps to accurately locate outliers and missing values.
[0029] In S2,Figure 3 As shown, outlier identification is performed on each hierarchical data, specifically including: S201: Normalize the first-level, second-level, and third-level hierarchical data; S202: Use the time feature data and environmental feature data corresponding to the target data as clustering features, and respectively form a first-level hierarchical clustering feature matrix, a second-level hierarchical clustering feature matrix, and a third-level hierarchical clustering feature matrix with the first-level, second-level, and third-level hierarchical data; each row in the clustering feature matrix contains a sample point, and the sample point has time features and environmental features; S203: Use the clustering algorithm to detect outliers in the first-level, second-level, and third-level hierarchical clustering feature matrices respectively, and mark the outliers in the first-level, second-level, and third-level hierarchical data correspondingly.
[0030] In S201, the normalization method is specifically:
[0031] In the formula, is the value after normalization, is the original hierarchical data, and represent the maximum and minimum values in the original hierarchical data.
[0032] In S202, use the time feature data and environmental feature data corresponding to the target data as clustering features, and respectively form a first-level hierarchical clustering feature matrix (F1), a second-level hierarchical clustering feature matrix (F2), and a third-level hierarchical clustering feature matrix (F3) with the first-level, second-level, and third-level hierarchical data.
[0033] For example, assume that it is necessary to analyze the building electricity consumption data of a certain shopping mall at different times of each day within a month. Select the electricity consumption data of a certain week as the key analysis object for this data cleaning. The building electricity measurement data of this week is the "target data". When detecting outliers in this part of the target data, its corresponding time feature data (such as the specific day of the week and the number of hours per day within this week) and environmental feature data (such as the outdoor temperature and weather conditions of each day within this week) will be used to form a clustering feature matrix with the first-level, second-level, and third-level hierarchical data, and then subsequent operations such as outlier identification will be carried out.
[0034] Among them, the feature matrix can be expressed as: X = [month, week, day, hour, outdoor temperature, indoor temperature, solar radiation intensity, sunny / cloudy / rainy state, hierarchical electricity consumption at the previous moment, hierarchical electricity consumption at the previous three moments, hierarchical electricity consumption at the corresponding moment of the previous day, hierarchical electricity consumption at the corresponding moment of the previous three days, hierarchical electricity consumption at the corresponding moment of the previous seven days].
[0035] Further, the specific representation forms of the F1, F2, and F3 matrices are as follows: each row contains the time features and environmental features of a sample.
[0036] Exemplarily, the time feature data uses the month, week, day, and hour numbers corresponding to each hierarchical data; the environmental feature data uses the outdoor temperature, indoor temperature, solar radiation intensity, and sunny / cloudy / rainy state.
[0037] In S203, the DBSCAN clustering algorithm is used to detect outliers in the F1, F2, and F3 matrices respectively, and the outliers of the first-level, second-level, and third-level hierarchical data are marked correspondingly.
[0038] Specifically, the specific method for identifying outliers in the F1, F2, and F3 matrices using the DBSCAN algorithm is as follows: For each pair of samples in the F1, F2, and F3 matrices, calculate the Euclidean distance between them according to the following formula to form a distance matrix:
[0039] In the formula, p and q are two n-dimensional sample points, and are their values on the i-th dimension respectively.
[0040] For each sample point, if Neps(p) ≥ min_samples, then the point p is a core point. Among them, Neps(p) represents the number of points within the neighborhood radius (eps) neighborhood of the point p, and min_samples is the minimum number of samples within the neighborhood.
[0041] Starting from a core point p, recursively find all core points that can be directly or indirectly reached, as well as the neighborhood points of these core points, to form a cluster.
[0042] Let , , …, represent all the clusters in the dataset, and each cluster is a set of points. Define an indicator function I(p ∈ Ci). If the point p belongs to the cluster Ci, then I(p ∈ Ci) = 1; otherwise, I(p ∈ Ci) = 0.
[0043] If ∀i ∈ {1, 2, …, k}, I(p ∈ Ci) = 0, then p is an outlier.
[0044] Further, the hyperparameter values of the DBSCAN algorithm are optimized based on the F1, F2, and F3 matrices respectively, including the eps value and the min_samples value. In this process, for each matrix, more than 20 typical outliers are determined manually according to the actual situation. The grid search technique is used to iteratively optimize the hyperparameter values until all the typical outliers determined manually are identified by the clustering results, and the optimal hyperparameter values are obtained. The optimal hyperparameters are applied to determine all the outliers in the first-level, second-level, and third-level hierarchical data.
[0045] In this embodiment, the DBSCAN algorithm is used to detect outliers in the hierarchical data of building electricity consumption. It can effectively identify clusters of any shape, which is in line with the complex and variable distribution characteristics of building electricity consumption data, and can accurately find outliers. Secondly, there is no need to pre-specify the number of clusters, avoiding the problem of affecting the detection results due to improper subjective setting of the number of clusters. Moreover, by setting hyperparameters such as the neighborhood radius and the minimum number of samples, the sensitivity of the detection can be flexibly adjusted according to the actual situation. And during the process of determining the hyperparameters, iterative optimization with the grid search technique can ensure the accuracy of the detection, effectively identify all the outliers in each level of hierarchical data, and provide a reliable data basis for subsequent data cleaning and building energy consumption analysis.
[0046] As an implementation method, an inter-layer cross-check is performed on the outliers in the first-level, second-level, and third-level hierarchical data to demarcate pseudo-outliers and true outliers, specifically including: S211: If the second-level hierarchical data and the first-level hierarchical data at the same moment are both identified as outliers, and the difference between the sum of the second-level hierarchical data and the first-level hierarchical data is within the error range of the metering instrument, then the outlier at this moment is determined to be a pseudo-outlier; S212: If the third-level hierarchical data and the second-level hierarchical data to which it belongs at the same moment are both identified as outliers, and the difference between the sum of the third-level hierarchical data and the second-level hierarchical data to which it belongs is within the error range of the metering instrument, then the outlier at this moment is determined to be a pseudo-outlier; S213: If only one layer of the second-level hierarchical data and the first-level hierarchical data at the same moment is identified as an outlier, then the outlier at this moment is determined to be a true outlier; S214: If only one layer of the third-level hierarchical data and the second-level hierarchical data to which it belongs at the same moment is identified as an outlier, then the outlier at this moment is determined to be a true outlier; S215: All other situations are determined to be true outliers.
[0047] In this embodiment, cross-layer verification is performed on the outliers of each layer, breaking through the limitation of traditional single-dimensional outlier judgment. Based on the characteristics of building electricity consumption hierarchical data, a multi-level correlation verification mechanism is constructed. By comparing the outliers between different levels of data and combining with the error range of the metering instruments, true and false outliers are accurately distinguished. This method fully considers the internal logical relationship between the data, effectively avoiding misjudgments caused by simple algorithm detection or accidental errors, and improving the accuracy and reliability of outlier judgment.
[0048] In S3, keep the original data of the false outliers unchanged and replace the true outliers with null values.
[0049] As a specific implementation, replacing the true outliers with null values specifically includes the following steps: Create a DataFrame containing timestamps and values based on Python. Define a judgment function using the above cross-layer verification to judge the true outliers, that is, represent the cross-layer verification rule in the form of a judgment function. Use the apply method to apply the judgment function to each hierarchical data column of the DataFrame and create a new column to store the results. Finally, use the loc method and boolean indexing to find all the rows marked as true outliers, replace the values of each hierarchical data column of these rows with NaN, and use the drop method to delete the auxiliary column used to mark the true outliers.
[0050] Furthermore, identify the true outliers replaced with null values as missing values.
[0051] As a specific implementation, identify the null values in the original data as missing values, specifically including: Create a complete time series. First, create a complete time index according to the time range and appropriate time intervals (such as days, hours, etc.) to ensure the continuity of the time series. Then, realign the target data to the complete time series according to the time index, using the reindex() method in Python for matching, so that the timestamps of the original data are aligned to the complete time axis. For the time points that are not matched, the data is filled with missing values (NaN).
[0052] In this embodiment, by creating a complete time series and realigning the target data with this time series, it is possible to comprehensively sort out the null value situation in the target data under a unified time dimension. The unified time framework provides a solid foundation for accurately identifying null values, and can avoid misjudgments of null values caused by discontinuous or inconsistent time. Based on the accurately identified null values, subsequent processing operations such as more accurate classification and filling of missing values can be carried out, thereby effectively improving the data quality and the reliability of subsequent analysis results.
[0053] Further, the missing values are classified into two categories: long-term missing values and local missing values.
[0054] When the power consumption data at N or more consecutive moments in each layer is a null value, it is determined as the long-term missing value; when the number of consecutive null values is less than N, it is determined as the local missing value; where N is a preset positive integer. In this embodiment, preferably N = 5.
[0055] In S4, as Figure 4 shown, the local missing values are filled by the interpolation method, and the long-term missing values are predictively filled by the support vector regression algorithm to obtain the cleaned power consumption data.
[0056] In this embodiment, the local missing values are filled by the following cubic spline interpolation method:
[0057] where f(x) is the interpolated data value, , , , are the interpolation coefficients solved by the least squares method, are the positions of each known hierarchical data point, and x is the interpolation point. The filling effect of the local missing values is as Figure 5 shown.
[0058] The process of predictively filling the long-term missing values by the support vector regression algorithm is as follows: The time feature data, environmental feature data corresponding to the target data, and the historical power consumption data of each layer are used to form the first-level hierarchical predictive filling feature matrix X1, the second-level hierarchical predictive filling feature matrix X2, and the third-level hierarchical predictive filling feature matrix X3; where the historical power consumption data is the first-level, second-level, and third-level hierarchical data at the previous moment, the previous three moments, the corresponding moments of the previous day, the corresponding moments of the previous three days, and the corresponding moments of the previous seven days.
[0059] For each hierarchical predictive filling feature matrix, an SVR model is constructed respectively. The objective function of SVR is:
[0060] where w is the weight vector; b is the bias term; is the linear model; C is the regularization parameter, which is used to control the complexity and fault tolerance of the model, is the insensitive loss function, is the sample corresponding true value, and is calculated in the following manner:
[0061] Map the input data to a high-dimensional space through a kernel function (such as a Gaussian kernel or a polynomial kernel):
[0062] in, is a kernel function, and the optional kernel functions include a linear kernel, a polynomial kernel, and a radial basis function (RBF) kernel. Those skilled in the art may select one according to actual needs, and the present invention does not impose any limitation on this; , is the Lagrange multiplier; b is the deviation term; are the support vectors.
[0063] Furthermore, the SVR model is trained using the X1, X2, and X3 matrices, and the optimal regularization parameter C and kernel function are selected through cross-validation.
[0064] For the missing power consumption data of each layer, the trained SVR model is used to make predictions as follows: According to the time points corresponding to the missing data, the corresponding time features, environmental features and historical power consumption features are extracted. The extracted features are normalized and encoded to meet the model input requirements; The preprocessed features are input into the SVR model to obtain the predicted data values of each layer; Fill the missing positions in each layer of data with the predicted values. The long-term predictive filling effect is as follows: Figure 6 shown.
[0065] In this embodiment, the number of continuous null values is used as the standard for classification, so that the missing value classification is more in line with the actual data characteristics, providing a reasonable basis for subsequent processing. For local missing values, the cubic spline interpolation method is used, which can accurately fit the known data point positions, smoothly fill in the gaps, and retain the local characteristics of the data. For long-term missing values, the support vector regression algorithm is used to build a model based on time, environment and historical power consumption characteristics. By cross-validating and optimizing parameters, the potential laws of the data can be effectively mined and accurate predictive filling can be achieved. The combination of the two improves the accuracy and adaptability of data filling and provides a high-quality data foundation for the analysis and application of building power consumption data.
[0066] In this specific embodiment, for outlier identification, based on the DBSCAN clustering algorithm, a clustering feature matrix is constructed by combining time and environmental features, which can effectively adapt to the complex distribution of building power consumption data, accurately identify outliers, and is different from traditional single-dimensional detection methods. In terms of inter-layer cross-verification, true and false outliers are delimited according to the data relationships at different levels and the error ranges of metering devices, breaking through the traditional single judgment mode, fully considering the internal logic of the data, reducing the false judgment rate, and providing a reliable basis for subsequent processing. In terms of missing value classification, continuous null value counts are used as the standard to divide into long-term and local missing values, making the filling strategy more targeted. For different types of missing values, interpolation methods and support vector regression algorithms are respectively used to improve the accuracy and rationality of data filling. The overall solution provides a systematic and innovative method for cleaning building power consumption data, which is of great significance to building energy consumption management.
[0067] Embodiment 2 This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in a method for hierarchical cleaning of building power consumption data as described in Embodiment 1 above.
[0068] Embodiment 3 This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a method for hierarchical cleaning of building power consumption data as described in Embodiment 1 above.
[0069] Embodiment 4 This embodiment provides a system for hierarchical cleaning of building power consumption data, including: A data stratification module, configured to obtain building electricity consumption measurement data and stratify the data according to the types of building electrical equipment, including primary stratification data, secondary stratification data, and tertiary stratification data; An outlier identification module, configured to identify outliers for each stratified data using a clustering algorithm; perform inter-layer cross-verification on the outliers in the primary, secondary, and tertiary stratified data to delimit false outliers and true outliers; A data classification module, configured to keep the original data of false outliers unchanged, replace true outliers with null values; for each stratified data, mark the null values after outlier replacement and the null values in the original data as missing values, and divide the missing values into long-term missing values and local missing values; A data cleaning module, configured to fill local missing values using interpolation methods and perform predictive filling on long-term missing values using support vector regression algorithms to obtain the cleaned power consumption data.
[0070] The steps or modules involved in the second to fourth embodiments above correspond to those in the first embodiment. For specific implementation manners, reference may be made to the relevant description part of the first embodiment. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and cause the processor to execute any method in the present invention.
[0071] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A layered cleaning method for building power consumption data, characterized in that: include: Acquire building electricity metering data, and stratify the data according to the type of building electricity equipment, including first-level stratified data, second-level stratified data, and third-level stratified data; Clustering algorithms are used to identify outliers for each stratified data; inter-stratum verification is performed on outliers in the primary, secondary, and tertiary stratified data to identify false outliers and true outliers; Keep the original data of false outliers unchanged and replace the true outliers with null values; for each layer of data, mark the null values after the outliers are replaced and the null values in the original data as missing values, and divide the missing values into long-term missing values and local missing values; Interpolation method is used to fill in local missing values, and support vector regression algorithm is used to predictively fill in long-term missing values to obtain the power consumption data after cleaning.
2. A layered cleaning method for building power consumption data according to claim 1, characterized in that: The first-level hierarchical data includes total building electricity consumption; The secondary hierarchical data includes electricity consumption of lighting sockets, electricity consumption of air conditioning systems, electricity consumption of power systems and electricity consumption of special systems; The three-level hierarchical data subordinate to lighting socket electricity consumption includes indoor lighting electricity consumption, indoor socket electricity consumption, public area socket electricity consumption and outdoor landscape lighting electricity consumption; The three-level hierarchical data belonging to the indoor socket electricity consumption includes the electricity consumption of the cold and heat source system, the electricity consumption of the air conditioning water system, the electricity consumption of the air conditioning wind system and the electricity consumption of the distributed air conditioning system; The three-level hierarchical data belonging to the power consumption of public area sockets includes elevator power consumption, water pump power consumption, fan power consumption and data room power consumption; The three-level hierarchical data belonging to special system electricity consumption includes kitchen and restaurant electricity consumption, car charging pile electricity consumption, machinery garage electricity consumption, production and operation electricity consumption, bathing electricity consumption and rental area electricity consumption.
3. A layered cleaning method for building power consumption data according to claim 1, characterized in that: The clustering algorithm is used to identify outliers for each layer of data, specifically including: Normalize the primary, secondary, and tertiary stratified data; The time feature data and the environmental feature data corresponding to the target data are used as clustering features, and are respectively combined with the first-level, second-level and third-level hierarchical data to form a first-level hierarchical clustering feature matrix, a second-level hierarchical clustering feature matrix and a third-level hierarchical clustering feature matrix; each row in the clustering feature matrix contains a sample point, and the sample point has a time feature and an environmental feature; The clustering algorithm is used to detect outliers on the primary, secondary and tertiary hierarchical clustering feature matrices, and the primary, secondary and tertiary hierarchical data outliers are marked accordingly.
4. A method for layered cleaning of building power consumption data as claimed in claim 3, characterized in that: The clustering algorithm is used to perform outlier detection on the primary, secondary and tertiary hierarchical clustering feature matrices respectively, and the primary, secondary and tertiary hierarchical data outliers are marked accordingly, wherein the clustering algorithm is the DBSCAN algorithm; specifically, it includes: Obtain sample points of the first-level, second-level, and third-level hierarchical clustering feature matrices, calculate the Euclidean distances between the sample points, and construct a distance matrix; Based on the distance matrix, if the number of points within the neighborhood radius of the sample point is not less than the minimum number of samples in the neighborhood, the sample point is regarded as the core point; Based on the core point, recursively find all core points that can be reached directly or indirectly, as well as their neighboring points, to form a cluster; According to the preset indicator function, if the sample point does not belong to the cluster, the sample point is an outlier.
5. A method for layered cleaning of building power consumption data as claimed in claim 1, characterized in that: The inter-layer verification of outliers in the primary, secondary and tertiary stratified data to define false outliers and true outliers specifically includes: If the secondary stratified data and the primary stratified data at the same time are both identified as outliers, and the difference between the sum of the secondary stratified data and the primary stratified data is within the error range of the meter, then the outlier at that time is determined to be a false outlier; If the third-level stratified data and the corresponding second-level stratified data at the same time are both identified as outliers, and the difference between the sum of the third-level stratified data and the corresponding second-level stratified data is within the error range of the meter, then the outlier at that time is determined to be a false outlier; If only one layer of the secondary stratified data and the primary stratified data at the same time is identified as an outlier, the outlier at that time is determined to be a true outlier; If only one layer of the three-level stratified data and the corresponding two-level stratified data at the same time is identified as an outlier, the outlier at that time is determined to be a true outlier; Other cases were judged as true outliers.
6. A method for layered cleaning of building power consumption data as claimed in claim 1, characterized in that: The missing values are divided into long-term missing values and local missing values. Specifically, when the power consumption data of each layer is null for N consecutive moments or more, it is determined to be a long-term missing value; when the number of null values at consecutive moments is less than N, it is determined to be a local missing value; N is a preset positive integer.
7. A method for layered cleaning of building power consumption data as claimed in claim 1, characterized in that: The support vector regression algorithm is used to predictively fill in long-term missing values, specifically including: According to the time points corresponding to the long-term missing values, the corresponding time features, environmental features and historical power consumption features are extracted, and the features are preprocessed; The preprocessed features are input into the pre-trained SVR model to obtain the predicted data values of each layer.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of a method for layered cleaning of building power consumption data as described in any one of claims 1 to 7 are implemented.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the method for layered cleaning of building power consumption data as described in any one of claims 1-7 are implemented.
10. A building power consumption data layered cleaning system, characterized in that: include: A data stratification module is configured to obtain building electricity metering data and stratify the data according to the type of building electricity equipment, including first-level stratified data, second-level stratified data and third-level stratified data; The anomaly identification module is configured to use a clustering algorithm to identify anomalies for each layer of data; perform inter-layer verification on the anomalies in the primary, secondary and tertiary layered data, and define false anomalies and true anomalies; The data partitioning module is configured to keep the original data of the false outliers unchanged and replace the true outliers with null values; for each layer of data, the null values after the outliers are replaced and the null values in the original data are marked as missing values, and the missing values are divided into long-term missing values and local missing values; The data cleaning module is configured to use the interpolation method to fill in the local missing values and the support vector regression algorithm to predictively fill in the long-term missing values to obtain the cleaned power consumption data.
Citation Information
Patent Citations
Power consumption data outlier detection and cleaning method based on DBSCAN and KNN algorithms
CN116089405A
Energy consumption metering statistics and energy flow presentation method, device and system and medium
CN117709582A
Big data assisted pollution and carbon reduction method and device based on power-economic data characteristics
CN118069632A
Scientific research building room energy consumption anomaly detection method based on ensemble learning and deep learning
CN118606861A
Power grid maintenance resource consumption prediction method and system
CN119005432A
Cited By
Data cleaning method combining spatio-temporal clustering and iterative threshold shrinkage algorithm
CN121188359A
A data cleaning method combining spatiotemporal clustering and iterative threshold shrinkage algorithm
CN121188359B
Park carbon emission accounting-oriented energy consumption data verification and completion method and system and medium
CN122155098A