Data interpolation method and system for data missing or error processing of intelligent electric meter

By adopting the optimal weighted average (OWA) data interpolation method in smart meter data, combining the advantages of linear interpolation and historical average interpolation, the problem of missing or errors of smart meter data is solved, and more accurate and reliable data interpolation is achieved, supporting the stable operation and optimized management of the power system.

CN120067550APending Publication Date: 2025-05-30JILIN POWER SUPPLY COMPANY STATE GRID JILIN ELECTRIC POWER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510091602.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Smart meter data is missing or errored due to technical failures, communication problems or installation errors, which seriously affects the analysis and decision-making process of the power system, especially in load prediction and grid status estimation.

Method used

The optimal weighted average (OWA) data interpolation method is adopted, combining the advantages of linear interpolation and historical average interpolation, and by optimizing the weight parameter α, the weighted average value of missing data is calculated to ensure the accuracy and reliability of the interpolation results.

Benefits of technology

It realizes more accurate and reliable data interpolation, can effectively handle missing or wrong data in smart meter data, and supports the stable operation and optimized management of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067550A_ABST
    Figure CN120067550A_ABST
Patent Text Reader

Abstract

The invention discloses a data interpolation method and system for data missing or error processing of an intelligent electric meter, and belongs to the field of electric power system measurement, and the method combines the advantages of linear interpolation and historical average interpolation, considers the continuity and mode regularity of load data, and is an optimal weighted average (OWA) load power data interpolation method. Therefore, more accurate and more reliable data interpolation can be realized. According to the method, additional explanatory variables are not needed, so that the method has wider applicability and simplicity and convenience in operation. The system comprises a computer readable storage medium and a processor. And the processor is used for reading the executable instruction stored in the computer readable storage medium and executing the data interpolation method for data missing or error processing of the intelligent electric meter. By using historical load power measurement data of the intelligent electric meter, missing or wrong data can be effectively interpolated in off-line and on-line environments, and support is provided for stable operation and optimal management of a power system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of power system measurement, and more specifically, relates to a data interpolation method and system for handling missing or incorrect data of smart meters. Background Art

[0002] With the development of smart grid technology, smart meters, as a key component of modern distribution systems, are increasingly widely deployed. Smart meters can provide high-frequency electricity consumption data, which is of great significance for the monitoring, management, and optimization of the power grid. However, despite the significant advantages brought by smart meter technology, the integrity and reliability of data still face challenges in practical applications. Specifically, smart meter data often suffers from missing or incorrect data due to technical failures, communication problems, or installation errors. These incomplete or inaccurate data fragments can seriously affect the analysis and decision-making processes of the power system, especially in scenarios where accurate load forecasting and power grid state estimation are required.

[0003] Traditional data interpolation methods, such as simply ignoring, linear interpolation, or using historical averages, although they can provide temporary solutions in some cases, often fail to fully capture the complexity and dynamics of load data. For example, although the linear interpolation method is computationally simple, its accuracy will significantly decrease for long missing data segments. And interpolation methods based on a single historical sample, such as only using data from the previous hour or the previous day, may lead to high variability of the estimation results due to the randomness of sample selection.

[0004] In addition, the characteristics of smart meter data are closely related to human activity patterns, showing obvious daily periodicity and seasonality. These data characteristics require interpolation methods to not only consider the continuity of data but also be able to capture the regularity of load patterns changing over time. However, existing methods often ignore these characteristics or are difficult to implement due to computational resources and data volume limitations. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a novel data interpolation method with high computational efficiency for estimating bad and missing load power measurement data in smart meter data.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] According to the first aspect of the present invention, a data interpolation method for handling missing or incorrect data of smart meters is proposed, including:

[0008] Step S1: Organize the smart meter data to be interpolated;

[0009] Step S2: Determine the time window of the smart meter data to be interpolated;

[0010] Step S3: Determine the interpolation principle for the smart meter data to be interpolated;

[0011] Step S4: Interpolate the smart meter data to be interpolated within the selected time window according to the interpolation principle;

[0012] Step S5: Repeat steps S2 - S4 to complete the interpolation and output the result;

[0013] Furthermore, the optimal weighted average OWA interpolation method is adopted as the interpolation principle in step S3, which is specifically as follows:

[0014]

[0015] Among them, is the estimated value obtained by linear interpolation LI, is the estimated value obtained by historical average HA, is the estimated value obtained by optimal weighted average OWA; the weight parameter w i is set to decay exponentially as d i > 0;

[0016]

[0017] Among them, α is the positive weight parameter, and α is optimized to obtain the optimal weight parameter α opt , and the optimal weight parameter α opt can minimize the error F(α) between the interpolation sample and the training data sample. The process of optimizing α to obtain the optimal weight parameter α opt is as follows:

[0018]

[0019] Among them, i represents the index of the missing data point, and N represents the length of the training data period or the number of samples; F i (α) represents the error between the interpolation sample of the i-th data missing point and the training data sample;

[0020] When using the squared error to calculate, F i (α) is given by the following formula:

[0021]

[0022] Among them represents the true value of the i-th data point, F′(α) = 0, and start iterating from the initial value α = α 0 ;

[0023]

[0024] where α k represents the value of the positive weight parameter in the k-th iteration, and α k+1 represents the updated value of the positive weight parameter in the next iteration. F′(α k ) represents the first-order derivative of the objective function F with respect to the positive weight parameter α at α k . F″(α k ) represents the second-order derivative of the objective function F with respect to the positive weight parameter α at α k ; until a selected convergence criterion is met.

[0025] Furthermore, in the data imputation method for missing or incorrect smart meter data, j ∈ H, |H| = N H , where N H represents the number of samples of the historical sample y j . To describe the set H, the concept of a week number WN is defined;

[0026]

[0027] where WD is a certain day of the week, WD ∈ {1, …, 7}; HH is the number of hours in a day, HH ∈ {1, …, 24}; MM is the number of minutes in an hour;

[0028] The set H is defined as including historical samples, and the day of the year DOY and the week number WN of the included historical samples are within a specific time range from the day of the year and the week number of the missing sample, with a range of ±8 days for the day of the year DOY and a range of ±1 hour and 1 minute for the week number WN.

[0029] Preferably, in the data imputation method for missing or incorrect smart meter data, α 0 = 0.001.

[0030] According to the second aspect of the present invention, a data imputation system for missing or incorrect smart meter data is proposed, including: a computer-readable storage medium and a processor;

[0031] The computer-readable storage medium is used to store executable instructions;

[0032] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the steps of the data imputation method for missing or incorrect smart meter data.

[0033] With the above design, the present invention can bring the following beneficial effects: The present invention proposes a new data interpolation method and system, namely the optimal weighted average (OWA) load power data interpolation method. This method aims to overcome the limitations of the prior art. By combining the advantages of linear interpolation and historical average interpolation, and taking into account the continuity and pattern regularity of load data, it can achieve more accurate and reliable data interpolation. The present invention does not require additional explanatory variables such as weather data, nor does it require customer-specific information, which makes this method have a wider applicability and operational simplicity. By using the historical load power measurement data of smart meters, the present invention can effectively interpolate missing or incorrect data in both offline and online environments, providing support for the stable operation and optimized management of power systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The following drawings are used to provide a further understanding of the present invention, and constitute a part of this application for the present invention. The schematic embodiments of the present invention and their descriptions are used to understand the present invention, and do not constitute an improper limitation of the present invention. In the drawings:

[0035] Figure 1 It is a flowchart of a data interpolation method for handling missing or incorrect data of smart meter data;

[0036] Figure 2 It is a data distribution diagram of data points with a 15-minute step size in an embodiment of the present invention;

[0037] Figure 3 It is a data distribution diagram of data points with a 15-second step size in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The technical solutions of the present invention will be further described below in conjunction with the drawings and embodiments, but are not limited thereto. Any modification or equivalent replacement of the technical solutions of the present invention without departing from the spirit scope of the technical solutions of the present invention shall be covered within the protection scope of the present invention.

[0039] The linear interpolation interpolation method is a commonly used processing method for estimating missing samples in a short time period from surrounding available samples. Usually, two methods, the nearest neighbor method and the interpolation method, are adopted. The operation of the nearest neighbor method is relatively simple, which is to directly set the missing sample as the nearest available sample, or take the average of several nearest samples. However, if the missing data time is slightly longer, it is more inclined to use the interpolation method because it can obtain an estimated value continuous with the existing data around. The data interpolation method proposed by the present invention selects linear interpolation. Compared with cubic interpolation or other more complex interpolation methods, linear interpolation performs more stably and consistently when dealing with missing data of different characteristics.

[0040] The linear interpolation (LI) interpolation method estimates the missing value through the nearest available values before and after.

[0041]

[0042] Among them, (x h , y h ) and (x j , y j ) are the coordinates of two adjacent non-missing data points, estimating the coordinates of the missing value by the linear interpolation imputation method.

[0043] The linear interpolation (LI) imputation method is simple and fast to operate, and only two available samples are required for imputation in each missing data period. However, as the length of the missing data period increases, the accuracy of LI imputation usually decreases.

[0044] The linear interpolation (LI) imputation often performs poorly in dealing with long-term missing data, and better estimation results can be obtained by selecting representative time periods from historical data. The simplest way to use historical data to fill in the missing values is to adopt the data samples of the previous hour, previous day, or previous month. However, using only a single sample may lead to large variations in the estimation results, and its accuracy depends largely on the time point of the missing sample. In the data imputation method proposed later, the present invention adopts the historical average (HA) imputation method, which estimates by calculating the average value of y i as N H representative historical samples y j , j ∈ H, |H| = N H .

[0045]

[0046] To describe the set H, a concept of "weeknum" (WN) is defined.

[0047]

[0048] Where: WD is the day of the week, WD ∈ {1, …, 7}; HH is the hour of the day, HH ∈ {1, …, 24}; MM is the minute of the hour;

[0049] Now, define the set H to contain those historical samples whose day of year (DOY) and week number (WN) are within a specific time range from the DOY and WN of the missing sample. In the present invention, a DOY range of ±8 days and a WN range of ±1 hour 1 minute (i.e., ±1 / 24 + 1 / (24×60)) are adopted. The DOY range ensures that the samples selected when calculating the historical average have similar seasonal characteristics in the present invention. The WN range ensures that the samples selected when calculating the historical average have similar day of the week and time of day. Such a definition can smooth the historical average profile of consecutive missing samples. If the present invention uses "rigid" time selection criteria, such as the same season, the same day of the week, and the same hour of the day, then when there are changes in seasons, weekdays, hours, etc., there will be mutations in the continuously interpolated samples.

[0050] The accuracy of historical average (HA) interpolation depends on the characteristics of the data and requires a clear historical repetition pattern. Based on these assumptions, when dealing with long-term missing data, it is expected that the average performance of HA interpolation will be better than that of linear interpolation (LI) interpolation.

[0051] Next, as Figure 1 shown, the present invention proposes a data interpolation method for missing or incorrect smart meter data, including the following: Step S1: Sort the smart meter data to be interpolated; Step S2: Determine the time window (i.e., the missing data period) of the smart meter data to be interpolated; Step S3: Determine the interpolation principle of the smart meter data to be interpolated; Step S4: Interpolate the smart meter data to be interpolated within the selected time window according to the interpolation principle; Step S5: Repeat Steps S2 to S4 to complete the interpolation and output the result; where the interpolation principle in Step S3 adopts the optimal weighted average (OWA) interpolation method, aiming to combine the accuracy of linear interpolation (LI) in short-term missing data and the accuracy of historical average (HA) in long-term missing data.

[0052] OWA interpolation estimates the missing data sample y by calculating the weighted average of the estimated value obtained by LI interpolation i .

[0053]

[0054] The weight parameter w i is set to decay exponentially as d i > 0, where d i refers to the positive distance (counted in the number of samples) to the nearest (previous or next) available sample.

[0055]

[0056] Here, α is a positive weight parameter. When d i is small (i.e., w i ≈ 1), the optimal weighted average (OWA) imputation value mainly depends on the imputation value of linear interpolation (LI) When d i is large (i.e., w i ≈ 0), the OWA imputation value mainly depends on the imputation value of historical average (HA) Next, how to select the value of α to optimize α will be elaborated in detail.

[0057] Optimal weight parameter α opt can minimize the error F(α) between the imputation samples and the training data samples.

[0058]

[0059] where i represents the index of the missing data point, N represents the length of the training data period or the number of samples; F i (α) represents the error between the imputation sample of the i-th data missing point and the training data sample; when calculated using the squared error, F i (α) is given by the following formula:

[0060]

[0061] Here and represent the true value of the i-th data point. To find the optimal solution α opt , a necessary condition is that the derivative F′(α) = 0. This so-called critical point can be found by Newton's method, starting from the initial value α = α 0 and iterating.

[0062]

[0063] where α k represents the value of the positive weight parameter in the k-th iteration, α k+1 represents the updated value of the positive weight parameter in the next iteration, F′(α k ) represents the first derivative of the objective function F with respect to the positive weight parameter α at α k , F″(α k ) represents the second derivative of the objective function F with respect to the positive weight parameter α at α k ; until a certain selected convergence criterion is met. For any set of training samples that result in F″(α) > 0 and interpolation samples and the error function F(α) are both non-convex. Therefore, if the initial value α 0 is not properly selected, the Newton's method may diverge. In actual operation, by choosing a small (but positive) initial value α 0 (e.g., α 0 = 0.001), good convergence can be obtained.

[0064] The optimal weight parameter α opt depends on the characteristics and length of the missing data period. Therefore, different values of α opt can be obtained according to the characteristics and length of different training data periods. By optimizing α for a series of randomly selected lengths and positions of training data periods, the distribution of α opt can be estimated. If the distribution of the missing data period length is known, the length of the missing data period can be sampled from it. The globally optimal α can be estimated from the mean (or median) of the obtained α opt sample distribution.

[0065] According to the method of the present invention, 15-minute data with 20 data points was randomly selected and optimized into 15-second granularity data with 1200 data points for comparison as Figure 2 and Figure 3 shown. The method proposed by the present invention effectively completes the 20 data points in Figure 2 to 1200 data points and effectively repairs the data.

[0066] In summary, the present invention proposes a new data interpolation method and system, namely the optimal weighted average (OWA) load power data interpolation method. This method aims to overcome the limitations of the prior art. By combining the advantages of linear interpolation and historical average interpolation, and taking into account the continuity and pattern regularity of load data, more accurate and reliable data interpolation can be achieved. The present invention does not require additional explanatory variables such as weather data, nor does it require customer-specific information, which makes the method have wider applicability and operational simplicity. By using the historical load power measurement data of smart meters, the present invention can effectively interpolate missing or incorrect data in offline and online environments, providing support for the stable operation and optimized management of power systems.

Claims

1. A data interpolation method for missing or erroneous data of a smart meter, comprising: Step S1: sorting the smart meter data to be interpolated; Step S2: Determine the time window for the smart meter data to be interpolated; Step S3: Determine the interpolation principle of the smart meter data to be interpolated; Step S4: interpolating the smart meter data to be interpolated within the selected time window according to the interpolation principle; Step S5: repeat steps S2 to S4 to complete interpolation and output the result; It is characterized in that the interpolation principle in step S3 adopts the optimal weighted average OWA interpolation method, which is as follows: in, is the estimated value obtained by linear interpolation LI, is the estimated value obtained by interpolation of the historical average HA, is the estimated value obtained by optimal weighted average OWA interpolation; the weight parameter w i Set to follow d i >0, exponential decay; Among them, α is a positive weight parameter, and α is optimized to obtain the optimal weight parameter α opt , the optimal weight parameter α opt It can minimize the error F(α) between the interpolation sample and the training data sample, and optimize α to obtain the optimal weight parameter α opt The process is as follows: Where i represents the index of the missing data point, N represents the length of the training data period or the number of samples; F i (α) represents the error between the interpolation sample of the i-th data missing point and the training data sample; When using squared error to calculate, F i (α) is given by the following formula: in represents the true value of the i-th data point, F′(α)=0, and iterates from the initial value α=α0; Among them, α k represents the value of the positive weight parameter in the kth iteration, α k+1 represents the updated value of the positive weight parameter in the next iteration, F′(α k ) indicates that the objective function F has a positive weight parameter α in α k The first-order derivative at k ) indicates that the objective function F has a positive weight parameter α in α k until a selected convergence criterion is met.

2. The data interpolation method for missing or erroneous data processing of smart meter according to claim 1, characterized in that: j∈H, |H|=N H , N H Represents historical sample y j In order to describe the set H, a concept of week number WN is defined; Where WD is the day of the week, WD∈{1,…,7}; HH is the hour of the day, HH∈{1,…,24}; MM is the minute of the hour; The set H is defined as containing historical samples, and the mid-year day DOY and week number WN of the included historical samples are within a specific time range with the mid-year day and week number of the missing samples, and the mid-year day DOY range is ±8 days and the week number WN range is ±1 hour and 1 minute.

3. The data interpolation method for missing or erroneous data processing of smart meter according to claim 1, characterized in that: α0=0.001。 4. A data interpolation system for processing missing or erroneous data of a smart meter, characterized in that: include: A computer readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the steps of the data interpolation method for processing missing or erroneous data of a smart meter as described in any one of claims 1 to 3.