A method and system for missing data imputation in virtual power plants

By calculating power consumption stability and similarity, and selecting reference users, and combining the temporal attention layer of the neural network model, the problem of missing data filling in the virtual power plant was solved, improving the accuracy of data filling and the accuracy of scheduling decisions in the virtual power plant.

CN122087288APending Publication Date: 2026-05-26STATE GRID SHANDONG ELECTRIC POWER CO MARKETING SERVICE CENT (MEASURING CENT)
View PDF 11 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID SHANDONG ELECTRIC POWER CO MARKETING SERVICE CENT (MEASURING CENT)
Filing Date
2026-04-23
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing data imputation methods fail to effectively consider the differences in missing electricity consumption data at different times and ignore individual user differences and group correlations, resulting in insufficient accuracy and timeliness of virtual power plant load forecasting and dispatching decisions.

Method used

By calculating power consumption stability, missing impact, and power consumption similarity, reference users are selected. Then, by combining the temporal attention layer of the neural network model and embedding prior weights, missing data is predicted and filled in.

Benefits of technology

It significantly improves the accuracy of missing data filling, making the filling results closer to the actual electricity consumption pattern, providing reliable data support for load forecasting and dispatching decisions of virtual power plants, and enhancing the accuracy and timeliness of peak shaving and valley filling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122087288A_ABST
    Figure CN122087288A_ABST
Patent Text Reader

Abstract

This invention proposes a method and system for missing data imputation in virtual power plants, belonging to the field of data processing technology. The method includes: acquiring electricity consumption data for each user on the power grid load side at various times within a preset neighboring period; calculating the electricity consumption stability of each local time period; calculating the missing impact and data importance of each user in each local time period; obtaining the electricity consumption similarity between any two users in each local time period to select reference users with similar electricity consumption behavior for each user; determining the group similarity and contribution of each user in each local time period, and setting prior weights for each time period accordingly, embedding these weights as attention bias terms in a time-series attention layer into a neural network model to predict and imput missing data. This invention enables the imputation results to closely approximate real electricity consumption patterns, improving the accuracy of missing data imputation, and thus providing reliable data support for load forecasting and scheduling decisions in virtual power plants.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data processing technology, and in particular relates to a method and system for filling in missing data in a virtual power plant. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Virtual power plants, as a key platform for aggregating distributed energy resources to participate in grid dispatch, rely on accurate perception and forecasting of user-side electricity load for operation. However, due to factors such as equipment aging, communication failures, and environmental interference, the collected data often has random gaps. These missing data not only affect the accuracy of load forecasting but may also lead to errors in dispatch decisions, thereby affecting the stable operation of the power grid.

[0004] Because electricity data typically exhibits distinct temporal characteristics, past electricity consumption patterns often influence future ones. However, existing data imputation methods fail to consider the varying impacts of missing electricity data at different times. Furthermore, due to the complex spatiotemporal heterogeneity, individual volatility, and group coordination inherent in user electricity consumption behavior, these methods neglect individual differences and group correlations, failing to effectively capture load change trends and periodic patterns. Consequently, the imputation results deviate from actual electricity consumption patterns, failing to provide reliable data support for load forecasting and dispatching decisions in virtual power plants, and impacting the accuracy and timeliness of peak shaving and valley filling in virtual power plants. Summary of the Invention

[0005] To overcome the shortcomings of the prior art, the present invention provides a method and system for filling missing data in a virtual power plant, which can make the filling results closer to the real power consumption pattern, improve the accuracy of missing data filling, and thus provide reliable data support for load forecasting and dispatching decisions of the virtual power plant.

[0006] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: The first aspect of the present invention provides a method for filling in missing data for a virtual power plant.

[0007] A method for missing data imputation in a virtual power plant, comprising: Acquire electricity consumption data for each user on the power grid load side at various times within a preset adjacent period; For each user, the volatility and trend of electricity consumption data within the corresponding local time period are analyzed in order to calculate the electricity consumption stability of each local time period; Based on the missing electricity consumption data within a local time period and the changing trend at the location of the missing data, the impact of the missing data on each user in each local time period is calculated; and combined with the electricity consumption stability, the data importance of each user in each local time period is determined. Based on the synergy of changes in electricity consumption data of any two users in each local time period, reference users with similar electricity consumption behavior are selected for each user by calculating the similarity of electricity consumption of any two users in each local time period. For each local time period, the consistency and similarity of electricity consumption trends of each user and corresponding reference user at a single moment are analyzed to determine the group similarity of each user in each local time period; combined with data importance, the contribution of each user in each local time period is determined, and the prior weight of each moment is set according to the contribution; by embedding the prior weight as an attention bias term of the temporal attention layer into the neural network model, missing data is predicted and imputed.

[0008] Furthermore, the power consumption stability for each local time period is calculated, including: For each user's local time period, obtain the extreme points of electricity consumption data for all moments within the internal time, and use the difference in electricity consumption data between two adjacent extreme points as the difference quantity; Curve fitting is performed on the electricity consumption data at all times within a local time period to calculate the slope of the tangent line at each time on the fitted curve, and the absolute values ​​of all tangent line slopes contained between two adjacent extreme points are positively merged as the change. The power stability is negatively correlated with all differences and all changes.

[0009] Furthermore, the impact of missing data for each user in each local time period is calculated, including: For local time periods, the corresponding time with missing data is marked as a missing time, and the time period consisting of consecutively adjacent missing times is defined as a missing time period; the time interval between each missing time and the time corresponding to the nearest extreme point is calculated; the changing trend of user data on the left and right sides of the missing time period where each missing time is located and the duration of the missing data are analyzed, and the potential criticality of each missing time is calculated in combination with the time interval. The percentage of missing moments within each local time period was calculated. The impact of the missing information is positively correlated with the proportion and the potential criticality.

[0010] Furthermore, the calculation of the potential criticality includes: The remaining time points, excluding the missing time points, are defined as non-missing time points. For each missing time point, the electricity consumption data of multiple consecutive non-missing time points before the first missing time point and multiple consecutive non-missing time points after the last missing time point within the missing time point are linearly fitted, and the slope of the fitted line is calculated. The duration of the missing time period to which each missing moment belongs is calculated, and a negative mapping is performed on it. The negative mapping result is positively fused with the absolute value of the slope to serve as the trend change of each missing moment. The potential criticality is negatively correlated with the time interval and positively correlated with the trend change.

[0011] Furthermore, by calculating the similarity of electricity consumption between any two users in each local time period, reference users with similar electricity consumption behavior are selected for each user, including: For each user's corresponding local time period, if each time period is missing, the label value is assigned to 1; otherwise, it is assigned to 0. When any two users in each local time period have the same label value at the same time, this time is defined as a synchronization time, and the time period consisting of consecutive adjacent synchronization times is defined as a synchronization time period. Calculate the correlation between all electricity consumption data of any two users in the synchronization time period. The electricity consumption similarity is the result of positively fusing the duration and correlation of all synchronous periods of any two users in each local time period; for each local time period, the remaining users whose electricity consumption similarity with each user is greater than or equal to a preset threshold are marked as reference users.

[0012] Furthermore, determining the group similarity of each user in each local time period includes: For each user and all corresponding reference users, users with positive tangent slopes and users with negative tangent slopes at the same time within each local time period are counted, forming the first user set and the second user set for each time period. The first user set and the second user set are labeled as user sets. The number of all users in a single user set is counted, and all users in the user set with the largest number are selected as the dominant user. The largest number is used as the intensity of the electricity consumption trend at each time period. Calculate the dispersion of the tangent slopes for all dominant users at each time point; Among them, the group similarity is positively correlated with the intensity of electricity consumption trend and the similarity of electricity consumption, and negatively correlated with the degree of dispersion.

[0013] Furthermore, the data importance is positively correlated with power consumption stability and negatively correlated with the impact of missing data; the contribution is the result of positive fusion of data importance and group similarity, and the prior weight is set as the contribution of each user to the local time period at each moment.

[0014] A second aspect of the present invention provides a missing data filling system for a virtual power plant.

[0015] A missing data imputation system for a virtual power plant includes: The data acquisition module is configured to acquire the electricity consumption data of each user on the power grid load side at each time within a preset adjacent period; The power consumption stability calculation module is configured to: analyze the volatility and trend of power consumption data in the corresponding local time period for each user, so as to calculate the power consumption stability of each local time period; The data importance calculation module is configured to: calculate the impact of missing data for each user in each local time period based on the missing data and the changing trend at the missing locations; and determine the data importance of each user in each local time period in conjunction with the power consumption stability. The reference user determination module is configured to: based on the synergy of changes in the electricity consumption data of any two users in each local time period, and by calculating the similarity of electricity consumption of any two users in each local time period, select reference users with similar electricity consumption behavior for each user; The missing data imputation module is configured to: determine the group similarity of each user in each local time period by analyzing the consistency and similarity of electricity consumption trends of each user and the corresponding reference user at a single moment; determine the contribution of each user in each local time period by combining data importance, and set the prior weight of each moment according to the contribution; and predict and impute missing data by embedding the prior weight as an attention bias term of the temporal attention layer into the neural network model.

[0016] A third aspect of the invention provides a computer-readable storage medium having a program stored thereon that, when executed by a processor, implements the steps of a missing data filling method for a virtual power plant as described in the first aspect of the invention.

[0017] A fourth aspect of the present invention provides an electronic device including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the steps of a missing data filling method for a virtual power plant as described in the first aspect of the present invention.

[0018] The above one or more technical solutions have the following beneficial effects: This invention comprehensively evaluates the importance of user data at each time period by calculating the stability and impact of missing data in each local time period, considering both intrinsic data quality and the degree of external damage. Simultaneously, it constructs a similarity score based on the synergy of electricity consumption changes among users, identifies reference users with similar electricity consumption behaviors, and further analyzes the consistency of electricity consumption trends between users and reference groups to determine group similarity. Finally, it combines data importance and group similarity to obtain prior weights for each time period, embedding them as attention bias terms in the temporal attention layer of the neural network model to guide the model to focus on high-quality, highly synergistic data sources. Compared to existing technologies, this invention effectively captures the temporal patterns of load changes and the spatiotemporal heterogeneity of user behavior, significantly improving the accuracy of missing data imputation and making the imputation results closer to real electricity consumption patterns. This provides reliable data support for load forecasting and dispatching decisions in virtual power plants, enhancing the accuracy and timeliness of peak shaving and valley filling.

[0019] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0020] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0021] Figure 1 This is a flowchart of a missing data filling method for a virtual power plant according to Embodiment 1 of the present invention.

[0022] Figure 2 This is a flowchart of the method for obtaining the group similarity of each user in each local time period in Embodiment 1 of the present invention. Detailed Implementation

[0023] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0024] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0025] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0026] Example 1 This embodiment discloses a method for filling in missing data for a virtual power plant.

[0027] like Figure 1 As shown, a method for imputing missing data in a virtual power plant includes: Step S1: Obtain the electricity consumption data of each user on the power grid load side at each time within a preset adjacent period; Step S2: For each user, analyze the volatility and trend of electricity consumption data within the corresponding local time period to calculate the electricity consumption stability of each local time period. Based on the missing electricity consumption data within a local time period and the changing trend at the location of the missing data, the impact of the missing data on each user in each local time period is calculated; and combined with the electricity consumption stability, the data importance of each user in each local time period is determined. Step S3: Based on the synergy of the changes in electricity consumption data of any two users in each local time period, the similarity of electricity consumption of any two users in each local time period is calculated, and reference users with similar electricity consumption behavior are selected for each user. Step S4: For each local time period, by analyzing the consistency and similarity of electricity consumption trends of each user and the corresponding reference user at a single moment, the group similarity of each user in each local time period is determined; combined with data importance, the contribution of each user in each local time period is determined, and the prior weight of each moment is set according to the contribution; by embedding the prior weight as an attention bias term of the temporal attention layer into the neural network model, the missing data is predicted and filled.

[0028] Based on the above process, this invention enables the filling results to closely resemble actual electricity consumption patterns, providing accuracy in filling missing data, and thus providing reliable data support for load forecasting and dispatching decisions of virtual power plants. To facilitate understanding of the technical solution of this invention, the specific implementation methods of this invention will be further explained and described below.

[0029] In step S1, the electricity consumption data of each user on the power grid load side at each time within a preset adjacent period is obtained.

[0030] A virtual power plant is a system that uses advanced information and automation technologies to centrally manage and optimize the dispatch of distributed energy resources scattered across different geographical locations. It can simulate the functions of a traditional power plant, coordinating and controlling the generation and storage of these distributed resources to perform peak shaving and valley filling, reducing the impact of distributed energy generation instability on the power grid, thereby achieving stability and reliability of power supply while improving energy utilization efficiency.

[0031] The efficient operation and precise control of virtual power plants heavily rely on the completeness and accuracy of data. However, in actual operation, factors such as sensor failures, electromagnetic interference, and transmission errors often lead to data gaps. Ignoring these missing data can cause deviations in the scheduling decisions of the virtual power plant, and may even trigger power grid accidents. In this embodiment, taking the electricity consumption data of load-side users in the power grid as an example, it is necessary to fill in the missing data to restore the true load baseline, enabling the virtual power plant to perform precise control.

[0032] Based on the above analysis, this embodiment uses smart meters deployed on the load side to collect electricity consumption data of different users in real time and uploads it to the centralized data management platform of the virtual power plant. From the centralized data management platform of the virtual power plant, the electricity consumption data of each user on the load side of the power grid at each moment in a preset adjacent period is extracted, and the collected data is normalized.

[0033] In this embodiment, the electricity consumption data of each user is the load power; the electricity consumption data of the user at all times in the most recent 7 days within the adjacent period. It is understood that the implementer can set the parameter adjustment factor according to the actual situation; secondly, the maximum and minimum value normalization method is used for normalization processing. The maximum and minimum value normalization method is a well-known technology and will not be described in detail here.

[0034] This allows us to obtain the electricity consumption data for each user at various times within a nearby period.

[0035] In step S2, for each user, the volatility and trend of electricity consumption data within the corresponding local time period are analyzed to calculate the electricity consumption stability of each local time period. Based on the missing electricity consumption data within the local time period and the trend of change at the missing locations, the impact of missing data for each user in each local time period is calculated; and combined with the electricity consumption stability, the data importance of each user in each local time period is determined.

[0036] When calculating the power consumption stability of each local time period, for each user's local time period, the extreme points of power consumption data at all times are obtained, and the difference in power consumption data between two adjacent extreme points is used as the difference quantity. Curve fitting is performed on the power consumption data at all times within the local time period to calculate the tangent slope of the fitted curve at each time. The absolute values ​​of all tangent slopes contained between two adjacent extreme points are positively merged as the change quantity. Among them, the power consumption stability is negatively correlated with all difference quantities and all change quantities.

[0037] When calculating the impact of missing data for each user in each local time period, for each local time period, the corresponding time with missing data is marked as a missing time, and the time period consisting of consecutively adjacent missing times is defined as a missing time period; the time interval between each missing time and the time corresponding to the nearest extreme point is calculated; the changing trend of user data on the left and right sides of the missing time period where each missing time is located and the duration of the missing data are analyzed, and the potential criticality of each missing time is calculated in combination with the time interval; the proportion of missing times in each local time period is calculated; the impact of missing data is positively correlated with the proportion and the potential criticality.

[0038] The calculation of potential criticality involves defining all times other than the missing time as non-missing times. For each missing time, the electricity consumption data of consecutive non-missing times before the first missing time and consecutive non-missing times after the last missing time within the missing time period are linearly fitted, and the slope of the fitted line is calculated. The duration of the missing time period to which each missing time belongs is statistically analyzed and negatively mapped. The negative mapping result is positively integrated with the absolute value of the slope to obtain the trend change of each missing time. Potential criticality is negatively correlated with the time interval and positively correlated with the trend change.

[0039] In the specific implementation process: One-dimensional convolutional neural network (CNN) models can effectively process time-series data. By extracting key features from the data and sliding convolutional kernels along the time dimension, they can capture local temporal correlations in the data, thus filling in missing data. However, since the core of a one-dimensional CNN model is learning the local temporal correlations of electricity consumption data, it requires training. Traditional one-dimensional CNN models assume that all time windows have equal importance during the convolution sliding process, and they fill in missing data by learning local patterns. However, in strong time-series data such as electricity load, the impact of missing data at different times varies. For example, data at key locations such as peak and valley times and trend inflection points have a much greater impact on the load baseline than data from stable periods. Treating them equally can lead to insufficient fitting of key temporal patterns by the model, and the filling results may have significant deviations during critical periods, thus misleading scheduling decisions.

[0040] Therefore, in order to improve the accuracy of missing data imputation, especially to ensure the accuracy of data restoration during critical load change periods, a temporal attention layer is embedded in a one-dimensional CNN model. This attention layer can assign a dynamic attention weight to the electricity consumption data at each moment. During training and imputation, the model will selectively and differentially learn, thereby improving the accuracy and reliability of imputing missing data.

[0041] Due to the influence of different users' electricity consumption habits, equipment usage, seasonal loads, and other factors, there are significant differences in electricity consumption patterns and demand characteristics among users. In order to improve the accuracy of filling in missing data, we analyze the electricity consumption characteristics and importance of each user in local time periods.

[0042] Based on this, the electricity stability is calculated by analyzing the fluctuations in each user's electricity consumption data over a local time period. Specifically: 1) When dividing all moments within a neighboring period into multiple local time periods, the APCA (Adaptive Piecewise Constant Approximation) segmentation method is used to divide the local time periods; when obtaining the extreme points of electricity consumption data for all moments within a local time period, the AMPD (Automatic multiscale-based peak detection) algorithm is used to obtain the extreme points. Both the APCA segmentation method and the AMPD algorithm are well-known techniques.

[0043] 2) The difference in electricity consumption data between two adjacent extreme points within a local time period is taken as the difference quantity; in this embodiment, the absolute value of the difference in electricity consumption data between two adjacent extreme points within a local time period is taken as the difference quantity.

[0044] 3) Perform curve fitting on the electricity consumption data at all times within the local time period and calculate the tangent slope of the fitted curve at each time. In this embodiment, the least squares method is used for curve fitting. The least squares method and the calculation of the tangent slope are well-known techniques and will not be described in detail here.

[0045] 4) The absolute values ​​of the tangent slopes of all times between two adjacent extreme points within a local time period are positively fused and used as the change quantity. In this embodiment, the specific process of positive fusion is as follows: the tangent slope is normalized using the maximum-minimum method, the absolute value of the normalized tangent slope is calculated, and the mean of the absolute values ​​of all times between two adjacent extreme points is used as the change quantity.

[0046] Among them, the electricity stability of each user in each local time period is negatively correlated with all differences and all changes; it should be noted that the negative correlation means that the dependent variable decreases as the independent variable increases, and increases as the independent variable decreases.

[0047] In this embodiment, the product of the difference and the change is calculated, and the sum of the products of the difference and the change between all adjacent extreme points within a local time period is calculated and negatively mapped to represent the power stability. The negative mapping process involves using an exponential function for negative mapping, assuming the sum is denoted as... Then The result, as a measure of power stability, is... This represents an exponential function with the natural constant as its base.

[0048] It should be noted that the smaller the difference, the smaller the difference in the user's electricity consumption data between peak and valley, and the more stable the electricity consumption fluctuation; the smaller the change, the more gradual the rate of change of electricity consumption data between peak and valley; the greater the obtained electricity consumption stability, the more stable the change of electricity consumption data is within the entire local time period, with small fluctuations and gradual changes.

[0049] Furthermore, since missing data can interfere with the effective determination of a user's electricity consumption stability, even if the electricity consumption stability is the same in two local time periods, the degree of impact from missing data in different local time periods will vary, and the importance of the data will also differ. For example, the load baseline of a virtual power plant is supported by data from key locations such as extreme points, peak and valley times, and trend inflection points. These locations are important reference points for judging electricity consumption patterns, and their absence will directly lead to baseline distortion. Therefore, this invention calculates potential criticality by assessing the criticality of the time corresponding to the missing data within a local time period, specifically: 1) Mark the time corresponding to the missing data within a local time period as the missing time, and define the time period consisting of consecutively adjacent missing times as the missing time period; 2) For a local time period, the time interval between each missing time and the time corresponding to the nearest extreme point is calculated. It should be noted that the time interval is measured by counting the number of times that are included between each missing time and the time corresponding to the nearest extreme point.

[0050] 3) Define all other times except for the missing time as non-missing times. For the missing time period to which each missing time belongs in the local time period, perform linear fitting on the electricity consumption data of multiple consecutive non-missing times before the first missing time and multiple consecutive non-missing times after the last missing time, and calculate the slope of the fitted line.

[0051] In this embodiment, the electricity consumption data of five consecutive non-missing moments before the first missing moment and five consecutive non-missing moments after the last missing moment are selected for linear fitting. The least squares method is used for linear fitting. The least squares method and the calculation of the slope are well-known techniques and will not be described in detail here.

[0052] It should be noted that if there are fewer than 5 non-missing moments to the left of the first missing moment or to the right of the last missing moment, then all consecutive non-missing moments are selected for fitting. For ease of understanding: assume that for a local time period, ;in, , , as well as Indicates missing data. , , , , , , as well as If it is not considered missing data, then the missing time period is... , as well as Time period Time period. (Targeting) , as well as The missing time period contains only data from 3 time points before the first missing time point. , , There are fewer than 5 missing data points. Therefore, we select data from the 3 time points before the first missing time point and data from the 5 time points after the last missing time point for linear fitting, i.e., for... , , , , , , as well as Perform linear fitting.

[0053] The duration of the missing time period to which each missing moment belongs is statistically analyzed, and a negative mapping is performed on it. This is then positively integrated with the absolute value of the slope to serve as the trend change measure for each missing moment.

[0054] In this embodiment, the specific process of negative mapping is as follows: the reciprocal of the duration is used as the negative mapping result. The specific process of positive fusion is as follows: the product between the negative mapping result and the absolute value of the slope is normalized to represent the trend change at each missing time point, wherein the maximum-minimum normalization method is used for normalization.

[0055] The potential criticality at each missing time point is negatively correlated with the time interval and positively correlated with the trend change. It should be noted that a positive correlation means that the dependent variable increases as the independent variable increases and decreases as the independent variable decreases. In this embodiment, the ratio of the trend change to the time interval is used as the potential criticality at each missing time point.

[0056] Understandably, the shorter the time interval and the steeper the slope, the closer the missing moment is to a critical location, and the more drastic the changes in electricity consumption data are at that point, making it more likely that important information on electricity consumption trends will be lost. Conversely, a longer duration results in a smaller negative mapping, indicating significant data loss within the missing time period. This makes the key trend in electricity consumption data represented by the slope of the fitted line less reliable. Conversely, a shorter duration makes the key trend in electricity consumption data represented by the slope of the fitted line more reliable. Introducing duration helps characterize the reliability of trend changes. Therefore, a higher potential criticality indicates that the missing moment occurs near a period of drastic change, and the electricity consumption data reflecting the missing moment is more important for analyzing electricity consumption trends.

[0057] Furthermore, the load baseline and dispatch decisions of the virtual power plant rely on the completeness of global electricity consumption patterns. The criticality of a single missing moment can only measure its local value, but there may be multiple missing data points within a local time period. Therefore, the impact of missing data is determined by the proportion of missing moments within a local time period and their potential criticality. Specifically: The percentage of missing moments within each local time period is statistically analyzed; the impact of each user's missing moments in each local time period is positively correlated with both the percentage and the potential criticality; in this embodiment, the sum of the potential criticality of all missing moments within a local time period is calculated, and its product with the percentage is used as the impact of each user's missing moments in each local time period.

[0058] It should be noted that the larger the percentage, the more missing data is contained in the local time period, the more severely incomplete the data is, and the poorer the overall data integrity. The larger the sum, the more important the trend information contained in the overall missing data is, and the greater the impact of the missing data, the more likely there is important missing data in that local time period, and the greater the impact of the missing data on the scheduling of the virtual power plant.

[0059] Furthermore, based on the impact of missing data and the stability of electricity consumption, the importance of data is determined. Specifically, the importance of data for each user in each local time period is positively correlated with the stability of electricity consumption, while it is negatively correlated with the impact of missing data.

[0060] In this embodiment, the ratio of power stability to the impact of missing data is used as the data importance. It should be noted that, to avoid a denominator of 0 when calculating the ratio, a parameter tuning factor is added to the denominator; the value range of the parameter tuning factor is... In this embodiment, the parameter tuning factor is set to 1. In other implementation methods, the implementer can set it according to the actual situation.

[0061] It should be noted that the greater the power consumption stability, the less the power consumption data is affected by random interference during that local period, and the stronger the regularity of data changes. Conversely, the lower the power consumption stability, the more drastic the data fluctuations. Although the data has high business value, it contains a lot of random noise, and its weight in the prediction should be reduced to prevent the model from overfitting noise. The greater the impact of missing data, the worse the data integrity. The greater the importance of the obtained data, the more typical and repeatable the stable power consumption state of the user during that period. The data can clearly reflect its basic load characteristics and is minimally affected by anomalies and missing data. When making predictions, such data should be given priority to accurately grasp the user's basic power consumption pattern and baseline load.

[0062] At this point, the data importance of each user in each local time period can be obtained.

[0063] In step S3, based on the synergy of the changes in electricity consumption data of any two users in each local time period, reference users with similar electricity consumption behavior are selected for each user by calculating the similarity of electricity consumption of any two users in each local time period.

[0064] For each user's corresponding local time period, if a time is missing, the label value is assigned 1; otherwise, it is assigned 0. When any two users in each local time period have the same label value at the same time, this time is defined as a synchronization time, and the time period consisting of consecutive adjacent synchronization times is defined as a synchronization time period. The correlation between all electricity consumption data of any two users in the synchronization time period is calculated. The electricity consumption similarity is the result of positively fusing the duration and correlation of all synchronization time periods of any two users in each local time period. For each local time period, the remaining users whose electricity consumption similarity with each user is greater than or equal to a preset threshold are marked as reference users.

[0065] In the specific implementation process: The core of virtual power plant dispatching is to ensure overall regional supply and demand balance, rather than focusing on the local patterns of individual users. If a user experiences prolonged data gaps, anomalies, or the addition of new users, relying solely on historical data to assess the importance of their electricity consumption behavior will be inaccurate. Therefore, it is necessary to consider the similarities in electricity consumption behavior across different users to determine the consistency between individual user behavior and the overall trend of the group, in order to more accurately assess the contribution of each user's electricity consumption data to the overall supply and demand balance.

[0066] Based on the above analysis, the electricity consumption similarity is calculated by analyzing the synchronicity of electricity consumption data of different users in a local time period. Specifically: For each user's corresponding local time period, if each time period is missing, the label value is assigned to 1; otherwise, it is assigned to 0. When any two users have the same tag value at the same time in each local time period, this time is defined as a synchronization time. The time period consisting of consecutively adjacent synchronization times is defined as a synchronization time period. Calculate the correlation between all electricity consumption data of any two users during the synchronization period.

[0067] In this embodiment, the Pearson correlation coefficient of all electricity consumption data between any two users during the synchronization period is calculated, and a positive mapping is performed on it as the degree of correlation. The calculation of the Pearson correlation coefficient is a well-known technique and will not be elaborated upon here. As other implementation methods, the implementer may also use other methods of the prior art, such as the reciprocal of the DTW distance, etc., and this embodiment does not impose any special restrictions on this. It should be noted that if the electricity consumption data of both users is missing during the synchronization period, the Pearson correlation coefficient is assigned a value of 1. Furthermore, the positive mapping process can be expressed as follows: the result of an exponential function with the natural constant as the base and the Pearson correlation coefficient as the exponent is used as the degree of correlation, and through the positive mapping process, the result of the degree of correlation is made non-negative.

[0068] The duration and correlation of all synchronous periods for any two users within each local time period are positively integrated to form the electricity consumption similarity of the two users within each local time period.

[0069] In this embodiment, the specific process of forward fusion is as follows: calculate the product of the duration of each synchronization period and the degree of correlation, and use the normalized result of the sum of the products of all synchronization periods in each local period as the power consumption similarity; as another implementation, the implementer may also use the normalized result of the mean of the products of all synchronization periods in each local period as the power consumption similarity; wherein, the maximum and minimum value normalization method is used for normalization processing.

[0070] It should be noted that the greater the correlation, the stronger the correlation between the electricity consumption data of the two users within the synchronized time period; the longer the synchronized time period, the longer the continuous synchronized moments between the two users within that local time period. Therefore, the more correlated the electricity consumption data of two users within the synchronized time period, and the longer the synchronized time period, the more consistent the electricity consumption patterns of the two users within the synchronized time period, that is, the greater the electricity consumption similarity, reflecting that the two users are more likely to belong to the same electricity consumption group or have similar electricity consumption behaviors.

[0071] For each local time period, other users whose electricity consumption similarity to each user is greater than or equal to a preset threshold are marked as reference users. In this embodiment, the preset threshold is obtained by calculating the upper quartile of the electricity consumption similarity between any two users, which is used as the preset threshold. The calculation of the upper quartile is a well-known technique and will not be described in detail here. In this embodiment, the preset threshold is set to 0.7, but implementers can also set it according to actual conditions.

[0072] Furthermore, each user shares similarities in electricity consumption behavior with reference users. Analyzing the electricity consumption data among these users with similar behaviors can reveal the patterns and characteristics of group electricity consumption in the virtual power plant. For example, in a residential community, household electricity consumption patterns may differ significantly between weekends and weekdays. By comparing these patterns, normal electricity consumption patterns under specific conditions can be identified, thereby improving the accuracy of missing data imputation. Therefore, analyzing the similarity between each user's electricity consumption and that of all reference users allows for the calculation of group similarity.

[0073] like Figure 2 As shown, the method for obtaining the group similarity of each user in each local time period specifically includes: For each user and all its reference users, users with positive tangent slopes and users with negative tangent slopes at the same time within each local time period are counted, forming the first user set and the second user set for each time period respectively. It should be noted that if the absolute value of the tangent slope is less than 0.01, it is considered as 0 and no positive or negative statistics are performed.

[0074] The first user set and the second user set are labeled as user sets. The number of all users in a single user set is counted. All users in the user set with the largest number are selected and defined as dominant users. The largest number is used as the intensity of the electricity consumption trend at each time. Then, the dispersion of the tangent slope of all dominant users at each time is calculated.

[0075] In this embodiment, the degree of dispersion is measured by calculating the coefficient of variation of the tangent slope of all dominant users at each time point. The coefficient of variation is a well-known technique and will not be elaborated here. As other implementation methods, implementers may also use other methods of the prior art, such as standard deviation. This embodiment does not impose any special restrictions on this.

[0076] Furthermore, the trend correlation degree for each user in each local time period is positively correlated with the intensity of the electricity consumption trend, and negatively correlated with the degree of dispersion. In this embodiment, the ratio of the intensity of the electricity consumption trend to the degree of dispersion at each moment is calculated, and the sum of the ratios at all moments within the local time period is taken as the trend correlation degree. When calculating the ratio, to avoid the denominator being 0, a parameter adjustment factor is added to the denominator, wherein the value range of the parameter adjustment factor is [missing value]. In this embodiment, the parameter tuning factor is set to 1. In other implementation methods, the implementer can set it according to the actual situation.

[0077] Furthermore, the group similarity, trend correlation, and electricity consumption similarity of each user in each local time period are all positively correlated. In this embodiment, the average electricity consumption similarity between each user and all reference users in each local time period is multiplied by the trend correlation and used as the group similarity.

[0078] It should be noted that the stronger the electricity consumption trend, the more consistent the electricity consumption trends of most users at that moment, reflecting a very clear and consistent group electricity consumption trend at that moment; the smaller the dispersion, the more similar the electricity consumption change rates of the dominant users, indicating highly synchronized behavior; the greater the trend correlation, the more the user's electricity consumption change trend follows a strong and unified group trend; the greater the group similarity, the more consistent the user's electricity consumption pattern is with the group's electricity consumption pattern in that local time period, reflecting the user's electricity consumption pattern as more representative and better able to reflect the group's electricity consumption pattern. This makes the user's data more valuable in analyzing and predicting group behavior, and thus requires higher weighting, allowing the model to focus more on these representative user data.

[0079] At this point, the group similarity of each user in each local time period is obtained.

[0080] In step S5, for each local time period, the consistency and similarity of electricity consumption trends of each user and the corresponding reference user at a single moment are analyzed to determine the group similarity of each user in each local time period; combined with data importance, the contribution of each user in each local time period is determined, and the prior weight of each moment is set according to the contribution; by embedding the prior weight as an attention bias term of the temporal attention layer into the neural network model, the missing data is predicted and filled.

[0081] For each user and all corresponding reference users, users with positive and negative tangent slopes at the same time point within each local time period are counted, forming a first user set and a second user set for each time point. The first and second user sets are labeled as user sets. The number of all users in a single user set is counted, and all users in the user set with the largest number are selected as dominant users. The largest number is used as the electricity consumption trend intensity at each time point. The dispersion of the tangent slopes of all dominant users at each time point is calculated. The group similarity is positively correlated with the electricity consumption trend intensity and electricity consumption similarity, and negatively correlated with the dispersion.

[0082] In the specific implementation process, the determination of contribution is as follows: The data importance and group similarity are positively fused to obtain the contribution of each user in each local time period. In this embodiment, the normalized result of the product of data importance and group similarity is used as the contribution of each user in each local time period. As another implementation, the implementer may also use the normalized result of the sum of data importance and group similarity as the contribution of each user in each local time period.

[0083] It should be noted that the greater the contribution, the higher the quality and representativeness of the user's electricity consumption data in that local time period. This data can clearly show individual patterns and reliably represent the commonalities of the group, allowing the model to learn better. By giving higher attention to high-contribution data, the model can be guided to focus its limited learning resources on the most valuable data.

[0084] The prior weights of each user at each moment in the adjacent period are set as the contribution of their respective local time period; a temporal attention layer is embedded in the neural network model, wherein the temporal attention layer uses the prior weights as attention bias terms to guide the model; based on the electricity consumption data of each user at all moments in the adjacent period, the neural network model with embedded temporal attention layer is used to predict and fill in the missing data. In this embodiment, the neural network model is a one-dimensional CNN model. Embedding a temporal attention layer in the one-dimensional CNN model is a well-known technique and will not be described in detail here. Additionally, a complete and uninterrupted segment of electricity consumption data from a nearby period is selected as the training set. Electricity consumption data from a portion of the time intervals within the training set are randomly selected and marked as missing. The true values ​​of these marked missing times are used as training labels. Based on the training set, the neural network model with the embedded temporal attention layer is trained. The training of the neural network model is a well-known technique and will not be described in detail here either.

[0085] Based on the above methods, the present invention achieves the following breakthroughs: 1) This invention quantifies the volatility and regularity of electricity consumption data by calculating the stability of electricity consumption in each local time period, providing a quantitative basis for evaluating the inherent reliability of data in different local time periods; 2) This invention calculates the impact of missing data for each user in each local time period, analyzes the location of the missing data and the drastic changes in the electricity consumption trend, as well as the duration of the missing data, and reflects the criticality of the data at different missing locations. It can effectively assess the impact of missing data in this local time period on the restoration of the real electricity consumption pattern. 3) This invention determines the data importance of each user in each local time period. By assessing the stability and impact of missing data in different local time periods, the data value of that local time period can be comprehensively evaluated from two dimensions: the intrinsic quality of the data reflected by stability and the degree of external damage reflected by the impact of missing data. 4) This invention obtains the similarity of electricity consumption between any two users in each local time period, and selects reference users with similar electricity consumption behavior for each user. It takes into account the synchronicity and similarity of electricity consumption behavior between different users, and can effectively reflect the consistency of electricity consumption patterns between two users, so as to select a group of reference users with similar electricity consumption behavior. 5) This invention determines the group similarity of each user in each local time period, takes into account the synchronicity of the electricity consumption trends of each user and the reference user, effectively captures the time-varying characteristics of the consistency between the user's electricity consumption behavior and the reference group to which they belong, and thus can better reflect the consistency between the user's electricity consumption pattern and the group's electricity consumption pattern. 6) This invention calculates the contribution of each user in each local time period. By comprehensively considering the data importance of individual users and their consistency with group electricity consumption behavior, it quantifies the important value of the user's electricity consumption data in a specific time period. Based on this, it sets the prior weights for each time moment and uses them as attention bias terms of the temporal attention layer to embed into the neural network model. This predicts and fills in missing data. By setting the prior weights for each time moment based on the contribution, the temporal attention layer is embedded in the neural network model, guiding the model's attention to focus on high-quality, highly collaborative information sources. This enhances the model's adaptability and robustness, greatly accelerates model convergence, and enables the model to simultaneously learn general temporal patterns and personalized electricity consumption characteristics. This improves the modeling ability for complex power temporal patterns and heterogeneous user behavior, thereby achieving accurate filling of missing data and providing strong data support for the efficient operation of virtual power plants.

[0086] Example 2 This embodiment discloses a missing data filling system for virtual power plants.

[0087] A missing data imputation system for a virtual power plant includes: The data acquisition module is configured to acquire the electricity consumption data of each user on the power grid load side at each time within a preset adjacent period; The power consumption stability calculation module is configured to: analyze the volatility and trend of power consumption data in the corresponding local time period for each user, so as to calculate the power consumption stability of each local time period; The data importance calculation module is configured to: calculate the impact of missing data for each user in each local time period based on the missing data and the changing trend at the missing locations; and determine the data importance of each user in each local time period in conjunction with the power consumption stability. The reference user determination module is configured to: based on the synergy of changes in the electricity consumption data of any two users in each local time period, and by calculating the similarity of electricity consumption of any two users in each local time period, select reference users with similar electricity consumption behavior for each user; The missing data imputation module is configured to: determine the group similarity of each user in each local time period by analyzing the consistency and similarity of electricity consumption trends of each user and the corresponding reference user at a single moment; determine the contribution of each user in each local time period by combining data importance, and set the prior weight of each moment according to the contribution; and predict and impute missing data by embedding the prior weight as an attention bias term of the temporal attention layer into the neural network model.

[0088] Example 3 The purpose of this embodiment is to provide a computer-readable storage medium.

[0089] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of a missing data filling method for a virtual power plant as described in Embodiment 1 of this disclosure.

[0090] Example 4 The purpose of this embodiment is to provide an electronic device.

[0091] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of a missing data filling method for a virtual power plant as described in Embodiment 1 of this disclosure.

[0092] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0093] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0094] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A missing data imputation method for a virtual power plant, characterized by, The method comprises the following steps: obtaining the power consumption data of each user at each time point in a preset adjacent period on the power grid load side; for each user, analyzing the volatility and change trend of the power consumption data in the corresponding local period to calculate the power consumption stability of each local period; according to the missing situation and the change trend at the missing position of the power consumption data in the local period, calculating the missing influence degree of each user in each local period; combining the power consumption stability, determine the data importance of each user in each local period; according to the change synergy of the power consumption data of any two users in each local period, calculate the power consumption similarity of any two users in each local period, and select the reference user with similar power consumption behavior for each user; for each local period, by analyzing the consistency of the power consumption change trend of each user and the corresponding reference user at a single time point and the power consumption similarity, determine the group similarity of each user in each local period; combine the data importance to determine the contribution degree of each user in each local period, and set the prior weight of each time point according to the contribution degree; by embedding the prior weight as the attention bias term of the time sequence attention layer into the neural network model, the missing data is predicted and filled.

2. The missing data imputation method for a virtual power plant of claim 1, wherein, The calculation of the power consumption stability of each local period comprises: for each local period of each user, obtain the extreme point of the power consumption data of all time points, and take the difference between the power consumption data of the adjacent two extreme points as the difference; curve fitting is performed on the power consumption data of all time points in the local period to calculate the tangent slope of each time point on the fitting curve, and the absolute values of all tangent slopes contained between the adjacent two extreme points are positively fused as the change; wherein, the power consumption stability is negatively correlated with all difference and all change.

3. The missing data imputation method for a virtual power plant of claim 1, wherein, The calculation of the missing influence degree of each user in each local period comprises: for the local period, the corresponding time point with missing data is marked as the missing time point, and the time period formed by the continuous adjacent missing time points is defined as the missing period; the time interval between each missing time point and the time point corresponding to the nearest extreme point is counted; the change trend of the user data on the left and right sides of the missing period where each missing time point is located and the duration of the missing are analyzed, and the potential key degree of each missing time point is calculated combined with the time interval; count the proportion of missing time points in each local period; wherein, the missing influence degree is positively correlated with the proportion and the potential key degree.

4. The missing data imputation method for a virtual power plant of claim 3, wherein, The calculation of the potential key degree comprises: define the remaining time points except the missing time points as non-missing time points, for each missing time point belonging to the missing period, linearly fit the power consumption data of the continuous multiple non-missing time points before the first missing time point and the continuous multiple non-missing time points after the last missing time point in the missing period, and calculate the slope of the fitting straight line; statistic the length of the missing period of each missing time point, and perform negative mapping, positively fuse the negative mapping result and the absolute value of the slope as the trend change of each missing time point; wherein, the potential key degree is negatively correlated with the time interval and positively correlated with the trend change.

5. The missing data imputation method for a virtual power plant of claim 1, wherein, By calculating the similarity of electricity consumption between any two users in each local time period, reference users with similar electricity consumption behavior are selected for each user, including: For each user's corresponding local time period, if each time period is missing, the label value is assigned to 1; otherwise, it is assigned to 0. When any two users in each local time period have the same label value at the same time, this time is defined as a synchronization time, and the time period consisting of consecutive adjacent synchronization times is defined as a synchronization time period. Calculate the correlation between all electricity consumption data of any two users in the synchronization time period. The electricity consumption similarity is the result of positively fusing the duration and correlation of all synchronous periods of any two users in each local time period; for each local time period, the remaining users whose electricity consumption similarity with each user is greater than or equal to a preset threshold are marked as reference users.

6. The missing data imputation method for a virtual power plant of claim 1, wherein, Determining the group similarity of each user in each local time period includes: For each user and all corresponding reference users, users with positive tangent slopes and users with negative tangent slopes at the same time within each local time period are counted, forming the first user set and the second user set for each time period. The first user set and the second user set are labeled as user sets. The number of all users in a single user set is counted, and all users in the user set with the largest number are selected as the dominant user. The largest number is used as the intensity of the electricity consumption trend at each time period. Calculate the dispersion of the tangent slopes for all dominant users at each time point; Among them, the group similarity is positively correlated with the intensity of electricity consumption trend and the similarity of electricity consumption, and negatively correlated with the degree of dispersion.

7. The missing data imputation method for a virtual power plant of claim 1, wherein, The data importance is positively correlated with power consumption stability and negatively correlated with the impact of missing data; the contribution is the result of positive fusion of data importance and group similarity, and the prior weight is set as the contribution of each user to the local time period at each moment.

8. A missing data imputation system for a virtual power plant, characterized by, include: The data acquisition module is configured to acquire the electricity consumption data of each user on the power grid load side at each time within a preset adjacent period; The power consumption stability calculation module is configured to: analyze the volatility and trend of power consumption data in the corresponding local time period for each user, so as to calculate the power consumption stability of each local time period; The data importance calculation module is configured to: calculate the impact of missing data for each user in each local time period based on the missing data and the changing trend at the missing locations; and determine the data importance of each user in each local time period in conjunction with the power consumption stability. The reference user determination module is configured to: based on the synergy of changes in the electricity consumption data of any two users in each local time period, and by calculating the similarity of electricity consumption of any two users in each local time period, select reference users with similar electricity consumption behavior for each user; The missing data imputation module is configured to: determine the group similarity of each user in each local time period by analyzing the consistency and similarity of electricity consumption trends of each user and the corresponding reference user at a single moment; determine the contribution of each user in each local time period by combining data importance, and set the prior weight of each moment according to the contribution; and predict and impute missing data by embedding the prior weight as an attention bias term of the temporal attention layer into the neural network model.

9. A computer-readable storage medium having stored thereon a program, characterized in that, When the program is executed by the processor, it implements the steps in a missing data imputation method for a virtual power plant as described in any one of claims 1-7.

10. An electronic device comprising a memory, a processor, and a program stored in the memory and capable of running on the processor, characterized by When the processor executes the program, it implements the steps in the missing data filling method for a virtual power plant as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Simulation data generation method and system for power consumption behavior prediction

    CN112560330A

  • Intelligent power consumption complementing method and system

    CN114611856A

  • Power grid intelligent scheduling method and system based on user demands

    CN118199049A

  • CNN-XGBoost power load prediction method based on attention mechanism

    CN119577598A

  • Power consumption data processing method and system based on time sequence correlation and trend fitting

    CN119719641A