Complementation method for daily load data missing value of power consumer

By obtaining historical electricity consumption data from the power trading center, extracting user characteristics, performing clustering and model training, the accuracy and reliability of filling missing values ​​of power users' daily load data is solved, and efficient missing value prediction and data completion are achieved.

CN120144936AActive Publication Date: 2025-06-13重庆玖奇科技有限公司

Patent Information

Application Number
CN202510350179.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-13
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

The prior art is difficult to accurately and reliably fill the missing values ​​in the daily load data of power users, especially in the case of excessive missing amount or lack of data, resulting in large prediction errors.

Method used

By obtaining historical electricity consumption data from the power trading center, performing data cleaning and standardization processing, extracting user's feature vector data, and dividing users into multiple aggregation classes based on the clustering algorithm, integrating meteorological and date information data, building a data completion model, performing model training and missing value filling.

Benefits of technology

Accurate and reliable filling of missing values ​​of daily load data of power users is achieved, the accuracy of missing value prediction and the robustness of data completion model are improved, and inaccurate completion situations caused by individual abnormal data or single user characteristics are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144936A_ABST
    Figure CN120144936A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power loads, and discloses a power consumer daily load data missing value complementing method, which comprises the following steps: acquiring meteorological data and date information data of each consumer in the same consumer aggregation class, integrating the meteorological data of each user, the date information data, the historical power consumption data of the user and the feature vector data to form training data set data corresponding to a user aggregation class; according to the training data set data corresponding to the same user aggregation class, training the data completion model corresponding to the user aggregation class, and after training is completed, outputting the corresponding data completion model and associating the data completion model with the corresponding user aggregation class; and inputting the meteorological information data, the date information data and the feature vector data of a certain user in the user aggregation class into a data completion model associated with the user aggregation class, outputting a missing value of the user, and filling the missing value into the historical power consumption data of the user to form the filled historical power consumption data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power load, and particularly relates to a method for filling missing values of daily load data of power users. Background Art

[0002] In the power industry, accurate and complete power consumption data of power users is of inestimable value to the operation and planning of power selling enterprises. Power selling enterprises need to accurately predict the future power consumption of users based on detailed historical power consumption data, so as to reasonably formulate marketing strategies, optimize power resource allocation and conduct cost control.

[0003] However, the acquisition of power consumption data of power users currently faces many severe challenges. On the one hand, the metering devices used by a large number of power users are relatively backward and cannot achieve automatic data acquisition. This makes the data acquisition work rely on manual operation, which is not only inefficient but also prone to human errors, resulting in the incompleteness and delay of data acquisition. On the other hand, the phenomenon of missed data collection occurs from time to time, further exacerbating the problem of missing power consumption data.

[0004] To solve the problem of missing values in historical power load data of power users, in the existing technical means, the similar day method or prediction algorithm is often used to fill in for a single user. However, the similar day method has significant defects. It is difficult to effectively capture the change trend of power consumption data in the short term before and after the date of missing values, resulting in the data after interpolation being unable to accurately reflect the actual power consumption situation, and further affecting the accuracy of the power selling enterprise's prediction of future power consumption. In the face of the problem of poor data filling effect caused by too large a missing amount or lack of data of a single user. When using the prediction algorithm, only a single user is filled in using the prediction algorithm, resulting in the problem of lack or too large a missing amount of data of a single user, so there is a large error in filling in the missing values, and there is still a large error even after the filling is completed.

[0005] Based on this, there is an urgent need for a method for filling missing values of daily load data of power users, which can accurately and reliably fill in the missing values of the power load data corresponding to users. Summary of the Invention

[0006] One of the purposes of the present invention is to provide a method for filling missing values of daily load data of power users, which can accurately and reliably fill in the missing values of the power load data corresponding to users.

[0007] To achieve the above purpose, a method for filling missing values of daily load data of power users is provided, including the following steps:

[0008] S1. Obtain the historical power consumption data corresponding to each user in a certain historical time period from the power trading center;

[0009] S2. Perform data cleaning operations and standardization processing on the historical electricity consumption data corresponding to each user in sequence;

[0010] S3. Extract the features corresponding to each user based on the historical electricity consumption data corresponding to each user after standardization processing, and generate the feature vector data corresponding to each user. The feature vector data includes the data domain feature data, time domain feature data, and frequency domain feature data of the user;

[0011] S4. Cluster each user based on the data domain feature data, time domain feature data, and frequency domain feature data corresponding to each user, based on a preset feature clustering strategy, to form multiple user aggregation classes;

[0012] S5. Obtain the meteorological data and date information data corresponding to each user in the same user aggregation class, and integrate the meteorological data and date information data corresponding to each user with the historical electricity consumption data and feature vector data corresponding to the user to form the training dataset data corresponding to the corresponding user aggregation class;

[0013] S6. Based on the training dataset data corresponding to the same user aggregation class, train the data completion model corresponding to the user aggregation class based on a pre-constructed data completion model and the model training strategy corresponding to the data completion model. After the training is completed, output the corresponding data completion model and associate it with the corresponding user aggregation class;

[0014] S7. Use the meteorological information data, date information data, and the feature vector data obtained by the user in S3 corresponding to a certain user in the user aggregation class as input data, input them into the data completion model associated with the user aggregation class, output the missing values corresponding to the user, and fill the missing values into the historical electricity consumption data of the user to form the filled historical electricity consumption data.

[0015] The technical principle and effect of this solution: In this solution, historical electricity consumption data is obtained from the power trading center. This is the basis of the entire solution, and comprehensive and accurate data acquisition provides raw materials for subsequent analysis. The electricity consumption data of different users in a certain historical period constitutes a complex dataset, which contains both useful information and may also have noise and incomplete parts.

[0016] Clean the obtained historical electricity consumption data to remove errors, duplicates, and outliers in the data to ensure the quality of the data. Standardization processing is to convert the data to a unified scale, eliminate the influence of the dimension between different features, and make subsequent data analysis and model training more accurate and effective.

[0017] Based on the standardized data, specific strategies are adopted to extract the features of users. Through these feature extractions, the original electricity consumption data is transformed into more representative feature vector data, providing more effective input for subsequent clustering and model training.

[0018] Using the feature vector data of each user, according to a certain clustering algorithm, users with similar features are grouped into one category, forming multiple user aggregation classes. The purpose of doing this is to gather users with similar electricity consumption behaviors together, so as to more accurately train and apply the data completion model for different categories of user characteristics in the follow-up.

[0019] Obtain relevant meteorological data (such as temperature, humidity, precipitation, etc.) and date information data (year, month, day of the week, whether it is a holiday, etc.) for each user aggregation class. There may be associations between these external data and users' electricity consumption behaviors. For example, in hot weather, the air conditioner electricity consumption of residential users may increase; during holidays, the electricity consumption patterns of commercial users may be different from those on weekdays. Integrating these meteorological data and date information data with users' historical electricity consumption data enriches the feature dimensions of the training dataset, enabling the model to learn more factors affecting electricity consumption data, thereby improving the accuracy of data completion.

[0020] Take the meteorological information data, date information data, and feature vector data of a certain user in the user aggregation class as input, and input them into the data completion model associated with this user aggregation class. The model predicts the missing electricity consumption data according to the rules learned from the previous training and outputs the missing values. Finally, fill these missing values into the user's historical electricity consumption data to complete the data completion.

[0021] In this solution, by extracting features and clustering the historical electricity consumption data of users, the classification of each user in the trading center is realized, and users with similar electricity consumption behaviors are grouped into the same aggregation class. This enables the model to learn and predict based on the commonalities of similar user groups during subsequent data completion, greatly improving the accuracy of predicting missing values compared with the similar-day method of a single user. The completion model constructed using the clustered multi-user data can better enhance the robustness and completion effect of the data completion model.

[0022] Integrate the meteorological data and date information data of each user in the same user aggregation class with the historical electricity consumption data and feature vector data to form a training dataset. Meteorological and date information have a significant impact on users' electricity consumption behaviors. Incorporating these multi-source data into the training enables the data completion model to learn more comprehensive influencing factors, thereby more accurately predicting missing values.

[0023] Based on the training data set corresponding to the user aggregation class, train based on the pre-constructed data completion model and model training strategy. Through a large amount of data training and optimization, the data completion model can adapt to the electricity consumption characteristics of different user aggregation classes, improving the generalization ability and stability of the model. Compared with the existing technology, this method reduces inaccurate completion situations caused by individual abnormal data or single user characteristics through multi-user aggregation and a large amount of data training, ensuring the reliability of data completion, filling data in combination with the data characteristics of a class of users, and reducing the error during data filling. That is, accurate and reliable filling of the missing values in the power load data corresponding to the users is achieved.

[0024] Further, the S2 includes:

[0025] S20. Based on the preset line-by-line traversal strategy, traverse each sub-electricity data in the historical electricity consumption data corresponding to each user line by line;

[0026] The preset line-by-line traversal strategy is:

[0027] Perform line-by-line identification and judgment on each sub-electricity data to determine whether there are any missing values in the sub-electricity data in the data row. If so, record the starting index of the missing value. When encountering the next data row with non-missing values, form a missing interval by combining the previously recorded starting index with the index before the current non-missing value data row, and record it. Then reset the starting index and continue to find the next missing interval until all sub-electricity data have been identified and judged, and output the set of missing intervals corresponding to the user;

[0028] According to the set of missing intervals corresponding to the user, obtain the electricity consumption data of the day before and the day after each missing interval, calculate the electricity consumption ratio between the electricity consumption data of the day after the missing interval and the electricity consumption data of the day before the missing interval, and determine whether the electricity consumption ratio is greater than the preset ratio. If so, it is determined that there is a data stacking situation in this missing interval; otherwise, it is determined that there is no data stacking situation in this missing interval;

[0029] S21. When the judgment result is that there is a data stacking situation in this missing interval, clean the data stacking situation in the corresponding missing interval;

[0030] S22. After completing the cleaning of the data stacking situation, exclude outliers from each sub-electricity data corresponding to each user;

[0031] S23. Perform standardization processing on each sub-electricity data corresponding to each user after excluding outliers.

[0032] Beneficial effects: By means of a preset line-by-line traversal strategy, the historical electricity consumption data is carefully checked. This strategy can accurately identify the missing values in the data rows and, through a clever indexing record method, precisely locate the missing intervals. This accurate positioning provides a reliable basis for subsequent analysis of the data before and after the missing intervals, avoiding incorrect judgments caused by inaccurate positioning of missing values. For example, in a complex electricity consumption dataset, it can accurately find the start and end positions of consecutive missing values, providing an accurate data range for judging data stacking situations. After determining the missing intervals, by calculating the ratio of the electricity consumption data of the day before and after the missing interval and comparing it with the preset ratio, it is possible to effectively judge whether there is data stacking. Data stacking will seriously affect the authenticity of the data and the reliability of the analysis results. By accurately identifying data stacking in this way, such data problems can be discovered and processed in a timely manner to ensure the accuracy of the data. For example, when the electricity consumption on the day after a certain missing interval is abnormally high, by calculating the ratio and comparing it with the preset ratio, it can be judged whether it is caused by data stacking, avoiding using such incorrect data for subsequent analysis.

[0033] After completing the cleaning of data stacking situations, outliers are excluded from the electricity consumption data of each user. By excluding outliers, the data distribution can be made more reasonable and the stability of the data can be enhanced.

[0034] Furthermore, the preset feature clustering strategy is as follows:

[0035] Step 1: According to the data domain feature data, time domain feature data, and frequency domain feature data corresponding to each user, the feature vector data corresponding to each user is integrated to form a feature matrix X corresponding to all users;

[0036]

[0037] where [x i1 ,x i2 ,…,x in is the feature vector corresponding to user i, and x in is the dimension value corresponding to the nth dimension in the corresponding feature vector data of user i;

[0038] Step 2: Determine the number of clusters K corresponding to this clustering;

[0039] Step 3: Based on the feature matrix corresponding to all users and the number of clusters K, randomly select K from m users as the cluster centers corresponding to this clustering;

[0040] Step 4: According to the feature matrix corresponding to the remaining users, based on the preset calculation formula for the user electricity consumption distance degree, calculate the user electricity consumption distance degree between the remaining users and the users corresponding to the K cluster centers;

[0041] The preset calculation formula for the user's power consumption distance degree is as follows:

[0042]

[0043] In the formula, d (x,u) is the user's power consumption distance degree between the user and the user corresponding to the cluster center, x ij is the dimension value corresponding to the j-th dimension corresponding to user i, and u j is the dimension value corresponding to the j-th dimension of the user corresponding to the cluster center;

[0044] Step 5: According to the user's power consumption distance degrees between each remaining user and the users corresponding to each cluster center, compare the user's power consumption distance degrees between the remaining users and the users corresponding to each cluster center, and classify the user corresponding to the cluster center with the smallest user's power consumption distance degree and the remaining user into the same class to form the initial aggregation classes corresponding to each cluster center;

[0045] Step 6: According to the user's power consumption distance degrees in each initial aggregation class, calculate the arithmetic mean of each dimension corresponding to the users in each corresponding initial aggregation class to obtain the center points of each initial aggregation class, calculate and compare the user's power consumption distance degrees between each user in each initial aggregation class and the center points of each initial aggregation class, and classify each user into the aggregation class corresponding to the center point with the smallest user's power consumption distance degree in turn;

[0046] Step 7: After completing the user adjustment of all cluster aggregation classes, output the users of the current round of aggregation classes corresponding to each cluster center to form the user aggregation classes corresponding to each cluster center;

[0047] Step 8: After forming the user aggregation classes corresponding to each cluster center corresponding to the current clustering, perform the next round of clustering and update the number of clusters corresponding to the next round of clustering, and re-execute Step 3 until the number of clusters reaches the preset number threshold;

[0048] Step 9: According to the user aggregation classes corresponding to each cluster center corresponding to each number of clusters, based on the preset loss function calculation formula, calculate the loss function values corresponding to each number of clusters, and determine and output the optimal number of clusters and the user aggregation classes corresponding to each cluster center corresponding to the optimal number of clusters.

[0049] Beneficial effects: By calculating the user's electricity consumption distance degree, the user with the smallest similarity and the cluster center are classified into the same category, which can reasonably divide users with similar electricity consumption behaviors into the same category and accurately identify different user groups. For example, in an electricity consumption dataset containing users from multiple industries, industrial users and commercial users can be accurately distinguished in this way because their electricity consumption characteristics are different in the feature matrix and can be correctly classified through similarity calculation.

[0050] By continuously updating the number of clusters and performing multiple rounds of clustering, and evaluating the clustering effect in combination with the loss function calculation formula, the optimal number of clusters and the corresponding user aggregation classes can be found. The loss function can quantify the quality of the clustering result. By iteratively adjusting the number of clusters, the model can continuously optimize the clustering result to make the clustering result more in line with the internal structure of the data. For example, in the analysis of electricity consumption behaviors of power users, the optimal clustering result found in this way can make the electricity consumption behaviors of users within each cluster more similar, while the differences in electricity consumption behaviors between different clusters are more obvious, thus improving the performance of subsequent data analysis and model training based on the clustering result.

[0051] In the initial clustering stage, although the user groups are divided by randomly selecting the cluster center and calculating based on the user's electricity consumption distance degree, due to the randomness of the cluster center selection and the complexity of the actual electricity consumption data, it is inevitable that some users will be misclassified. For example, the electricity consumption behavior of some users may be similar to that of a certain cluster center in some features, but from the overall multi-dimensional features, they are more suitable to belong to other clusters. Without user elimination and secondary division, these misclassified users will interfere with the consistency of the data within the cluster, making the clustering result unable to accurately reflect the true electricity consumption behavior pattern of users.

[0052] By taking the arithmetic mean of the feature vectors of all users in the current cluster to generate the center point, it can comprehensively reflect the overall behavior pattern of the users in the cluster. As a virtual user, the characteristics of the center point are the average of the characteristics of all users in the cluster, so it is more universal. By iteratively adjusting the center point in multiple rounds, the algorithm can gradually optimize the clustering result and reduce the local optimum problem caused by the initial random selection of the cluster center..

[0053] Further, the S6 includes:

[0054] S60. Divide the dataset information corresponding to the same user aggregation class into a training set and a validation set according to a preset ratio;

[0055] S61. Based on the pre-constructed data completion model, determine the parameter search range corresponding to each hyperparameter of the data completion model;

[0056] S61. Randomly generate a number of hyperparameter groups according to the parameter search ranges corresponding to each hyperparameter, and form data completion models corresponding to each hyperparameter group;

[0057] S62. Use the meteorological data, date information data, and feature vector data in the training set as input data, and input them into the data completion models corresponding to each hyperparameter group to predict the prediction data corresponding to each hyperparameter group as the predicted values, and use the corresponding electronic data in the training set as the target values;

[0058] S63. According to the predicted values and corresponding target values corresponding to each hyperparameter group, calculate the loss values of the data completion models corresponding to each hyperparameter group based on a preset loss value calculation formula;

[0059] The preset loss value calculation formula is:

[0060]

[0061] In the formula, MSE is the loss value, y i is the target value, is the predicted value, and n is the number of samples in the dataset information;

[0062] S64. According to the loss values of the data completion models corresponding to each hyperparameter group, determine whether each loss value is less than a preset first loss threshold. If not, mutate all the hyperparameter groups to form multiple new hyperparameter groups and corresponding data completion models, and re-execute S62;

[0063] If so, select the corresponding hyperparameter group, mutate the hyperparameter group to form multiple new hyperparameter groups and corresponding data completion models, re-execute S62 to S63, and determine whether the loss value calculated in S63 is less than the loss value corresponding to this hyperparameter group. If not, this hyperparameter group is the optimal hyperparameter group, and output the data completion model corresponding to this hyperparameter group; if so, the corresponding new hyperparameter group is the optimal hyperparameter group, and output the data completion model corresponding to this new hyperparameter group;

[0064] S65. Use the meteorological data, date information data, and feature vector data obtained by the user in S3 in the validation set as input data, input them into the data completion model corresponding to the output corresponding hyperparameter group, output the corresponding prediction data, and calculate the corresponding prediction error based on the user sub-data in the validation set. Determine whether the prediction error is less than a preset error threshold. If so, the data completion model corresponding to the output corresponding hyperparameter group is a feasible model, otherwise it is an infeasible model, and re-execute S61.

[0065] Beneficial effects: When the loss values of the data completion models corresponding to each hyperparameter group are calculated for the first time, if no loss value is less than the preset first loss threshold, it indicates that the current hyperparameter groups may not have reached an optimal state. At this time, mutating all hyperparameter groups can generate new hyperparameter combinations. Each hyperparameter can potentially produce new values within its search range, and the recombination of different hyperparameter values greatly expands the search space of hyperparameters. Originally, the search for the optimal solution might be confined to certain local regions. After mutation, more hyperparameter combination regions that were not previously explored can be discovered, thus significantly increasing the likelihood of finding the optimal hyperparameter group that can minimize the model's loss value.

[0066] Even if there are hyperparameter groups with loss values less than the preset first loss threshold, it is still very necessary to mutate them. Because the seemingly optimal hyperparameter groups initially found may only be local optimal solutions, that is, they perform well in a certain local region of the model parameter space but are not the global optimum. By mutating this hyperparameter group, the current local optimal state can be broken, and other potentially more optimal parameter combinations can be explored. This enables the model to break free from the limitations of local optimality and adapt to more complex and variable data characteristics and patterns. Brief Description of the Drawings

[0067] Figure 1 It is a flowchart of the method for completing missing values in the daily load data of power users in the first embodiment of the present invention;

[0068] Figure 2 It is a graph showing the change of the number of clusters and the loss function value in the first embodiment of the present invention.

[0069] Figure 3 It is an effect diagram of missing value completion in the first embodiment of the present invention. Detailed Embodiments

[0070] The following is a further detailed description through specific embodiments:

[0071] Embodiment 1

[0072] A method for completing missing values in the daily load data of power users is basically as Figure 1 shown, and includes the following steps:

[0073] S1. Obtain the historical electricity consumption data corresponding to each user in a certain historical time period from the power trading center; for example, obtain the historical electricity consumption data of users at the trading center, with a total of 1147 user numbers and a time length from January 1, 2023 to October 31, 2024.

[0074] S2. Perform data cleaning operations and standardization processing on the obtained historical electricity consumption data corresponding to each user in sequence;

[0075] The said S2 includes:

[0076] S20. Based on a preset line-by-line traversal strategy, perform a line-by-line traversal on each sub-power consumption data in the historical power consumption data corresponding to each user;

[0077] The preset line-by-line traversal strategy is as follows:

[0078] Perform a line-by-line identification and judgment on each sub-power consumption data. Judge whether there is any missing value in each sub-power consumption data in the data row. If so, record the starting index of the missing value. When encountering the data row of the next non-missing value, form a missing interval by combining the previously recorded starting index and the index before the current non-missing value data row, and record it. Then reset the starting index and continue to find the next missing interval until all sub-power consumption data have been identified and judged, and output the set of missing intervals corresponding to the user;

[0079] According to the set of missing intervals corresponding to the user, obtain the power consumption data of the day before and the day after each missing interval, calculate the power consumption ratio between the power consumption data of the day after the missing interval and the power consumption data of the day before the missing interval, and judge whether the power consumption ratio is greater than the preset ratio. If so, judge that there is a data stacking situation in this missing interval; otherwise, judge that there is no data stacking situation in this missing interval;

[0080] S21. When the judgment result is that there is a data stacking situation in this missing interval, perform cleaning on the data stacking situation of the corresponding missing interval;

[0081] S22. After completing the cleaning of the data stacking situation, exclude outliers from each sub-power consumption data corresponding to each user;

[0082] S23. Perform standardization processing on each sub-power consumption data corresponding to each user after excluding outliers.

[0083] S3. Extract the features corresponding to each user based on the historical electricity consumption data of each user after standardization processing, and generate the feature vector data corresponding to each user. The feature vector data includes the data domain feature data, time domain feature data, and frequency domain feature data of the user. The data domain feature data includes the data of the district or county where the user is located, industrial and commercial information, and the ratio data between demand charge and capacity charge. Among them, demand charge is a way of calculating electricity charges based on the actual maximum electricity demand of the user. Capacity charge is a way of calculating electricity charges based on the rated capacity of the user's transformer. When extracting features, first convert some of the original data into numerical labels that can represent the user's electricity consumption behavior for subsequent machine learning analysis. Map the name of the district or county where the user is located to a unique code through address classification as the feature of the district or county information. Define the user type, where 1 represents commercial electricity consumption and 2 represents industrial electricity consumption to represent the industrial and commercial types. Define the charging method, where 3 represents demand charge and 4 represents capacity charge.

[0084] The time domain feature data includes the maximum electricity consumption value, the maximum absolute value of electricity consumption, the minimum electricity consumption value, the average electricity consumption value, the peak-to-peak value of electricity consumption, the absolute square value of electricity consumption, the root mean square value of electricity consumption, the amplitude of electricity consumption, the standard deviation of electricity consumption, the kurtosis of electricity consumption, the skewness of electricity consumption, the margin index, the impulse index, and the peak index corresponding to the user. The calculation formulas corresponding to the time domain feature data are shown in Table 1:

[0085]

[0086]

[0087] Table 1

[0088] In this embodiment, when the user's power consumption is stable, the maximum value, the maximum absolute value (which can also be regarded as the peak value), or the minimum value has a small change range and is basically stable below a threshold. However, once the maximum value, the maximum absolute value, or the minimum value becomes abnormally large or small, it can be basically considered that there has been a change in the user's power consumption situation. If it becomes too large or too small to a certain extent, there must be a certain statistical error. The mean value reflects the change of the vibration signal generated due to the change of the power consumption situation during the user's power consumption process. The root mean square value, also known as the effective value, reflects the energy intensity and stability of the vibration signal. Power load forecasting usually focuses most on this indicator. If this indicator becomes abnormally large, it means that the power consumption characteristics are strongly interfered by some external factors. Kurtosis reflects the impact characteristics of the vibration signal. Kurtosis is sensitive to impacts. Generally, the kurtosis value should be around 3 because the kurtosis of the normal distribution is equal to 3. If it deviates too much from 3, it indicates that the power consumption behavior habits of power users are easily interfered by external factors. Skewness reflects the asymmetry of the vibration signal. Usually, the vibration signal is symmetric about the x-axis, and at this time, the skewness should approach 0. If the periodic average value of the power consumption behavior moves up or down, the skewness will become larger. The margin index is used to characterize the change rate characteristics of the power consumption periodic change. The impulse index and the peak index are both used to detect whether there are impacts in the signal

[0089] The frequency domain characteristic data includes the power consumption center frequency, the root mean square of the power consumption frequency, the average power consumption frequency, and the variance of the power consumption frequency. The calculation formulas for the respective data corresponding to the frequency domain characteristic data are shown in Table 2:

[0090]

[0091] Table 2

[0092] In the table, f k is the frequency value of the k-th component, and S(k) is the weight value of the k-th component.

[0093] S4. According to the data domain characteristic data, time domain characteristic data, and frequency domain characteristic data corresponding to each user, based on a preset feature clustering strategy, perform clustering processing on each user to form multiple user aggregation classes;

[0094] The preset feature clustering strategy is as follows:

[0095] Step 1. According to the data domain characteristic data, time domain characteristic data, and frequency domain characteristic data corresponding to each user, perform data integration on the feature vector data corresponding to each user to form a feature matrix X corresponding to all users;

[0096]

[0097] In the formula, [x i1 ,x i2,…,x in is the feature vector corresponding to user i, and x in is the dimension value corresponding to the nth dimension in the corresponding feature vector data of user i; in this embodiment, the dimensions of the data corresponding to the data domain feature data, time domain feature data, and frequency domain feature data are determined. For example, the data of the district or county where the user is located is the first dimension, and the industrial electricity consumption is the second dimension.

[0098] Step 2: Determine the number of clusters K corresponding to this clustering;

[0099] Step 3: Based on the feature matrix corresponding to all users and the number of clusters K, randomly select K from m users as the cluster centers corresponding to this clustering;

[0100] Step 4: Based on the preset user electricity consumption distance calculation formula and the feature matrix corresponding to the remaining users, calculate the user electricity consumption distance between the remaining users and the users corresponding to the K cluster centers;

[0101] The preset user electricity consumption distance calculation formula is:

[0102]

[0103] In the formula, d (x,u) is the user electricity consumption distance between the user and the user corresponding to the cluster center, x ij is the dimension value corresponding to the jth dimension of user i, and u j is the dimension value corresponding to the jth dimension of the user corresponding to the cluster center;

[0104] Step 5: According to the user electricity consumption distances between each remaining user and the users corresponding to each cluster center, compare the user electricity consumption distances between the remaining users and the users corresponding to each cluster center, and classify the user corresponding to the cluster center with the smallest user electricity consumption distance and the remaining user into the same class to form the initial aggregation classes corresponding to each cluster center;

[0105] Step 6: According to the user electricity consumption distances in each initial aggregation class, calculate the arithmetic mean of each dimension corresponding to the users in each initial aggregation class to obtain the center points of each initial aggregation class, calculate and compare the user electricity consumption distances between each user in each initial aggregation class and the center points of each initial aggregation class, and classify each user into the aggregation class corresponding to the center point with the smallest user electricity consumption distance in turn;

[0106] Step 7: After completing the user adjustment of all cluster aggregation classes, output the users of the current round of aggregation classes corresponding to each cluster center to form the user aggregation classes corresponding to each cluster center;

[0107] Step 8: After the formation of the user aggregation classes corresponding to each cluster center for the current clustering is completed, perform the next round of clustering, update the number of clusters corresponding to the next round of clustering, and re-execute Step 3 until the number of clusters reaches the preset number threshold;

[0108] Step 9: Based on each user aggregation class corresponding to each cluster center under each number of clusters, calculate the loss function values corresponding to each number of clusters according to the preset loss function calculation formula, and determine and output the optimal number of clusters and each user aggregation class corresponding to the optimal number of clusters and each cluster center.

[0109] The preset loss function calculation formula is:

[0110]

[0111] In the formula, J is the loss function value, c (d) is the cluster category to which the data point x (d) belongs. Here, d ranges from 1 to m, corresponding to m data points, and each data point is assigned to one of the K clusters. c (d) is the label used to identify the cluster to which the data point belongs; m is the total number of data points, and in this clustering scenario, it is the number of all user feature vectors participating in the clustering. It is used as the divisor when calculating the average loss value to measure the overall data scale. μ 1 , …, μ K , which are the center vectors of the K clusters respectively. μ K (k = 1, 2, ..., K) represents the center of the kth cluster, which is the mean of the feature vectors of all data points within the cluster and reflects the central tendency of the data in the cluster. x (d) is the feature vector of the ith data point, corresponding to the feature vector of each user after integration, and it contains multi-dimensional information such as data domain features, time domain features, and frequency domain features. is the square of the Euclidean distance from the data point x (d) to its corresponding cluster center . In this embodiment, the elbow method is used to determine the optimal number of clusters in the classification of different clusters, that is, the number of clustering clusters is determined through the elbow point. As Figure 2 shown, the number of clusters K is on the x-axis, and the loss function value is on the y-axis. Among the classification parameters of different clusters, an "elbow point" appears at K = 5, and the reduction rate of the within-cluster sum of squares significantly slows down. Therefore, the optimal number of K-means clusters is determined to be 5.

[0112] For example, when forming 5 clusters, the user situation corresponding to each cluster is shown in Table 3:

[0113] cluster total number of users number of users with complete data number of users with missing values 1 261 198 63 2 644 473 171 3 341 289 52 4 284 261 23 5 596 440 156

[0114] Table 3

[0115] S5. Obtain the meteorological data and date information data corresponding to each user in the same user aggregation class, and integrate the meteorological data and date information data corresponding to each user with the historical electricity consumption data and feature vector data corresponding to the user to form the training dataset data corresponding to the corresponding user aggregation class; in this embodiment, the meteorological data includes the temperature, humidity, and precipitation information corresponding to the user's location information, and the date information data includes the year, month, the number of weeks in a month, the day of the month, the day of the week, whether it is a legal holiday, and whether it is an adjusted holiday.

[0116] S6. Based on the training dataset data corresponding to the same user aggregation class, train the data completion model corresponding to the user aggregation class based on the pre-constructed data completion model and the model training strategy corresponding to the data completion model. After the training is completed, output the corresponding data completion model and associate it with the corresponding user aggregation class;

[0117] The S6 includes:

[0118] S60. Divide the dataset information corresponding to the same user aggregation class into a training set and a validation set according to a preset ratio;

[0119] S61. Based on the pre-constructed data completion model, determine the parameter search range corresponding to each hyperparameter of the data completion model;

[0120] S61. According to the parameter search range corresponding to each hyperparameter, randomly generate several hyperparameter groups to form the data completion models corresponding to each hyperparameter group;

[0121] S62. Use the meteorological data, date information data, and feature vector data in the training set as input data, input them into the data completion models corresponding to each hyperparameter group, predict the predicted data corresponding to each hyperparameter group as the predicted value, and use the corresponding electronic data in the training set as the target value;

[0122] S63. According to the predicted value and the corresponding target value corresponding to each hyperparameter group, calculate the loss value of the data completion model corresponding to each hyperparameter group based on the preset loss value calculation formula;

[0123] The preset loss value calculation formula is:

[0124]

[0125] In the formula, MSE is the loss value, y i is the target value, is the predicted value, and n is the number of samples in the dataset information;

[0126] S64. Complete the loss value of the model according to the data corresponding to each hyperparameter group, and determine whether each loss value is less than a preset first loss threshold. If not, mutate all the hyperparameter groups to form multiple new hyperparameter groups and corresponding data completion models, and re - execute S62;

[0127] If so, select the corresponding hyperparameter group, mutate the hyperparameter group to form multiple new hyperparameter groups and corresponding data completion models, re - execute S62 to S63, and determine whether the loss value calculated in S63 is less than the loss value corresponding to the hyperparameter group. If not, the hyperparameter group is the optimal hyperparameter group, and output the data completion model corresponding to the hyperparameter group; if so, the corresponding new hyperparameter group is the optimal hyperparameter group, and output the data completion model corresponding to the new hyperparameter group;

[0128] S65. Take the meteorological data, date information data in the validation set, and the feature vector data obtained by the user in S3 as input data, input them into the data completion model corresponding to the output corresponding hyperparameter group, output the corresponding prediction data, calculate the corresponding prediction error based on the user sub - data in the validation set, and determine whether the prediction error is less than a preset error threshold. If so, the data completion model corresponding to the output corresponding hyperparameter group is a feasible model, otherwise it is an infeasible model, and re - execute S61.

[0129] S7. Take the meteorological information data, date information data, and feature vector data corresponding to a certain user in the user aggregation class as input data, input them into the data completion model associated with the user aggregation class, output the missing values corresponding to the user, and fill the missing values into the historical electricity consumption data of the user to form the filled historical electricity consumption data. As Figure 3 shown, use the users with complete data from June 1, 2024 to October 31, 2024 as the validation set, intercept some users as the test set, randomly place missing values, and then fill them through the above steps and merge them with the validation set for comparison. R = ratio (true value / test value). In the algorithm filling result, the deviation rate is controlled within ±2%.

[0130] The above are only embodiments of the present invention. Specific structures and characteristics and other common knowledge in the art are described in detail here. Those of ordinary skill in the art know all the common general technical knowledge in the technical field to which the invention belongs before the filing date or the priority date, can know all the prior art in this field, and have the ability to apply the conventional experimental means before this date. Those of ordinary skill in the art can, under the inspiration given in this application, combine their own abilities to improve and implement this solution. Some typical well-known structures or well-known methods should not become an obstacle for those of ordinary skill in the art to implement this application. It should be noted that for those skilled in the art, without departing from the structure of the present invention, several deformations and improvements can also be made, and these should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent. The protection scope required by this application should be based on the content of its claims, and the specific implementation manners and other records in the specification can be used to interpret the content of the claims.

Claims

1. A method for completing missing values ​​of daily load data of power users, characterized by: The following steps are involved: S1. Obtain the historical electricity consumption data corresponding to each user in a certain historical period from the power trading center; S2. Perform data cleaning and standardization on the historical electricity consumption data corresponding to each user in turn; S3. Extract the features corresponding to each user according to the historical electricity consumption data corresponding to each user after the standardized processing, and generate feature vector data corresponding to each user, wherein the feature vector data includes the data domain feature data, time domain feature data and frequency domain feature data of the user; S4, clustering each user according to the data domain feature data, time domain feature data and frequency domain feature data corresponding to each user based on a preset feature clustering strategy to form multiple user aggregation classes; S5, obtaining meteorological data and date information data corresponding to each user in the same user aggregation class, and integrating the meteorological data and date information data corresponding to each user with the historical electricity consumption data and feature vector data corresponding to the user to form a training data set data corresponding to the corresponding user aggregation class; S6. According to the training data set data corresponding to the same user aggregation class, based on the pre-built data completion model and the model training strategy corresponding to the data completion model, the data completion model corresponding to the user aggregation class is trained. After the training is completed, the corresponding data completion model is output and associated with the corresponding user aggregation class; S7. The meteorological information data, date information data corresponding to a user in the user aggregation class and the feature vector data obtained by the user in S3 are used as input data and input into the data completion model associated with the user aggregation class, the missing values ​​corresponding to the user are output, and the missing values ​​are filled into the historical electricity consumption data of the user to form the filled historical electricity consumption data.

2. A method for completing missing values ​​of daily load data of power users according to claim 1, characterized in that: The S2 includes: S20, based on a preset line-by-line traversal strategy, traverse line by line each of the electricity consumption data in the historical electricity consumption data corresponding to each user; The preset line-by-line traversal strategy is: Identify and judge each electronic data row by row to determine whether each electronic data in the data row is missing. If so, record the starting index of the missing value. When encountering the next data row with non-missing values, combine the previously recorded starting index with the previous index of the current data row with non-missing values ​​to form a missing interval, record it, reset the starting index, and continue to search for the next missing interval until all electronic data are identified and judged, and output the missing interval set corresponding to the user; According to the missing interval set corresponding to the user, the electricity consumption data of the previous day and the next day corresponding to each missing interval is obtained, the electricity consumption ratio between the electricity consumption data of the day after the missing interval and the electricity consumption data of the day before the missing interval is calculated, and it is determined whether the electricity consumption ratio is greater than a preset ratio. If so, it is determined that there is data overlap in the missing interval, otherwise, it is determined that there is no data overlap in the missing interval; S21, when the result of the judgment is that there is data overlap in the missing interval, the corresponding missing interval is cleaned of data overlap; S22, after the data stacking is cleaned, outliers are eliminated from each electronic data corresponding to each user; S23. Standardize the electronic data corresponding to each user after the outliers are eliminated.

3. A method for completing missing values ​​of daily load data of power users according to claim 2, characterized in that: The preset feature clustering strategy is: Step 1: According to the data domain feature data, time domain feature data and frequency domain feature data corresponding to each user, the feature vector data corresponding to each user is integrated to form a feature matrix X corresponding to all users; In the formula, [x i1 ,x i2 ,…,x in ] is the feature vector corresponding to user i, x in is the dimension value corresponding to the nth dimension in the corresponding feature vector data of user i; Step 2: Determine the number of clusters K corresponding to this clustering; Step 3: According to the feature matrix corresponding to all users and based on the number of clusters K, randomly select K users from the m users as the cluster centers corresponding to this clustering; Step 4: According to the feature matrix corresponding to the remaining users and based on the preset user power usage distance calculation formula, calculate the user power usage distance between the remaining users and the users corresponding to the K cluster centers; The preset user power usage distance calculation formula is: Where, d (x,u) is the user electricity distance between the user and the user corresponding to the cluster center, x ij is the dimension value corresponding to the jth dimension corresponding to user i, u j is the dimension value corresponding to the j-th dimension of the user corresponding to the cluster center; Step 5: According to the user power usage distances between each remaining user and the user corresponding to each cluster center, the user power usage distances between the remaining users and the users corresponding to each cluster center are compared, and the user corresponding to the cluster center with the smallest user power usage distance is classified into the same class as the remaining users, thereby forming an initial aggregation class corresponding to each cluster center; Step 6: According to the electricity usage distance of each user in the initial aggregation class, the arithmetic mean of each dimension corresponding to the user in each initial aggregation class is calculated to obtain the center point of each initial aggregation class, and the user electricity usage distance between each user in each initial aggregation class and the center point of each initial aggregation class is calculated and compared and judged, and each user is sequentially divided into the aggregation class corresponding to the center point with the smallest user electricity usage distance; Step 7: After completing the adjustment of users of all cluster aggregation classes, output the users of this round of aggregation classes corresponding to each cluster center to form user aggregation classes corresponding to each cluster center; Step 8: After the formation of the user aggregation classes corresponding to the cluster centers corresponding to the current clustering is completed, the next round of clustering is performed and the number of clusters corresponding to the next round of clustering is updated, and step 3 is re-executed until the members in the cluster no longer change; Step 9: According to the user aggregation classes corresponding to the cluster centers under each number of clusters, based on the preset loss function calculation formula, calculate the loss function value corresponding to each number of clusters, and determine and output the optimal number of clusters and the user aggregation classes corresponding to the cluster centers corresponding to the optimal number of clusters.

4. A method for completing missing values ​​of daily load data of power users according to claim 3, characterized in that: The S6 includes: S60, dividing the data set information corresponding to the same user aggregation class into a training set and a validation set according to a preset ratio; S61. Based on the pre-built data completion model, determine the parameter search range corresponding to each hyperparameter corresponding to the data completion model; S61. According to the parameter search range corresponding to each hyperparameter, a number of hyperparameter groups are randomly generated to form a data completion model corresponding to each hyperparameter group; S62, using the meteorological data, date information data and feature vector data in the training set as input data, inputting them into the data completion model corresponding to each hyperparameter group, predicting the predicted data corresponding to each hyperparameter group as the predicted value, and using the corresponding electronic data in the training set as the target value; S63, according to the predicted value and the corresponding target value corresponding to each hyperparameter group, based on the preset loss value calculation formula, calculate the loss value of the data completion model corresponding to each hyperparameter group; The preset loss value calculation formula is: In the formula, MSE is the loss value, y i is the target value, is the predicted value, n is the number of samples in the data set information; S64, according to the loss value of the data completion model corresponding to each hyperparameter group, determine whether each loss value is less than a preset first loss threshold, if not, mutate all hyperparameter groups to form multiple new hyperparameter groups and corresponding data completion models, and re-execute S62; If so, select the corresponding hyperparameter group, mutate the hyperparameter group, form multiple new hyperparameter groups and corresponding data completion models, re-execute S62 to S63, and determine whether the loss value calculated by S63 is smaller than the loss value corresponding to the hyperparameter group. If not, the hyperparameter group is the optimal hyperparameter group, and the data completion model corresponding to the hyperparameter group is output; if so, the corresponding new hyperparameter group is the optimal hyperparameter group, and the data completion model corresponding to the new hyperparameter group is output; S65. Use the meteorological data, date information data in the verification set, and the feature vector data obtained by the user in S3 as input data, and input them into the data completion model corresponding to the corresponding hyperparameter group of the output, output the corresponding prediction data, and calculate the corresponding prediction error based on the user sub-data in the verification set, and determine whether the prediction error is less than the preset error threshold. If so, the data completion model corresponding to the corresponding hyperparameter group of the output is a feasible model, otherwise it is an infeasible model, and re-execute S61.

Citation Information

Patent Citations

  • A cleaning method for power utilization time sequence data

    CN112732694A

  • A method for complementing frozen data of an electric energy meter in a missing day

    CN113239029A

  • TransCNN medical eye fundus image classification algorithm based on hyper-parameter optimization

    CN115965807A

  • User electrical load data prediction method and system, terminal and storage medium

    CN118017503A

  • Deep learning-based short-term power load prediction method

    CN119627865A

Cited By

  • Time series data quality enhancement method and device based on prediction large model

    CN120975286A

  • Electric vehicle battery intelligent charging control method and system and intelligent charger

    CN121157703A