A method suitable for multiple types of mass power user data completion

By employing time-series power data decomposition, delayed embedding transformation, and tensor parallel factor decomposition, the missing data problem in the power data acquisition process is solved, achieving high-precision completion of massive power user data of various types. This method is applicable to power systems and power markets and has significant practical engineering application value.

CN119782290BActive Publication Date: 2026-03-03TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-06
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing technologies cannot effectively address the problem of missing data during power data acquisition, especially the lack of data integrity caused by database write failures and hardware malfunctions during high-load periods. Furthermore, machine learning methods are not effective in cases of missing high-dimensional data, and low-rank tensor decomposition models do not perform well in repairing one-dimensional data.

Method used

We employ time-series power data decomposition, delay embedding transform, and tensor parallel factor decomposition. We verify the periodic terms through autocorrelation calculation and fast Fourier transform. Combining time-series smoothing constraints and periodic term constraints, we use alternating least squares method to solve the comprehensive tensor completion objective function and output complete data.

Benefits of technology

It achieves high-precision completion of massive amounts of electricity user data of various types, is applicable to one-dimensional and multi-dimensional scenarios, has low computational complexity, is suitable for user electricity consumption data in power systems and electricity markets, and has good scalability and completion effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119782290B_ABST
    Figure CN119782290B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of suitable for multiple types of mass power user data completion method, comprising the following steps: step 1, the power consumption data of multiple types of power users is collected;Step 2, the power consumption data of single different user is decomposed into the sum of trend term and periodic term;Step 3, periodic term verification is carried out;Step 4, the power consumption data of single user collected in step 1 and the periodic term, trend term time series vector after decomposition in step 2 are carried out standard delay transformation, obtain the hankel form tensor;Step 5, establish time series smoothing constraint and periodic term constraint;Step 6, establish comprehensive tensor completion objective function;Step 7, solve the comprehensive tensor completion objective function obtained in step 6 and output complete data by anti-MDT technology.The present application can be applied to one-dimensional and multidimensional data scene, and the completion effect is good.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power user data completion technology, and relates to a power user data completion method, especially a method applicable to massive amounts of multi-type power user data completion. Background Technology

[0002] Electricity data plays a vital role in modern society. It is not only the cornerstone of stable power grid operation and efficient management, but also a key element in promoting energy transformation and achieving green and low-carbon development. By collecting and analyzing electricity data in real time, we can accurately predict electricity demand, optimize resource allocation, reduce waste, and ensure a safe and reliable power supply. At the same time, electricity data provides scientific basis for enterprises to conserve energy, reduce emissions, and control costs, helping them achieve sustainable development. In the construction of smart cities, electricity data is an important component of smart energy systems, promoting the formation of the energy internet and improving urban energy utilization efficiency and management levels. Therefore, the accuracy, timeliness, and security of electricity data have immeasurable value for economic and social development.

[0003] Currently, power data in the power system encompasses infrastructure data such as power grid lines, substations, and switching equipment, as well as information on the grid's operational status and security. This data is crucial for power grid monitoring, dispatching, and fault diagnosis. In the power market, data also includes market transaction data, electricity price data, and market participant information. This data reflects the operational status and trading activities of the power market and is of great significance for its regulation and control.

[0004] However, due to: aging or malfunctions of data acquisition devices such as electricity meters leading to missing data; communication disruptions or malicious network attacks during power data transmission causing data loss; the irreversible nature of power time-series data acquisition and its high real-time requirements, leading to data loss when data writing operations occur during periods of high load and the database cannot handle all concurrent write requests; and system hard drive failures, memory failures, or other hardware failures during data storage potentially causing data write failures and resulting in data loss, data integrity cannot be guaranteed, leading to poor performance of data-driven analysis, prediction, and identification models.

[0005] Furthermore, current machine learning-based data completion methods cannot effectively address high-dimensional missing data situations and require a large amount of non-missing data for pre-training. While low-rank tensor and ordinary tensor decomposition models can effectively handle high-dimensional data, their repair performance is far lower than that of machine learning-based methods when dealing with one-dimensional data repair problems.

[0006] Therefore, this invention proposes a method for completing massive amounts of power user data of various types, which can solve the above problems.

[0007] A search revealed no publicly available literature of the same or similar prior art as this invention. Summary of the Invention

[0008] The purpose of this invention is to overcome the shortcomings of the prior art and propose a method for completing massive amounts of power user data of various types. It is applicable to both one-dimensional and multi-dimensional data scenarios and has good completion effect.

[0009] The present invention solves its practical problem by adopting the following technical solution:

[0010] A method for data completion applicable to massive amounts of electricity user data of various types includes the following steps:

[0011] Step 1: Collect electricity consumption data from multiple types of electricity users and extract the electricity consumption data of any single user.

[0012] Step 2: Based on the individual user electricity consumption data extracted in Step 1, decompose the individual different user electricity consumption data into a sum of trend terms and periodic terms based on the time series power data decomposition method;

[0013] Step 3: Verify the periodic terms using autocorrelation and fast Fourier transform calculations on the periodic term data obtained in Step 2;

[0014] Step 4: Using the delayed embedding transformation technique, perform a standard delayed transformation on the individual user electricity consumption data collected in Step 1 and the decomposed periodic and trend time series vectors in Step 2 to obtain the Hankel form tensor.

[0015] Step 5: Based on the trend and periodic data obtained in Step 2, establish time-series smoothing constraints and periodic constraints;

[0016] Step 6: For the Hankel form tensor of the trend term data obtained in Step 4, tensor parallel factor decomposition is used to obtain the tensor parallel factor decomposition model. Combined with the two constraints in Step 5, a comprehensive tensor completion objective function is established, which is applicable to different types of power users.

[0017] Step 7: Solve the comprehensive tensor obtained in Step 6 using Alternating Least Squares (HALS) and output the complete data using inverse MDT.

[0018] Furthermore, the specific steps of step 2 include:

[0019] The user's electricity consumption data is decomposed into a weighted sum of trend and periodic components using a moving average.

[0020] Perform additive decomposition on the user electricity consumption data extracted in step 1:

[0021]

[0022] In the formula: α=2N+1 is the order of smoothness, x t For electricity user electricity consumption data, s t For trend term and e t It is a periodic term.

[0023] Furthermore, the specific method for step 3 is as follows:

[0024] Periodic terms were validated using autocorrelation and fast Fourier transform calculations:

[0025]

[0026] In the formula: e t Let T be the mean of the periodic terms after sequence decomposition, and T be the lag order.

[0027] Furthermore, the specific method for step 4 is as follows:

[0028] For single-user electricity consumption data: Assuming a basic cycle of one day, selecting n days of data with a data length of L, we can obtain L×n vectors, where the electricity consumption on the i-th (i=1,…,n) day can be represented as x i ∈x L×1 ;

[0029] Then, using the delayed embedding transformation technique, the time series vector is subjected to a standard delayed transformation, also known as Hankelization, to obtain the Hankel form matrix;

[0030]

[0031] Where L is the length of the data vector, τ is the size of the delay window, and fold (τ,L-τ+1) :R τ(L-τ+1) →R τ×(L-τ+1) It is a folding operator, δ∈{0,1} τ(L-τ+1)×L It is a repeating matrix that satisfies:

[0032] vec(H τ (v))=δx (4)

[0033] Stacking the Hankel matrices established above along the third dimension forms a Hankel tensor.

[0034] Furthermore, the specific method for step 5 is as follows:

[0035] Using the trend and periodic data obtained from step 2, establish time-series smoothing constraints: Constraints on periodic terms:

[0036] Furthermore, the objective function and constraints for the comprehensive tensor completion in step 6 are as follows:

[0037]

[0038]

[0039] Among them: E to The value represents all values ​​within the period containing the missing element, where m is the number of delays, T is the period, and w is the weighting coefficient (w t-mT +…+w t+mT =1), λ e It is the regularization coefficient. This represents the complete Hankel output tensor. Use 0 and 1 to represent missing and available elements, The Kronecker product, in Hankel form, is represented by a tensor for missing terms. To fill:

[0040]

[0041] in: ρ=[ρ (1) ,ρ (2) ,...,ρ (N) ] T It is a smoothness parameter vector, L1∈R (In-1)×In , where In∈R, is the number of elements in the factor vector. The first derivative operator is defined as:

[0042]

[0043] In the established objective function:

[0044] Represents observation terms and reconstruction tensor The root mean square error (RMSE) of the values ​​is used to estimate the difference between the two, where

[0045] Time-series smoothing regularization term: After decomposing the electricity consumption data of electricity users participating in spot trading in the first part, the smoothing trend part is... Furthermore, by incorporating the special structure of the Hankel tensor and adding temporal smoothness constraints (i.e., the values ​​of adjacent elements are close), this invention employs a quadratic variation (QV) constraint. This regularization term is a formal expression of the temporal smoothness constraint, defined as:

[0046]

[0047] (3) This indicates a strong correlation between the periodicity of the data. However, due to the irregular noise in the electricity consumption data and the fact that the data does not strictly follow the periodic terms after the autocorrelation coefficient calculation, a suitable set of parameters W is found based on the periodic values ​​of each non-missing value in the data to fit the missing values ​​in the periodic terms.

[0048] R is the rank of the parallel factorization of the trend term. The choice of R needs to balance accuracy and computational complexity. Too low an R value can reduce computational complexity, but will produce a large truncation error, affecting model accuracy; too high an R value will increase computational complexity.

[0049] The optimization of R is achieved by conducting a truncation error experiment. The specific steps are as follows:

[0050] (1) The upper limit of R max Related to the size of the tensor; For example, R max The values ​​are shown in equation (9):

[0051] R max ≤min(I×J,I×K,J×K) (9)

[0052] In the formula, I, J, and K represent the rows, columns, and number of tensors of the tensor, respectively.

[0053] (2) Use equation (10) to determine the cutoff error Q of parallel factorization. CP :

[0054]

[0055] In the formula: Let this represent the non-missing index tensor reconstructed from the r groups of CP decomposition components. The truncation error Q is selected. CP The rank corresponding to the moment when it gradually decreases and becomes relatively stable within a certain range is the rank of the tensor parallel factorization in this paper.

[0056] Furthermore, the specific method for step 7 is as follows:

[0057] The Alternating Least Squares (HALS) method is used to solve the optimization problem. According to the HALS algorithm, consider minimizing the following local objective function:

[0058]

[0059] in:

[0060] in:

[0061] Optimize the above equation in the following order, updating the parameter matrix W and the factor vector. Adaptive factor g r ,:

[0062]

[0063] in: It is a subset of unit vectors;

[0064] (1) Update W: Optimize the weight coefficients W to simplify calculations and achieve the goal of optimizing parameters. p (W) simplifies to:

[0065]

[0066] The parameters are updated according to the following rules:

[0067]

[0068] In the formula: η is the learning rate. Is the function about W k The gradient.

[0069] Find a suitable set of weighting coefficients, based on E t The expression calculates the missing values ​​for the periodic terms. The completed periodic sequence is then transformed into a Hankel tensor using a delayed embedding transformation.

[0070] (2) Update u: To prevent non-differentiable cases during optimization, this invention patent selects a quadratic variational constraint (QV) and adopts a gradient-normalized update form. To simplify the notation, u is used. k Indicate u r (n) kth update, v k The gradient of the objective function

[0071]

[0072] The above equation assumes that the gradient is updated on the hypersphere. To facilitate the representation of the gradient of the objective function, It can be simplified to:

[0073]

[0074] in:

[0075]

[0076] The gradient of the objective function is then:

[0077]

[0078] (3) Update g r The problem of updating the adaptive factor is an unconstrained quadratic optimization problem, and the unique solution can be obtained analytically. The objective function can be simplified to:

[0079]

[0080] So g r The update rules are as follows:

[0081]

[0082] The above three steps are iteratively solved using HALS until the error between two updates of each variable is less than a set threshold ε, typically set to 0.001. The completed Hankel tensor is then obtained. Output complete data through inverse delay transform:

[0083]

[0084] in This is the Moore–Penrose pseudo-inverse transform, unfold (τ,L-τ+1) fold (τ,L-τ+1) The inverse operation.

[0085] Advantages and beneficial effects of the present invention:

[0086] 1. This invention proposes a data completion method applicable to massive amounts of electricity user data of various types. First, the data is decomposed into trend terms and periodic terms, and then transformed into Hankel tensors. The data is then completed in one-dimensional or multi-dimensional form by combining the tensor parallel factor decomposition model.

[0087] 2. This invention proposes a data completion method applicable to massive amounts of electricity user data of various types. Addressing the issue of missing electricity consumption data for different users, it achieves high-precision data completion through time-series electricity consumption data decomposition, delayed embedding transformation, and tensor parallel factor decomposition. This invention is applicable to completion of data for different electricity user types, not only for data collected in power systems but also for user electricity consumption data in the electricity market, exhibiting good scalability and low computational complexity. Specifically, the acquired data is first classified to determine the electricity user type. Then, the data undergoes a novel time-series decomposition. Next, the decomposed one-dimensional sequence data is transformed into Hankel tensors. For the trend and periodic terms, corresponding Hankel tensors are also transformed, and corresponding constraints are established. A tensor parallel factor decomposition model is established for the trend term. Then, a comprehensive completion model applicable to different electricity user types is established by combining the constraints of the trend and periodic terms with the established tensor parallel factor decomposition model. Finally, hierarchical alternating least squares solution and inverse delayed transformation are used to output the complete data. A set of evaluation indicators is used to compare the completion effect of the proposed method with existing methods. The results show that the present invention is applicable to one-dimensional and multi-dimensional data scenarios, and the completion effect is good, proving the good practical application significance of the method. Attached Figure Description

[0088] Figure 1 This is a flowchart of the processing of the present invention;

[0089] Figure 2 This is a data decomposition diagram of the present invention;

[0090] Figure 3 This is a diagram of the Hankel tensor transformation of the present invention;

[0091] Figure 4 This is a diagram illustrating the effect of completing single-user power consumption data according to the present invention.

[0092] Figure 5 This is a diagram illustrating the power consumption completion effect of the load aggregator (multi-user) according to the present invention. Detailed Implementation

[0093] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings:

[0094] A data completion method applicable to massive amounts of electricity user data of various types is proposed. Through multiple steps including data decomposition, periodic term constraints, smoothing trend term constraints, and establishment of a comprehensive objective function, it achieves high-precision completion of missing electricity consumption values ​​for different electricity user types. The method includes the following steps:

[0095] Step 1: Collect electricity consumption data from multiple types of electricity users and extract the electricity consumption data of any single user.

[0096] The residential electricity consumption data, coal mine user electricity consumption data from a certain province in China, and commercial electricity consumption data from a load participating in the Chinese electricity market obtained by this invention represent three different types of electricity users: ordinary residential users, industrial and commercial users, and comprehensive users including multiple entities.

[0097] Analysis of different user types reveals the following characteristics: Ordinary residential users consume relatively small amounts of electricity, exhibiting a certain degree of periodicity and minimal fluctuations in electricity consumption data; Industrial and commercial users consume large amounts of electricity, with significant fluctuations due to the start-up, shutdown, or withdrawal of high-power generating units, and also exhibiting strong periodicity; Load aggregators can aggregate multiple small and medium-sized electricity users together, and an aggregator typically includes two or more electricity users, with the electricity consumption data changes of each electricity user exhibiting similar characteristics to those of industrial and commercial users.

[0098] Step 2: Based on the individual user electricity consumption data extracted in Step 1, decompose the individual different user electricity consumption data into a sum of trend terms and periodic terms based on the time series power data decomposition method;

[0099] The specific method for step 2 is as follows:

[0100] The electricity consumption data of individual users is decomposed into a sum of trend and periodic terms using the moving average method;

[0101] In this embodiment, step 2 includes the following specific steps:

[0102] A novel time series data decomposition method is proposed: user electricity consumption data is decomposed into a weighted sum of trend and periodic terms using moving averages, as shown in the figure. Figure 2 As shown.

[0103] Perform additive decomposition on the user electricity consumption data extracted in step 1:

[0104]

[0105] In the formula: α=2N+1 is the order of smoothness, x t For electricity user electricity consumption data, s t For trend term and e t It is a periodic term.

[0106] Step 3: Verify the periodic terms using autocorrelation and fast Fourier transform calculations on the periodic term data obtained in Step 2;

[0107] Step 3 can verify the periodic term in step 2 through autocorrelation calculation and fast Fourier transform. The specific method of step 3 is as follows:

[0108] Periodic terms were validated using autocorrelation and fast Fourier transform calculations:

[0109]

[0110] In the formula: e t Let T be the mean of the periodic terms after sequence decomposition, and T be the lag order.

[0111] Step 4: Use the delayed embedding transformation technique to perform a standard delayed transformation, also known as Hankel transformation, on the individual user electricity consumption data collected in Step 1 and the time series vectors of the periodic and trend terms decomposed in Step 2 to obtain the Hankel form tensor.

[0112] The specific method for step 4 is as follows:

[0113] For the individual user electricity consumption data collected in step 1 and the trend and periodic terms obtained from step 2, the Delayed Embedding Transform (MDT) technique is used to transform them into Hankel tensors, such as... Figure 3 As shown, the specific steps are as follows:

[0114] For single-user electricity consumption data: Assuming a basic cycle of one day, selecting n days of data with a data length of L, we can obtain L×n vectors, where the electricity consumption on the i-th (i=1,…,n) day can be represented as x i ∈x L×1 ;

[0115] Then, by using the delayed embedding transformation technique, the time series vector is subjected to a standard delayed transformation, also known as Hankelization, to obtain a Hankel form matrix, which can give the data good properties such as low rank or smoothness.

[0116]

[0117] Where L is the length of the data vector, τ is the size of the delay window, and fold (τ,L-τ+1) :R τ(L-τ+1) →R τ×(L-τ+1) It is a folding operator, δ∈{0,1} τ(L-τ+1)×L It is a repeating matrix that satisfies:

[0118] vec(H τ (v))=δx (4)

[0119] Whether it is residential electricity consumption data in low-voltage distribution areas or electricity consumption data of industrial and commercial users on other medium and high voltage sides, the electricity consumption data of a single user has multiple periods and temporal and spatial correlations; the electricity consumption data of users included in the load aggregator in the electricity market has the characteristics of periodic changes, consistent collection frequency, and multiple participants in the electricity market transaction.

[0120] Stacking the Hankel matrices established above along the third dimension forms a Hankel tensor.

[0121] Step 5: Based on the trend and periodic data obtained in Step 2, establish time-series smoothing constraints and periodic constraints;

[0122] The specific method for step 5 is as follows:

[0123] Using the trend and periodic data obtained from step 2, establish time-series smoothing constraints: Constraints on periodic terms:

[0124] Step 6: For the Hankel form tensor of the trend term data obtained in Step 4, tensor parallel factor decomposition is used to obtain the tensor parallel factor decomposition model. Combined with the two constraints in Step 5, a comprehensive tensor completion objective function is established, which is applicable to different types of power users.

[0125] The objective function and constraints for the comprehensive tensor completion in step 6 are as follows:

[0126]

[0127] Among them: E to The value represents all values ​​within the period containing the missing element, where m is the number of delays, T is the period, and w is the weighting coefficient (w t-mT +…+w t+mT =1), λ e It is the regularization coefficient. This represents the complete Hankel output tensor. Use 0 and 1 to represent missing and available elements, The Kronecker product, in Hankel form, is represented by a tensor for missing terms. To fill:

[0128]

[0129] in: ρ=[ρ (1) ,ρ (2) ,...,ρ (N) ] T It is a smoothness parameter vector, L1∈R (In-1)×In , where In∈R, is the number of elements in the factor vector. The first derivative operator is defined as:

[0130]

[0131] In the established objective function:

[0132] Represents observation terms and reconstruction tensor The root mean square error (RMSE) of the values ​​is used to estimate the difference between the two, where

[0133] Time-series smoothing regularization term: After decomposing the electricity consumption data of electricity users participating in spot trading in the first part, the smoothing trend part is... Furthermore, by incorporating the special structure of the Hankel tensor and adding temporal smoothness constraints (i.e., the values ​​of adjacent elements are close), this invention employs a quadratic variation (QV) constraint. This regularization term is a formal expression of the temporal smoothness constraint, defined as:

[0134]

[0135] (3) This indicates a strong correlation between the periodicity of the data. However, due to the irregular noise in the electricity consumption data and the fact that the data does not strictly follow the periodic terms after the autocorrelation coefficient calculation, a suitable set of parameters W is found based on the periodic values ​​of each non-missing value in the data to fit the missing values ​​in the periodic terms.

[0136] R is the rank of the parallel factorization of the trend term. The choice of R needs to balance accuracy and computational complexity. Too low an R value can reduce computational complexity, but will produce a large truncation error, affecting model accuracy; too high an R value will increase computational complexity.

[0137] This invention optimizes R by conducting truncation error experiments. The specific steps are as follows:

[0138] (1) The upper limit of R max Related to the size of the tensor; For example, R max The values ​​are shown in equation (9):

[0139] R max ≤min(I×J,I×K,J×K) (9)

[0140] In the formula, I, J, and K represent the rows, columns, and number of tensors of the tensor, respectively.

[0141] (2) Use equation (10) to determine the cutoff error Q of parallel factorization. CP :

[0142]

[0143] In the formula: Let this represent the non-missing index tensor reconstructed from the r groups of CP decomposition components. The truncation error Q is selected. CPThe rank corresponding to the moment when it gradually decreases and becomes relatively stable within a certain range is the rank of the tensor parallel factorization in this paper.

[0144] Step 7: Solve the comprehensive tensor completion objective function obtained in Step 6 using Alternating Least Squares (HALS) and output the complete data using inverse MDT technique;

[0145] The specific method for step 7 is as follows:

[0146] The Alternating Least Squares (HALS) method is used to solve the optimization problem. According to the HALS algorithm, consider minimizing the following local objective function:

[0147]

[0148] in:

[0149] in:

[0150] Optimize the above equation in the following order, updating the parameter matrix W and the factor vector. Adaptive factor g r ,:

[0151]

[0152] in: It is a subset of unit vectors;

[0153] (1) Update W: Optimize the weight coefficients W to simplify calculations and achieve the goal of optimizing parameters. p (W) simplifies to:

[0154]

[0155] The parameters are updated according to the following rules:

[0156]

[0157] In the formula: η is the learning rate. Is the function about W k The gradient.

[0158] Find a suitable set of weighting coefficients, based on E t The expression calculates the missing values ​​for the periodic terms. The completed periodic sequence is then transformed into a Hankel tensor using a delayed embedding transformation.

[0159] (2) Update u: To prevent non-differentiable cases during optimization, this invention patent selects a quadratic variational constraint (QV) and adopts a gradient-normalized update form. To simplify the notation, u is used. k Indicate u r (n) kth update, v k The gradient of the objective function

[0160]

[0161] The above equation assumes that the gradient is updated on the hypersphere. To facilitate the representation of the gradient of the objective function, It can be simplified to:

[0162]

[0163] in:

[0164]

[0165] The gradient of the objective function is then:

[0166]

[0167] (3) Update g r The problem of updating the adaptive factor is an unconstrained quadratic optimization problem, and the unique solution can be obtained analytically. The objective function can be simplified to:

[0168]

[0169] So g r The update rules are as follows:

[0170]

[0171] The above three steps are iteratively solved using HALS until the error between two updates of each variable is less than a set threshold ε, typically set to 0.001. The completed Hankel tensor is then obtained. Output complete data through inverse delay transform:

[0172]

[0173] in This is the Moore–Penrose pseudo-inverse transform, unfold (τ,L-τ+1) fold (τ,L-τ+1) The inverse operation.

[0174] Step 8: Verify the effectiveness of the data completion method proposed in this invention using evaluation metrics such as relative recovery error (RRE) and mean absolute percentage error (MAPE).

[0175] The specific method for step 8 is as follows:

[0176] The model recovery performance was evaluated using relative recovery error, mean absolute percentage error, symmetric mean absolute percentage error, and root mean square error, respectively, for different missing types and different missing proportions.

[0177]

[0178] In the formula: Complete the missing electricity consumption data for electricity users; This provides the actual power quality data for the missing areas. Figure 4 Provide a graphical representation of the completed user's electricity consumption data. Figure 5 The diagram shows the effect of simultaneously completing the information for three electricity users (load aggregators), demonstrating the effectiveness of the text method for different user types.

[0179] Table 1 shows the RRE results of the method of the present invention for missing value interpolation, tensor CP decomposition completion, low-rank matrix completion, and support vector machine regression.

[0180] Table 1. Comparison of RRE methods (data from a single industrial user)

[0181]

[0182]

[0183] Table 2 shows the MAPE results of the method of the present invention compared with tensor CP decomposition completion, low-rank matrix completion, and support vector machine regression for missing value interpolation.

[0184] Table 2 Comparison of different methods for MAPE (resident user dataset)

[0185]

[0186] Table 3 shows the SMAPE results of the method of the present invention for missing value interpolation, along with tensor CP decomposition completion, low-rank matrix completion, and support vector machine regression.

[0187] Table 3 Comparison of different SMAPE methods (resident user dataset)

[0188]

[0189] Table 4 shows the RMSE results of the method of the present invention for missing value interpolation compared with tensor CP decomposition completion, low-rank matrix completion, and support vector machine regression.

[0190] Table 4. Comparison of RMSE using different methods (resident user dataset)

[0191]

[0192] As can be seen from the experimental results in Tables 1 to 4, the method proposed in this invention has a good completion effect.

[0193] It should be emphasized that the embodiments described in this invention are illustrative rather than limiting. Therefore, this invention includes, but is not limited to, the embodiments described in the specific implementation. Any other implementations derived by those skilled in the art based on the technical solutions of this invention are also within the scope of protection of this invention.

Claims

1. A method for data completion suitable for multi-type mass power users, characterized in that: It comprises the following steps: Step 1, collecting power consumption data of multiple types of power users, and extracting any single user power consumption data; Step 2, based on the single user power consumption data extracted in step 1, decomposing the single user power consumption data into the sum of trend items and periodic items based on the time series power data decomposition method; Step 3, using autocorrelation calculation and fast Fourier calculation to verify the periodic item data obtained in step 2; Step 4, using delay embedding transformation technology to perform standard delay transformation on the single user power consumption data collected in step 1 and the periodic item and trend item time series vectors decomposed in step 2, to obtain a Hankel form tensor; Step 5, based on the trend item and periodic item data obtained in step 2, establishing time series smoothing constraints and periodic item constraints; Step 6, for the trend item data Hankel form tensor obtained in step 4, using tensor parallel factor decomposition to obtain a tensor parallel factor decomposition model and combining the two constraints in step 5 to establish a comprehensive tensor completion objective function suitable for different types of power users; Step 7, using alternating least squares (HALS) to solve the comprehensive tensor completion objective function obtained in step 6 and using reverse MDT technology to output complete data; The specific steps of step 2 include: Decompose the user power consumption data into the weighted sum of trend items and periodic items through moving average; Step 1, additively decompose the user power consumption data extracted in step 1: where: a = 2N + 1 is the order of smoothing, x t is the power user consumption data, s t is the trend term and e t is the periodic term; The specific method of step 5 is: The trend and cycle data obtained by step 2 are decomposed to establish time series smoothing constraints: and cycle constraints: The comprehensive tensor completion objective function and constraint conditions of step 6 are: where: E to is the value of all elements in the period where the missing element is located, m is the number of time delays, T is the period, w is the weight coefficient (w t-mT +…+w t+mT =1), λ e is the regularization coefficient, H(X) represents the complete Hankel output tensor, 0 and 1 represent missing and available elements, denotes the Kronecker product, and the missing term in the Hankel form is filled with the tensor H(Z): where: p = [p (1) , p (2) ,..., p (N) ] T is a smoothness parameter vector, L1 e R (In-1)×In , In e R, is the number of elements in the factor vector First derivative operator is defined as: In the established objective function: The root mean square error (RMSE) of the values of the observed term H(X) and the reconstructed tensor H(Z) is indicative of the difference between the two, where H(Z) == H(S) + H(E); The timing smoothness regular term: after decomposing the power consumption data of the power users participating in the spot transaction in the first part, the smooth trend part And combined with the special structure of the Hankel tensor, the timing smoothness constraint is added, that is, the values between adjacent elements are close, and the application adopts quadratic variation (QV) constraint, which is a formal expression of the timing smoothness constraint, and is defined as: (3) The periodic strong correlation of the data is represented, but because of the irregular noise of the power consumption data, and after the autocorrelation coefficient calculation, the data is not strictly changed according to the periodic term, so a set of suitable parameters W is found according to the periodic numerical value of each non-empty value of the data, which is used to fit the numerical value of the empty value in the periodic term. R is the rank of the trend item parallel factor decomposition, and the selection of R needs to fully balance the accuracy and computational complexity; too low R value can reduce the computational complexity, but will produce large truncation error, affecting the model accuracy; too high R value will increase the computational complexity; Optimize R through truncation error experiment, the specific steps are as follows: (1) the upper limit R of R max related to the size of the tensor; H(S) e I×J×K , R max takes the value as shown in equation (9): R max ≤ min(I x J, I x K, J x K) (9) Where I, J, and K represent the rows, columns, and number of tensors, respectively; (2) using formula (10) to determine the truncation error Q of the parallel factorization CP : where H(S') is the Hessian of S' at S; and Ω is the reconstructed tensor from the r CP-decomposition components; the truncation error Q is selected CP The rank size corresponding to the moment when the error gradually decreases and becomes relatively stable in a certain range is the rank of the tensor parallel factorization in this paper.

2. The method according to claim 1, wherein the method is suitable for multi-type mass power user data completion. The specific method of step 3 is: Verify the periodic item data using autocorrelation calculation and fast Fourier calculation: where: e t is the mean of the decomposed period term, and T is the order of the lag.

3. The method of claim 1, wherein the method is suitable for multiple types of mass power user data completion. The specific method of step 4 is: For single-user electricity consumption data: assuming a day as a basic cycle, select n days of data, the data length is L, and L x n vectors can be obtained, where the electricity of the i(i=1,…,n) day can be expressed as x i ∈x L×1 ; Then use delay embedding transformation technology to perform standard delay transformation on the time series vector, that is, Hankelization, to obtain a Hankel form matrix; where L is the data vector length, τ is the delay window size, fold (τ,L-τ+1) :R τ(L-τ+1) →R τ×(L-τ+1) is a folding operator, δ ∈ {0, 1} τ(L-τ+1)×L is a repetition matrix satisfying: vec(H τ (v)) = δx (4) Stack the above established Hankel matrix in the third dimension to form a Hankel tensor.

4. The method of claim 1, wherein the method is suitable for multiple types of mass power user data completion. The specific method of step 7 is: Use alternating least squares (HALS) to solve the optimization problem, according to the HALS algorithm, consider minimizing the following local objective function: wherein: wherein: The parameter matrix W is updated in the following order, optimizing the above equation, factor vector Adaptive factor g r , : wherein: is a subset of unit vectors; (1) Update W: Optimization is performed on the weight coefficient W, and the objective function F p (W) is simplified as: Update the parameters, the update rule is as follows: where: η is the learning rate, is the gradient of the function with respect to W k ; find a set of suitable weight coefficients, according to E t The expression calculates the missing values of the periodic term; the completed periodic sequence is delayed embedded into a Hankel tensor H(E); (2) Update u: In order to prevent the occurrence of non-differentiable conditions in the optimization process, the patent selects the quadratic variation constraint (QV), and adopts the gradient normalization update form, and in order to simplify the symbol, u k represents u r (n) The kth update, v k represents the gradient of the objective function The above formula considers updating the gradient on a hypersphere. In order to facilitate the representation of the gradient of the objective function, which can be simplified as: Where: The gradient of the objective function is: (3) Update g r The update problem for the adaptive factor is a unconstrained quadratic optimization problem, and the unique solution can be obtained analytically, with the objective function simplifying to: So g r The update rule for g is: Use HALS to alternately iterate the above 3 steps until the error of each variable calculated by twice updating is less than the set threshold ε, which is 0.001; finally obtain the completed Hankel tensor H(X') == H(S')+H(E'), and output the complete data through delay inverse transformation: wherein is the Moore-Penrose pseudo-inverse transform, unfold ( is the inverse operation of fold (τ,L-τ+1) .

Citation Information

Patent Citations

  • Ubiquitous power Internet of Things perception data missing restoration method based on matrix filling

    CN110705762A

  • User power consumption behavior analysis method considering missing data supplementation

    CN114048200A