A power user power consumption behavior profiling method considering power consumption sensitivity

By combining weighted Euclidean distance and dynamic time curvature distance with the improved CRITIC method, the electricity consumption behavior of power users is profiled, and the base load and sensitive load are decomposed. This solves the problem of insufficient analysis in the existing technology, and achieves accurate characterization of user electricity consumption behavior and improvement of service level.

CN120470430BActive Publication Date: 2026-08-04HARBIN INST OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HARBIN INST OF TECH
Filing Date
2025-04-17
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing methods for profiling electricity users fail to effectively distinguish between base load and sensitive load, resulting in insufficiently detailed analysis of electricity consumption behavior and an inability to accurately grasp the characteristics of user electricity consumption.

Method used

An improved CRITIC method combining weighted Euclidean distance and dynamic time curvature distance was used to perform preliminary clustering of load curves, decomposing them into basic loads and sensitive loads, and performing cluster analysis separately to generate accurate profiles.

Benefits of technology

It enables detailed analysis of users' electricity consumption behavior, accurately grasps users' basic electricity consumption patterns and sensitivity to external factors, and improves the service level of power companies and the data support for the formulation of demand response strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470430B_ABST
    Figure CN120470430B_ABST
Patent Text Reader

Abstract

The application provides a power user power consumption behavior portrait method considering power consumption sensitivity, and belongs to the technical field of industrial data processing. The method solves the problem that in most power user portrait methods, sensitive loads and basic loads influenced by other factors are mixed for analysis, which makes the portrait result not fine enough and unable to accurately grasp the power consumption behavior characteristics of users. The method comprises the following steps: obtaining a power load data set, drawing a load curve, and performing preliminary clustering; labeling the preliminary clustering cluster; inputting the preliminary clustered load data set into an STL model for decomposition and division into basic loads and sensitive loads; performing clustering analysis on the basic loads and the sensitive loads respectively, defining a basic load label library and an external factor sensitive load label library according to the clustering results; and labeling each user with two types of labels according to the basic load label library and the external factor sensitive load label library, combining the labels of the initial clustering results, and generating a precise portrait of the power user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for profiling the electricity consumption behavior of power users that takes into account their sensitivity to electricity consumption, and belongs to the field of industrial data processing technology. Background Technology

[0002] With the continuous improvement of the national economic development level, the demand for electricity in society is increasing, and the electricity consumption behavior of different industries is showing diverse and complex characteristics. At the same time, the widespread adoption of various intelligent metering terminals on the user side has brought massive amounts of user-side data to the power system. How to effectively utilize this data and extract valuable information has become an important research issue. User profiling, a data analysis tool that uses user data to define labels to characterize users and develop precise marketing strategies for target users, can help power companies understand users' electricity consumption behavior and provide convenient conditions for intelligent management of power supply companies. In the context of power market reform, research on electricity user behavior profiling technology can help power companies understand user behavior, improve marketing capabilities, and also serve as a basis for future real-time electricity pricing.

[0003] Against the backdrop of rapid development of smart grids, the application of machine learning technology in power grids is gradually increasing, and research on user profiling in the power system field is also deepening. Power systems can uncover the demand characteristics of electricity users by profiling their behavior, thereby implementing differentiated marketing strategies and improving the service level of the power system. Currently, user profiling methods mainly focus on three aspects: improving the user feature system, enhancing the quality of feature sets, and improving the efficiency of clustering algorithms. To achieve accurate profiling of electricity user behavior, existing research has proposed various analysis methods based on electricity consumption characteristics. Among them, some studies use the maximum correlation minimum redundancy criterion combined with the k-means clustering algorithm to extract user electricity consumption characteristics and construct behavioral profiling models. Other scholars use the fuzzy C-means clustering algorithm to perform clustering analysis on industry electricity consumption data, forming a refined electricity consumption labeling system covering four dimensions by deconstructing load curve characteristics. In residential user research, some studies have established a load characteristic labeling system based on big data platforms, combining typical daily load curves of different seasons to analyze user load volatility and demand response potential, forming a variable time-scale profiling model that reflects the temporal patterns and elastic characteristics of electricity consumption. Furthermore, to address the need for user behavior tracking, some scholars have proposed profiling methods based on fine-grained data from non-in-home terminals and an improved k-means algorithm. However, these methods still rely on expert experience for support during the label definition process. Overall, existing methods mainly revolve around clustering algorithms, constructing dynamic and scenario-based user profile systems through multi-dimensional electricity consumption characteristic analysis.

[0004] Current research on electricity user profiling focuses excessively on constructing user characteristics and classifying user types, while research on user electricity consumption behavior remains insufficient. Many studies profile the overall electricity consumption behavior characteristics of users across different industries, but few decompose user loads to separate base loads and sensitive loads, and then analyze the basic electricity consumption attributes and sensitivity to different influences in different seasons. Because user-side electricity consumption characteristics under different seasons are the result of the combined effects of the user group's inherent electricity consumption characteristics and those influenced by external environmental factors, analyzing sensitive loads influenced by other factors together with base loads results in imprecise user profiling and an inability to accurately grasp the characteristics of user electricity consumption behavior.

[0005] In the clustering of electricity load curves, researchers have proposed various methods to improve clustering performance and reduce reliance on the selection of initial cluster centers. One method is to obtain initial cluster centers based on density peaks, which can more accurately identify natural groupings in the data. Furthermore, some studies have introduced the concepts of inter-cluster optimization and intra-cluster optimization, using local convergence and global diffusion to not only reduce the algorithm's requirements for initial cluster center selection but also improve clustering quality. To further improve clustering results, some studies have used characteristic indicators such as daily load factor and daily peak-to-valley difference rate for data dimensionality reduction, thereby improving clustering quality. Other works utilize kernel functions to map load curve data to a high-dimensional space, enhancing data separability and making clustering more effective. Some studies combine deep learning techniques for data dimensionality reduction and then use the fuzzy C-means algorithm for clustering; this method can capture more complex patterns and improve clustering performance. Additionally, some scholars have proposed using quantile and difference algorithms to evaluate the dynamic characteristics of load curves and assign different feature values ​​based on the rate of change; this method helps to better understand and classify different types of load curves. A review of current research reveals that most studies focus on initial cluster center selection and dimensionality processing of daily load curve data, with few in-depth investigations into similarity measurement methods. The mainstream similarity measurement method—Euclidean distance—only considers the numerical distribution characteristics of daily load curves at corresponding time points, failing to reflect dynamic characteristics such as curve shape and slope changes. This leads to significant deviations in measurement accuracy under extreme climbing conditions and is equally affected by all local differences, resulting in a lack of key distinguishing features between loads. Furthermore, the time intervals between load points on load curves are increasingly shorter, diminishing the significance of simply calculating the Euclidean distance between corresponding load points. Dynamic curvature distance has also been increasingly used in similarity measurement in recent years, but it emphasizes the dynamic characteristics of the curve, to some extent ignoring the impact of the curve's numerical distribution characteristics on similarity measurement, and suffers from measurement bias due to excessive curvature. Summary of the Invention

[0006] This invention addresses the problem that most electricity user profiling methods mix sensitive loads affected by other factors with base loads for analysis, resulting in insufficiently detailed profiling results and an inability to accurately grasp the characteristics of users' electricity consumption behavior. Therefore, this invention proposes an electricity user behavior profiling method that considers electricity consumption sensitivity.

[0007] The technical solution adopted by the present invention to solve the above problems is as follows: The present invention includes the following steps:

[0008] Step 1: Obtain the power load dataset and plot the load curve;

[0009] Step 2: For any two given time series X = [x1, x2, ..., x] in the load curve n ] and Y = [y1, y2, ..., y n The weighted Euclidean distance, load change trend and DTW distance of the sampling points in the load curve are calculated. The weights are selected based on the improved CRITIC method. The load curve is initially clustered and the initial clusters are labeled.

[0010] Step 3: Input the initially clustered load dataset into the STL model for decomposition, and divide the decomposed load data into basic load and sensitive load;

[0011] Step 4: Perform cluster analysis on the basic load and sensitive load respectively, and define the label library for the basic load and the label library for the external factor sensitive load according to the clustering results;

[0012] Step 5: Based on the label library of basic load and the label library of loads sensitive to external factors, assign two types of labels to each user. Combine the labels from the initial clustering results to generate an accurate profile of the power user.

[0013] Furthermore, step 2 specifically includes:

[0014] Step 2.1: Based on the two time series X = [x1, x2, ..., x] given in the load curve... n ] and Y = [y1, y2, ..., y n ] Calculate the standard deviation σ(X) of the load value in the i-th time period. i ) and average value μ(X) i ), based on standard deviation σ(X) i ) and average value μ(X) i The load fluctuation index w was calculated. i Based on the load volatility index w i The normalized weight w' of the load value in the i-th time period is calculated. i This allows us to obtain the weighted Euclidean distance between the same sampling point in curves X and Y;

[0015] Step 2.2: Calculate the time series X = [x1, x2, ..., x] using the DTW algorithm. n ] and Y = [y1, y2, ..., y n DTW distance;

[0016] Step 2.3: For discrete load curves X = [x1, x2, ..., x n Transform it into a morphological sequence X' = [x'1, x'2, ..., x'] of length n-1. n-1To reflect the trend information of the load curve in each time period and describe the local dynamic characteristics of the curve, the similarity of the curve is described by weighted Euclidean distance, load change trend and DTW distance;

[0017] Step 2.4: Perform preliminary clustering on the curves after calculating similarity, output the best preliminary clustering results, and label the preliminary clusters;

[0018] Load fluctuation index w i The calculation formula is:

[0019]

[0020] The load value normalized weight w' for the i-th time period i The calculation formula is:

[0021]

[0022] The expression for weighted Euclidean distance is:

[0023]

[0024] Further, step 2.2.1: Based on the time series X = [x1, x2, ..., x...] n ] and Y = [y1, y2, ..., y n Construct an n×m distance matrix D∈R n×m ;

[0025] Step 2.2.2: Define the set of each pair of adjacent elements in matrix D as the curved path P, where P = {p1, p2, ..., p...} s ,…,p k}, where k is the total number of elements in the path, and element p s Let p be the coordinates of the s-th point on the path. s = (i,j);

[0026] Step 2.2.3: Set two constraints for the curved path. Constraint one is that the curved path must start from the lower left corner p1 = (1,1) and end at the upper right corner p1. k = (n,m) ends, constraint two is that each point in the curved path must match its adjacent point, that is, if p s = (i,j), then p s+1 = (a, b) must satisfy 0 ≤ ai ≤ 1, 0 ≤ bj ≤ 1

[0027] Step 2.2.4: Based on the constraints in Step 2.2.3, add constraints on the number of consecutive bends;

[0028] Step 2.2.5: After completing the bending path constraint, calculate the bending path P that minimizes the total bending cost of time series X and Y by constructing the cumulative cost matrix L, and then obtain the DTW distance D of time series X and Y. dtw (X,Y)=L n,m ;

[0029] Distance matrix D∈R n×m Element D i,j The calculation formula is:

[0030]

[0031] The expression for the constraint on the number of consecutive bends is:

[0032]

[0033] In formula (5), r x and r y These represent the number of consecutive bends of the path along the x-axis and y-axis, respectively; r x,max and r y,max These represent the maximum number of consecutive bends allowed on the x-axis and y-axis, respectively.

[0034] The formula for calculating the bending path P with the minimum total cost is:

[0035]

[0036] In formula (6), D(p) s () represents the cumulative distance of the curved path;

[0037] The expression for the cumulative cost matrix L is:

[0038] L i,j =D i,j +min(L i-1,j-1 ,L i,j-1 ,L i-1,j (7);

[0039] In formula (7), i = 1, 2, ..., n; j = 1, 2, ..., m, L i,j Let L be the element in the i-th row and j-th column of matrix L, where L 0,0 =0,L i,0 =L 0,j =+∞.

[0040] Furthermore, step 2.3 specifically includes:

[0041] Step 2.3: For discrete load curves X = [x1, x2, ..., x n Transform it into a morphological sequence X' = [x'1, x'2, ..., x'] of length n-1.n-1 To reflect the trend information of the load curve in each time period and describe the local dynamic characteristics of the curve, the similarity of the curve is described by weighted Euclidean distance, load change trend and DTW distance;

[0042] Step 2.3.2: Using weighted Euclidean distance and DTW distance, combined with the transformed curves Y' and X', calculate the similarity between load curves X and Y;

[0043] The expression for element xi′ in sequence X' is:

[0044]

[0045] In formula (8), Δt is the time interval;

[0046] The formula for calculating the similarity between load curves X and Y is:

[0047]

[0048] In formula (9), D2'(X,Y) is the weighted Euclidean distance between curves X and Y, used to measure the similarity of the numerical distributions of the curves at corresponding time points, that is, the similarity of the overall distribution characteristics of the curves. dtw (X,Y) represents the DTW distance between curves X and Y, used to measure the overall shape similarity of the curves, i.e., the overall dynamic similarity of the curves. dtw (X',Y') represents the DTW distance corresponding to the local dynamic characteristics of curves X and Y described by formula (8); α, β, and γ are the weights, and D all The smaller the (X,Y) value, the higher the similarity between the two curves.

[0049] Furthermore, step 2.4 specifically includes:

[0050] Step 2.4.1: Preprocess the data in the loading curve after similarity calculation, setting the initial number of clusters to L = L min ;

[0051] Step 2.4.2: Set the initial cluster center curve;

[0052] Step 2.4.3: Configure similarity weights based on the improved CRITIC method;

[0053] Step 2.4.4: Combine weighted Euclidean distance, DTW distance, and K-means algorithm for aggregation;

[0054] Step 2.4.5: Calculate the clustering quality index I DBI ;

[0055] Step 2.4.6: Determine if L is equal to Lmax If not equal, let L = L + 1, and repeat steps 2.4.4-2.4.5 until L = L. max ;

[0056] Step 2.4.7: Select the L corresponding to the minimum DBI value of the clustering index and output the corresponding best clustering result;

[0057] The expression for configuring similarity weights is:

[0058]

[0059]

[0060] In formulas (10) and (11), σ j Let r be the standard deviation of the j-th indicator. ij Let C be the correlation coefficient between the i-th indicator and the j-th indicator. j Let W represent the total amount of information represented by the j-th evaluation indicator. j To determine the objective weight of the j-th evaluation index obtained through normalization, The average value of the j-th indicator;

[0061] Clustering quality index I DBI The expression is:

[0062]

[0063] In formulas (12) and (13), S w and R w All are parameters, R w S is used to measure the similarity between the center data points of class w and class q; w Used to measure the clustering of data points in the w-th class; M wq This represents the degree of dispersion between the data points of the w-th class center and the q-th class center.

[0064] Furthermore, step 3 specifically includes:

[0065] The load dataset after preliminary clustering is input into the STL model. The time-series load data is decomposed using the addition principle to obtain the trend component value, periodic component value, and residual component value at the corresponding time. The decomposed trend component and periodic component are used as the basic load, and the residual component is used as the sensitive load.

[0066] The expression for decomposing time-series load data is:

[0067] Y t =T t +S t +I t (14);

[0068] In formula (18), Y t Let T be the observation value at time t; t S t I t These represent the trend component value, periodic component value, and residual component value at time t, respectively.

[0069] The expression for the base load is:

[0070] Z t =T t +S t (15);

[0071] In formula (15), Z t It is the basic load component.

[0072] Furthermore, step 4 specifically includes:

[0073] Step 4.1: Determine the number of clusters k for the basic load using the clustering effectiveness index control method, thereby obtaining the number of basic load types, and perform cluster analysis on the basic load data. Define electricity consumption behavior labels and assign scores based on the clustering results.

[0074] Step 4.2: Perform correlation analysis on sensitive loads and extract the correlation coefficients of various factors on each type of load.

[0075] Using the same clustering effectiveness index as in step 4.1, the number of clusters k affected by other factors was determined, and cluster analysis was performed to obtain different types of influence, and labels were defined for each.

[0076] Furthermore, steps 4.1 and 4.2, which use the clustering validity index control method to determine the number of clusters k and perform cluster analysis, include:

[0077] The sum of squared errors is used as the evaluation index for the number of clusters, and the true number of clusters is set as k. * When the number of clusters k <k * When k increases, SSE decreases rapidly; when k ≥ k * As k increases, the optimal number of clusters k is determined based on the changes in SSE, and clustering is performed.

[0078] Assuming that both the basic load data and the sensitive load data follow a Gaussian distribution, the probability density functions of k classes of Gaussian distributions and the weight α of each class are estimated using the training data of the Gaussian distribution probability model. k Then, substitute each data point into k Gaussian distributions to calculate the probability P(y) of belonging to each class. i ) k The data is then grouped into the category with the highest probability value to complete the cluster analysis.

[0079] The expression for the sum of squared errors is:

[0080]

[0081] In formula (16), k is the number of clusters, and C i For a cluster in the clustering, u i C i The mean of all data points in the cluster, where x is a point in the cluster;

[0082] The probability model of the Gaussian distribution is:

[0083]

[0084] In formula (17), y represents the sample; α represents the sample. k As the weight, φ(y|θ) k Let θ be the probability density function of a Gaussian distribution. k The parameters of the probability density include μ. k and

[0085] The expression for the probability density function is:

[0086]

[0087] In formula (18), σ k Let μ be the standard deviation of sample y. k Let y be the mean of the sample y;

[0088] Probability P(y) i ) k The calculation formula is:

[0089]

[0090] In formula (19), y i is the corresponding i-th data; k is the k-th Gaussian distribution.

[0091] Furthermore, step 4.1, which defines electricity consumption behavior labels and assigns scores based on the basic load clustering results, specifically includes:

[0092] The electricity consumption behavior characteristics of electricity users are described by calculating the load factor k1, peak-valley difference k2, and peak-to-average power ratio k3. The scoring system is set with a full score of 10 points, and the electricity consumption behavior characteristics of all users are scored.

[0093] The formula for calculating the load factor k1 is:

[0094]

[0095] The formula for calculating the peak-to-valley difference rate k2 is:

[0096]

[0097] The formula for calculating the peak-to-average power ratio k3 is:

[0098]

[0099] The expression for scoring applied electrical behavior features is:

[0100]

[0101] In formula (23), k i For feature categories, C j For the j-th type of electricity consumption behavior, For all j-th type of electricity consumption behavior, k i The average value of the feature k for all users i The minimum value of the feature. k for all users i The maximum value of the feature.

[0102] Furthermore, step 4.2, which involves correlation analysis of sensitive loads, specifically includes:

[0103] The mesh partitioning method is used to calculate the inter-variable correlation (MI) between random variables in the sensitive load dataset. Based on the MI, the MIC value between two variables in the sensitive load dataset is calculated. The closer the MIC value is to 1, the stronger the correlation between the two variables.

[0104] The formula for calculating the inter-variable randomness (MI) is:

[0105]

[0106] In formula (24), X and Y are random variables in the sensitive load dataset, p(x,y) is the joint probability density between X and Y, and p(x) and p(y) are the marginal probability densities between X and Y, respectively.

[0107]

[0108] In formula (25), B is the sample size, N is the sample variable, and I(x,y) is the MI value between variables x and y.

[0109] The beneficial effects of this invention are:

[0110] 1. This invention proposes a load curve clustering method based on weighted Euclidean distance, dynamic time-curvature distance, and an improved CRITIC method. This method combines DTW distance and weighted Euclidean distance to comprehensively determine curve similarity, and uses this comprehensive similarity as the basis for clustering. It considers three similarity characteristics of daily load curves: overall distribution characteristics, local dynamic characteristics, and overall dynamic characteristics, and fully takes into account the different criticality levels at different times. The improved CRITIC method is used to weight the similarity of these curve characteristics, and the k-means algorithm is applied to conduct preliminary clustering analysis on the daily load curves of power grid users. This effectively avoids misclassification of curves during clustering, resulting in high-quality clustering and providing more accurate cluster center curves for the analysis of load curves from multiple users.

[0111] 2. This invention considers simultaneously analyzing the basic electricity consumption behavior patterns of users as a whole and the sensitivity of users to other factors. It proposes to use the STL algorithm to decompose the original load data into basic load and load affected by other factors. Then, these two types of loads are clustered and analyzed to form a label library for the two types of loads. Finally, the two types of labels are combined with the initial clustering labels to form an accurate profile of the electricity user.

[0112] 3. The method proposed in this invention, which combines direct clustering analysis of raw load data with separate clustering analysis of load decomposition to create user profiles, enables more refined and comprehensive user behavior analysis. It retains the original information while improving the accuracy of the profiles, helping power companies to accurately grasp user behavior, improve their service levels, and provide data support for the formulation of future demand response strategies. It has significant application value. Attached Figure Description

[0113] Figure 1 A flowchart illustrating a method for profiling electricity user behavior that considers electricity sensitivity, provided by the present invention;

[0114] Figure 2 A flowchart illustrating the preliminary clustering algorithm provided by this invention;

[0115] Figure 3 A flowchart illustrating the electricity user behavior profile provided by this invention;

[0116] Figure 4 This is a schematic diagram of the DTW sequence provided by the present invention. Detailed Implementation

[0117] Combination Figure 1-4 This implementation method is described as follows: Figure 1 As shown, the steps of the method for profiling electricity user behavior considering electricity sensitivity according to this embodiment include:

[0118] S1: Obtain the power load dataset and plot the load curve;

[0119] S2: Perform preliminary clustering of the load curves and label the preliminary clusters;

[0120] For the corresponding load curves, this implementation uses the following three characteristics to describe the similarity between load curves: ① The distance between the same sampling point of each curve in a cluster reflects the overall similarity of the curves, which is called the "overall distribution characteristic"; ② The shape characteristics of the curves throughout the entire sampling period reflect the overall trend of the curves, which is called the "overall dynamic characteristic"; ③ The degree of drastic change of the curves at each sampling time point (or sampling time interval) is called the "local dynamic characteristic". The preliminary clustering steps include:

[0121] S201: Calculate the weighted Euclidean distance;

[0122] Given two time series X = [x1, x2, ..., x...] n ] and Y = [y1, y2, ..., y n If we use the Minkowski distance as a measure, its expression is:

[0123]

[0124] In formula (1), W is the distance coefficient of the Minkowski distance. Depending on its value, it can represent different distance measurement methods. When W = 2, D W This is the Euclidean distance.

[0125] Traditional clustering methods fail to differentiate between load variations at different time points, thus not fully reflecting the characteristics of different time periods and resulting in inaccurate clustering results. To address the shortcomings of traditional load clustering, this implementation method uses weighted Euclidean distance as an alternative when comparing similarity. Weights are assigned to the n-dimensional data based on the importance of different time periods, strengthening the influence of key time periods in clustering. The formula for weighted Euclidean distance is:

[0126]

[0127] Load fluctuation index w i The calculation formula is:

[0128]

[0129] The load value normalized weight w' for the i-th time period i The calculation formula is:

[0130]

[0131] In formulas (2)-(4), wi w' is an indicator of load volatility. i The normalized weight for the load value in the i-th time period is σ(X). i ) represents the standard deviation of the load value in the i-th time period, μ(X) i ) represents the average load value in the i-th time period.

[0132] S202: Calculate DTW distance;

[0133] S20201: Obtain the DTW path;

[0134] Unlike Euclidean distance, which strictly calculates distance based on the values ​​of corresponding time points in two time series, DTW employs dynamic programming. It adjusts the relationships between corresponding elements at different time points in the time series to obtain an optimal curved path that minimizes the distance between time series along that path. This algorithm effectively measures the overall shape similarity between time series. For example... Figure 4 The diagram shows two time series X and Y. The points connected by the black line between the two time series are points of similarity between them. When n points in one time series correspond to one point in another time series, it indicates that the time series exhibits curvature. The sum of the distances between all these similar points is the DTW distance. The DTW algorithm uses this as a basis to measure the similarity between time series.

[0135] S20202: Calculate the time series X = [x1, x2, ..., x] using the DTW algorithm. n ] and Y = [y1, y2, ..., y n DTW distance;

[0136] For any two given time series X = [x1, x2, ..., x...] n ] and Y = [y1, y2, ..., y m Construct an n×m distance matrix D∈R n×m , where element D i,j As shown in the following formula:

[0137]

[0138] Formula (5) represents two time series points x i and y j The Euclidean distance. The set of each pair of adjacent elements in matrix D is called the "curved path," denoted as P = {p1, p2, ..., p...}. s ,…,p k}, where k is the total number of elements in the path, and element p s Let p be the coordinates of the s-th point on the path. s= (i,j). Meanwhile, the curved path must also satisfy the following constraints: ① The selected path must start from the lower left corner and end at the upper right corner, i.e., p1 = (1,1), p k = (n, m); ② Each point must match its adjacent points, that is, if p s = (i,j), then p s+1 = (a, b) must satisfy 0 ≤ ai ≤ 1, 0 ≤ bj ≤ 1. Furthermore, to avoid excessive bending caused by multiple consecutive bends in the same horizontal or vertical direction (i.e., to avoid one point in one time series corresponding to too many points in another time series), a constraint on the number of consecutive bends is added to the existing constraints, namely:

[0139]

[0140] In formula (6), r x and r y These represent the number of consecutive bends of the path along the x-axis and y-axis, respectively; r x,max and r y,max These represent the maximum number of consecutive bends allowed on the x-axis and y-axis, respectively.

[0141] There are multiple paths P mentioned above. The goal of DTW is to find an optimal curved path that minimizes the total cost of the curvature of sequences X and Y.

[0142]

[0143] In formula (7), D(p) s () represents the cumulative distance of the curved path;

[0144] To solve equation (7), this implementation method constructs a cumulative cost matrix L using dynamic programming, the element expression of which is:

[0145] l i,j =D i,j +min(L i-1,j-1 ,L i,j-1 ,L i-1,j (8);

[0146] In formula (8), i = 1, 2, ..., n; j = 1, 2, ..., m, L i,j Let L be the element in the i-th row and j-th column of matrix L, where L 0,0 =0,L i,0 =L 0,j =+∞.

[0147] S203: Description of curve similarity;

[0148] For discrete load curves X = [x1, x2, ..., x...] nIn this embodiment, it is converted into a morphological sequence X' = [x'1, x'2, ..., x'] of length n-1. n-1 This reflects the trend information of the load curve over various time periods and describes the local dynamic characteristics of the curve. The element xi′ is given by the following formula:

[0149]

[0150] In formula (9), Δt is the time interval;

[0151] Based on the characteristics of the load curve described above, this implementation method combines weighted Euclidean distance and DTW to propose a similarity measurement method that takes into account the overall distribution characteristics, overall dynamic characteristics, and local dynamic characteristics of the daily load curve, as shown in the following formula:

[0152]

[0153] In formula (10), D2'(X,Y) is the weighted Euclidean distance between curves X and Y, used to measure the similarity of the numerical distributions at corresponding time points between the curves, that is, the similarity of the overall distribution characteristics of the curves. dtw (X,Y) represents the DTW distance between curves X and Y, used to measure the overall shape similarity of the curves, i.e., the overall dynamic similarity of the curves. dtw (X',Y') represents the DTW distance corresponding to the local dynamic characteristics of curves X and Y described by formula (8); α, β, and γ are the weights, and D all The smaller the (X,Y) value, the higher the similarity between the two curves.

[0154] After completing the load curve similarity calculation, the load curves are then preliminarily clustered. The preliminary clustering process is as follows: Figure 2 As shown.

[0155] S204: Weight selection based on the improved CRITIC method;

[0156] To more accurately characterize the daily load curve, this invention employs an improved CRITIC method to determine the weights of three similarity indicators. The CRITIC method is a type of objective weighting method, its basic principle being to determine the objective weights of indicators based on their discriminative power and conflict among them. The magnitude of the change in an indicator's value across different samples reflects its discriminative power; the greater the change, the stronger its discriminative power. The correlation between indicators reflects their conflict. Therefore, standard deviation and correlation coefficient are generally used to measure the magnitude and direction of the indicator's discriminative power and conflict, as shown in the following formula:

[0157]

[0158] In formulas (11) and (12), σ j Let r be the standard deviation of the j-th indicator. ij Let C be the correlation coefficient between the i-th indicator and the j-th indicator. j Let W represent the total amount of information represented by the j-th evaluation indicator. j To determine the objective weight of the j-th evaluation index obtained through normalization, The average value of the j-th indicator;

[0159] However, the dimensionless standard deviation σ j It is not possible to directly compare the discriminative power of different indicators, and the correlation coefficient r ij The expression can be either positive or negative, and therefore does not fully reflect the conflict between indicators. Therefore, this implementation method adopts an improved CRITIC method, as shown in the following formula:

[0160]

[0161]

[0162] In formulas (13) and (14), The average value of the j-th indicator;

[0163] S205: Aggregation is performed by combining weighted Euclidean distance, DTW distance, and K-means algorithm;

[0164] S206: Calculate clustering quality indicators;

[0165] The general requirement for clustering is that individuals within the same category have high similarity, while individuals between different categories have high dissimilarity. Clustering quality assessment metrics are used to evaluate clustering quality using indicators such as the sum of squared errors, the CHI (Calinski Harabasz index), and the DBI (Davies-Bouldin index). The DBI index... DBI This value represents the ratio of the sum of intra-class distances to the inter-class distances. A smaller value indicates better clustering. Furthermore, this index is more suitable for clustering power system load curves compared to other indices. Its formula is simple, has a small range of variation, and is easy to apply. Therefore, this implementation uses index I. DBI The indicator is calculated using the following formula:

[0166]

[0167] In formulas (15) and (16), S w and R w All are parameters, R w S is used to measure the similarity between the center data points of class w and class q; w Used to measure the clustering of data points in the w-th class; Mwq This represents the degree of dispersion between the data points of the w-th class center and the q-th class center.

[0168] Will I DBI The minimum number of clusters is set as the optimal number of clusters. Figure 2 The initial number of clusters L min =2, final number of clusters L max =0.5N, where N is the number of effective daily load curves of the clustered database.

[0169] This invention combines DTW distance and weighted Euclidean distance to comprehensively determine curve similarity. Clustering is then performed based on this comprehensive similarity, taking into account three similarity characteristics: overall distribution characteristics, local dynamic characteristics, and overall dynamic characteristics of the daily load curves. It also fully considers the varying degrees of criticality at different times, using an improved CRITIC method to weight the similarity of these curve characteristics. The k-means algorithm is then applied to conduct preliminary clustering analysis of the daily load curves of power grid users. This effectively avoids misclassification of curves during clustering, resulting in high-quality clustering and providing more accurate cluster center curves for the analysis of electricity load curves from multiple users.

[0170] S3: Input the initially clustered load dataset into the STL model for decomposition, and divide the decomposed load data into basic load and sensitive load;

[0171] User-side load typically consists of base load and load affected by other factors. The former reflects the user's basic electricity consumption attributes when unaffected by external factors, while the latter reflects the user's sensitivity to external factors. By performing cluster analysis on both, the former can analyze the user's long-term electricity consumption patterns and habits, while the latter can analyze the user's response characteristics to factors such as electricity prices and weather. Combining the initial clustering results, a more accurate and practical electricity user profile can be constructed.

[0172] Seasonal and trend decomposition using Loess (STL) is an algorithm for decomposing time series data based on Loess's locally weighted regression smoothing estimation technique. It can handle any type of time series data, dividing it into trend components, periodic components, and residual components. The STL algorithm is simple to operate, has few parameters, is fast in computation, provides interpretable components, eliminates the need to worry about the choice of component count, is insensitive to outliers, and has good robustness. STL typically uses the addition principle to decompose time series data, as expressed in the following expression:

[0173] Y t =T t +S t +I t (17);

[0174] In formula (17), Y t Let T be the observation value at time t; t S t I t These represent the trend component value, periodic component value, and residual component value at time t, respectively.

[0175] In this implementation, the sum of the trend component and the periodic component after STL decomposition of the original load data is regarded as the basic load, and the residual component is regarded as the sensitive load, that is:

[0176] Z t =T t +S t (18);

[0177] In formula (15), Z t It is the basic load component.

[0178] S4: Perform cluster analysis on the basic load and sensitive load respectively, and define the label library for the basic load and the label library for the external factor sensitive load according to the clustering results;

[0179] After separating the base load and the load sensitive to external factors through load decomposition, clustering models are constructed separately. Following analysis, label libraries are defined for the base load and the load sensitive to external factors. Finally, these two types of labels are used to characterize user electricity consumption behavior, achieving a precise profile of user electricity consumption behavior. The profiling process is as follows: Figure 3 As shown.

[0180] S401: The clustering effectiveness index control method is used to determine the number of clusters k for the basic load, thereby obtaining the number of basic load types. Cluster analysis is performed on the basic load data, and electricity consumption behavior labels are defined and scored based on the clustering results.

[0181] The electricity consumption behavior characteristics of electricity users can usually be described using electricity consumption characteristics derived from the user's load curve, such as load factor, peak-to-valley ratio, peak-to-average power ratio, etc.

[0182] The load factor k1 can be used to describe the usage of a user's electrical equipment, and is expressed as the daily average load P. av With maximum load P max The ratio:

[0183]

[0184] The peak-to-valley difference rate k2 can be used to describe the volatility of electricity consumption behavior, and is expressed as the peak-to-valley difference (P). 峰 -P 谷 ) and peak load P 峰 The ratio:

[0185]

[0186] The peak-to-average power ratio k3 can be used to describe the user's electricity consumption level during peak hours, and is expressed as the peak load P. 峰 With average load P av The ratio:

[0187]

[0188] To more vividly illustrate the level of different characteristics of a certain type of electricity consumption within the overall picture, this implementation method can be described using a scoring system. The maximum score is 10 points, and the score for this characteristic among all users can be obtained using the following formula:

[0189]

[0190] In formula (22), k i For feature categories, C j For the j-th type of electricity consumption behavior, For all j-th type of electricity consumption behavior, k i The average value of the feature k for all users i The minimum value of the feature. k for all users i The maximum value of the characteristic. S402: Perform correlation analysis on sensitive loads to extract the correlation coefficients of various factors on each type of load.

[0191] Using the same clustering effectiveness index as in step 4.1, the number of clusters k affected by other factors was determined, and cluster analysis was performed to obtain different types of influence, and labels were defined for each.

[0192] In correlation analysis, the maximal information coefficient (MIC) is a correlation analysis algorithm based on mutual information. It uses a grid partitioning method to measure the strength of the association between two variables, whether linear or nonlinear. It is commonly used for feature selection and has good universality, fairness, and symmetry. Let X = [x1, x2, ..., x...]. n ] and Y = [y1, y2, ..., y n Let X and Y be random variables in the dataset, and n be the number of samples. Then the MI between X and Y is:

[0193]

[0194] In formula (23), X and Y are random variables in the sensitive load dataset, p(x,y) is the joint probability density between X and Y, and p(x) and p(y) are the marginal probability densities between X and Y, respectively.

[0195] MIC overcomes the difficulty of MI in calculating the joint probability density function of continuous variables, and can find the correlation between two variables to the greatest extent. The formula for calculating MIC is:

[0196]

[0197] In formula (24), B is the sample size, N is the sample variable, and I(x,y) is the MI value between variables x and y. The closer the MIC value between two variables is to 1, the stronger their correlation.

[0198] The specific steps for determining the number of clusters k and performing cluster analysis using the clustering validity index control method in S401 and S402 are as follows:

[0199] The Gaussian Mixture Model (GMM) clustering algorithm must first determine the number of clusters, k. Based on the principles of the GMM algorithm, the sum of squared errors (SSE) can be used as an evaluation metric for determining the number of clusters. The formula for SSE is as follows:

[0200]

[0201] In formula (25), k is the number of clusters, and C i For a cluster in the clustering, u i C i The mean of all data points in the cluster, where x is a point in the cluster;

[0202] Assume the actual number of clusters is k * When the number of clusters k <k * When k increases, SSE decreases rapidly; when k ≥ k * As k increases, the decreasing trend of SSE slows down significantly, so the optimal number of clusters can be determined based on the changes in SSE.

[0203] To analyze user behavior of different types in different seasons, it is necessary to cluster the basic load and sensitive load. This invention uses the GMM algorithm for clustering, with the aim of classifying similar loads into the same category for analysis.

[0204] After determining the number of clusters k using SSE, it is assumed that each cluster follows a Gaussian distribution, where the probability model of the Gaussian distribution is:

[0205]

[0206] In formula (26), y represents the sample; α represents the sample. k As the weight, φ(y|θ) k Let θ be the probability density function of a Gaussian distribution. k The parameters of the probability density include μ.k and The probability density expression is:

[0207]

[0208] In formula (27), σ k Let μ be the standard deviation of sample y. k Let y be the mean of the sample y;

[0209] The probability density functions of these k classes of Gaussian distributions and the weight α of each class were estimated using the training data of the Gaussian model. k Next, calculate the probability of each data point appearing in each of the k Gaussian distributions, that is, substitute the data point into each of the k Gaussian distributions to find the probability P(y) of belonging to each class. i ) k :

[0210]

[0211] In formula (28), y i is the corresponding i-th data; k is the k-th Gaussian distribution.

[0212] Finally, by comparing the results, the sample is assigned to the class with the highest probability value, thus completing the cluster analysis.

[0213] S5: Based on the label library of basic load and the label library of loads sensitive to external factors, assign two types of labels to each user. Combine the labels from the initial clustering results to generate an accurate profile of the power user.

[0214] In summary, this invention considers simultaneously analyzing the overall basic electricity consumption behavior patterns of users and their sensitivity to other factors. It proposes using the STL algorithm to decompose the original load data into basic load and loads affected by other factors. Then, these two types of loads are clustered and analyzed separately to form label libraries for each type. Finally, the two types of labels are combined with the initial clustering labels to form a precise profile of the electricity user. This invention's method of combining direct clustering analysis of original load data with separate clustering analysis of load decomposition for profiling enables more refined and comprehensive user behavior analysis. It retains original information while improving the accuracy of the profile, helping power companies accurately grasp user behavior, improve their service levels, and provide data support for future demand response strategy development. It has significant application value.

[0215] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent substitutions, and improvements made to the above embodiments without departing from the scope of the present invention, based on the technical essence of the present invention and within the spirit and principles of the present invention, shall still fall within the protection scope of the present invention.

Claims

1. A method for profiling electricity user behavior considering electricity sensitivity, characterized in that, The steps of the method for profiling electricity user behavior that considers electricity sensitivity include: Step 1: Obtain the power load dataset and plot the load curve; Step 2: For any two given time series in the load curve and The weighted Euclidean distance, load change trend and DTW distance of the sampling points in the load curve are calculated. The weights are selected based on the improved CRITIC method. The load curve is initially clustered and the initial clusters are labeled. Step 2 specifically includes: Step 2.1: Based on the two time series given in the load curve and Calculate the first i Standard deviation of time period load values and average Based on standard deviation and average The load fluctuation index was calculated. Based on load fluctuation index Calculate the first i Normalized weights of load values ​​for different time periods to obtain the curve X and curve Y The weighted Euclidean distance of the same sampling point in the sample; Step 2.2: Calculate the time series using the DTW algorithm and DTW distance; Step 2.3: For discrete load curves This is transformed into a sequence of shapes of length n - 1. To reflect the trend information of the load curve in each time period and describe the local dynamic characteristics of the curve, the weighted Euclidean distance, load change trend and DTW distance are used to describe the curve similarity. Step 2.4: Perform preliminary clustering on the curves after calculating similarity, output the best preliminary clustering results, and label the preliminary clusters; Step 2.4 specifically includes: Step 2.4.1: Preprocess the data in the loading curve after similarity calculation, setting the initial number of clusters to... ; Step 2.4.2: Set the initial cluster center curve; Step 2.4.3: Configure similarity weights based on the improved CRITIC method; Step 2.4.4: Combine weighted Euclidean distance, DTW distance, and K-means algorithm for aggregation; Step 2.4.5: Calculate cluster quality indices ; Step 2.4.6: Determine L Is it equal to If not equal, in order L=L+ 1, and repeat steps 2.4.4-2.4.5 until L= ; Step 2.4.7: Select Class Quality Indicators The value corresponding to the minimum L And output the corresponding best clustering result; The expression for configuring similarity weights is: (10); (11); In formulas (10) and (11), For the first The standard deviation of the indicator For the first Item and the first The correlation coefficient between the indicators For the first The evaluation indicators represent a comprehensive amount of information. The first one obtained through normalization operation The objective weight of each evaluation indicator For the first The average value of each indicator; Clustering quality indicators The expression is: (12); (13); In formulas (12) and (13), and All are parameters. Used to measure the first w Class and the q Similarity of class center data points; Used to measure the first w The degree of clustering of data points within a class; For the first w Class center and the first q The degree of dispersion of the class center data points; Load fluctuation index The calculation formula is: (1); No. i Normalized weights of load values ​​for different time periods The calculation formula is: (2); The expression for weighted Euclidean distance is: (3); Step 3: Input the initially clustered load dataset into the STL model for decomposition, and divide the decomposed load data into basic load and sensitive load; Step 4: Perform cluster analysis on the basic load and sensitive load respectively, and define the label library for the basic load and the label library for the external factor sensitive load according to the clustering results; Step 5: Based on the label library of basic load and the label library of loads sensitive to external factors, assign two types of labels to each user. Combine the labels from the initial clustering results to generate an accurate profile of the power user.

2. The method for profiling electricity user behavior considering electricity sensitivity according to claim 1, characterized in that, Step 2.2 specifically includes: Step 2.2.1: Based on time series and Build Type distance matrix ; Step 2.2.2: Convert the matrix D Each set of adjacent elements is defined as a curved path. , ,in k The total number of elements in the path, element For the first on the path The coordinates of each point, i.e. ; Step 2.2.3: Set two constraints for the curved path. Constraint one is that the curved path must start from the lower left corner. Start, go to the top right corner End. Constraint two is that each point in the curved path must match its adjacent point, that is, if... ,but Must meet Step 2.2.4: Based on the constraints in Step 2.2.3, add constraints on the number of consecutive bends; Step 2.2.5: After completing the curved path constraint, construct the cumulative cost matrix. L Calculate the time series and The bending path with the minimum total bending cost P, This leads to the time series. X and Y DTW distance ; Distance matrix Middle elements The calculation formula is: (4); The expression for the constraint on the number of consecutive bends is: (5); In formula (5), and The paths are respectively in x shaft and y Number of consecutive bends on the axis; and They are respectively in x shaft and y The maximum number of consecutive bends allowed on the shaft; Minimum total cost curved path P The calculation formula is: (6); In formula (6), This represents the cumulative distance along the curved path. Cumulative cost matrix L The expression is: (7); In formula (7), ; , For matrix L The Middle line, number Column elements, where .

3. The method for profiling electricity user behavior considering electricity sensitivity according to claim 1, characterized in that, Step 2.3 specifically includes: Step 2.3.1: Convert the discrete load curves Convert to a set of length n - 1 morphological sequence To reflect the trend information of the load curve in different time periods, the load curve The same process was performed to obtain ; Step 2.3.2: Use weighted Euclidean distance and DTW distance, combined with the transformed curve. and Calculate the load curve and Similarity; sequence Middle elements xi The expression for ′ is: (8); In formula (8), For time intervals; Load curve and The formula for calculating the similarity is: (9); In formula (9), For curves X and Y The weighted Euclidean distance between the curves is used to measure the similarity of their numerical distributions at corresponding time points, i.e., the similarity of their overall distribution characteristics. For curves X and Y The DTW distance between curves is used to measure the overall shape similarity of the curves, that is, the overall dynamic similarity of the curves. The curve described by formula (8) X and Y The DTW distance corresponding to local dynamic characteristics; As weight, The smaller the value, the higher the similarity between the two curves.

4. The method for profiling electricity user behavior considering electricity sensitivity according to claim 1, characterized in that, Step 3 specifically includes: The load dataset after preliminary clustering is input into the STL model. The time-series load data is decomposed using the addition principle to obtain the trend component value, periodic component value, and residual component value at the corresponding time. The decomposed trend component and periodic component are used as the basic load, and the residual component is used as the sensitive load. The expression for decomposing time-series load data is: (14); In formula (14), for The observed value at time; , , They are respectively Trend component value, periodic component value, and residual component value at any given time; The expression for the base load is: (15); In formula (15), It is the basic load component.

5. The method for profiling electricity user behavior considering electricity sensitivity according to claim 1, characterized in that, Step 4 specifically includes: Step 4.1: Determine the number of clusters for the basic load using the clustering effectiveness index control method. This leads to the number of basic load types, and cluster analysis is performed on the basic load data. Based on the clustering results, electricity consumption behavior labels are defined and scored. Step 4.2: Perform correlation analysis on sensitive loads and extract the correlation coefficients of various factors on each type of load. The clustering effectiveness index, consistent with step 4.1, was used to determine the number of clusters affected by other factors. Cluster analysis was performed to obtain different impact types, and labels were defined for each type.

6. The method for profiling electricity user behavior considering electricity sensitivity according to claim 5, characterized in that, In steps 4.1 and 4.2, the clustering validity index control method is used to determine the number of clusters. The steps involved in cluster analysis include: Using the sum of squared errors as the evaluation index for the number of clusters, the true number of clusters is set as... When the number of clusters hour, k As SSE increases, it decreases rapidly; when hour, k Increase the number of clusters and determine the optimal number of clusters based on the changes in SSE. k And perform clustering; Both the base load data and the sensitive load data are assumed to follow a Gaussian distribution. The data is estimated using the training data of a probability model based on the Gaussian distribution. k The probability density function of a Gaussian-like distribution and the weights of each class. And substitute each data point into k Find the probability of belonging to each class from a Gaussian distribution. The data is then grouped into the category with the highest probability value to complete the cluster analysis. The expression for the sum of squared errors is: (16); In formula (16), k The number of clusters, For a cluster in the clustering, for The mean of each data point in the middle. x For a point in the cluster; The probability model of the Gaussian distribution is: (17); In formula (17), sample; As weight, Let be the probability density function of a Gaussian distribution. The parameters of the probability density include and ; The expression for the probability density function is: (18); In formula (18), For the sample y standard deviation For the sample y The mean; probability The calculation formula is: (19); In formula (19), For the corresponding number i One data point; For the first A Gaussian distribution.

7. The method for profiling electricity user behavior considering electricity sensitivity according to claim 5, characterized in that, Step 4.1, which defines electricity consumption behavior labels and assigns scores based on the baseline load clustering results, specifically includes: By calculating the load factor Peak-valley difference rate Peak-to-average power ratio To describe the electricity consumption behavior characteristics of electricity users, a scoring system with a maximum score of 10 points is set, and the electricity consumption behavior characteristics of all users are scored. Load factor The calculation formula is: (20); Peak-valley difference The calculation formula is: (21); Peak-to-average power ratio The calculation formula is: (22); The expression for scoring applied electrical behavior features is: (23); In formula (23), For feature categories, For the first Types of electricity consumption behavior, For all the In similar electricity consumption behaviors The average value of the feature For all users The minimum value of the feature. For all users The maximum value of the feature.

8. The method for profiling electricity user behavior considering electricity sensitivity according to claim 5, characterized in that, Step 4.2, which involves correlation analysis of sensitive loads, specifically includes: The mesh partitioning method is used to calculate the inter-variable correlation (MI) between random variables in the sensitive load dataset. Based on the MI, the MIC value between two variables in the sensitive load dataset is calculated. The closer the MIC value is to 1, the stronger the correlation between the two variables. The formula for calculating the inter-variable randomness (MI) is: (24); In formula (24), X and Y For random variables in the sensitive load dataset, Let X be the joint probability density between X and Y. and These are the marginal probability densities between X and Y, respectively; (25); In formula (25), B is the sample size and N is the sample variable. For variables and The MI value between.