Improved Gaussian mixture analysis method for accurately identifying electricity utilization rule and electricity utilization abnormity of user
Through the improved Gaussian hybrid analysis method, the complexity of electricity load analysis and abnormal electricity identification problems are solved during holidays, more accurate electricity usage rules and abnormal detection are achieved, and the reliability of power grid management is improved.
Patent Information
- Application Number
- CN202510086750.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-30
AI Technical Summary
Holiday power load analysis is complex and difficult to accurately capture. Traditional models are difficult to deal with the volatility and irregularity of holiday data, and it is difficult to identify and predict abnormal power usage events.
The improved Gaussian hybrid analysis method is used to strictly analyze the grid load data, extract the user's electricity usage behavior characteristics, and cluster it using the Gaussian hybrid model to determine the user's daily electricity usage rules and abnormal electricity usage patterns.
It improves the accuracy and reliability of load analysis, can more accurately identify user electricity usage patterns and abnormalities, and helps power grid companies formulate effective energy management and planning strategies.
Smart Images

Figure CN120067850A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electricity load analysis, and in particular to an improved Gaussian mixture analysis method for accurately identifying user electricity consumption patterns and electricity anomalies. Background Art
[0002] Accurately analyzing the regularity and anomaly of holiday load data of electricity users is crucial for power enterprises to formulate reasonable power dispatching and energy efficiency management plans. Compared with daily load analysis, the electricity load during holidays has more complexity and variability. First of all, the electricity load during holidays is usually affected by people's work, leisure and consumption patterns, and these factors show great uncertainty among different holidays. Secondly, the electricity demand during holidays often shows periodic changes, but at the same time, there will also be large fluctuations and irregularities, making it difficult for traditional load analysis models to accurately capture these changes.
[0003] The difficulties in holiday electricity load analysis mainly focus on the following aspects: First, there are obvious periodic characteristics in electricity consumption behavior during holidays, but due to the diversity and uncertainty of holidays, the periodic patterns are often disrupted, increasing the complexity of analysis; Secondly, the amount of holiday electricity load data is small and the volatility is large, which makes it difficult for traditional statistical models to effectively analyze using limited historical data; Finally, during holidays, people's consumption habits and activity types may change drastically, and the electricity demands of industries such as commerce, entertainment, and tourism usually show strong suddenness and volatility, and traditional models are difficult to comprehensively consider these complex factors.
[0004] In addition, there are abnormal electricity consumption phenomena in holiday electricity loads, mainly referring to sudden or atypical electricity demands, such as extreme electricity consumption events on specific holidays or under specific weather conditions. Such abnormal electricity consumption events usually do not conform to the daily electricity consumption pattern, and their occurrence and influence range are often relatively limited, but they pose a serious threat to the safety and dispatching of the power grid. Traditional load analysis methods are often difficult to identify and predict these abnormal fluctuations, resulting in situations where the power system may have untimely resource allocation or over-dispatching when dealing with these sudden demands. Therefore, accurately identifying and handling abnormal electricity consumption events is an important task in holiday load analysis.
[0005] Considering comprehensively the periodic volatility of holiday electricity loads and the suddenness of user abnormal electricity consumption behavior, and combining the Gaussian mixture model for analyzing user electricity consumption patterns and electricity anomalies can effectively improve the accuracy and reliability of load analysis, and help power grid companies formulate effective energy management and planning strategies under the national policy of promoting orderly power consumption. Summary of the Invention
[0006] The present invention provides an improved Gaussian mixture analysis method for accurately identifying users' electricity consumption patterns and electricity consumption anomalies. By performing strict data analysis on grid load data, the basic behavioral characteristics of users and the sources of behavioral variability considering accurate matching of holidays are determined; these behavioral characteristics are analyzed to appropriately select clustering algorithms and their parameters, which will be extracted from each time series and then used in an improved Gaussian mixture modeling algorithm for clustering; after clustering, the daily electricity consumption patterns of users are determined and characterized; after clustering, the posterior probability, log-likelihood probability, and existence of sparse clustering are calculated to verify the results of user electricity consumption anomaly detection, so as to solve the problems in the background technology.
[0007] The technical solution of the present invention is as follows:
[0008] An improved Gaussian mixture analysis method for accurately identifying users' electricity consumption patterns and electricity consumption anomalies, comprising:
[0009] S1: First, screen the annual data to obtain the time series load data of weekdays, holidays, daily changes, and seasonal differences, standardize the daily electricity consumption average value of the whole year's data, process the missing values, and perform normalization processing;
[0010] S2: Extract behavioral characteristics from each time series data, calculate the characteristic indicators based on the whole year's data, and calculate the correlation between the characteristic indicators through the Pearson algorithm;
[0011] S3: Select a Gaussian mixture model to cluster the behavioral characteristics extracted in S2;
[0012] S4: After clustering, use principal component analysis to obtain the visualized final clusters to characterize the daily electricity consumption patterns of users;
[0013] S5: Calculate the posterior probability, log-likelihood probability, and existence of the sparse clustering clusters to verify the results of user electricity consumption anomaly detection.
[0014] Further, in S2, the characteristic indicators include the relative average power during the morning peak, the relative average power during the day, the relative average power during the evening peak, the relative average power at night, the holiday indicator, the relative average standard deviation, and the season indicator;
[0015] Due to the working system, users in the living area go to work in the morning, get off work in the evening, and rest at night; the daily load curve of the living day often shows the characteristic of "double peaks"; so four of the selected characteristic values are the relative average power during the morning peak, the evening peak, during the day, and at night. This also applies to the other three types of electricity consumption patterns (high energy consumption, industrial, commercial): the morning peak - day and day - evening peak are the electricity consumption peaks, and the day and night rest are the electricity consumption troughs.
[0016] Due to seasonal influences, electricity consumption is usually highest in summer, especially in hot regions. The use of cooling equipment such as air conditioners and fans increases significantly, leading to a notable rise in the electricity demand for residential and commercial use; peak loads mostly occur from noon to evening because the temperature is highest at this time and the load on cooling equipment is large. In winter in cold regions, electricity consumption may also increase significantly, mainly due to the use of heating equipment (such as electric heaters, heat pumps, etc.); peak loads may occur in the early morning and evening because the temperature is lowest at these times and the heating demand is high. In spring and autumn, due to relatively mild climates, the demand for cooling and heating is low, and electricity consumption is usually lower than in summer and winter; the electricity load is relatively stable with small fluctuations throughout the day. Therefore, one of the selected characteristic values is the season index.
[0017] The frequency of holidays is low, and the load characteristics are affected by various factors such as the year and holiday arrangements, so the amount of available data is small; in addition, the regularity of holidays is weak, and the holiday times, policies, social and economic conditions, etc. may vary from year to year, which will lead to obstacles in load analysis. The volatility and uncertainty of these holidays will both cause obvious changes in the distribution characteristics of electricity loads in residential areas and those in commercial and industrial areas, resulting in load imbalance. Therefore, in order to better obtain the electricity consumption pattern, it is necessary to conduct in-depth analysis of holidays, and the holiday index can be selected as a characteristic value for subsequent clustering.
[0018] Furthermore, the correlation between characteristic indicators is represented by the Pearson correlation coefficient formula, and the calculation formula is as follows:
[0019]
[0020] In the formula, n represents the sample size, x i and y i respectively represent the variable characteristics, x and y respectively represent the means of the variable characteristics, r represents the Pearson correlation coefficient, and its value range is [-1, 1].
[0021] When r = 1, it indicates that the two variables are completely positively linearly correlated, that is, y i increases strictly according to a linear relationship as x i increases;
[0022] When r = -1, it indicates that the two variables are completely negatively linearly correlated, that is, y i decreases strictly according to a linear relationship as x i increases;
[0023] When r = 0, it indicates that there is no linear correlation between the two variables.
[0024] Further, in S3, the expectation maximization algorithm and the log-likelihood function are used to determine the parameters of the Gaussian mixture model, including initializing the parameters of the expectation maximization algorithm. The expectation step and the maximization step in the expectation maximization algorithm are executed alternately. After each execution, the log-likelihood function is calculated.
[0025] Further, the process of determining the parameters of the Gaussian mixture model is as follows:
[0026] S3.1: Determine the number of components k of the Gaussian mixture model and initialize the parameters w k , u k and ∑ k ;
[0027] S3.2: Expectation step. For each feature index data point x i , calculate its posterior probability γ k (x i ), and the calculation formula is:
[0028]
[0029] In the formula, N represents the total number of feature index data points, i represents the time period of the corresponding data point, and N(x i |u k , ∑k) represents the Gaussian probability density function;
[0030] S3.3: Maximization step. According to the posterior probability γ k (x i ) calculated in the expectation step, update the parameters w k , u k and ∑ k , and the calculation formulas are as follows:
[0031]
[0032] In the formula, u k is the mean value, ∑k is the covariance matrix, and w k is the weight, that is, the mixing coefficient;
[0033] S3.4: Repeat S3.2 and S3.3 until the Gaussian mixture model converges. The convergence judgment criterion is confirmed by the value of the log-likelihood function, and the calculation formula is as follows:
[0034]
[0035] When the change amount of the calculated value of the log-likelihood function is less than the preset range, stop the iteration to obtain the parameters of the Gaussian mixture model.
[0036] Further, in S4, principal component analysis maps the original high-dimensional data to a low-dimensional space through linear transformation and retains the key information of the data. Its mapping steps include:
[0037] S4.1: First, standardize the data to avoid the influence of feature scale differences on the results. The standardization calculation formula is as follows:
[0038]
[0039] In the formula, x' represents the standardized original data, x represents the original data, and σ i is the standard deviation of the feature index;
[0040] S4.2: For the original data matrix X, calculate its covariance matrix δ lm , and the calculation formula is as follows:
[0041]
[0042] In the formula, δ lm represents the covariance of feature l and feature m, and T represents the transpose;
[0043] S4.3: Obtain the eigenvalues and eigenvectors of the covariance matrix through eigenvalue decomposition. The calculation formula is as follows:
[0044] δ lm = λv
[0045] In the formula, the eigenvalue λ is used to measure the variance of the principal component data, and the eigenvector v represents the direction of the principal component;
[0046] S4.4: Sort the eigenvalues λ from largest to smallest, and select the corresponding eigenvectors according to the sorting. Select the first q eigenvectors to form the projection matrix W, denoted as W = [v 1 , v 2 ,..., v q , project the original data matrix X into the principal component space, denoted as Z = XW. Z represents the data matrix after dimensionality reduction. By adjusting the size of q, the degree of dimensionality reduction can be controlled;
[0047] S4.5: Through data mapping, realize the visualization of users' electricity consumption patterns.
[0048] Further, in S5, the sparse clustering cluster represents an abnormal situation. The likelihood probability and weighted logarithmic probability are used as measures of abnormality, and the thresholds of the likelihood probability and weighted logarithmic probability are calculated. The standard deviation method is used to determine the standard deviation value lower than the cluster average to distinguish abnormal users from regular users;
[0049] The calculation formula for the likelihood probability of the sparse clustering cluster is as follows:
[0050]
[0051] Weighted logarithmic probability:
[0052]
[0053] In the above formula, π represents the normalization used with the total number of data points for the probability density function, ensuring that the total probability in the entire data space is 1. represents the data point x i the difference vector with the mean u k transpose, and together with (x i - u k ) is used to calculate the Mahalanobis distance. exp represents the exponential function, and exp is used to calculate the exponential form of the Mahalanobis distance between the data point x i and the mean u k .
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] 1. Through the clustering analysis of users' electricity consumption behavior characteristics, the present invention determines and characterizes the daily (e.g., weekly load curve of working five days and resting two days or daily load curve with two peaks) or abnormal electricity consumption patterns of different users (high energy consumption, industrial, residential, commercial), and at the same time determines from the user load data factors such as the repeatedly occurring peak demand time periods, holidays and non-holidays, and seasonally induced variability, and groups and characterizes the electricity consumption rules of different users during holidays according to the peak demand characteristics. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 is a schematic flowchart of an improved Gaussian mixture analysis method for accurately matching users' electricity consumption rules and electricity consumption anomalies considering holidays provided by the present invention;
[0057] Figure 2 is a heat map of the correlation of characteristic indicators of an improved Gaussian mixture analysis method for accurately matching users' electricity consumption rules and electricity consumption anomalies considering holidays provided by the present invention;
[0058] Figure 3 is a comparison chart of the clustering results of holiday and non-holiday load data in a certain area in January 2023 and February 2024 in the implementation case of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] The following further describes in detail the embodiments of the present invention with reference to the drawings and examples. The following examples are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.
[0060] Example:
[0061] In this embodiment, Figure 1 It is a flow schematic diagram provided by the present invention to illustrate the analysis method for accurately matching the user's electricity consumption pattern and electricity consumption anomaly considering holidays in this embodiment, specifically including:
[0062] S1: First, screen the annual data to obtain the time series load data of weekdays, holidays, daily changes, and seasonal differences. Standardize the daily electricity consumption average value of the annual data, process the missing values, and perform normalization processing.
[0063] S2: Extract the behavioral characteristics from each time series data, calculate the characteristic indicators based on the annual data, and calculate the correlation between the characteristic indicators through the Pearson algorithm;
[0064] The characteristic indicators include the relative average power of the morning peak, the relative average power of the daytime, the relative average power of the evening peak, the relative average power of the night, the holiday indicator, the relative average standard deviation, and the season indicator;
[0065] The above seven characteristic indicators are calculated in the following manner:
[0066]
[0067] In the formula, MP represents Morning Peak, that is, the average consumption during the morning peak period relative to the total average consumption (relative average value). u 早 represents the average power during the morning peak period in a year. In this embodiment, the morning peak period is 6:00 - 8:00, and the average power of the user in a year is
[0068]
[0069] In the formula, DT represents Day Time, that is, the average consumption during the daytime period relative to the total average consumption (relative average value), u 白 represents the average power during the daytime period in a year. In this embodiment, the daytime period is 8:00 - 18:00.
[0070]
[0071] In the formula, NP represents Night Peek, that is, the average consumption during the evening peak period relative to the total average consumption (relative average value), u 晚 represents the average power during the evening peak period in a year. In this embodiment, the evening peak period is 18:00 - 20:00.
[0072]
[0073] In the formula, NT represents Night Time, i.e., night time. The average consumption during the night time period relative to the total average consumption (relative average value) is u 夜 represents the average power during the night time period in a year. In this embodiment, the night time period is from 20:00 to 6:00 the next day.
[0074]
[0075] In the formula, MRSD represents Mean Relative Standard Deviation. Considering the relative standard deviations in different time periods to measure the variability and irregularity of consumers, it can be used as a feature for anomaly detection. The standard deviation in each time period of a year is σ i , where i = (1, 2, 3, 4), representing four time periods in a day.
[0076]
[0077] In the formula, SSCORE represents Seasonal Score. Through this value, it is shown that the difference in consumption between summer and winter is proportional to the average annual demand. represents the average efficacy in winter for each time period, represents the average efficacy in summer for each time period.
[0078]
[0079] In the formula, HDSCORE represents Holiday vs Weekday Score. The difference in consumption between holidays and weekdays is proportional to the average annual demand. According to this characteristic value, the difference between weekdays and holidays can be measured.
[0080] It should be noted that these features are calculated based on the annual data under the condition of considering the measurements of different time periods, seasons, weekdays, holiday behaviors, and the daily variations and fluctuations of each user. These features can be applied to different data sets with different time resolutions and are easier to calculate.
[0081] Considering different combinations of feature indicators, such as the seasonality in each time period, the intra-day entropy, the day-time entropy in different time periods, the day-time entropy in different seasons, and the index of holidays and non-holidays. However, it is found that these feature indicators are highly correlated with the seasonality, and the index of holidays and weekdays; the correlation between feature indicators is represented by the Pearson correlation coefficient formula, and the calculation formula is as follows:
[0082]
[0083] where n represents the number of samples, and x i and y i represent variable features respectively, and represent the means of the variable features respectively, r represents the Pearson correlation coefficient, and its value range is [-1, 1].
[0084] Combined with Figure 2 , Figure 2 which is a heat map of the correlation of feature indicators provided by the present invention, to illustrate the correlations of the seven feature indicators of the improved Gaussian mixture analysis method for accurately matching the user's electricity consumption pattern and electricity consumption anomaly considering holidays provided in this embodiment. These correlations are calculated through the Pearson correlation coefficient formula, indicating that the correlations between most features are very weak. Therefore, it is feasible to cluster through these seven features. However, there is a certain correlation between the average relative standard deviation and the weekend and weekday scores and the seasonal scores. This is because the average relative standard deviation represents the daily fluctuations of the demands of users (industrial, high-energy-consuming, residential, commercial), and such fluctuations are affected by seasonality, date type, and other factors. Therefore, it is an important parameter for capturing the daily electricity consumption changes of each user. The high correlation between the daytime relative average power and the nighttime relative average power is due to the work shifts of users in the living area, that is, users who do not consume electricity during the day will increase their electricity consumption at night, and vice versa. The same is true for users in the commercial area, high-energy-consuming, and industrial areas, except that users in these areas consume electricity during the day and reduce their electricity consumption at night.
[0085] S3: Select a Gaussian mixture model to cluster the behavior features extracted in S2;
[0086] Optimize the fit between the observed data and the mathematical model from the perspective of probability theory to visualize the underlying probability distribution of the features used for clustering; based on the above S2, it is finally obtained that the seven extracted features follow a Gaussian distribution, and a Gaussian mixture-based clustering algorithm is selected to process the selected data;
[0087] Use the expectation-maximization algorithm and the log-likelihood function to determine the parameters of the Gaussian mixture model, including initializing the parameters of the expectation-maximization algorithm, and alternately executing the expectation step and the maximization step in the expectation-maximization algorithm. After each execution, calculate the log-likelihood function.
[0088] The process of determining the parameters of the Gaussian mixture model is as follows:
[0089] S3.1: Determine the number of components k of the Gaussian mixture model and initialize the parameters w k 、u k and ∑ k ;
[0090] S3.2: Expectation step, for each feature index data point x i , calculate its posterior probability γ k (x i ) for each Gaussian component, and the calculation formula is:
[0091]
[0092] where N represents the total number of feature index data points, i represents the time period corresponding to the data point, and N(x i |u k ,∑k) represents the Gaussian probability density function;
[0093] S3.3: Maximization step, according to the posterior probability γ k (x i ) calculated in the expectation step, update the parameters w k , u k and ∑ k of each Gaussian component, and the calculation formulas are as follows:
[0094]
[0095] where u k is the mean value, ∑k is the covariance matrix, and w k is the weight, i.e., the mixing coefficient;
[0096] S3.4: Repeat S3.2 and S3.3 until the Gaussian mixture model converges. The convergence criterion is confirmed by the value of the log-likelihood function, and the calculation formula is as follows:
[0097]
[0098] Stop the iteration when the change amount of the calculated value of the log-likelihood function is less than the preset range, and obtain the parameters of the Gaussian mixture model.
[0099] Furthermore, find the best number of clusters k that is most suitable for the given data set. By using five clustering validity indices, namely, Arbelaitz, Silhouette, Calinski Harabasz, Davies-Bouldin, AIC (Akaike Information Criterion), and BIC (Bayesian Information Criterion), improve and select the best model suitable for all reasonable values through comprehensive comparison.
[0100] The five clustering validity indices are calculated as follows:
[0101]
[0102] where AIC k(Akaike Information Criterion) represents the Akaike information criterion coefficient, which is the maximum value of the log-likelihood function of the model. Gaussian mixture is performed by maximizing the log-likelihood as a suitable clustering validation index for clustering, and the optimal k is selected according to the lower score to maximize the likelihood.
[0103]
[0104] In the formula, BIC k (Bayesian Information Criterion) represents the Bayesian information criterion coefficient. Similar to AIC, except for the penalty term, it depends on the sample size in BIC, and f is the number of observations.
[0105]
[0106] In the formula, SS k (Silhouette Score) represents the silhouette coefficient, which captures the separation and compactness of each object and returns the average of the scaled differences between separation and compactness. a g is the distance between the g-th clustering object and all other objects in the same cluster, and b g is the average pairwise distance between the g-th object and all other objects in its nearest cluster.
[0107]
[0108] In the formula, CH k (Calinski Harabasz) represents the CH index, also known as the variance ratio criterion, which is used to measure the degree of definition of clustering. This score is calculated from the ratio of the average between-cluster and within-cluster sum of squares. f h is the number of observations in the h-th cluster, ‖ρ h -u‖ is the Euclidean distance between the centroid of the h-th cluster and the mean of all objects, and ‖x - ρ h ‖ is the Euclidean distance between each object and the cluster centroid.
[0109]
[0110] In the formula, DB k (Davies Bouldin) represents the DB index, which captures the separation and compactness of all data cluster pairs and returns a system-wide similarity measure of each cluster compared to its most similar neighbor. θ(x g , x j ) is the inter-cluster distance calculated using the Minkowski distance, Δ(xg ) and Δ(x j ) are the within - cluster distances of clusters g and j respectively.
[0111] S4: Clustering is completed, and principal component analysis is used to obtain the final visualized clusters, which characterize the user's daily electricity consumption pattern;
[0112] Principal component analysis maps the original high - dimensional data to a low - dimensional space through linear transformation and retains the key information of the data. Its mapping steps include:
[0113] S4.1: First, standardize the data to avoid the influence of feature scale differences on the results. The standardization calculation formula is as follows:
[0114]
[0115] In the formula, x' represents the standardized original data, x represents the original data, and σ i is the standard deviation of the feature index;
[0116] S4.2: For the original data matrix X, calculate its covariance matrix δ lm , and the calculation formula is as follows:
[0117]
[0118] In the formula, δ lm represents the covariance of feature l and feature m, and T represents the transpose;
[0119] S4.3: Obtain the eigenvalues and eigenvectors of the covariance matrix through eigenvalue decomposition. The calculation formula is as follows:
[0120] δ lm = λv
[0121] In the formula, the eigenvalue λ is used to measure the variance of the principal component data, and the eigenvector v represents the direction of the principal component;
[0122] S4.4: Sort the eigenvalues λ from largest to smallest, and select the corresponding eigenvectors according to the sorting. Select the first q eigenvectors to form the projection matrix W, denoted as W = [v 1 , v 2 ,..., v q , project the original data matrix X into the principal component space, denoted as Z = XW. Z represents the data matrix after dimensionality reduction. By adjusting the size of q, the degree of dimensionality reduction can be controlled;
[0123] S4.5: Through data mapping, the visualization of the user's electricity consumption pattern is realized.
[0124] S5: Calculate the posterior probability, log-likelihood probability, and existence of the sparse clustering clusters to verify the user's abnormal power consumption detection results;
[0125] It should be noted that: in the cluster analysis of specific user load data, the regularity of power consumption behavior is greatly affected by date types, holiday effects, daily volatility, and seasonal changes. The load data of some users show different peak demand time periods during different holidays, seasons, or date types. To characterize the power consumption regularity of different user groups, the mean of the relative mean standard deviation index is used as a feature for cluster analysis. However, after in-depth analysis of the power consumption behavior of each cluster, it is found that multiple factors jointly affect the peak demand regularity that appears within a specific time period. These factors include the power consumption difference between holidays and non-holidays, seasonal consumption differences, and the volatility of daily power consumption behavior. The means of the relative mean standard deviation, seasonal index, and holiday index of different user clusters can accurately reflect the regularity characteristics of their power consumption behavior. Clusters with obvious sparse members (for example: some users have power consumption behaviors that do not conform to the cycle law due to special time points such as holidays in their originally regular cycle power consumption patterns.) represent abnormal profiles, while other clusters (high-energy-consuming, residential, industrial, commercial clusters) represent the daily power consumption habits commonly adopted by different users. The average relative standard deviation value of the sparse profile (abnormal power consumption cluster) is very high because their behavior varies greatly throughout the year.
[0126] Based on the fact that the same daily power consumption behavior has the same characteristics, the clusters are further divided, and the similarities and differences between the divided clusters are analyzed.
[0127] To explain and verify the above-mentioned "clusters with obvious sparse members represent abnormal profiles", additional anomaly detection is added. The likelihood probability and weighted log probability are regarded as measures of anomalies; and the thresholds of the posterior and weighted log probabilities are estimated, and finally the thresholds of the posterior and weighted log probabilities are determined, using the standard deviation method to determine the standard deviation value lower than the average outlier; to distinguish abnormal users from regular users.
[0128] The formula for calculating the likelihood probability of the sparse clustering cluster is as follows:
[0129]
[0130] Weighted log probability:
[0131]
[0132] In the above formula, π represents being used together with the total number of data points to normalize the probability density function, ensuring that the sum of probabilities in the entire data space is 1. represents the data point x i and the mean uk The transpose of the difference vector, and is used together with (x i - u k ) to calculate the Mahalanobis distance. exp represents the exponential function, and exp is used to calculate the exponential form of the Mahalanobis distance between the data point x i and the mean u k .
[0133] Based on the above steps, finally, only the load data of a certain residential area in 2023 and 2024 are used to verify the algorithm, and the clustering results ( Figure 3 ) are obtained.
[0134] Combined with Figure 3 ; Figure 3 This is the comparison chart of the clustering results of the holiday and non - holiday load data in a certain area in January 2023 and February 2024 in the implementation case of the present invention.
[0135] It should be noted that: Holidays (blue line, yellow line):
[0136] Early morning period (0:00 - 8:00): The load is lower. It may be that the residents in this area have a more flexible schedule and get up later.
[0137] Morning period (8:00 - 12:00): The load increases more slowly during holidays, and the lunch peak time is postponed.
[0138] Evening peak (18:00 - 20:00): The load during holidays is usually higher, reflecting an increase in residents' activities at home.
[0139] Non - holidays (red line, purple line): Have a more obvious bimodal feature (morning peak and evening peak). The electricity consumption peak is more concentrated in the working hours (8:00 - 17:00), indicating that the daily life and schedule during weekdays are more regular. Through the clustering results, it can be analyzed that the electricity consumption behavior of this type of user conforms to residential electricity consumption.
[0140] The electricity consumption during holidays is more random and dispersed, and the electricity consumption peak is not as concentrated as that during non - holidays.
[0141] The electricity consumption during non - holidays shows a more regular weekday pattern and is strongly restricted by daily life and work.
[0142] Comparing the curves of January 2023 (winter) and February 2024 (end of winter), it can be found that:
[0143] The winter heating equipment and household electricity load are relatively large, resulting in a relatively high overall level of the curve.
[0144] Through the clustering effect diagram, the basic behavioral characteristics and sources of behavioral variability of users with accurate holiday matching can be clearly obtained.
[0145] In summary: The electrical load is more dispersed and fluctuates more during holidays, while the load on non-holiday days is more concentrated and has strong regularity. Precise matching of the electricity consumption pattern during holidays can not only optimize power grid dispatching and resource utilization, but also improve the user experience and promote the development of smart grids and green energy.
[0146] The embodiments of the present invention are given for the purposes of illustration and description. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. An improved Gaussian mixture analysis method for accurately identifying user electricity consumption patterns and anomalies, characterized in that: include: S1: First, annual data are screened to obtain time series load data for weekdays, holidays, daily changes, and seasonal differences. The daily average power consumption of the annual data is standardized, and missing values are processed for normalization; S2: Extract behavioral features from each time series data, calculate feature indicators based on the full-year data, and calculate the correlation between feature indicators using the Pearson algorithm; S3: Select the Gaussian mixture model to cluster the behavioral features extracted in S2; S4: Clustering is completed, and the final cluster is visualized using principal component analysis to represent the user's daily electricity consumption pattern; S5: Calculate the posterior probability, log-likelihood probability and existence of the sparse clustering clusters to verify the user power consumption anomaly detection results.
2. The improved Gaussian mixture analysis method for accurately identifying user electricity usage patterns and anomalies as claimed in claim 1 is characterized by: In S2, the characteristic indicators include relative average power during the morning peak, relative average power during the day, relative average power during the evening peak, relative average power during the night, holiday indicators, relative average standard deviation, and seasonal indicators.
3. The improved Gaussian mixture analysis method for accurately identifying user electricity usage patterns and anomalies as claimed in claim 2 is characterized by: The Pearson correlation coefficient formula is used to express the correlation between characteristic indicators. The calculation formula is as follows: In the formula, n represents the number of samples, x i and i They represent variable characteristics respectively, x and y represent the means of variable characteristics respectively, and r represents the Pearson correlation coefficient, which ranges from [-1,1].
4. The improved Gaussian mixture analysis method for accurately identifying user electricity usage patterns and anomalies as claimed in claim 1 is characterized by: In S3, the expectation maximization algorithm and the log-likelihood function are used to determine the parameters of the Gaussian mixture model, including initializing the parameters of the expectation maximization algorithm. The expectation step and the maximization step in the expectation maximization algorithm are executed alternately, and the log-likelihood function is calculated after each execution.
5. The improved Gaussian mixture analysis method for accurately identifying user electricity usage patterns and anomalies as claimed in claim 4 is characterized by: The parameter determination process of the Gaussian mixture model is as follows: S3.1: Determine the number of components k of the Gaussian mixture model and initialize the parameters w of each component k 、u k and∑ k ; S3.2: Expectation step, for each feature index data point x i , calculate the posterior probability γ that belongs to each Gaussian component k (x i ), the calculation formula is: In the formula, N represents the total number of characteristic index data points, i represents the time period corresponding to the data point, and N(x i |u k ,∑k) represents the Gaussian probability density function; S3.3: Maximization step, based on the posterior probability γ calculated in the expectation step k (x i ), update the parameters w of each Gaussian component k 、u k and∑ k , the calculation formula is as follows: In the formula, u k is the mean, ∑k is the covariance matrix, w k is the weight or mixing coefficient; S3.4: Repeat S3.2 and S3.3 until the Gaussian mixture model converges. The convergence criterion is confirmed by the log-likelihood function value, and the calculation formula is as follows: When the change in the calculated value of the log-likelihood function is less than a preset range, the iteration is stopped to obtain the parameters of the Gaussian mixture model.
6. The improved Gaussian mixture analysis method for accurately identifying user electricity usage patterns and anomalies as claimed in claim 5 is characterized by: In S4, principal component analysis maps the original high-dimensional data to a low-dimensional space through linear transformation and retains the key information of the data. The mapping steps include: S4.1: First, standardize the data to avoid the influence of feature scale differences on the results. The standardization calculation formula is as follows: In the formula, x' represents the original data after standardization, x represents the original data, σ i is the standard deviation of the characteristic index; S4.2: For the original data matrix X, calculate its covariance matrix δ lm , the calculation formula is as follows: In the formula, δ lm represents the covariance of feature l and feature m, and T represents transposition; S4.3: The eigenvalues and eigenvectors of the covariance matrix are obtained by eigendecomposition. The calculation formula is as follows: d lm =λv In the formula, the eigenvalue λ is used to measure the variance of the principal component data, and the eigenvector v represents the direction of the principal component; S4.4: Sort the eigenvalues λ from large to small, and select the corresponding eigenvectors according to the sorting. Select the first q eigenvectors to form the projection matrix W, expressed as W = [v1, v2, ..., v q ], project the original data matrix X into the principal component space, expressed as Z = XW, where Z represents the data matrix after dimensionality reduction. By adjusting the size of q, the degree of dimensionality reduction can be controlled; S4.5: Visualize users’ electricity usage patterns through data mapping.
7. The improved Gaussian mixture analysis method for accurately identifying user power consumption patterns and power consumption anomalies as claimed in claim 5 is characterized by: In S5, sparse clustering clusters represent abnormal situations, and the likelihood probability and weighted log probability are used as measures of abnormality. The thresholds of the likelihood probability and weighted log probability are calculated, and the standard deviation method is used to determine the standard deviation value below the cluster mean to distinguish abnormal users from regular users; The likelihood probability calculation formula for sparse clustering is as follows: Weighted log probability: In the above formula, π is used together with the total number of data points to normalize the probability density function to ensure that the sum of the probabilities in the entire data space is 1. i -u k ) T Represents data point x i With the mean u k The transpose of the difference vector and (x i -u k ) are used together to calculate the Mahalanobis distance, exp represents the exponential function, To calculate the data point x i With the mean u k The exponential form of the Mahalanobis distance between .
Citation Information
Cited By
Power consumer clustering method and device, computer equipment and storage medium
CN120596959A