A user power consumption behavior data analysis method based on a CCASM algorithm
By combining the CCASM and CCA algorithms, the accuracy problem of clustering algorithms in user electricity consumption behavior analysis is solved, enabling comprehensive analysis of user electricity consumption behavior and proactive predictive maintenance, thereby improving the predictability and stability of the power grid.
Patent Information
- Application Number
- CN202211556407.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-06
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2042-12-06
AI Technical Summary
Existing technologies lack label references in user electricity consumption behavior analysis, resulting in inaccurate clustering algorithms, incomplete analysis of user electricity consumption correlations, and difficulty in achieving proactive predictive maintenance.
The CCASM algorithm is used to extract and analyze feature indicators of power grid data. The analytic hierarchy process and correlation matrix method are combined for primary and secondary classification. The CCA algorithm is used for multivariate correlation analysis. The optimal analysis result is selected through the optimal result algorithm and kernel CCA method to achieve comprehensive data collection and analysis of user electricity consumption behavior.
It enables comprehensive data collection and analysis of user electricity consumption behavior, improves the accuracy of electricity consumption behavior prediction and the timeliness of proactive maintenance, and enhances the predictability and stability of power grid operation.
Smart Images

Figure SMS_31 
Figure SMS_37 
Figure SMS_86
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of user power consumption, in particular to a user power consumption behavior data analysis method based on a CCASM algorithm. BACKGROUND
[0002] In recent years, with the rapid development of high-tech such as artificial intelligence, big data, smart grid, the industry prospect of big data has become a severe test and valuable opportunity for each power enterprise. Under the vigorous development of information technology, power big data fusion has a wide range of applications in power market, residential power consumption, power system safety evaluation, power grid disaster warning and other fields. It is necessary to analyze and invent the fusion technology of power grid user power big data. The key technology of big data technology in smart grid is data fusion technology. In the current state grid, whether it is the use information of power transmission and distribution or the specific information of user power consumption, it can be collected into the big data information database because of the computerization of office. Using data mining and data processing technology, these technologies can be quickly completed to maximize the normal operation of the power grid. Predictive analysis of user power consumption using big data, proactive maintenance, and avoidance of traditional post-fault maintenance to prevent sudden failures that cause life and production difficulties.
[0003] Traditional user power consumption behavior adopts clustering algorithm, but without labels as reference, it brings difficulties to classification prediction model in feature engineering and effect evaluation, so other methods need to be fused to better analyze user power consumption behavior.
[0004] A large building user behavior analysis method based on improved clustering fusion is disclosed in Chinese patent CN 109064353, which includes the following steps: (1) obtaining total load data and sub-metering data of the large building user to be analyzed; (2) constructing a clustering effect comprehensive evaluation index and selecting multiple high-quality clustering methods; (3) using the selected high-quality clustering method to cluster the total load data of the large building user to be analyzed to obtain different clustering results; (4) fusing the clustering results obtained by the high-quality clustering method to obtain the final power consumption mode; but this invention only analyzes and processes user power consumption itself, uses clustering fusion method, and the model is not accurate enough, the relevance of user power consumption is not considered, and the analysis of user power consumption behavior is not comprehensive.
[0005] The Chinese patent with publication number CN 114065819 A discloses a power consumption behavior analysis method and system based on multi-feature fusion and improved clustering, which includes: data cleaning of power consumption data; extracting power consumption features from the cleaned power consumption data based on load characteristic curve, signal processing and load feature construction; feature selection of power consumption features by recursive feature elimination, and feature fusion of selected power consumption features; classification of different power consumption behaviors based on the feature subsets obtained by fusion using an improved spectral clustering model; the improved spectral clustering model includes constructing an adjacency matrix of the feature subset based on an enhanced Gaussian kernel function, and classifying based on the adjacency matrix, and the enhanced Gaussian kernel function is an edge weight obtained according to the distance between sample points in the feature subset and a preset positive parameter, and the adjacency matrix is constructed based on the edge weight; although the patent adds feature extraction and fusion to analyze and process user power consumption behavior based on the clustering method, the feature extraction is also for data preparation for clustering. SUMMARY
[0006] The application provides a user power consumption behavior data analysis method based on a CCASM algorithm, which is used for correlation analysis of user data and non-power factors, realizes analysis of user power consumption behavior, and timely discovers relevant inspection problems, and converts passive inspection after failure into active predictive inspection.
[0007] To achieve the above purpose, the application adopts the following technical solutions:
[0008] A user power consumption behavior data analysis method based on a CCASM algorithm, comprising the following steps:
[0009] Step one: extracting and analyzing feature indicators of power grid data;
[0010] Step two: secondary classification of power consumption load based on feature indicators;
[0011] Step three: lever analysis of user load characteristics;
[0012] Step four: analyzing power consumption correlation by using a CCASM algorithm.
[0013] Further, the step one collects internal data of the power grid and related data, including user power consumption, power consumption behavior, power consumption frequency, causality of power consumption time, device operation data and normal device operation data when the user causes transformer and other equipment to trip and fail, forms heterogeneous data flow, and extracts and fuses the data by using an open source tool.
[0014] Further, the step two classifies the user load based on the characteristic index, selects the analytic hierarchy process to evaluate the first level index system, uses the correlation matrix method to evaluate the influence factors and the second level index system, determines the weight between each index and the influence factor, and performs clustering analysis on each type of user load based on the initial classification.
[0015] Further, the step three finds the classification containing the most users in the secondary classification results based on the user electricity load curve directly or indirectly obtained according to the user electricity characteristics, obtains the benchmark, compares with the load curves of other users, and evaluates the electricity habit according to the difference between the benchmark and the load curve.
[0016] Further, the step four uses the correlation analysis CCASM algorithm to solve the correlation analysis of two or more than two user electricity behaviors, and the steps are as follows:
[0017] 1) Multivariate clustering is performed on the electricity data X, and K categories of are obtained according to the clustering results. Then, the non-power data Y and Z are also divided into K categories, and K groups of are obtained.
[0018] 2) Correlation analysis is performed, and CCA analysis is performed on each group of , and then the unique and optimal single-day correlation coefficient and the weight matrix of each variable are obtained through the optimal result algorithm. ;
[0019] 3) All single-day correlation coefficients and weight matrices are arranged and summarized in time sequence to obtain the daily correlation and variable weight data between X, Y and Z for N consecutive days, and the optimal selection is performed on the analysis results.
[0020] Wherein, the multivariate set is the N-day electricity load data. is the daily load data with p variables, and ;
[0021] is the load data of a kind of electricity equipment.
[0022] is the daily load data of a kind of electricity equipment, and ;
[0023] The multivariate set is any non-power factor of the user.
[0024] is daily data of the factor with q variables;
[0025] is load data of the factor;
[0026] is daily data of a sub-factor of the factor, and ;
[0027] and completely correspond to each other, and represent daily data of the same day.
[0028] Compared with the prior art, the present application has the following beneficial effects:
[0029] 1) The method of the present application analyzes the correlation between user data and non-power factors, realizes comprehensive data collection and analysis of user power consumption behavior, and
[0030] 2) The CCA algorithm is used to analyze the user power consumption behavior from multiple angles, so that the user power consumption behavior is more comprehensive, and the proactive predictive maintenance is more timely. DETAILED DESCRIPTION
[0031] The specific embodiments of the present application will be further described below:
[0032] The user power consumption behavior data analysis method based on the CCASM algorithm comprises the following steps:
[0033] Step 1: Extract and analyze the feature indicators of the power grid data;
[0034] Step 2: Secondary classification of power consumption load based on feature indicators;
[0035] Step 3: Lever analysis of user load characteristics;
[0036] Step 4: Analysis of power consumption correlation using the CCASM algorithm.
[0037] Further, the step 1 collects internal data of the power grid and related data, including user power consumption, power consumption behavior, power consumption frequency, causality of power consumption time, device operation data and normal operation data of the device when the user causes the transformer and other equipment to trip, forms a heterogeneous data stream, and uses an open source tool to extract and fuse the data.
[0038] For the user power consumption of the low-carbon project, first, the actual user power consumption load and the user power consumption correlation behavior are collected;
[0039] Then, the feature index of user classification is obtained according to the user electricity correlation behavior; 15 influence factors are extracted, and the 15 influence factors are divided into three categories: hardware condition, energy saving and environmental protection consciousness and electricity habit, wherein the hardware condition includes: C1 member number, C2 house ownership, C3 working at home, C4 room number, C5 bedroom proportion, C6 appliance number, the energy saving and environmental protection consciousness includes: C7 energy saving consciousness, C8 environmental protection consciousness, C9 LED proportion, the electricity habit includes: C10 heating mode, C11 heating state, C12 price influence use, C13 cheap time use, C14 appliance power consumption, C15 timing device, the 15 influence factors uniformly constitute a primary index, the hardware condition, the energy saving and environmental protection consciousness and the electricity habit constitute three secondary indexes; the user extraction range is further expanded, 1024 user 1 year 0.5h electricity consumption data are extracted, one primary index, two secondary indexes and 8 influence factors are extracted, the primary index is actual electricity feature, the two secondary indexes are statistical feature and shape feature, wherein the statistical feature includes: daily average load, maximum load utilization hour, daily average peak load and daily average base load, the shape feature includes: skewness, kurtosis, rising time and falling time;
[0040] The maximum load utilization hour calculation formula is as follows:
[0041] (1)
[0042] Wherein, h is the maximum load utilization hour, L is the total daily load of the user, is the daily average peak load;
[0043] The skewness calculation formula is as follows:
[0044] (2)
[0045] Wherein, is the skewness, is the third central distance, is the standard deviation;
[0046] The kurtosis calculation formula is as follows:
[0047] (3)
[0048] Wherein, is the kurtosis, n is the data number, is the average value, is each data, D is the standard deviation;
[0049] The actual electricity load rising time and falling time are: five o'clock from the electricity load 0.2kWh starts to rise steadily, fluctuates from eleven o'clock to three o'clock in the afternoon, starts to fall at eight o'clock in the evening and continues to five o'clock in the morning.
[0050] Further, the step two is based on the characteristic index to firstly classify the user load, and the user load is classified by using the analytic hierarchy process to evaluate the first level index system, and the correlation matrix method is used to evaluate the influence of the second level index and the influence factors, and the correlation matrix is as shown in Table 1:
[0051]
[0052] Wherein, is the weight, A is the specific evaluation user, N is the characteristic index for evaluating the user; A is the sub first level index / evaluation scheme; is the evaluation result of a certain scheme for a specific index; is the comprehensive result of the evaluation result of each index of a certain individual; the correlation matrix method is used to calculate the weight relationship between the influence factors corresponding to each second level index, and finally the index value between them is obtained;
[0053] Secondly, the weight between each index and the influence factor is determined, the AHP method is used to divide the index into 3 categories, a hierarchical structure is formed, the relative importance of each element in the same level is determined through pairwise comparison, and then the corresponding weight value is obtained, and the formula is as follows:
[0054] (4)
[0055] Wherein, Z is the upper level index, is the relative weight of the sub first level criterion, the weight value of each level for the upper level is calculated, and thus the user characteristic is obtained;
[0056] According to the classification result of the user characteristic, the actual user load corresponding to each classification is statistically operated, the average load under each category is obtained, and the working day and the non-working day are analyzed from two angles, and the analysis result is as follows:
[0057] When the load is divided into 3 categories, 5 categories or 8 categories, the working day load exists the phenomenon that the loads are overlapped together, and for the non-working day load, the loads in the peak load period are not overlapped, which indicates that according to this classification method, the peak load of the non-working day can be distinguished, and with the increase of the user characteristic value, the load curve moves upward as a whole, which verifies that the user characteristic value is greatly affected by the actual power consumption information of the user; through the analysis of the classification results of the three user characteristics, the method of dividing into 3 categories is adopted, and three categories are divided into Table 2:
[0058]
[0059] Finally, on the basis of the initial classification, the user load of each class is subjected to cluster analysis; the user characteristic values of the experimental users conform to normal distribution, the k-means clustering algorithm is selected, each type under the initial classification is further classified, according to the calculation of the CH clustering evaluation function, the clustering result obtained by the application is set to 5, and the calculation formula is as follows:
[0060] (5)
[0061] Wherein, n is the number of clusters, i is the current divided class, trB(i) is the trace of the inter-class dispersion matrix, and trM(i) is the trace of the intra-class dispersion matrix.
[0062] Further, step three is directed to the user power consumption characteristics, the user power consumption load curve is taken as a reference, in the secondary classification result, the classification containing the most users is found, a benchmark is obtained, and the load curves of other users are compared, and the power consumption habits are evaluated according to the difference between the benchmark and the load curves;
[0063] According to the results of the previous two cluster analyses, the second class in the 5-class results is selected as the benchmark "user" representing the class, the specific load of the users contained is statistically calculated, a "virtual" benchmark is obtained, and in this way, each class under the initial classification is subjected to cluster analysis in turn, and the best benchmark corresponding to each class is found.
[0064] For the benchmark of the second class, the specific power consumption load of the workday and the non-workday can be ignored, only in summer, the workday load exceeds the power consumption load of the non-workday. And in the non-workday, the power consumption load is a gentle curve, and there is no fluctuation, which indicates that the users of the second class often go out on weekends.
[0065] Further, step four uses the correlation analysis CCASM algorithm to solve the correlation analysis of two or more than two user power consumption behaviors, and the steps are as follows:
[0066] 1) Multivariate clustering is performed on the power consumption data X, and K classes are obtained according to the clustering result , then the non-power data Y and Z are also divided into K classes, and K groups are obtained ;
[0067] 2) Correlation analysis is performed, and CCA analysis is performed on each group in , and then the unique and optimal single-day correlation coefficient and the weight matrix of each variable are obtained by the optimal result algorithm.
[0068] 3) All the single-day correlation coefficients and weight matrices are arranged and summarized in time sequence to obtain the daily correlation and variable weight data between X, Y and Z for N consecutive days, and the optimal selection is made for the analysis results.
[0069] wherein the multivariate set is the N-day electricity load data; is the daily load data with p variables, and ;
[0070] is the load data of an electricity-using device;
[0071] is the daily load data of an electricity-using device, and ;
[0072] the multivariate set is any non-power factor of the user;
[0073] is the daily data with q variables;
[0074] is the load data of the factor;
[0075] is the daily data of a sub-factor in the factor, and ;
[0076] and completely correspond to each other, representing the daily data of the same day.
[0077] The above is subjected to a fine-grained correlation analysis, based on two or more multivariate data sets, a typical correlation analysis method CCA is adopted in combination with an optimal result selection mechanism to analyze the correlation influence of multiple non-power factors on the user's electricity-using behavior, and this stage includes two steps: first, CCA or kernel CCA is performed on the single-day data in each category to obtain multiple sets of results; then, an optimal result selection algorithm is performed based on the category data to select the optimal CCA analysis result, and the process is as follows:
[0078] For each pair of single-day data , the CCA can find the linear relationship between and , assuming that there is a corresponding pair of projection weight matrices and , each variable can be projected into the typical space, and the formula is as follows:
[0079] (6)
[0080] (7)
[0081] (8)
[0082] (9)
[0083] where, and are the canonical weights, projecting and are the canonical components, is the transpose of matrix ;
[0084] The correlation between and is maximized, and the formula is as follows:
[0085] (10)
[0086] where, the maximum of and is the maximum canonical correlation coefficient, is the covariance matrix of , and is the autocovariance matrix of ;
[0087] In order to maximize the correlation coefficient , the denominator is fixed when solving, and the numerator is maximized, and the second team canonical weight formula is as follows:
[0088] (11)
[0089] (12)
[0090]
[0091] According to the optimization condition, the generalized eigenvalue is calculated by using the Lagrange algorithm:
[0092] (13)
[0093] (14)
[0094] At this time, the obtained is the second largest eigenvalue, and the canonical weight is obtained by performing CCA, and the number of canonical weights is constrained by the number of variables contained in and , which is less than or equal to , otherwise it will lead to overfitting, and the correlation coefficient of the three multivariate data sets is a The total number of the canonical weights obtained by CCA must be less than or equal to ;
[0095] The CCA is combined with the and function method, the data is projected to a high-dimensional feature space through the kernel of the polynomial, and then the CCA is performed in the new feature space, therefore, the objective function formula is as follows:
[0096] (15)
[0097] The corresponding optimization problem is as follows:
[0098] (16)
[0099] (17)
[0100] If the kernel function is reversible, the regularization is used to control the overfitting problem, when the data dimension is high, the objective function of the regularized CCA is as follows:
[0101] (18)
[0102] The objective function of the regularized kernel CCA is as follows: (19)
[0103] Finally, the optimal selection is made on the analysis result, and the formula is as follows:
[0104] (20)
[0105] The above examples are implemented on the premise of the technical scheme of the application, and detailed implementation modes and specific operation processes are given, but the protection scope of the application is not limited to the above examples. The methods used in the above examples are all conventional methods unless otherwise specified.
Claims
1. A user electricity consumption behavior data analysis method based on a CCASM algorithm, characterized in that, Comprise the following steps: Step one: feature index extraction and analysis of power grid data; collect internal data and related data, including user power consumption, power consumption behavior, power consumption frequency, power consumption time causality, user reason leading to transformer and other equipment tripping fault device operation data and device normal operation data, form heterogeneous data stream, adopt open source tool to extract and fuse data; Step two: secondary classification of power load based on feature index; Step three: leverage analysis of user load characteristics; Step four: use CCASM algorithm to analyze power consumption correlation: Use correlation analysis CCASM algorithm to solve the correlation analysis of two or more than two user power consumption behaviors, the steps are as follows: 1) Multivariate clustering on the electricity data X, and get K clusters of Then, non-power data Y and Z are also divided into K categories, and K groups of , where ; 2) correlation analysis is performed on each group of CCA analysis is performed, and then the unique and optimal single-day correlation coefficient is obtained by the optimal result algorithm and the weight matrix of each variable ; 3) arrange and summarize all single-day correlation coefficients and weight matrices in time sequence to obtain the daily correlation and variable weight data between X, Y and Z for N consecutive days, and make optimal selection for the analysis results; wherein the multivariate set is the N-day electricity load data; is daily load data with p variables, and ; load data of an electric device; is daily load data of an electric device, and ; Multivariate set For any non-electric factor of the user; is daily data with q variables; a load data for this factor; day data for one of the factors, and ; and completely corresponds, indicates the same day data.
2. The method for analyzing user's electricity consumption behavior data based on CCASM algorithm according to claim 1, characterized in that, The step two is based on the feature index to classify the user load, the analytic hierarchy process is used to evaluate the first level index system, the correlation matrix method is used to evaluate the influence factors and the second level index system, the weight between each index and the influence factor is determined, and the clustering analysis is carried out on each kind of user load based on the initial classification.
3. The method of claim 1, wherein the method is characterized by: The step three is directly or indirectly obtained from the user power consumption characteristics, the user power consumption load curve is taken as the benchmark, in the secondary classification result, the most common classification of the user is found, the benchmark is obtained, and the load curve of other users is compared, the power consumption habit is evaluated according to the difference between the benchmark.
Citation Information
Patent Citations
Power consumption behavior analysis method and system based on multi-feature fusion and improved spectral clustering
CN114065819A
Method for analyzing power consumption behavior of user
CN108510165A
Method for mining dominant influence factors of power utilization behavior of user
CN108596227A