Subspace clustering-CatBoost-based electrical load analysis method, system and device, and storage medium

Through the power load analysis method combined with subspace clustering and CatBoost algorithm, the problems of low clustering accuracy and low computing efficiency in the existing technology are solved, and efficient and accurate load mode classification and analysis are achieved, which improves the intelligence level of electricity prediction and energy utilization efficiency.

CN120235344APending Publication Date: 2025-07-01INNER MONGOLIA ELECTRIC POWER TRADING CENT CO LTD

Patent Information

Application Number
CN202510303944.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing power load analysis methods have low clustering accuracy, low computing efficiency, and insufficient ability to identify load fluctuations, which makes it difficult to meet the needs of fast response in smart grid environments.

Method used

The method of combining subspace clustering and CatBoost algorithm is used to standardize the electrical load data, and cluster labels are generated using the subspace clustering algorithm, and classified through the CatBoost classifier to analyze the fluctuations of electrical loads.

Benefits of technology

It improves the accuracy and calculation efficiency of power load analysis, enhances the interpretability and intelligence of the model, and can more accurately characterize complex load modes, optimize power consumption management, and improve energy utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235344A_ABST
    Figure CN120235344A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of analysis of load fluctuation behaviors, and particularly relates to an electrical load analysis method, system and device based on subspace clustering-CatBoost and a storage medium. Electrical load data are collected and subjected to standardization processing; clustering the standardized data through a subspace clustering algorithm to generate a clustering label; the clustering labels are classified through a CatBoost classifier, and the fluctuation behavior of the electrical load is analyzed; by performing standardization processing on the electric load data, the inconsistency of data scales is eliminated, the stability of data input is improved, different types of load modes are automatically identified through a subspace clustering technology, the classification of the load data is more accurate, meanwhile, the calculation efficiency of data processing is improved, and through a CatBoost classifier, the calculation efficiency of data processing is improved. Different types of load modes are accurately classified, the intelligent level of electricity consumption prediction is improved, and the interpretability of the model is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of analyzing load fluctuation behavior, and particularly relates to a method, system, device and storage medium for analyzing electricity load based on subspace clustering - CatBoost. Background Art

[0002] In modern power systems, load forecasting and analysis are key links in power dispatching, demand - side management, energy efficiency optimization, and power market transactions. With the development of smart grids, traditional load forecasting methods based on statistical analysis have gradually been replaced by methods based on machine learning and big data analysis. These emerging methods rely on large - scale data collection, intelligent perception, and computer modeling capabilities, and can more accurately depict users' electricity consumption behavior patterns and achieve efficient power resource allocation.

[0003] Although the above - mentioned load analysis technologies have played a certain role in applications such as power dispatching and market forecasting, there are still the following problems: Traditional clustering methods often assume that data has a uniform distribution in the entire feature space, and it is difficult to effectively capture the complex relationships of load data in different dimensions. When facing non - linear and multi - dimensional complex load data, it is often difficult to accurately capture the internal laws of load fluctuations. Especially in the case of drastic load changes or anomalies, deep learning methods perform well in load forecasting tasks, but their high computational complexity makes the model training and inference processes require a large amount of computing resources, which limits their real - time application in large - scale power load analysis. Especially in the smart grid environment, the power system needs to be able to quickly respond to load fluctuations, and models based on complex neural networks are difficult to meet the actual application requirements.

[0004] In view of the above problems, the present invention combines subspace clustering and the CatBoost algorithm to propose a more refined and accurate method for analyzing electricity load. First, the load data is divided by the subspace clustering method, so that different types of load patterns can be more accurately identified in different subspaces. Secondly, the CatBoost classifier is used to further classify the clustering results to improve the ability to identify complex load fluctuation patterns. Summary of the Invention

[0005] In view of the problems existing in the above - mentioned prior art, the present invention is proposed.

[0006] Therefore, the technical problem solved by the present invention is: Existing electricity load analysis methods have low clustering accuracy, low computational efficiency, insufficient ability to identify load fluctuation behavior, and the problem of how to achieve efficient and accurate load pattern classification and analysis.

[0007] To solve the above technical problems, the present invention provides the following technical solution, a power consumption load analysis method based on subspace clustering - CatBoost, including: collecting power load data and performing normalization processing on the data; clustering the normalized data through a subspace clustering algorithm to generate clustering labels; classifying the clustering labels through a CatBoost classifier to analyze the power consumption load fluctuation behavior.

[0008] As a preferred solution of the power consumption load analysis method based on subspace clustering - CatBoost described in the present invention, wherein: the normalization processing of the data includes converting the collected power load data into a matrix format and representing it as:

[0009] X ∈ R n×d

[0010] where X is the power consumption load matrix, R is the set of real numbers, n is the number of power consumption load curves, and d is the number of time nodes for detecting power load data;

[0011] Sparsely represent the power consumption load curve data, and represent it as:

[0012]

[0013] where A ij is the sparse representation matrix of the power consumption load data, x i is the i-th power consumption load curve data, and x j is the j-th power consumption load curve data;

[0014] By adding a norm to measure the sparsity of the data matrix, it is represented as:

[0015]

[0016] where ||·|| F is the F-norm, and λ is the regularization parameter.

[0017] As a preferred solution of the power consumption load analysis method based on subspace clustering - CatBoost described in the present invention, wherein: the clustering of the normalized data through a subspace clustering algorithm includes constructing a correlation matrix through a sparse matrix and representing it as:

[0018] S = |A| + |A T |

[0019] where S is the correlation matrix, and the element S ij in S represents the similarity between the i-th power consumption load curve data and the j-th power consumption load curve data, A is the sparse representation matrix, and A T is the transpose matrix of A;

[0020] The subspace clustering algorithm is used to partition the data points of the correlation matrix and classify the electricity load curves.

[0021] As a preferred solution of the electricity load analysis method based on subspace clustering - CatBoost according to the present invention, wherein: the subspace clustering algorithm includes extracting the low - dimensional feature representation of the subspace, clustering the low - dimensional representation, and constructing the adjacency matrix expressed as:

[0022]

[0023] where W ij is the degree - similarity per unit electricity between vertices s i and s j , s i and s j are the i - th and j - th columns of the correlation matrix of the electricity load data, and dis(s i , s j ) is the similarity measure between vertices;

[0024] Construct the degree matrix expressed as:

[0025]

[0026] where D is the degree matrix;

[0027] Calculate the graph Laplacian matrix expressed as:

[0028] L = D - W

[0029] where L is the graph Laplacian matrix.

[0030] As a preferred solution of the electricity load analysis method based on subspace clustering - CatBoost according to the present invention, wherein: the clustering of the low - dimensional representation includes calculating the eigenvalues and eigenvectors of the graph Laplacian matrix to obtain the eigenmatrix expressed as:

[0031] Lu m = λ m u m , m = 1, 2, …, k1

[0032]

[0033] where Lu m is the graph Laplacian matrix of the eigenvector u m , u m is the eigenvector of the electricity load pattern, λ m is the eigenvalue corresponding to the eigenvector, U is the eigenmatrix, and k1 is the number of eigenvalues;

[0034] Through k-means clustering, the row vectors of the feature matrix are divided into k2 clusters to obtain clustering labels.

[0035] As a preferred solution of the method for analyzing electricity load based on subspace clustering-CatBoost according to the present invention, wherein: the classification of the clustering labels by the CatBoost classifier includes constructing a strong learner by accumulating weak learners, which is expressed as:

[0036]

[0037] Wherein, F(x) is the predicted value of the electricity load, K is the number of iterations of the CatBoost classifier, and f k (x,θ k is the predicted value of the kth gradient boosting decision tree on the input x;

[0038] During the iteration of the learner, the negative gradient of the optimized loss function is expressed as

[0039]

[0040] Wherein, g i is the negative gradient, y i is the clustering label, F t-1 (x i ) is the cumulative predicted value of the electricity load after the (t-1)th iteration, and L(.) is the electricity load loss function;

[0041] Fitting the learner with the negative gradient, the updated model is expressed as:

[0042] F t (x) = F t-1 (x) + ηf t (x)

[0043] Wherein, F t (x) is the predicted value of the electricity load pattern after the tth iteration, η is the learning rate, and f t (x) is the decision tree model during the tth training process.

[0044] As a preferred solution of the method for analyzing electricity load based on subspace clustering-CatBoost according to the present invention, wherein: the analysis of the electricity load fluctuation behavior includes using the trained classifier to standardize the electricity load within the preset time of the region and then sending it into the classifier for classification, and generating a portrait of the electricity load fluctuation behavior on the electricity side of the region by analyzing the frequency distribution of the classification results.

[0045] As a preferred solution of the power consumption load analysis system based on subspace clustering - CatBoost according to the present invention, it includes: a data standardization processing module, a standardized data clustering module, and a module for analyzing the power consumption load fluctuation behavior; the data standardization processing module includes a data acquisition module and a data standardization module. The data acquisition module is used to extract the power consumption load data of the previous 24 hours from the power consumption dataset, and the data standardization module is used to construct a power consumption load data matrix and find the sparse representation of the power consumption load data matrix; the standardized data clustering module includes a low-dimensional feature representation module and a clustering label generation module. The low-dimensional feature representation module is used to extract the low-dimensional feature representation of the subspace through the correlation matrix, and the clustering label generation module is used to calculate the eigenvalues and eigenvectors of the graph Laplacian matrix to obtain a feature matrix, divide the row vectors of the feature matrix into k clusters to obtain clustering labels; the module for analyzing the power consumption load fluctuation behavior includes a clustering label classification module and a power consumption load fluctuation behavior analysis module. The clustering label classification module is used to construct a strong learner by accumulating weak learners, and approximate the target distribution by iteratively optimizing the loss function through the CatBoost classifier. The power consumption load fluctuation behavior analysis module is used to use the trained classifier to analyze the frequency distribution of the classification results and generate a portrait of the power consumption side load fluctuation behavior of the region.

[0046] A computer device includes a memory and a processor. The memory stores a computer program. It is characterized in that when the processor executes the computer program, it implements the steps of any one of the methods in the power consumption load analysis method based on subspace clustering - CatBoost.

[0047] A computer-readable storage medium stores a computer program. It is characterized in that when the computer program is executed by a processor, it implements the steps of any one of the methods in the power consumption load analysis method based on subspace clustering - CatBoost.

[0048] Advantages of the present invention: By standardizing the electrical load data, the inconsistency of data scales is eliminated, the stability of data input is improved, providing a more reliable basis for subsequent subspace clustering and classification. This not only ensures the accuracy of data analysis but also enhances the portability of the entire method, enabling it to be applicable to different scales of electricity users. Through subspace clustering technology, different categories of load patterns are automatically identified, making the classification of load data more accurate. At the same time, the computational efficiency of data processing is improved. Compared with traditional clustering methods, this method can more accurately depict complex load patterns, ensuring the efficiency and reliability of classification tasks. Through the CatBoost classifier, different categories of load patterns are accurately classified, and the fluctuation behavior of electrical loads is deeply analyzed, improving the intelligent level of electricity consumption prediction. This not only improves the recognition accuracy but also enhances the interpretability of the model, enabling the power dispatching department to optimize electricity management based on the classification results and improve energy utilization efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0050] Figure 1 Schematic flowchart of the electricity load analysis method based on subspace clustering - CatBoost provided by the first embodiment of the present invention.

[0051] Figure 2 Schematic diagram for determining the optimal number of clusters by the elbow method of the electricity load analysis method based on subspace clustering - CatBoost provided by the second embodiment of the present invention.

[0052] Figure 3 Comparison chart of cluster center curves of the electricity load analysis method based on subspace clustering - CatBoost provided by the second embodiment of the present invention.

[0053] Figure 4 Distribution diagrams of load curves for each category of the electricity load analysis method based on subspace clustering - CatBoost provided by the second embodiment of the present invention.

[0054] Figure 5 Confusion matrix diagram of the CatBoost classifier of the electricity load analysis method based on subspace clustering - CatBoost provided by the second embodiment of the present invention.

[0055] Figure 6ROC curve graphs of various categories of the power consumption load analysis method based on subspace clustering - CatBoost provided for the second embodiment of the present invention.

[0056] Figure 7 Typical power consumption behavior portrait graph of the power consumption load analysis method based on subspace clustering - CatBoost provided for the second embodiment of the present invention.

[0057] Figure 8 User power consumption behavior analysis radar graph of the power consumption load analysis method based on subspace clustering - CatBoost provided for the second embodiment of the present invention. Detailed implementation manners

[0058] To make the above - mentioned objects, features, and advantages of the present invention more obvious and understandable, the following provides a detailed description of the specific implementation manners of the present invention in conjunction with the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0059] Example 1, referring to Figure 1 , which is the first embodiment of the present invention. This embodiment provides a power consumption load analysis method based on subspace clustering - CatBoost, including:

[0060] S1: Collect power load data and perform standardization processing on the data.

[0061] It should be noted that performing standardization processing on the data includes converting the collected power load data into a matrix format, expressed as:

[0062] X ∈ R n×d

[0063] where X is the power consumption load matrix, R is the set of real numbers, n is the number of power consumption load curves, and d is the number of time nodes for detecting power consumption load data;

[0064] Sparsely represent the power consumption load curve data, expressed as:

[0065]

[0066] where A ij is the sparse representation matrix of the power consumption load data, x i is the i - th power consumption load curve data, and x j is the j - th power consumption load curve data;

[0067] Measure the sparsity of the data matrix by adding a norm, expressed as:

[0068]

[0069] Among them, ||·|| F is the F-norm, and λ is the regularization parameter.

[0070] It should also be noted that the load data for the previous 24 hours is extracted from the electricity consumption dataset, and the load data is represented in matrix format. The number of time nodes d is 24 for one day, and each row represents a data point. Assuming that the data points can be distributed in multiple low-dimensional subspaces, and there is a certain similarity between the data points within the same subspace. Through the sparse representation method, the internal correlation of the load data is mined, so that each data point can be expressed by a small number of key data points, thereby reducing data redundancy. For a given load curve data x i , an attempt is made to represent it as a sparse linear combination of other load data, A ij is the sparse representation matrix of the load data, indicating how the load curve data x i is linearly represented by other data points. To ensure the sparsity of this representation, in this embodiment, it is obtained by adding a norm to the optimization problem. ||·||1 in the formula represents the 1-norm, which is used to ensure the sparsity of the metric matrix.

[0071] S2: Cluster the standardized data through the subspace clustering algorithm to generate clustering labels.

[0072] It should be noted that clustering the standardized data through the subspace clustering algorithm includes constructing a correlation matrix through a sparse matrix, which is expressed as:

[0073] S = |A| + |A T |

[0074] Among them, S is the correlation matrix, and the element S ij in S represents the similarity between the i-th electricity load curve data and the j-th electricity load curve data. A is the sparse representation matrix, and A T is the transpose matrix of A;

[0075] Perform subspace partitioning on the data points of the correlation matrix through the subspace clustering algorithm to classify the electricity load curves.

[0076] It should also be noted that after the sparse representation calculation of the load data is completed, the load data correlation matrix S is an important tool for capturing the mutual relationship between data points. Through the sparse matrix A, the correlation matrix is constructed to reflect the similarity between data points. |A| represents the absolute value matrix of the sparse matrix, which eliminates the influence of negative weights on similarity. The introduction of |A T | ensures that S is symmetric, providing support for subsequent spectral clustering.

[0077] Further, the subspace clustering algorithm includes extracting the low-dimensional feature representation of the subspace, clustering the low-dimensional representation, and constructing an adjacency matrix expressed as:

[0078]

[0079] where W ij is the per-degree similarity between vertices s i and s j , s i and s j are the i-th and j-th columns of the correlation matrix of the electricity load data, and dis(s i , s j ) is the similarity measure between vertices;

[0080] Construct the degree matrix expressed as:

[0081]

[0082] where D is the degree matrix;

[0083] Calculate the graph Laplacian matrix expressed as:

[0084]

[0085] where L is the graph Laplacian matrix.

[0086] It should also be noted that after constructing the load data correlation matrix S, the spectral clustering method can be used to partition the data points into subspaces, that is, to classify the load curves (or data points). Spectral clustering extracts the low-dimensional feature representation of the subspace through the correlation matrix and then clusters the low-dimensional representation.

[0087] First, a graph G=(S, E) needs to be constructed from the data points, where each data point corresponds to a vertex s∈S, and the edge e∈E is defined based on a certain similarity measure, such as the Euclidean distance or the Gaussian kernel function, to construct an adjacency matrix, whose element W ij represents the similarity between vertices, and the graph Laplacian matrix reflects the topological structure of the graph, especially the connection relationship between vertices.

[0088] Furthermore, clustering the low-dimensional representation includes calculating the eigenvalues and eigenvectors of the graph Laplacian matrix to obtain the eigenmatrix expressed as:

[0089] Lu m =λ m u m , m = 1, 2, …, k1

[0090]

[0091] where Lum is the eigenvector u m of the graph Laplacian matrix, where u m is the eigenvector of the electricity load pattern, and λ m is the eigenvalue corresponding to the eigenvector, U is the eigenmatrix, and k1 is the number of eigenvalues;

[0092] Through k-means clustering, the row vectors of the eigenmatrix are divided into k2 clusters to obtain cluster labels.

[0093] Calculate the eigenvalues and eigenvectors of the graph Laplacian matrix L. The result of eigenvalue decomposition is an eigenvalue vector and the corresponding eigenvector matrix. Select the eigenvectors corresponding to the first k1 smallest eigenvalues of L to form a new eigenmatrix U. The eigenvectors provide a representation of the data in the low-dimensional space and retain the structural information of the original graph. The selection of eigenvectors and the number k1 are usually determined based on the specific requirements of the problem or through cross-validation. Use the k-means clustering algorithm to cluster the row vectors of the eigenmatrix U. Each row vector represents the position of a data point in the low-dimensional space. Through k-means clustering, these row vectors can be divided into k2 clusters, thereby realizing the clustering of the original data points. The final clustering result reflects the distribution and similarity of the data points in the original space. Through the above steps in this embodiment, the subspace clustering method can effectively group the data to obtain cluster labels.

[0094] S3: Classify the cluster labels through the CatBoost classifier to analyze the electricity load fluctuation behavior.

[0095] It should be noted that classifying the cluster labels through the CatBoost classifier includes constructing a strong learner by accumulating weak learners, which is expressed as:

[0096]

[0097] where F(x) is the predicted value of the electricity load, K is the number of iterations of the CatBoost classifier, and f k (x, θ k ) is the predicted value of the k-th gradient boosting decision tree on the input x;

[0098] During the iteration of the learner, the negative gradient of the optimized loss function is expressed as

[0099]

[0100] where g i is the negative gradient, y i is the cluster label, and F t-1 (x iis the cumulative predicted value of electricity consumption load after the (t - 1)-th round of iteration, and L(.) is the electricity consumption load loss function;

[0101] Fit the learner to the negative gradient, and the updated model is expressed as:

[0102] F t F(x) = F t-1 (x) + ηf t (x)

[0103] where, F t (x) is the predicted value of the electricity consumption load pattern after the t-th round of iteration, η is the learning rate, and f t (x) is the decision tree model during the t-th round of training.

[0104] It should also be noted that by classifying the clustering labels through the CatBoost classifier, accurate identification of different types of electricity consumption load patterns can be achieved. CatBoost can automatically capture the non-linear features in the load data, making the classification more accurate. CatBoost adopts a class balance processing method, which can avoid the model from biasing towards certain data classes and improve the classification stability. Through multiple rounds of iterative training, CatBoost can continuously adjust the weights, enabling the model to maintain a high classification accuracy in different electricity consumption load environments. Using the Gradient Boosting Decision Tree (GBDT) for classification, the loss function is optimized in each round of training, and the model is updated through the negative gradient. Each round of decision tree learns the misclassified data to improve the prediction ability of the overall model. CatBoost adopts a regularization method, making the learner in each round not overfit the training data and enhancing the stability of the model.

[0105] The negative gradient is used to indicate the direction that needs to be optimized in each round of training, enabling the model to converge to the optimal solution faster. Compared with the method of directly updating the weights, the negative gradient method can optimize the model parameters more efficiently, accelerate the model training, and reduce the computational redundancy, so that a better classification effect can still be obtained with fewer computational resources. The learning rate η ensures that the model will not fluctuate violently during the update process, thereby improving the stability of the classifier. Through step-by-step optimization, the model can continuously adjust the classification boundary, making the final prediction result more accurate. In different power users or time periods, the load pattern may change. The gradually updated model can better adapt to this change and improve the generalization ability of the classifier.

[0106] Further, analyze the electricity consumption load fluctuation behavior, including using the trained classifier to standardize the electricity consumption load within the preset time in the region and then sending it into the classifier for classification. By analyzing the frequency distribution of the classification results, a portrait of the electricity consumption side load fluctuation behavior in the region is generated.

[0107] It should also be noted that after standardizing the electricity consumption load of enterprises in a certain area over a period of time using the trained classifier, it is sent to the classifier for classification. By analyzing the frequency distribution of the classification results, a load fluctuation behavior portrait of the electricity consumption side in this area is generated. This portrait can reveal the characteristics and changing trends of the electricity consumption patterns of enterprises in this area, providing data support for power dispatching and demand response.

[0108] Example 2, referring to Figures 2 - 8 , which is an embodiment of the present invention, provides a method for analyzing electricity consumption load based on subspace clustering - CatBoost. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0109] To verify the effectiveness of the method, measured load curves on the electricity consumption side are extracted from a certain area for case analysis. There are a total of 10,000 daily load curves (using the maximum value of all daily load curves as the normalization benchmark, and implementing maximum value normalization for the daily load curves). The first 3,000 data are extracted for subspace clustering to obtain daily load pattern labels.

[0110] For example, according to the elbow method, as Figure 2 shown, after obtaining the classification categories, assuming it is 6, then the first 3,000 load curves are divided into 6 types of typical curves using subspace clustering, as Figure 3 shown, and each group of load curves is labeled. If a load curve belongs to the third category, then its label is 3. Through the subspace clustering process, these 3,000 data become labeled data, and the clustering result is as Figure 4 shown.

[0111] Then, the CatBoost algorithm is used for these 3,000 labeled data, and training is carried out in the ratio of 8:2 for the training set and the test set. For example, there are 2,400 data in the training set and 600 data in the test set. Training is carried out using the CatBoost algorithm to obtain training weights and save the trained model; the remaining 7,000 data are discriminated using the trained model. For example, when the 3001st load curve model is input and fed into the model, it returns 2, indicating that the 3001st data belongs to the second type of typical curve, and all the remaining load curves are operated on in turn. The reason for such a design is that when the data is too large, the clustering effect is not good, so it is segmented, and the trained model (i.e., the classifier) can also be reused multiple times, and the training results are as Figure 5 、 Figure 6 shown.

[0112] After subspace clustering and CatBoost classification, 10,000 labeled load data are obtained. By using subspace clustering and CatBoost classification, user behavior portraits are constructed. For example, these 10,000 data may be the data of 100 enterprises in 100 days. For a certain enterprise A, there are 100 load curves in total. Among these 100 load curves, 20, 30, 10, 10, 10, and 20 belong to the 1st, 2nd, 3rd, 4th, 5th, and 6th categories respectively. Then, the vector [0.2, 0.3, 0.1, 0.1, 0.1, 0.2], that is, the frequency of each type of load, is used to represent the behavior portrait of enterprise A in these 100 days, as Figure 7 shown.

[0113] Finally, a frequency vector is obtained for each enterprise to represent its behavior portrait. For example, in the above example, there are 100 enterprise behavior portraits, that is, 100 frequency vectors. Then, subspace clustering is used again for these 100 frequency vectors. Assuming that the optimal clustering is 4 (obtained by the elbow method), the overall typical four types of electricity consumption behaviors in this area are obtained, as Figure 8 shown.

[0114] Example 3 is an embodiment of the present invention, which provides an electricity load analysis system based on subspace clustering - CatBoost, including a data normalization processing module 100, a normalized data clustering module 200, and an analysis module 300 for electricity load fluctuation behavior.

[0115] Among them, S4: The data normalization processing module 100 includes a data acquisition module 101 and a data normalization module 102. The data acquisition module 101 is used to extract the electricity load data of the previous 24 hours from the electricity dataset. The data normalization module 102 is used to construct an electricity load data matrix and find the sparse representation of the electricity load data matrix.

[0116] It should also be noted that the data normalization module 102 transfers the normalized data matrix to the normalized data clustering module 200. The sparse representation matrix is used to calculate the correlation between data points and provide data support for constructing a low - dimensional feature representation.

[0117] S5: The normalized data clustering module 200 includes a low - dimensional feature representation module 201 and a clustering label generation module 202. The low - dimensional feature representation module 201 is used to extract the low - dimensional feature representation of the subspace through the correlation matrix. The clustering label generation module 202 is used to calculate the eigenvalues and eigenvectors of the graph Laplacian matrix to obtain a feature matrix, and divide the row vectors of the feature matrix into k clusters to obtain clustering labels.

[0118] It should also be noted that the data clustering module 200 takes the standardized data matrix as input to obtain clustering labels, and the clustering label module 202 passes the clustering labels to the module 300 for analyzing the fluctuation behavior of electricity load.

[0119] S6: The module 300 for analyzing the fluctuation behavior of electricity load includes a clustering label classification module 301 and an electricity load fluctuation behavior analysis module 302. The clustering label classification module 301 is used to construct a strong learner by accumulating weak learners and approximate the target distribution by iteratively optimizing the loss function through a CatBoost classifier. The electricity load fluctuation behavior analysis module 302 is used to analyze the frequency distribution of the classification results by using the trained classifier and generate a portrait of the electricity load fluctuation behavior on the user side of the region.

[0120] It should also be noted that the clustering label classification module 301 takes the clustering labels as input, trains the model, and passes the trained classification model to the electricity load fluctuation behavior analysis module 302. The electricity load fluctuation behavior analysis module 302 depends on the classification results of the clustering label classification module 301 to calculate the distribution of different electricity load patterns in different time periods and form a portrait of the load behavior.

[0121] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.

[0122] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.

[0123] More specific examples (a non-exhaustive list) of computer-readable media include the following: electrical connections (electronic devices) having one or more wirings, portable computer diskettes (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.

[0124] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0125] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. The power load analysis method based on subspace clustering-CatBoost is characterized by: include: Collect electric load data and standardize the data; Cluster the standardized data using the subspace clustering algorithm to generate cluster labels; The cluster labels are classified by CatBoost classifier to analyze the power load fluctuation behavior.

2. The power load analysis method based on subspace clustering-CatBoost according to claim 1, characterized in that: The data standardization process includes converting the collected electric load data into a matrix format as follows: X∈R n×d Among them, X is the power load matrix, R is a real number set, n is the number of power load curves, and d is the number of time nodes for detecting power load data; The power load curve data is sparsely represented as follows: Among them, A ij is the sparse representation matrix of the power load data, x i is the i-th power load curve data, x j is the j-th electricity load curve data; The sparsity of the data matrix is ​​measured by adding the norm, expressed as: Among them, ||·|| F is the F-norm, and λ is the regularization parameter.

3. The power load analysis method based on subspace clustering-CatBoost according to claim 1 or 2, characterized in that: The method of clustering the standardized data by using a subspace clustering algorithm includes constructing a correlation matrix by using a sparse matrix, which is expressed as: S=|A|+|A T | Among them, S is the correlation matrix, and the element S in S ij represents the similarity between the i-th power load curve data and the j-th power load curve data, A is a sparse representation matrix, A T is the transposed matrix of A; The subspace clustering algorithm is used to divide the data points of the correlation matrix into subspaces and classify the power load curves.

4. The power load analysis method based on subspace clustering-CatBoost as claimed in claim 3, characterized in that: The subspace clustering algorithm includes extracting low-dimensional feature representations of subspaces, clustering the low-dimensional representations, and constructing an adjacency matrix representation as follows: Among them, W ij For vertex s i and j The electrical similarity between i and j is the i-th and j-th columns of the power load data correlation matrix, dis(s i ,s j ) is the similarity measure between vertices; The constructed degree matrix is ​​expressed as: Where D is the degree matrix; The computation graph Laplacian matrix is ​​expressed as: L=DW Where L is the graph Laplacian matrix.

5. The power load analysis method based on subspace clustering-CatBoost according to claim 1 or 4, characterized in that: The clustering of the low-dimensional representation includes calculating the eigenvalues ​​and eigenvectors of the graph Laplacian matrix, and obtaining a characteristic matrix represented as: Lu m =λ m you m ,m=1,2,…,k1 Among them, Lu m is the feature vector u m The graph Laplacian matrix, u m is the characteristic vector of the power load pattern, λ m is the eigenvalue of the corresponding eigenvector, U is the characteristic matrix, and k1 is the number of eigenvalues; Through k-means clustering, the row vectors of the feature matrix are divided into k2 clusters to obtain cluster labels.

6. The method for analyzing power load based on subspace clustering-CatBoost according to claim 5, characterized in that: The classification of cluster labels by CatBoost classifier includes constructing a strong learner by accumulating weak learners, which is expressed as: Among them, F(x) is the predicted value of power load, K is the number of iterations of CatBoost classifier, and f k (x,θ k ) is the predicted value of the kth gradient boosting decision tree on the input x; During the learner iteration process, the negative gradient of the optimized loss function is expressed as, Among them, g i is the negative gradient, y i is the cluster label, F t-1 (x i ) is the cumulative predicted value of power load after the t-1th iteration, and L(.) is the power load loss function; Fitting the learner with negative gradient, the updated model is expressed as: F t (x)=F t-1 (x)+ηf t (x) Among them, F t (x) is the predicted value of the power load pattern after the tth iteration, η is the learning rate, f t (x) is the decision tree model in the tth round of training.

7. The power load analysis method based on subspace clustering-CatBoost according to claim 1 or 6, characterized in that: The analysis of power load fluctuation behavior includes using a trained classifier to standardize the power load in a region within a preset time, and then sending the standardized power load to the classifier for classification, and generating a power load fluctuation behavior portrait of the region by analyzing the frequency distribution of the classification results.

8. The power load analysis system based on subspace clustering-CatBoost is characterized by: It includes a data standardization processing module (100), a standardized data clustering module (200), and a power load fluctuation behavior analysis module (300); The data standardization processing module (100) comprises a data acquisition module (101) and a data standardization module (102), wherein the data acquisition module (101) is used to extract the power load data of the previous 24 hours from the power consumption data set, and the data standardization module (102) is used to construct a power load data matrix and find a sparse representation of the power load data matrix; The standardized data clustering module (200) comprises a low-dimensional feature representation module (201) and a clustering label generation module (202), wherein the low-dimensional feature representation module (201) is used to extract the low-dimensional feature representation of the subspace through the correlation matrix, and the clustering label generation module (202) is used to calculate the eigenvalues ​​and eigenvectors of the graph Laplacian matrix to obtain a feature matrix, and divide the row vectors of the feature matrix into k clusters to obtain clustering labels; The power load fluctuation behavior analysis module (300) comprises a clustering label classification module (301) and a power load fluctuation behavior analysis module (302). The clustering label classification module (301) is used to construct a strong learner by accumulating weak learners, and to approximate the target distribution by iteratively optimizing the loss function through the CatBoost classifier. The power load fluctuation behavior analysis module (302) is used to use the trained classifier to analyze the frequency distribution of the classification results and generate a load fluctuation behavior portrait of the power consumption side of the region.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Load curve data visualization method based on combination of supervised and unsupervised algorithms

    CN110321390A

  • Aerial target activity rule prediction method

    CN114330509A

  • Unbalanced industrial load identification method

    CN116561658A

  • Data management method and system for reducing computing power occupation

    CN119440776A

Cited By

  • Vehicle auxiliary driving method based on driver personalized fatigue detection

    CN121354073A