A power load behavior clustering method and system based on a deep clustering coupled model
Patent Information
- Application Number
- CN202611282768.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-24
- Publication Date
- 2026-09-18
AI Technical Summary
该发明用以解决现有技术中负荷数据降维导致的信息丢失及噪声样本造成的负荷类型识别效果不够理想问题
(1)传统基于深度学习的聚类方法通常采用先利用自编码器提取特征、再独立执行聚类的两阶段方法,容易产生特征与聚类任务脱节的问题,导致聚类结果的紧凑性与区分度较低,最终预测得到的日负荷行为模式精度与稳定性较差;于是本发明将训练完成的时间序列自编码器TS-Mixer与UNSEEN DCN深度聚类模块进行耦合训练,构建了端到端的统一学习框架,能够有效提升日负荷行为模式的聚类精度与稳定性;
Smart Images

Figure CN122778085A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system technology, and in particular to a method and system for clustering power load behavior based on a deep clustering coupling model. Background Technology
[0002] With the continuous advancement of new power system construction and the large-scale integration of smart meters and other intelligent sensing devices, the spatiotemporal scope of power load data is constantly expanding. The monitoring scope of power load is gradually extending from distribution substations to the user side and terminal equipment, driving rapid growth in the scale and increasing richness of load data types. Therefore, fully exploring the potential value of power load data has become one of the important issues in the construction of new power systems and the digital development of the power industry.
[0003] Load clustering is a crucial foundation for power big data analysis, applicable to various aspects such as electricity consumption behavior research, electricity pricing, and load modeling. Accurate load clustering also contributes to optimized power system operation and improved demand-side management. Currently, power load behavior clustering technology has shifted from traditional clustering methods to deep learning-based methods. The core idea is to automatically extract deep features from the data using deep learning models, then combine this with clustering algorithms to achieve accurate feature classification, effectively overcoming the limitations of manual feature extraction.
[0004] Autoencoders, as a classic unsupervised deep learning model, can map high-dimensional load data into low-dimensional, compact features through an encoder, and reconstruct the original data through a decoder, achieving effective dimensionality reduction and information preservation of features, making them a commonly used tool for load data feature extraction. However, traditional autoencoders are prone to gradient vanishing and incomplete capture of time-series features when processing long-term power load data.
[0005] Most clustering methods do not employ end-to-end joint optimization of feature extraction and clustering processes, which may result in low-dimensional features extracted by autoencoders being insufficiently adapted to the clustering task, affecting the final clustering results. Furthermore, most existing deep clustering models require pre-determining the number of clusters and lack the ability to adaptively estimate the number of clusters.
[0006] The invention disclosed in CN118551242A presents a load type identification method based on deep embedding clustering, belonging to the field of power system load characteristic analysis and data mining. The method includes data preprocessing, data dimensionality reduction, load clustering, and load classification. The method preprocesses the collected daily load curve data to ensure data integrity and accuracy; it uses a one-dimensional residual convolutional autoencoder to extract features from the input data, achieving data dimensionality reduction to address the problem of excessively high data dimensionality; it employs a clustering algorithm to divide the low-dimensional features, constructing a reliable sample database to mitigate the impact of noisy samples; and it uses the reliable sample database to train a convolutional neural network (CNN) classification model to obtain the load type identification result. This invention aims to solve the problems of information loss due to load data dimensionality reduction and the unsatisfactory load type identification effect caused by noisy samples in existing technologies. However, this scheme uses a two-stage method of first extracting features using an autoencoder and then independently performing clustering, which suffers from a disconnect between features and the clustering task, resulting in poor final clustering results. Summary of the Invention
[0007] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a power load behavior clustering method and system based on a deep clustering coupling model.
[0008] The objective of this invention can be achieved through the following technical solutions: A power load behavior clustering method based on a deep clustering coupling model includes: Obtain and preprocess the load data of electricity users; input the preprocessed data into a pre-trained deep clustering coupling model to obtain the clustering results of electricity load behavior; The training process of the deep clustering coupling model includes: acquiring historical load data and preprocessing it; constructing a daily load curve dataset based on the preprocessed data; inputting the daily load curve dataset into the deep clustering coupling model to extract features and obtain low-dimensional features; and performing initial clustering based on the low-dimensional features to obtain initial sample clusters, sample cluster centers, and labels with embedded features. Based on the initial sample clusters, sample cluster centers, and labels of the embedded features, a total loss function including reconstruction loss, cluster assignment loss, and nearest neighbor constraint loss is calculated. Backpropagation is performed based on the calculated loss function values, and the model parameters of the deep clustering coupling model are optimized until training converges, resulting in a trained deep clustering coupling model.
[0009] Furthermore, the deep clustering coupling model includes a pre-trained time series autoencoder and a deep clustering model; the time series autoencoder is used to extract features from the daily load curve dataset and obtain low-dimensional features; during the training process of the deep clustering coupling model, the number of samples in the initial sample cluster is counted, and the initial sample clusters with a sample number lower than a preset threshold are removed.
[0010] Furthermore, the pre-training process of the time series autoencoder includes: Historical electricity load data of various types of electricity users are acquired; based on the historical electricity load data, a daily load curve dataset is constructed; the daily load curve dataset is input into the time series autoencoder to perform feature extraction and feature transformation in the time dimension to obtain feature vectors; the feature vectors are restored and output constraints are applied to obtain reconstructed curves; based on the reconstructed curves, the reconstruction loss is calculated; based on the reconstruction loss, backpropagation is performed and the parameters of the time series autoencoder are updated until training converges, resulting in a trained time series autoencoder.
[0011] Furthermore, for feature extraction along the time dimension, the corresponding feature extraction expression is: in, It is an activation function; Drop(·) indicates the Dropout layer. This indicates a projection with equal time intervals. Indicates the selected feature matrix of the th All rows of data in the column, This represents the output feature vector of the timing transformation module, and Norm() represents layer normalization or batch normalization.
[0012] Furthermore, the feature transformation, and the corresponding transformation expression, are as follows: in, This represents the feature vector after the intermediate transformation. It is an activation function; Drop(·) indicates the Dropout layer. The weight matrix is the first-level linear transformation. This represents the cross-layer mapping features from the encoder side C to the decoder side D. This represents all column data in the j-th row of the selected feature matrix. This represents the bias vector of the direct mapping path. This represents the weight matrix of the direct mapping path. This represents the bias vector for the second-level linear transformation. b1 represents the weight matrix of the second-level linear transformation, Norm(·) represents the layer normalization or batch normalization, and b2 represents the bias vector of the first-level linear transformation.
[0013] Furthermore, based on the initial sample cluster of the embedded features, the reconstruction loss is calculated, and the corresponding calculation formula is as follows: in, To reconstruct the loss, This refers to the number of samples within the batch. For input dimensions, Indicates sample d represents the input dimension d, To reconstruct the output, This is the original input.
[0014] Furthermore, based on the embedded feature sample cluster centers and labels, the clustering assignment loss is calculated, and the corresponding calculation formula is as follows: in, For the sample The latent features, For the sample The center of the cluster to which it belongs, For the sample Hard clustering labels, This refers to the number of samples within the batch. Assign loss to clustering.
[0015] Furthermore, based on the embedded feature sample cluster centers and labels, the nearest neighbor constraint loss is calculated using the following formula: Where K is the number of currently active clusters, and N is the number of samples in the batch. For nearest neighbor constraint loss, This represents the number of nearest neighbor samples. Indicates the first One sample, It is a sample The The latent features of the nearest neighbor samples, Indicates the first The latent features of each sample.
[0016] Furthermore, the reconstruction loss, clustering assignment loss, and nearest neighbor constraint loss are weighted and summed to obtain the total loss function, the corresponding function expression of which is: in, For the total loss function, , These are the weighting coefficients corresponding to the loss term. To reconstruct the loss, Assign loss to clustering The loss is due to the nearest neighbor constraint.
[0017] The present invention also provides a system for a power load behavior clustering method based on a deep clustering coupling model, comprising a memory and a processor, wherein the memory stores a computer program, and the processor calls the computer program to execute the steps of the method described above.
[0018] Compared with the prior art, the present invention has the following advantages: (1) Traditional deep learning-based clustering methods usually adopt a two-stage approach of first extracting features using an autoencoder and then performing clustering independently. This approach is prone to the problem of features being disconnected from the clustering task, resulting in low compactness and discriminability of the clustering results, and ultimately poor accuracy and stability of the predicted daily load behavior patterns. Therefore, this invention couples the trained time series autoencoder TS-Mixer with the UNSEEN DCN deep clustering module to construct an end-to-end unified learning framework, which can effectively improve the clustering accuracy and stability of daily load behavior patterns. Furthermore, a comprehensive total loss function was constructed, comprising reconstruction loss, clustering assignment loss, and nearest neighbor constraint loss. The reconstruction loss restores the low-dimensional features extracted by the encoder to the original daily load curve shape, ensuring that the low-dimensional feature space retains the key physical morphological information in the original data. The clustering assignment loss causes sample points in the low-dimensional feature space to continuously move towards their corresponding cluster centers, thereby improving the intra-class compactness of load curves for users of the same type and the inter-class separation between different categories. The nearest neighbor constraint loss fully utilizes the continuous characteristics of power load data in the time dimension, avoiding the problem of category mutation caused by abnormal daily data. This allows the finally trained total loss function to make predictions based on key physical information and continuous characteristics such as peaks, valleys, and their duration in the original data, with high recognition accuracy. The final predicted daily load behavior pattern clustering accuracy and stability are good.
[0019] (2) This invention employs an autoencoder based on the TS-Mixer architecture to perform bidirectional hybrid modeling in both the time and feature channel dimensions. By combining downsampling and upsampling compression and reconstruction mechanisms, it achieves efficient feature extraction of the daily load curve. Compared with traditional autoencoders, ordinary convolutional autoencoders, and other shallow autoencoders, this method can more accurately and comprehensively capture the peaks, valleys, and their durations of the daily load curve, effectively characterizing the multi-peak morphology and temporal evolution of the load, thereby providing more discriminative representation information for subsequent load pattern clustering.
[0020] (3) This invention introduces a dying cluster threshold judgment mechanism, which counts the number of samples allocated to each cluster in real time during the model iterative training process, automatically identifies and removes invalid small clusters with a long-term low sample ratio and sparse distribution, and dynamically adjusts the number of effective clusters based on the actual data distribution. Compared with traditional K-means, DCN and other clustering methods that directly preset a fixed K value, this mechanism can effectively avoid redundant clusters, fragmented clusters and unreasonable partitioning caused by improper prior setting of the number of clusters, reduce the interference of noisy clusters on the overall clustering results, make the clustering results more in line with the actual electricity consumption behavior of users, and make the load grouping structure more reasonable and more meaningful in engineering practice. Attached Figure Description
[0021] Figure 1 This is a flowchart of a power load behavior clustering method based on a deep clustering coupling model provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a TS-mixer encoder, which is a power load behavior clustering method based on a deep clustering coupling model, provided in an embodiment of the present invention. Figure 3 This is a schematic diagram of the reconstruction curve of a user based on a power load behavior clustering method based on a deep clustering coupling model provided in an embodiment of the present invention. Figure 4 This is a typical load curve of a user based on a power load behavior clustering method based on a deep clustering coupling model provided in an embodiment of the present invention. Figure 5 This is a calendar heatmap of a user's electricity consumption behavior based on a power load behavior clustering method using a deep clustering coupling model, as provided in this embodiment of the invention. Figure 6 This is a visualization of a user clustering method based on a deep clustering coupling model for power load behavior provided in an embodiment of the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0023] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0024] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0025] Example 1 like Figure 1 As shown, this embodiment provides a power load behavior clustering method based on a deep clustering coupling model. The method includes the following steps: S1: Obtain and preprocess the load data of electricity users; input the preprocessed data into a pre-trained deep clustering coupling model to obtain the power load behavior clustering results; S2: The training process of the deep clustering coupling model includes: acquiring historical load data and preprocessing it; constructing a daily load curve dataset based on the preprocessed data; inputting the daily load curve dataset into the deep clustering coupling model, extracting features and obtaining low-dimensional features; performing initial clustering based on the low-dimensional features to obtain the initial sample clusters, sample cluster centers and labels embedded with the features; S201: Data Acquisition and Preprocessing: Acquire historical load data of electricity users and preprocess the historical load data; Historical electricity load information from various electricity users is collected, and then the raw load data undergoes systematic preprocessing. After preprocessing, the processed data is integrated using individual electricity users as the basic unit to construct a standardized daily load curve dataset, which serves as both the input data and the reconstruction output target for the autoencoder.
[0026] Preferred, Preprocessing includes data cleaning, completion, noise reduction, normalization, etc., and generates standardized daily load curve datasets according to the user. S202: Construction and Training of Time Series Autoencoders: Building a time series autoencoder to learn the key change patterns of the daily load curve through compression and reconstruction, mapping high-dimensional daily load data to a low-dimensional latent feature space; Preferred, The construction and training of the time series autoencoder adopts the TS-Mixer structure, and the specific steps are as follows: The normalized daily load data is organized into a matrix form and used as the input and reconstruction target of the autoencoder; The autoencoder encoder is based on the TSMixer residual block, and performs linear transformation and residual connection in the time dimension and channel dimension in sequence. The 96 time points are compressed into a shorter potential time length by linear downsampling in the time dimension, and the compressed temporal features are mapped to a low-dimensional potential vector space by a fully connected layer. The decoding end starts from this latent vector, first restores it to the compressed time axis through a fully connected layer, then uses linear upsampling in the time dimension to restore the sequence to 96 time points, and finally outputs the reconstructed daily load curve through an activation function; During training, mean squared error is used as the reconstruction loss function, and the model parameters are iteratively updated within the mini-batch gradient descent framework, in conjunction with the Adam optimizer and the adaptive learning rate decay strategy.
[0027] Preferred, The encoder, belonging to the feature extraction part of TS-Mixer, consists of multiple layers of TS-Mixer residual blocks, including temporal dimension mixing and feature dimension mixing. The feature extraction formula for the temporal dimension is: in, , It's the activation function, Drop is Dropout, and Norm can be layer normalization or batch normalization. (Encoder input, L is the backtracking window, C is the number of features). This is a projection with equal time intervals.
[0028] Feature mixing acts on the input matrix The blocks are designed to model feature transformations and are shared across all rows of the input matrix. Its core formula is: in, , .
[0029] After the construction is completed, the time axis is compressed by time dimension downsampling, and then the downsampled tensor is flattened and mapped to a low-dimensional latent space to obtain the feature vector.
[0030] Build a decoder: The feature vectors in the latent space are restored to their original lengths, and the output is constrained to between 0 and 1 using the Sigmoid function to obtain the reconstructed curve.
[0031] Preferred, The reconstruction curve obtained from forward propagation is used to calculate the reconstruction loss using the reconstruction loss function. Backpropagation is then performed to calculate the gradients of the parameters of the entire autoencoder, and all parameters are updated based on these gradients. Multiple iterations are implemented to gradually converge the encoder, providing input features for subsequent UNSEENDCN deep clustering.
[0032] S203: Coupling of autoencoder and deep clustering model: The trained time series autoencoder is coupled with the improved UNSEEN DCN deep clustering module to form an end-to-end deep clustering framework, ensuring that the low-dimensional features output by the autoencoder can be directly used for subsequent clustering training. Preferred, The trained TSMixer autoencoder is encapsulated to conform to the interface specifications of the deep clustering framework, so that the low-dimensional latent vector obtained by mapping the load curve through the encoder can be directly used as the embedding input of DCN.
[0033] S3: Based on the initial sample clusters, sample cluster centers, and labels of the embedded features, calculate the total loss function, including reconstruction loss, cluster assignment loss, and nearest neighbor constraint loss; perform backpropagation based on the calculated loss function value, and optimize the model parameters of the deep clustering coupling model until training converges, thus obtaining the trained deep clustering coupling model.
[0034] S301: Joint Loss Function Optimization: Global optimization training is performed on the entire deep clustering framework, and a joint loss function is constructed using self-supervised reconstruction loss, cluster assignment loss and nearest neighbor constraint loss; Preferred, First, perform deep clustering initialization: using latent features as input, perform an initial K-means clustering to obtain initial sample labels, initial clusters, and initial cluster numbers. The results serve as the starting point for subsequent end-to-end training.
[0035] Preferred, The original DCN retains the mechanism of joint optimization using self-supervised reconstruction loss and cluster assignment loss. It constrains the representation quality of the latent space by minimizing the autoencoder reconstruction error, updates the sample cluster labels and cluster centers according to the distance between the current embedding and the cluster center, and introduces a nearest neighbor constraint loss to impose additional penalties on the local neighborhood structure in the embedding space.
[0036] S302: Adaptive Clustering Training: When implementing end-to-end clustering training for low-dimensional latent features, a dyingcluster threshold mechanism is introduced to automatically filter and remove clusters that are too small during the model training process, thereby achieving adaptive adjustment of the number of clusters.
[0037] Preferred, UNSEENDCN integrates a dying cluster threshold mechanism, which dynamically monitors the sample size of each cluster during training, automatically removes "dead clusters" whose size is consistently below a set ratio, and re-estimates the number of effective clusters and cluster centers accordingly, thereby achieving adaptive adjustment of the number of clusters.
[0038] After several iterations, UNSEENDCN will count the current number of samples in each cluster and compare it with the initial size. If the number of samples in a cluster is lower than the set threshold for a long period of time, it will be regarded as a "dead cluster" and removed from the cluster center list.
[0039] The labels of the affected samples are re-predicted, and the number of effective clusters and cluster centers are re-counted. During this process, the latent features are still output by the autoencoder, and the clustering module and the encoder continue to be trained together.
[0040] Preferred, The entire model training process is an iterative process of alternating optimization, and its specific steps are as follows: In each training batch, all loss terms are first calculated through forward propagation; Then, the joint loss function is calculated using the backpropagation algorithm. The gradient relative to all trainable parameters of the autoencoder and clustering layers; Finally, the optimizer is used to synchronously update the parameters of the entire network based on the calculated gradients.
[0041] Convergence and Final Clustering Output: After multiple rounds of iterative training, when the total loss and various metrics tend to stabilize, the end-to-end framework is considered to have converged, and the cluster centers and labels learned by UNSEENDCN are stable. Finally, a final K-means clustering can be performed on the embedding space to obtain the final cluster labels and cluster centers for subsequent evaluation and business analysis.
[0042] Preferred, The clustering assignment loss and nearest neighbor loss are calculated based on the sample cluster centers and labels of the embedded features, and then added to the reconstruction loss to obtain the total loss. The formula for the total loss is: Next, we proceed with backpropagation: The gradients of the trainable parameters in the TS-Mixer and UNSEENDCN models are calculated together, and the parameters of the autoencoder and clustering module are updated using the Adam optimizer.
[0043] This embodiment also provides a system for a power load behavior clustering method based on a deep clustering model, including a memory and a processor. The memory stores a computer program, and the processor calls the computer program to execute the steps of the method described above.
[0044] Example 2 This embodiment provides a specific implementation method and effectiveness analysis of a power load behavior clustering method based on a deep clustering coupling model. The overall technical process includes: N1. Data Acquisition and Preprocessing N101. Collect historical load data of different electricity users in the study area, and then systematically preprocess the acquired historical load data. The preprocessing process includes key steps such as data cleaning, data completion, data denoising, and data normalization.
[0045] N102. After preprocessing, the processed data is integrated to generate a standardized daily load curve dataset for each individual electricity user. Each sample is a time series with a length of 96, which serves as the input and reconstruction target of the autoencoder.
[0046] N2, Construction and Training of Time Series Autoencoders This step aims to perform feature reduction on the raw power load data.
[0047] N201, Setting up the TS-Mixer autoencoder N2011, Building the encoder The encoder is part of the feature extraction section of TS-Mixer, consisting of multiple layers of TS-Mixer residual blocks, including temporal dimension mixing and feature dimension mixing. A schematic diagram of the TS-Mixer encoder structure is shown below. Figure 2 As shown, the time dimension captures the overall shape change of the load over time within a day, and its output is then added to the input as a residual. The feature dimension is equivalent to further processing the amplitude at each time point, and the residuals are also added. The feature extraction formula for the time dimension is: in, , It's the activation function, Drop is Dropout, and Norm can be layer normalization or batch normalization. (Encoder input, L is the backtracking window, C is the number of features). This is a projection with equal time intervals.
[0048] Feature mixing acts on the input matrix The blocks are designed to model feature transformations and are shared across all rows of the input matrix. Its core formula is: in, , .
[0049] By cascading multiple residual blocks, the encoder can extract more stable and discriminative temporal features layer by layer. The formula for multi-layer stacked Mixer is as follows: After the construction is completed, the time axis is compressed by time dimension downsampling, and then the downsampled tensor is flattened and mapped to a low-dimensional latent space to obtain the feature vector.
[0050] N2012, Building a decoder The feature vectors in the latent space are restored to their original lengths, and the output is constrained to between 0 and 1 using the Sigmoid function to obtain the reconstructed curve.
[0051] N202, Constructing a training loop The reconstruction curve obtained from forward propagation is used to calculate the reconstruction loss based on the reconstruction loss function. Backpropagation is then performed to calculate the gradients of the parameters of the entire autoencoder, and the Adam optimizer is used to update all parameters based on the gradients. Multiple iterations are performed based on these steps to gradually converge the encoder, providing input features for subsequent UNSEENDCN deep clustering. In one embodiment, a user's reconstruction curve is as follows: Figure 3 As shown.
[0052] N3, End-to-end deep clustering We construct a model that co-optimizes the TSMixer autoencoder and UNSEENDCN within the same training framework to achieve end-to-end deep clustering. The features output by the trained encoder become the input to the model.
[0053] N301, Initialize deep clustering Using latent features as input, we first perform an initial K-means clustering to obtain initial sample labels, initial clusters, and initial cluster numbers. The results serve as the starting point for subsequent end-to-end training.
[0054] N302, Define the loss function The loss function includes the supervised reconstruction loss function, the clustering assignment loss function, and the nearest neighbor constraint loss. The formula for the supervised reconstruction loss function is: in, This refers to the number of samples within the batch. For input dimensions, To reconstruct the output, This is the original input.
[0055] The formula for the clustering assignment loss function is: in, For the sample The latent features, For the sample The center of the cluster to which it belongs, For the sample Hard clustering labels.
[0056] The nearest neighbor constraint loss formula is: in, (K is the number of currently active clusters, and N is the number of samples in the batch.) It is a sample The The latent features of the nearest neighbor samples.
[0057] N303, End-to-End Joint Training N3031 Total Loss The clustering assignment loss and nearest neighbor loss are calculated based on the sample cluster centers and labels of the embedded features, and then added to the reconstruction loss to obtain the total loss. The formula for the total loss is: N3032 Backpropagation The gradients of the trainable parameters in the TS-Mixer and UNSEENDCN models are calculated together, and the parameters of the autoencoder and clustering module are updated using the Adam optimizer.
[0058] N3033 dyingcluster mechanism: dynamically adjusts the number of clusters. After several iterations, UNSEENDCN counts the current number of samples in each cluster and compares it with the initial size. If the number of samples in a cluster remains below a set threshold for an extended period, it is considered a "dead cluster" and removed from the cluster center list. Labels are re-predicted for the affected samples, and the number of valid clusters and cluster centers are re-counted. During this process, latent features are still output by the autoencoder, and the clustering module and encoder continue training together. The judgment formula is: , For clusters that are dead, perform deletion. in, For the first In the training rounds, the first A sample set of clusters, for The current size, For the initial size, This is the cluster extinction threshold.
[0059] N3034 convergence and final clustering result output After multiple rounds of training, when the total loss and various metrics stabilize, the end-to-end framework is considered converged, and the cluster centers and labels learned by UNSEENDCN are stable. Finally, a final K-means clustering can be performed on the embedding space to obtain the final cluster labels and cluster centers for subsequent evaluation and business analysis.
[0060] To verify the effectiveness of the method in this embodiment, an experiment was conducted using an electricity load dataset.
[0061] Data set: Daily load data of multiple electricity users in a certain region for one year. The daily load sampling frequency is 15 minutes, with 96 collection points per day.
[0062] The TS-Mixer+UNSEEN model proposed in this embodiment is compared with multiple benchmark models, including a single k-means model and different combinations of deep clustering models (CAE+k-means, CAE+DEC, etc.). Profile dilution (SC), Davidsonburgin index (DBI), and Kalinsky-Hallabas index (CH) are used as evaluation metrics.
[0063] Given the large number of electricity users, it would be meaningless to display them all individually. Therefore, we selected three representative user groups—User A (large enterprises), User B (small enterprises), and User C (commercial users)—for display and result analysis.
[0064] Table 1: Comparison of Clustering Indicators for User A (Large Enterprise) across Various Clustering Models Table 2: Comparison of Clustering Indicators for User B (Small Business) across Clustering Models Table 3: Comparison of Clustering Indicators for User C (Business User) across Clustering Models The clustering results for different users are shown in Tables 1, 2, and 3. The proposed TS-Mixer+UNSEEN model achieved the best results for all user types. Analysis of the results shows that compared to the inaccurate clustering results obtained by simply performing k-means clustering on the original data, introducing a deep network for feature dimensionality reduction of long-term series significantly improves the performance of various data types. Replacing the convolutional network in the CAE autoencoder with the TS-Mixer network significantly improves the results of the deep clustering model. Furthermore, the added cluster elimination mechanism avoids the strong dependence of the clustering results on the initial embedding quality caused by pre-setting the number of clusters. The results also show that adding the cluster elimination mechanism further improves the clustering results. This fully demonstrates that the proposed model has powerful clustering capabilities and can adapt to power load clustering tasks with different data characteristics.
[0065] A typical load curve after clustering a user (business user) is as follows: Figure 4 As shown, the clustering model divides the user's daily load data over the past year into six categories, and combines them with... Figure 5 The user's calendar heatmap shown can be used to analyze the user's annual electricity consumption behavior. From a load pattern perspective, six types of electricity consumption curves differentiate into diverse electricity consumption patterns: such as... Figure 6 As shown, Cluster 3 has the highest peak load, with prominent midday and evening peaks, representing typical examples of the main business formats operating during midday and evening. Cluster 1 has a similar peak shape to Cluster 3, but with slightly lower peak values, together constituting the core peak electricity consumption during midday and evening. Both clusters account for a very high proportion from June to September, indicating that during the high-temperature summer season in this region, air conditioning and other cooling equipment operate at high loads for extended periods, significantly increasing the proportion of cooling energy consumption. This results in high electricity consumption combined with holidays / midday and evening consumption periods, making these two electricity consumption patterns the main electricity consumption curves. Cluster 0 exhibits a mild double peak with stable load, corresponding to basic electricity consumption for daily operations and public areas. Its distribution shows that Cluster 0 is concentrated in peak seasons, appearing densely during holidays and peak consumption periods, but its proportion is low in summer and off-seasons. Clusters 5 and 2 have relatively small peak-to-valley differences and are evenly distributed during non-summer periods, representing stable auxiliary electricity consumption. The difference between the two is that Cluster 5, with its higher electricity consumption, is mostly distributed on weekends. Cluster 3 represents stable electricity consumption in winter, maintaining a high proportion only in winter, with unstable distribution throughout the year, exhibiting strong seasonality. Cluster 2 represents the weakest electricity demand, with a uniform distribution throughout the year. Cluster 4 shows the largest peak-to-valley difference, indicating significant fluctuations in electricity load within a single day. Such clusters are few in number and relatively dispersed, potentially representing special cases of electricity consumption. The t-SNE plot visualized after clustering is shown in the figure. It can be seen that the six clusters exhibit a clear banded distribution in two-dimensional space, with distinct boundaries and minimal overlap between clusters, indicating a good clustering effect and effective capture of the characteristic differences in different electricity consumption patterns.
[0066] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A power load behavior clustering method based on a deep clustering coupling model, characterized in that, include: Acquire and preprocess load data from electricity users; The preprocessed data is input into a pre-trained deep clustering coupling model to obtain the power load behavior clustering results; The training process of the deep clustering coupling model includes: acquiring historical load data and preprocessing it; constructing a daily load curve dataset based on the preprocessed data; inputting the daily load curve dataset into the deep clustering coupling model to extract features and obtain low-dimensional features; and performing initial clustering based on the low-dimensional features to obtain initial sample clusters, sample cluster centers, and labels with embedded features. Based on the initial sample clusters, sample cluster centers, and labels of the embedded features, a total loss function including reconstruction loss, cluster assignment loss, and nearest neighbor constraint loss is calculated. Backpropagation is performed based on the calculated loss function values, and the model parameters of the deep clustering coupling model are optimized until training converges, resulting in a trained deep clustering coupling model.
2. The power load behavior clustering method based on a deep clustering coupling model according to claim 1, characterized in that, The deep clustering coupling model includes a pre-trained time series autoencoder and a deep clustering model; the time series autoencoder is used to extract features from the daily load curve dataset and obtain low-dimensional features; during the training process of the deep clustering coupling model, the number of samples in the initial sample cluster is counted, and the initial sample clusters with a sample number lower than a preset threshold are removed.
3. The power load behavior clustering method based on a deep clustering coupling model according to claim 2, characterized in that, The pre-training process of the time series autoencoder includes: Historical electricity load data of various types of electricity users are acquired; based on the historical electricity load data, a daily load curve dataset is constructed; the daily load curve dataset is input into the time series autoencoder to perform feature extraction and feature transformation in the time dimension to obtain feature vectors; the feature vectors are restored and output constraints are applied to obtain reconstructed curves; based on the reconstructed curves, the reconstruction loss is calculated; based on the reconstruction loss, backpropagation is performed and the parameters of the time series autoencoder are updated until training converges, resulting in a trained time series autoencoder.
4. The power load behavior clustering method based on a deep clustering coupling model according to claim 3, characterized in that, The feature extraction for the time dimension corresponds to the following feature extraction expression: in, It is an activation function; Drop(·) indicates the Dropout layer. This indicates a projection with equal time intervals. Indicates the selected feature matrix of the th All rows of data in the column, This represents the output feature vector of the timing transformation module, and Norm() represents layer normalization or batch normalization.
5. The power load behavior clustering method based on a deep clustering coupling model according to claim 3, characterized in that, The feature transformation corresponds to the following transformation expression: in, This represents the feature vector after the intermediate transformation. It is an activation function; Drop(·) indicates the Dropout layer. The weight matrix is the first-level linear transformation. This represents the cross-layer mapping features from the encoder side C to the decoder side D. This represents all column data in the j-th row of the selected feature matrix. This represents the bias vector of the direct mapping path. This represents the weight matrix of the direct mapping path. This represents the bias vector for the second-level linear transformation. b1 represents the weight matrix of the second-level linear transformation, Norm(·) represents the layer normalization or batch normalization, b2 represents the bias vector of the first-level linear transformation, K is the number of currently active clusters, and N is the number of batch samples.
6. The power load behavior clustering method based on a deep clustering coupling model according to claim 1, characterized in that, Based on the initial sample cluster of the embedded features, the reconstruction loss is calculated, and the corresponding calculation formula is as follows: in, To reconstruct the loss, This refers to the number of samples within the batch. For input dimensions, Indicates sample d represents the input dimension d, To reconstruct the output, This is the original input.
7. The power load behavior clustering method based on a deep clustering coupling model according to claim 1, characterized in that, Based on the cluster centers and labels of the embedded feature samples, the clustering assignment loss is calculated using the following formula: in, For the sample The latent features, For the sample The center of the cluster to which it belongs, For the sample Hard clustering labels, This refers to the number of samples within the batch. Assign loss to clustering.
8. The power load behavior clustering method based on a deep clustering coupling model according to claim 1, characterized in that, Based on the embedded feature sample cluster centers and labels, the nearest neighbor constraint loss is calculated using the following formula: Where K is the number of currently active clusters, and N is the number of samples in the batch. For nearest neighbor constraint loss, This represents the number of nearest neighbor samples. Indicates the first One sample, It is a sample The The latent features of the nearest neighbor samples, Indicates the first The latent features of each sample.
9. A power load behavior clustering method based on a deep clustering coupling model according to claim 1, characterized in that, The reconstruction loss, clustering assignment loss, and nearest neighbor constraint loss are weighted and summed to obtain the total loss function, whose corresponding function expression is: in, For the total loss function, , These are the weighting coefficients corresponding to the loss term. To reconstruct the loss, Assign loss to clustering The loss is due to the nearest neighbor constraint.
10. A system for clustering power load behavior based on a deep clustering coupling model, characterized in that, It includes a memory and a processor, the memory storing a computer program, the processor invoking the computer program to perform the steps of the method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Load type identification method based on deep embedded clustering
CN118551242A