Industrial user load curve classification method, device and system

By using indicator rules, LSTM, autoencoder, PSO algorithm, K-means algorithm and LSTM-CNN-KAN model in the industrial user load curve classification method, the problems of insufficient feature extraction, difficult label acquisition and poor interpretability in the existing technology are solved, and efficient and accurate load curve classification is achieved.

CN120670927APending Publication Date: 2025-09-19NANJING SUYI IND +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510582018.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The existing industrial user load curve classification methods based on the combination of unsupervised and supervised learning have problems such as insufficient feature extraction of long time series samples, difficulty in obtaining accurate category labels of training sets under conditions of unbalanced data volume, and poor interpretability.

Method used

The indicator rule is used to obtain labeled training samples, and the LSTM and autoencoder technologies are combined for feature extraction and dimensionality reduction. The PSO algorithm and K-means algorithm are used for classification, and finally the LSTM-CNN-KAN model is used for training and classification.

Benefits of technology

It achieves efficient classification of load curves of massive industrial users, improves the accuracy, robustness and interpretability of classification results, and supports the refined management of power systems and the improvement of their intelligence level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670927A_ABST
    Figure CN120670927A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial user load curve classification method, device and system. The method comprises the following steps: acquiring a certain number of labeled training samples by adopting an index rule; secondly, combining LSTM with an auto-encoder technology, performing feature extraction and dimension reduction on a training set sample, and solving the problems of possible information loss and damage in the process; then, a PSO algorithm and a Kmeans algorithm are combined, residual unclassified training samples are classified, a complete training set with labels is obtained, and the problem of initial value sensitivity existing when the Kmeans algorithm is directly used can be solved through introduction of the PSO algorithm; and finally, the LSTM-CNN-KAN model is trained by using the training set with the label, and the test set is classified, so that the classification of the massive load curves is completed, and the introduction of the KAN model improves the interpretability of the classification algorithm. Therefore, the mass load curve classification solution with precision, robustness and interpretability can be provided for energy efficiency management of industrial users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method, device and system for classifying industrial user load curves, and belongs to the technical field of industrial user load classification. Background Art

[0002] Analyzing user electricity usage behavior is a crucial component of power grid analysis and planning. With the widespread adoption of intelligent data collection devices, user electricity usage can be sampled through smart meters and represented in load curves and other forms. This data is characterized by high volume, high velocity, and low value density. Developing efficient load curve classification methods based on the characteristics of user load data can help power companies uncover potential user electricity usage patterns from massive amounts of electricity consumption data, which is crucial for load forecasting, demand response, and electricity pricing decisions.

[0003] Currently, load curve classification methods primarily include unsupervised clustering, supervised classification, and a combination of unsupervised and supervised methods. In recent years, this combination has garnered significant attention. Load curve data, as unlabeled data, can be classified using unsupervised clustering to obtain class labels. This allows for the training of supervised learning models for classification, combining the strengths of unsupervised and supervised methods to achieve efficient classification of massive amounts of load data.

[0004] The difficulties faced in combining unsupervised and supervised methods include: 1) Since the number of loads in some categories is far less than that in other categories, the efficiency of using unsupervised learning to obtain accurate category labels for the training set is low under such conditions of data imbalance; 2) As a long time series, the load curve has a high dimension and contains time series features. If the data compression and feature extraction methods are inappropriate, the time series features may be destroyed and lost. Without data compression, the classification efficiency is difficult to improve, thus facing a bottleneck. Summary of the Invention

[0005] The purpose of this invention is to provide a method, device and system for classifying industrial user load curves. By classifying massive industrial user load curves, the refined management and intelligent level of the power system can be significantly improved, and the safe and economic operation of the power grid and the realization of the dual carbon goals can be promoted.

[0006] In order to achieve the above objectives / solve the above technical problems, the present invention is implemented by adopting the following technical solutions.

[0007] In one aspect, the present invention provides a method for classifying industrial user load curves, comprising:

[0008] Obtain historical load curve data of target industrial users and divide the load curve data into training sets and test sets in proportion;

[0009] Calculate the index vector corresponding to the historical load curve data;

[0010] Based on the indicator vector, label classification is performed on the historical load curve data in the training set;

[0011] Perform dimensionality reduction on the historical load curve data in the training set, and concatenate the reduced-dimensional historical load curve data with its corresponding indicator vector to obtain a dimensionality-reduced and concatenated sample vector that corresponds one-to-one to the historical load curve data.

[0012] Obtaining the initial cluster center vector of each category after the historical load curve data is classified using the PSO algorithm according to the sample vector;

[0013] The sample vector and the initial cluster center vector are processed by the K-means algorithm to construct a sample vector with classification labels;

[0014] Decode the sample vector with classification labels to obtain the historical load curve of the original dimension, and merge it with the classification label to form the labeled historical load curve data;

[0015] The historical load curve data with classification labels is input into the LSTM-CNN-KAN classification model for training until the relative change rate of the cross entropy loss function value is lower than the given threshold, and a trained LSTM-CNN-KAN classification model is obtained;

[0016] The test set is input into the trained LSTM-CNN-KAN classification model for classification processing to obtain the classification prediction results.

[0017] Furthermore, the calculation of the index vector corresponding to the historical load curve data specifically includes:

[0018] The index vector includes: daily peak-to-valley difference rate, daily load rate, peak load rate and valley load rate;

[0019] Calculate the daily peak-to-valley difference rate using the following expression:

[0020] ;

[0021] in, is the daily peak-to-valley difference rate of load curve j, represents the load power value of load curve j in time period t, represents the maximum power of load curve j, represents the minimum power of load curve j;

[0022] Calculate the daily load rate using the expression:

[0023] ;

[0024] in, is the daily load rate of load curve j;

[0025] Calculate the peak load rate using the expression:

[0026] ;

[0027] in, is the peak load rate of load curve j, T Pj represents the peak period set of load curve j, card(T Pj ) represents the set T Pj Total number of time periods;

[0028] Calculate the valley load rate using the expression:

[0029] ;

[0030] in, is the valley load rate of load curve j, represents the valley period set of load curve j, Representative Set Total number of time periods.

[0031] Furthermore, the label classification of the historical load curve data in the training set according to the indicator vector specifically includes:

[0032] when - >H PV1 , then the label of the historical load curve data is peak type, where H PV1 is the first threshold of the peak-valley load rate difference;

[0033] when - >H PV1 , then the label of the historical load curve data is peak avoidance type;

[0034] When | - |>H PV2 , >H PVD , <H L , then the label of the historical load curve data is high load rate type, where H PV2 is the second threshold of the peak-valley load rate difference, H PVD is the threshold of the peak-to-valley difference rate, H L is the threshold value of daily load rate;

[0035] When | — | <H PV2 , <H PVD , >H L , the label of the historical load curve data is continuous.

[0036] Furthermore, the dimensionality reduction of the historical load curve data in the training set or the decoding of the sample vectors with classification labels are both implemented based on the autoencoder, specifically including:

[0037] The autoencoder comprises: an input layer, a hidden layer and an output layer; wherein: the hidden layer comprises an encoder and a decoder;

[0038] The dimensionality reduction of the historical load curve data in the training set is performed as follows:

[0039] The historical load curve data of the training set obtained from the input layer is fed into two consecutive LSTM layers of the encoder to fully extract the features of the load data. The extracted features are then fed into the Dense layer of the encoder and output as the dimensionality-reduced vector of the historical load curve data of the training set.

[0040] Decode the sample vector with classification labels, specifically:

[0041] The sample vector with the indicator vector deleted is input into the Reshape layer of the decoder to restore the vector dimension. The restored vector is then sent to two consecutive LSTM layers of the decoder for decoding. The output is the restored historical load curve data, which is re-spliced ​​with the corresponding classification label to obtain the labeled training set historical load curve data.

[0042] Furthermore, the initial cluster center vector of each category after the historical load curve data is classified by the PSO algorithm according to the sample vector specifically includes:

[0043] Calculate the sum of the squares of the Euclidean distances from all sample vectors to the cluster center vector of each PSO particle. The expression is:

[0044] ;

[0045] Where: f i is the sum of squared Euclidean distances between the four cluster center vectors and the sample vector contained in the i-th PSO particle; m j is the jth sample vector, c ik is the kth cluster center vector contained in the i-th PSO particle, is m j to c ik The square of the Euclidean distance; M is the number of historical load curves in the training set;

[0046] The fitness function is constructed based on the Euclidean distance squared sum, and the expression is:

[0047] ;

[0048] Where: 2 is the fitness variance, is the set number of PSO particles;

[0049] The initial cluster center vector of each category after the historical load curve data classification is obtained by combining the PSO algorithm with the fitness function, specifically including:

[0050] When setting σ 2 When it is lower than a given threshold or the iteration reaches the maximum number T, the PSO optimization process is stopped, and the four cluster center vectors contained in the global optimal PSO particle at this time are used as the initial cluster center vectors.

[0051] Furthermore, the expression of the cross entropy loss function is:

[0052] ;

[0053] in, is the cross entropy loss function value, is the number of historical load curves in the training set, is the true label of load curve j belonging to category k, is the probability that the output load curve j belongs to category k;

[0054] The expression of the relative rate of change of the loss function value is:

[0055] ;

[0056] Among them, L d and L d -τ are the cross entropy loss function values ​​after the dth and d-τth rounds of training, and ε is the threshold for the relative rate of change of the loss function value.

[0057] Furthermore, the pre-trained LSTM-CNN-KAN classification model performs a classification process, specifically including:

[0058] Extract spatial and temporal features of target industrial user load curve data.

[0059] Perform tensor splicing on spatial features and temporal features to generate fused feature vectors, perform dimensionality reduction on the fused feature vectors, and then perform detection on the fused feature vectors after dimensionality reduction;

[0060] The extraction of spatial features and temporal features specifically includes: converting the data dimension of the load curve data of the target industrial user, performing convolution processing on the converted data to extract spatial features, and pooling the extracted spatial features to obtain simplified spatial features;

[0061] The intrinsic time series features of the target industrial user load curve data are extracted twice continuously to obtain the simplified time series features.

[0062] In a second aspect, the present invention provides an industrial user load curve classification device, comprising:

[0063] An acquisition module is used to obtain historical load curve data of target industrial users and divide the load curve data into training sets and test sets in proportion;

[0064] Index module, used to calculate the index vector corresponding to the historical load curve data;

[0065] The label module is used to classify the historical load curve data in the training set according to the indicator vector;

[0066] A dimensionality reduction and splicing module is used to reduce the dimensionality of the historical load curve data in the training set, and splice the reduced dimensionality historical load curve data with its corresponding indicator vector to obtain a dimensionality-reduced and spliced ​​sample vector that corresponds one-to-one to the historical load curve data;

[0067] A construction module is used to construct a fitness function based on the sum of squares of the Euclidean distances from the sample vector to the cluster center vector of each PSO particle, and obtain the initial cluster center vector of each category after the historical load curve data is classified by combining the PSO algorithm with the fitness function;

[0068] The construction module is used to process the sample vector and the initial cluster center vector through the K-means algorithm to construct the sample vector with classification labels;

[0069] A decoding module is used to decode the sample vector with classification labels, restore the sample vector to the historical load curve of the original dimension, and merge it with the classification label to form the labeled historical load curve data;

[0070] The training module is used to input the historical load curve data with classification labels into the LSTM-CNN-KAN classification model for training until the relative change rate of the cross entropy loss function value is lower than a given threshold, thus forming a trained LSTM-CNN-KAN classification model;

[0071] The classification module is used to input the test set into the pre-trained LSTM-CNN-KAN classification model for classification processing to obtain classification prediction results.

[0072] In a third aspect, the present invention provides an industrial user load curve classification system, comprising:

[0073] Memory, used to store computer programs / instructions;

[0074] A processor is configured to execute the computer program / instructions to implement the steps of the above-mentioned industrial user load curve classification method.

[0075] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program / instruction stored thereon, which, when executed by a processor, implements the steps of the above-mentioned industrial user load curve classification method.

[0076] Compared with the existing technology, the beneficial effects achieved by the present invention are as follows: the present invention solves the technical bottlenecks in the existing industrial user load curve classification method based on the combination of unsupervised and supervised learning, mainly including insufficient feature extraction of long time series samples, difficulty in obtaining accurate category labels of the training set under the condition of unbalanced data volume, and poor interpretability; the present invention uses indicator rules to label some samples in the training set; then, the training set samples are reduced in dimension, and the PSO-Kmeans algorithm is used to classify the reduced training set to obtain a labeled training set, wherein the initial particle swarm of the PSO algorithm is generated by the labeled samples obtained in the training set using the indicator rules, thereby solving the problems caused by the classification results falling into local optimality and unbalanced load data; the labeled training set is then used to train the LSTM-CNN-KAN model, and finally the trained LSTM-CNN-KAN model is used to classify the industrial user load curves, thereby obtaining the classification results of massive industrial user load curves.

[0077] The present invention first classifies the acquired load curves of industrial users through the trained LSTM-CNN-KAN model, which can be used to accurately identify the electricity consumption patterns of different users, and provide data support for load forecasting, grid planning, demand-side management and electricity price design of the power system, thereby optimizing power generation scheduling, reducing peak-to-valley differences, and improving energy utilization efficiency. It can provide a massive load curve classification solution with both accuracy, robustness and interpretability for the energy efficiency management of industrial users.

[0078] In addition, it can assist in electricity market transactions, tap the potential of adjustable loads, promote the consumption of green electricity, and provide a basis for distribution network upgrades, fault diagnosis and energy efficiency management, thereby significantly improving the refined management and intelligence level of the power system, and promoting the safe and economic operation of the power grid and the realization of dual carbon goals. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 This is a flow chart of a method for classifying massive industrial user load curves based on a KAN improved neural network according to an embodiment of the present invention;

[0080] Figure 2 This is a structural diagram of the LSTM-CNN-KAN model in the massive industrial user load curve classification method based on the KAN improved neural network described in an embodiment of the present invention. DETAILED DESCRIPTION

[0081] It should be noted that:

[0082] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0083] The term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Additionally, the character " / " generally indicates an "or" relationship between the related objects.

[0084] Example 1

[0085] like Figure 1 An embodiment shown in FIG. 1 provides a method for classifying industrial user load curves, including:

[0086] Step 1: Obtain historical load curve data of target industrial users and divide the load curve data into training and test sets in a ratio of 3:7;

[0087] Step 2: Calculate the index vector that corresponds to the historical load curve data, including:

[0088] Calculate the daily peak-to-valley difference rate using the following expression:

[0089] ;

[0090] in, is the daily peak-to-valley difference rate of load curve j, represents the load power value of load curve j in time period t, represents the maximum power of load curve j, represents the minimum power of load curve j;

[0091] Calculate the daily load rate using the expression:

[0092] ;

[0093] in, is the daily load rate of load curve j;

[0094] Calculate the peak load rate using the expression:

[0095] ;

[0096] in, is the peak load rate of load curve j, T Pj represents the peak period set of load curve j, card(T Pj ) represents the set T Pj Total number of time periods;

[0097] Calculate the valley load rate using the expression:

[0098] ;

[0099] in, is the valley load rate of load curve j, represents the valley period set of load curve j, Representative Set Total number of time periods.

[0100] Step 3: Based on the indicator vector, label classification is performed on the historical load curve data in the training set, specifically including:

[0101] when - >H PV1 , then the label of the historical load curve data is peak type, where H PV1 is the first threshold of the peak-valley load factor difference, which is 60%;

[0102] when - >H PV1 , then the label of the historical load curve data is peak avoidance type;

[0103] When | - |>H PV2 , >H PVD , <H L , then the label of the historical load curve data is high load rate type, where H PV2 The second threshold value of the peak-valley load rate difference is 30%, H PVD is the threshold of the peak-to-valley difference rate, which is 30%, and H L is the threshold of daily load rate, which is 90%;

[0104] When | — | <H PV2 , <HPVD , >H L , the label of the historical load curve data is continuous.

[0105] Step 4: Reduce the dimension of the historical load curve data in the training set, and concatenate the reduced-dimensional historical load curve data with its corresponding indicator vector to obtain a sample vector that corresponds one-to-one to the historical load curve data after dimension reduction and concatenation. Specifically, the following steps are performed:

[0106] Step 4.1: Input the historical load curve data in the training set into the autoencoder and pass it through two LSTM layers with 64 and 32 neurons respectively to achieve full feature extraction;

[0107] Step 4.2: Input the features extracted by the two LSTM layers into a fully connected layer (Dense layer) with 8 neurons to reduce the feature dimensionality of the load curve and output the reduced dimensionality vector.

[0108] Step 4.3: Concatenate the dimension-reduced historical load curve data with its corresponding indicator vector to obtain a dimension-reduced and concatenated sample vector that corresponds one-to-one to the historical load curve data.

[0109] Step 5: Construct a fitness function based on the sum of the squares of the Euclidean distances from the sample vector to the cluster center vector of each PSO particle, and obtain the initial cluster center vector of each category after the historical load curve data is classified by combining the PSO algorithm with the fitness function, specifically including:

[0110] In order to prevent the subsequent Kmeans algorithm from falling into local optimality, the PSO algorithm is required to provide it with the initial cluster center; therefore, the position x of the i-th PSO particle in the PSO algorithm is i There are 4 cluster center vectors c ik (k=1, 2, 3, 4), the dimension of each cluster center vector is consistent with the dimension of the sample vector after dimensionality reduction and splicing of the training set load curve in step 4;

[0111] All PSO particles independently search for the optimal solution in the given space, regard the individual optimal value as the current independent individual extreme value, and then share their own optimal position with other particles, compare the individual optimal position of each particle group as the global optimal solution, and iterate continuously. When the particle group is initialized before the iteration begins, since each class has obtained a certain number of labeled load curve samples in step 3, the c of the i-th particle in the initial particle group is ik , should be randomly sampled from the labeled dimensionality reduction-splicing samples obtained for the kth class. This operation can prevent the poor classification effect caused by unbalanced data (i.e., the number of loads in some categories is much smaller than that in other categories);

[0112] Then, the PSO iteration process officially begins; all particles adjust their individual speed and position by comparing their individual optimal position with the global optimal position of the entire particle swarm. The calculation expression is:

[0113] ;

[0114] in, and are the velocities of the i-th particle after the t-th and t+1-th iterations, respectively. and are the positions of the i-th particle after the t-th and t+1-th iterations, respectively. is the optimal position currently searched by the i-th particle, G t is the global optimal position shared by the entire particle swarm; rand() represents the generation of a sample of random numbers that obey a uniform distribution in the interval [0,1]; w t+1 is the weight coefficient at the t+1th iteration, c 1,t+1 is the individual learning factor of each particle in the t+1th iteration; c 2,t+1 is the social learning factor of each particle in the t+1th iteration;

[0115] ;

[0116] Among them, w max and w min w t+1 The maximum and minimum values ​​of c 1,max and c 1,min c 1,t+1 The maximum and minimum values ​​of c 2,max and c 2,min c 2,t+1 The maximum and minimum values ​​of , T is the upper limit of the number of iterations.

[0117] Among them: the sum of the squares of the Euclidean distances from the sample vector to the cluster center vector of each PSO particle is calculated, and the expression is:

[0118] ;

[0119] Where: f i is the sum of squared Euclidean distances from the sample vector to the cluster center vector after dimension reduction and splicing of the i-th load curve; m j is the sample vector after dimension reduction and splicing of the j-th load curve, c ik is the kth cluster center vector in the i-th PSO particle, is m j to c ik The square of the Euclidean distance; M is the number of historical load curves in the training set;

[0120] The fitness function is constructed based on the Euclidean distance squared sum, and the expression is:

[0121] ;

[0122] Where: 2 is the fitness variance, is the set number of PSO particles;

[0123] The initial cluster center vector of each category of the historical load curve data is obtained by combining the PSO algorithm with the fitness function, including:

[0124] When setting σ 2 When it is lower than a given threshold or the maximum number of iterations T is reached, the PSO optimization process is stopped, and the four cluster center vectors contained in the global optimal PSO particle at this time are used as the initial cluster center vectors, and the Kmeans algorithm is used for local fast optimization to complete the classification of all remaining samples in the training set and affix category labels.

[0125] Step 6: Process the sample vector and the initial cluster center vector through the K-means algorithm to construct a sample vector with classification labels;

[0126] Step 7: Decode the sample vector with classification labels through the decoding part of the autoencoder, restore the sample vector to the historical load curve of the original dimension, and merge it with the classification label to form the labeled historical load curve data, specifically:

[0127] When decoding a sample vector with a classification label, the indicator vector in the sample vector is first deleted. The resulting vector is then fed into the decoder's Reshape layer (with the number of neurons set to 8) to restore the vector dimension. The restored vector is then fed into two consecutive LSTM layers of the decoder (with the number of neurons set to 32 and 64, respectively) for decoding. The output is the restored historical load curve, which is reconstructed with the corresponding classification label to obtain the labeled training set historical load curve.

[0128] Step 8: Input the historical load curve data with classification labels into the LSTM-CNN-KAN classification model for training until the relative change rate of the cross entropy loss function value is lower than the given threshold, forming a trained LSTM-CNN-KAN classification model. Specifically, the following steps are performed:

[0129] The expression of the cross entropy loss function is:

[0130] ;

[0131] in, is the cross entropy loss function value, is the number of historical load curves in the training set, is the true label of load curve j belonging to category k (1 if it belongs, 0 if it does not), is the probability that the output load curve j belongs to category k;

[0132] The expression of the relative rate of change of the loss function value is:

[0133] ;

[0134] Among them, L d and L d -τ are the cross entropy loss function values ​​after the dth round and the d-τth round of training respectively, ε is the given threshold for the relative rate of change of the loss function value, and τ can be 5.

[0135] Step 9: Input the test set into the pre-trained LSTM-CNN-KAN classification model for classification processing to obtain the classification prediction results, including:

[0136] The trained LSTM-CNN-KAN classification model mainly includes: CNN submodule, LSTM submodule and KNN classification module;

[0137] Spatial features are extracted based on the CNN submodule: First, the original load curve data enters the Reshape layer in the module for data dimension conversion; then, the converted data enters the one-dimensional convolution layer with the Relu activation function for spatial feature extraction; then, the extracted features are sent to the maximum pooling layer for feature reduction; finally, the reduced feature data is sent to the one-dimensional convolution layer and the global pooling layer again to obtain fully reduced spatial features;

[0138] Extracting time series features based on the LSTM submodule: A two-layer LSTM network is configured, with each layer containing 64 memory units. The ReLU nonlinear activation mechanism is used to perform two consecutive intrinsic time series feature extractions on the original load curve data through the two-layer LSTM. This is used to mine the dynamic time dependencies in the data stream and obtain simplified time series features.

[0139] Classification detection based on the KAN classification module: After the above two modules complete feature extraction, the spatial features and the temporal features are tensor-spliced ​​in the feature splicing layer of the module, and the spliced ​​one-dimensional fused feature vector is output; then, the fused feature vector is sent to the first fully connected layer of the module with 32 neurons and Relu activation function to achieve dimensionality reduction of the fused feature vector; finally, the reduced fusion feature vector is sent to the second KAN fully connected layer of the module with the number of neurons equal to the number of load curve categories (4 categories in this embodiment) and Softmax activation function. The output of this layer is the probability of each load curve belonging to each of the 4 categories. The category with the largest probability is selected as the classification label of the curve, and the classification prediction results of all samples in the test set are finally obtained.

[0140] Example 2

[0141] This embodiment provides an industrial user load curve classification device, including:

[0142] An acquisition module is used to obtain historical load curve data of target industrial users and divide the load curve data into training sets and test sets in proportion;

[0143] Index module, used to calculate the index vector corresponding to the historical load curve data;

[0144] The label module is used to classify the historical load curve data in the training set according to the indicator vector;

[0145] A dimensionality reduction and splicing module is used to reduce the dimensionality of the historical load curve data in the training set, and splice the reduced dimensionality historical load curve data with its corresponding indicator vector to obtain a dimensionality-reduced and spliced ​​sample vector that corresponds one-to-one to the historical load curve data;

[0146] A construction module is used to construct a fitness function based on the sum of squares of the Euclidean distances from the sample vector to the cluster center vector of each PSO particle, and obtain the initial cluster center vector of each category after the historical load curve data is classified by combining the PSO algorithm with the fitness function;

[0147] The construction module is used to process the sample vector and the initial cluster center vector through the K-means algorithm to construct the sample vector with classification labels;

[0148] A decoding module is used to decode the sample vector with classification labels, restore the sample vector to the historical load curve of the original dimension, and merge it with the classification label to form the labeled historical load curve data;

[0149] The training module is used to input the historical load curve data with classification labels into the LSTM-CNN-KAN classification model for training until the relative change rate of the cross entropy loss function value is lower than a given threshold, thus forming a trained LSTM-CNN-KAN classification model;

[0150] The classification module is used to input the test set into the pre-trained LSTM-CNN-KAN classification model for classification processing to obtain classification prediction results.

[0151] Example 3

[0152] This embodiment provides an industrial user load curve classification system, including:

[0153] Memory, used to store computer programs / instructions;

[0154] A processor is configured to execute the computer program / instructions to implement the steps of the industrial user load curve classification method of the first embodiment.

[0155] Example 4

[0156] This embodiment provides a computer-readable storage medium on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the steps of the above-mentioned embodiment 1 are implemented.

[0157] The purpose of the present invention is to solve the technical bottlenecks in the existing industrial user load curve classification method based on the combination of unsupervised and supervised learning, mainly including insufficient feature extraction of long time series samples, difficulty in obtaining accurate category labels for training sets under conditions of unbalanced data volume, and poor interpretability. The present invention first adopts indicator rules to obtain a certain number of labeled training samples; then combines LSTM with autoencoder technology to extract features and reduce the dimension of the training set samples to solve the possible information loss and damage problems in the process; then combines the PSO algorithm and the Kmeans algorithm to classify the remaining unclassified training samples to obtain a complete labeled training set, wherein the introduction of the PSO algorithm can solve the initial value sensitivity problem that exists when the Kmeans algorithm is used directly; finally, the LSTM-CNN-KAN model is trained using the labeled training set to classify the test set, thereby completing the classification of massive load curves. In particular, the introduction of the KAN model improves the interpretability of the classification algorithm; therefore, the present invention can provide a massive load curve classification solution with both accuracy, robustness and interpretability for industrial user energy efficiency management.

[0158] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0159] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0160] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0162] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present invention and the claims, which are all protected by the present invention.

Claims

1. A method for classifying industrial user load curves, characterized in that: include: Obtain historical load curve data of target industrial users and divide the load curve data into training sets and test sets in proportion; Calculate the index vector corresponding to the historical load curve data; Based on the indicator vector, label classification is performed on the historical load curve data in the training set; Perform dimensionality reduction on the historical load curve data in the training set, and concatenate the reduced-dimensional historical load curve data with its corresponding indicator vector to obtain a dimensionality-reduced and concatenated sample vector that corresponds one-to-one to the historical load curve data. Obtaining the initial cluster center vector of each category after the historical load curve data is classified using the PSO algorithm according to the sample vector; The sample vector and the initial cluster center vector are processed by the K-means algorithm to construct a sample vector with classification labels; Decode the sample vector with classification labels to obtain the historical load curve of the original dimension, and merge it with the classification label to form the labeled historical load curve data; The historical load curve data with classification labels is input into the LSTM-CNN-KAN classification model for training until the relative change rate of the cross entropy loss function value is lower than the given threshold, and a trained LSTM-CNN-KAN classification model is obtained; The test set is input into the trained LSTM-CNN-KAN classification model for classification processing to obtain the classification prediction results.

2. The industrial user load curve classification method according to claim 1, characterized in that: The calculation of the index vector corresponding to the historical load curve data specifically includes: Calculate the daily peak-to-valley difference rate using the following expression: ; in, is the daily peak-to-valley difference rate of load curve j, represents the load power value of load curve j in time period t, represents the maximum power of load curve j, represents the minimum power of load curve j; Calculate the daily load rate using the expression: ; in, is the daily load rate of load curve j; Calculate the peak load rate using the expression: ; in, is the peak load rate of load curve j, T Pj represents the peak period set of load curve j, card(T Pj ) represents the set T Pj Total number of time periods; Calculate the valley load rate using the expression: ; in, is the valley load rate of load curve j, represents the valley period set of load curve j, Representative Set Total number of time periods.

3. The industrial user load curve classification method according to claim 2, characterized in that: The label classification of the historical load curve data in the training set according to the indicator vector specifically includes: when - >H PV1 , then the label of the historical load curve data is peak type, where H PV1 is the first threshold of the peak-valley load rate difference; when - >H PV1 , then the label of the historical load curve data is peak avoidance type; When | - |>H PV2 , >H PVD , <H L , then the label of the historical load curve data is high load rate type, where H PV2 is the second threshold of the peak-valley load rate difference, H PVD is the threshold of the peak-to-valley difference rate, H L is the threshold value of daily load rate; When | — | <H PV2 , <H PVD , >H L , the label of the historical load curve data is continuous.

4. The industrial user load curve classification method according to claim 1, characterized in that: The dimensionality reduction of the historical load curve data in the training set or the decoding of the sample vector with classification labels are both implemented based on the autoencoder, specifically including: The autoencoder comprises: an input layer, a hidden layer and an output layer; wherein: the hidden layer comprises an encoder and a decoder; The dimensionality reduction of the historical load curve data in the training set is performed as follows: The historical load curve data of the training set obtained from the input layer is fed into two consecutive LSTM layers of the encoder to fully extract the load data features. The extracted features are then fed into the Dense layer of the encoder, and the output is the dimensionally reduced vector of the historical load curve data of the training set. Decode the sample vector with classification labels, specifically: The sample vector with the indicator vector deleted is input into the Reshape layer of the decoder to restore the vector dimension. The restored vector is then sent to two consecutive LSTM layers of the decoder for decoding. The output is the restored historical load curve data, which is re-spliced ​​with the corresponding classification label to obtain the labeled training set historical load curve data.

5. The industrial user load curve classification method according to claim 1, characterized in that: The initial cluster center vector of each category after the historical load curve data is classified by the PSO algorithm according to the sample vector specifically includes: Calculate the sum of the squares of the Euclidean distances from all sample vectors to the cluster center vector of each PSO particle. The expression is: ; Where: f i is the sum of squared Euclidean distances between the four cluster center vectors and the sample vector contained in the i-th PSO particle; m j is the jth sample vector, c ik is the kth cluster center vector contained in the i-th PSO particle, is m j to c ik The square of the Euclidean distance; M is the number of historical load curves in the training set; The fitness function is constructed based on the Euclidean distance squared sum, and the expression is: ; Where: 2 is the fitness variance, is the set number of PSO particles; The initial cluster center vector of each category after the historical load curve data classification is obtained by combining the PSO algorithm with the fitness function, specifically including: When setting σ 2 When it is lower than a given threshold or the iteration reaches the maximum number T, the PSO optimization process is stopped, and the four cluster center vectors contained in the global optimal PSO particle at this time are used as the initial cluster center vectors.

6. The industrial user load curve classification method according to claim 1, characterized in that: The expression of the cross entropy loss function is: ; in, is the cross entropy loss function value, is the number of historical load curves in the training set, is the true label of load curve j belonging to category k, is the probability that the output load curve j belongs to category k; The expression of the relative rate of change of the loss function value is: ; Among them, L d and L d -τ are the cross entropy loss function values ​​after the dth and d-τth rounds of training, and ε is the threshold for the relative rate of change of the loss function value.

7. The industrial user load curve classification method according to claim 1, characterized in that: The pre-trained LSTM-CNN-KAN classification model performs a classification process, specifically including: Extract spatial and temporal features of target industrial user load curve data. Perform tensor splicing on spatial features and temporal features to generate fused feature vectors, perform dimensionality reduction on the fused feature vectors, and then perform detection on the fused feature vectors after dimensionality reduction; The spatial feature and time series feature extraction specifically includes: performing data dimension conversion on the load curve data of the target industrial user, performing convolution processing on the converted data to extract spatial features, and performing pooling processing on the extracted spatial features to obtain simplified spatial features; The intrinsic time series features of the target industrial user load curve data are extracted twice continuously to obtain the simplified time series features.

8. An industrial user load curve classification device, characterized in that: include: An acquisition module is used to obtain historical load curve data of target industrial users and divide the load curve data into training sets and test sets in proportion; Index module, used to calculate the index vector corresponding to the historical load curve data; The label module is used to classify the historical load curve data in the training set according to the indicator vector; A dimensionality reduction and splicing module is used to reduce the dimensionality of the historical load curve data in the training set, and splice the reduced dimensionality historical load curve data with its corresponding indicator vector to obtain a dimensionality-reduced and spliced ​​sample vector that corresponds one-to-one to the historical load curve data; A construction module is used to obtain an initial cluster center vector of each category after the historical load curve data is classified using a PSO algorithm according to the sample vector; The construction module is used to process the sample vector and the initial cluster center vector through the K-means algorithm to construct the sample vector with classification labels; The decoding module is used to decode the sample vector with classification labels to obtain the historical load curve of the original dimension and merge it with the classification label to form the labeled historical load curve data; The training module is used to input the historical load curve data with classification labels into the LSTM-CNN-KAN classification model for training until the relative change rate of the cross entropy loss function value is lower than a given threshold, thereby obtaining a trained LSTM-CNN-KAN classification model; The classification module is used to input the test set into the trained LSTM-CNN-KAN classification model for classification processing to obtain classification prediction results.

9. An industrial user load curve classification system, characterized in that: include: Memory, used to store computer programs / instructions; A processor is configured to execute the computer program / instructions to implement the steps of the industrial user load curve classification method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instruction is executed by a processor, the steps of the industrial user load curve classification method described in any one of claims 1-8 are implemented.