Clustering Method, Device, Electronic Device, and Storage Medium for Load Data
By calculating the clustering effect and feature extraction of the load data set, combined with the calculation of multiple sets of feature weights and distance gain, the problem of poor clustering effect of load data is solved, and a better clustering effect of load data is achieved.
Patent Information
- Application Number
- CN202210211046.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-04
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-03-04
AI Technical Summary
The prior art causes poor clustering effect due to the low signal-to-noise ratio, nonlinear, non-stationary and non-normal characteristics when processing load data.
By calculating the first distance that characterizes the clustering effect of each load data set, feature extraction is performed and multiple sets of feature weights are set, the load data is clustered into n subload data sets, and the distance gain is calculated to finally determine the best clustering characteristics and results.
The clustering effect of load data is improved, and it can adapt to different types of load data well, achieving more accurate clustering results.
Smart Images

Figure CN114580538B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a clustering method, device, electronic device, and storage medium for load data. Background Art
[0002] The electricity reform in China is in full swing, and the number of domestic power sales companies is increasing significantly. Each province is also actively preparing for the reform of the spot market. Under the background of the coming of the spot market, important business modules of power sales companies such as signing strategies, quotation strategies, trading strategies, and economic calculations for individual users are all based on the results of load forecasting. And for the accuracy of load forecasting, the first step is to perform clustering analysis on the common electricity loads of each electricity-consuming enterprise and unit.
[0003] Since load data generally has the characteristics of low signal-to-noise ratio, non-linearity, non-stationarity, and non-normality, the use of existing clustering methods for clustering often has poor effects. Summary of the Invention
[0004] The present invention provides a clustering method, device, electronic device, and storage medium for load data to improve the clustering effect.
[0005] The present invention solves the above technical problems through the following technical solutions:
[0006] In a first aspect, a clustering method for load data is provided, including:
[0007] Calculating a first distance representing the clustering effect of each load data set; the load data set contains multiple data samples; each data sample represents the load data of a load curve;
[0008] Performing feature extraction on the data samples included in each load data set and setting multiple groups of feature weights;
[0009] Respectively clustering the load data sets corresponding to each group of feature weights into n sub-load data sets and calculating the distance gain corresponding to each group of feature weights; where n≥2; the distance gain represents the deviation between the first distance and the second distance, and the second distance represents the clustering effect of the n sub-load data sets;
[0010] Determining the n sub-load data sets corresponding to the feature weight with the largest distance gain as the clustering result of this round of clustering;
[0011] Judging whether this round of clustering meets the clustering stop condition;
[0012] In the case where the judgment result is no, returning to the step of calculating the first distance of each load data set; in the case where the judgment result is yes, generating a clustering result.
[0013] Optionally, the first distance is determined according to the distances of the respective data samples included in the load data set to the cluster center;
[0014] The second distance is determined according to the distances of the n sub-load data sets, and the distances of the respective sub-load data sets are determined according to the distances of the respective data samples included in the sub-load data sets to the cluster center.
[0015] Optionally, the clustering stop condition includes at least one of the following:
[0016] The number of clustering times is greater than the number threshold;
[0017] There is at least one sub-load data set among the n sub-load data sets corresponding to the feature weights with the largest distance gain that contains a number of data samples less than the number threshold;
[0018] The distance gain is less than the gain threshold.
[0019] Optionally, clustering the load data sets corresponding to each group of feature weights into n sub-load data sets includes:
[0020] Clustering the load data sets corresponding to each group of feature weights into n sub-load data sets according to the K-medoids algorithm.
[0021] Optionally, before performing feature extraction on the data samples included in each load data set, it includes:
[0022] Performing preprocessing on the data samples included in each load data set;
[0023] And / or, performing normalization processing on the data samples included in each load data set.
[0024] In a second aspect, a clustering device for load data is provided, including:
[0025] A calculation module, configured to calculate a first distance characterizing the clustering effect of each load data set; the load data set includes a plurality of data samples; each data sample characterizes the load data of a load curve;
[0026] An extraction module, configured to perform feature extraction on the data samples included in each load data set and set multiple groups of feature weights;
[0027] A clustering module, configured to respectively cluster the load data sets corresponding to each group of feature weights into n sub-load data sets and calculate the distance gain corresponding to each group of feature weights; where n≥2; the distance gain characterizes the deviation between the first distance and the second distance, and the second distance characterizes the clustering effect of the n sub-load data sets;
[0028] A determination module, configured to determine n sub-load data sets corresponding to the feature weights with the largest distance gain as the clustering result of the current round of clustering;
[0029] A judgment module, configured to judge whether the current round of clustering meets the clustering stop condition; in the case where the judgment result is no, call the calculation module; in the case where the judgment result is yes, call the generation module;
[0030] The generation module is configured to generate a clustering result.
[0031] Optionally, the first distance is determined according to the distances from the respective data samples included in the load data set to the clustering center;
[0032] The second distance is determined according to the distances of the n sub-load data sets, and the distances of the respective sub-load data sets are determined according to the distances from the respective data samples included in the sub-load data set to the clustering center.
[0033] Optionally, the clustering stop condition includes at least one of the following:
[0034] The number of clustering times is greater than the number threshold;
[0035] There is at least one sub-load data set among the n sub-load data sets corresponding to the feature weights with the largest distance gain that contains a number of data samples less than the number threshold;
[0036] The distance gain is less than the gain threshold.
[0037] Optionally, the clustering module is specifically configured to:
[0038] Cluster the load data sets corresponding to each group of feature weights into n sub-load data sets according to the K-medoids algorithm.
[0039] Optionally, it further includes:
[0040] A preprocessing module, configured to preprocess the data samples included in each load data set; and / or, perform normalization processing on the data samples included in each load data set.
[0041] In a third aspect, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the clustering method of the load data described in any one of the above is implemented.
[0042] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored, and characterized in that when the computer program is executed by a processor, the clustering method of the load data described in any one of the above is implemented.
[0043] The positive and progressive effects of the present invention are as follows: In the embodiments of the present invention, by traversing the clustering effects under different feature weights, a set of optimal feature weights for each round of clustering can be selected, and a set of optimal clustering features can be determined, so as to well adapt to the clustering of different types of load data and have good clustering effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a flowchart of a clustering method for load data provided by an exemplary embodiment of the present invention;
[0045] Figure 2 It is a flowchart of another clustering method for load data provided by an exemplary embodiment of the present invention;
[0046] Figure 3 It is a schematic diagram of the usage scenario of a clustering method for load data provided by an exemplary embodiment of the present invention;
[0047] Figure 4a It is a schematic diagram of a clustering result obtained by using the clustering method for load data provided by the embodiments of the present invention;
[0048] Figure 4b It is a schematic diagram of another clustering result obtained by using the clustering method for load data provided by the embodiments of the present invention;
[0049] Figure 4c It is a schematic diagram of another clustering result obtained by using the clustering method for load data provided by the embodiments of the present invention;
[0050] Figure 5 It is a schematic diagram of the modules of a clustering device for load data provided by an exemplary embodiment of the present invention;
[0051] Figure 6 It is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present invention;
[0052] Figure 7 It is a schematic diagram of the comparison of the clustering effects of the clustering method for load data in the embodiments of the present invention with different n. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] The present invention will be further described below by way of embodiments, but the present invention is not limited to the scope of the described embodiments.
[0054] Figure 1 It is a flowchart of a clustering method for load data provided by an exemplary embodiment of the present invention, and the clustering method includes the following steps:
[0055] Step 101, calculate the first distance characterizing the clustering effect of each load data set.
[0056] In the embodiments of the present invention, through multiple rounds of clustering on the load data set to be clustered, a clustering result is finally obtained. For each round of clustering, the load data set is clustered into n load data sets. It can be understood that after the first round of clustering, the load data set to be clustered is clustered into n load data sets, and in the second round of clustering, all or part of the n load data sets are clustered respectively, and so on. Therefore, the number of load data sets for calculating the first distance in step 101 is related to the number of clustering rounds.
[0057] The load data set contains multiple data samples; each data sample represents the load data of a load curve. The load curve can be an active power load curve, a reactive power load curve, or other types of load curves, and the embodiments of the present invention do not make special limitations on this.
[0058] In one embodiment, the first distance is determined according to the distances from the respective data samples included in the load data set to the clustering center, and can be the sum of the distances from the respective data samples to the clustering center, or the average value or standard deviation of the distances from the respective data samples to the clustering center, etc., or can also be the weighted result of at least two of the above.
[0059] Taking the load data set containing M data samples and using the average value of the distances from the respective data samples to the clustering center as the first distance as an example, the calculation process of the first distance is introduced below. The calculation formula of the first distance of the load data set can be but is not limited to being expressed as follows:
[0060]
[0061]
[0062]
[0063] Among them, each data sample is represented by N sampling points sampled in time sequence on the load curve; D 1 represents the first distance; d j represents the distance from the jth data sample in the load data set to the clustering center (cluster center) ; x ij represents the ith sampling point of the jth data sample.
[0064] Before the first round of clustering, there is only one load data set. When calculating the first distance, the load data features included in this load data set are extracted as a point set, and a target point is selected. The sum of the distances from this target point to all other points in the point set is the smallest, and this target point is used as the center point. This characteristic center point is mapped to the corresponding load curve, and this load curve is used as the cluster center for calculating the first distance before the first round of clustering. Compared with directly finding the cluster center on the load curve, in the embodiment of the present invention, the cluster center is selected on the features corresponding to the load, effectively reducing the dimension and improving the calculation speed.
[0065] In one embodiment, before calculating the first distance, the following processing is performed on the load data set: preprocessing the data samples included in each load data set; and / or, normalizing the data samples included in each load data set. Through preprocessing, abnormal data can be eliminated to ensure the smooth progress of subsequent calculations and the accuracy of subsequent calculations. After normalization processing, the order of magnitude of the features of the load data can be offset, facilitating the selection of weights.
[0066] Step 102: Extract the features of the data samples included in each load data set and set multiple groups of feature weights.
[0067] The features of the data samples can include but are not limited to at least one of the maximum value, minimum value, mean value, integral value, standard deviation, variance, skewness, kurtosis, etc. of the load. Different features correspond to different characteristics of the load distribution. For example, skewness can describe the asymmetry of the load distribution. Kurtosis can describe the steepness of the load distribution.
[0068] The number of groups of feature weights can be set according to the actual situation, and each group of feature weights is different from each other. The number of weights in each group of feature weights corresponds to the number of features extracted. For example, if 3 features of the data samples are extracted, namely the maximum value, integral, and standard deviation, then each group of feature weights includes 3 weights, namely the weight of the maximum value, the weight of the integral, and the weight of the standard deviation. The weights of each feature can be set according to actual needs. For example, the weight is positively correlated with the importance of the feature.
[0069] It should be noted that the execution order of step 101 and step 102 is not limited to first calculating the first distance and then performing feature extraction as shown in the figure. Step 101 and step 102 can be executed synchronously, that is, calculating the first distance and feature extraction are executed synchronously; or step 102 can be executed first and then step 101, that is, first performing feature extraction and then calculating the first distance.
[0070] Step 103: Cluster the load data sets corresponding to each group of feature weights into n sub-load data sets respectively, and calculate the distance gain corresponding to each group of feature weights.
[0071] Among them, n≥2, and n can be set according to actual needs.
[0072] For example, when n is equal to 2, after each round of clustering, for each group of feature weights, the load data set is clustered into 2 sub-load data sets; when n is equal to 3, after each round of clustering, the load data set is clustered into 3 sub-load data sets.
[0073] When the number of groups of feature weights is h, each round of clustering will obtain h clustering results. When performing clustering, all groups of feature weights are traversed, and the corresponding load data sets are clustered. For example, if 3 groups of feature weights are set, the load data set needs to be clustered 3 times, and correspondingly, 3 distance gains can be obtained.
[0074] The distance gain characterizes the deviation between the first distance and the second distance, and the second distance characterizes the clustering effect of n sub-load data sets.
[0075] In one embodiment, the deviation between the first distance and the second distance is characterized by the difference between the first distance and the second distance, and the distance gain formula is expressed as follows:
[0076] D g = D 1 - D 2 ;
[0077] In one embodiment, the deviation between the first distance and the second distance is characterized by the ratio of the first distance to the second distance, and the distance gain formula is expressed as follows:
[0078] D g = D 1 / D 2 ;
[0079] In one embodiment, the deviation between the first distance and the second distance can also be characterized by the standard deviation, and the specific implementation method will not be elaborated here.
[0080] In one embodiment, the second distance is determined according to the distances of n sub-load data sets. For example, the second distance is the sum of the distances of n sub-load data sets, or the second distance is the average value or standard deviation of the distances of n sub-load data sets. The second clustering can also be the weighted result of the above items; the distance of each sub-load data set is determined according to the distances of each data sample included in the sub-load data set to the clustering center. For example, the distance of each sub-load data set is the sum or average value or standard deviation of the distances of each data sample to the clustering center, and it can also be the weighted result of at least two of the above items.
[0081] Taking the sum of the distances of n sub-load data sets as the second distance and the average of the distances from each data sample to the cluster center as the distance of the sub-load data set as an example, the calculation process of the second distance is introduced below. The calculation formula of the second distance of the load data set can be but not limited to expressed as follows:
[0082]
[0083]
[0084] Among them, D 2 represents the second distance; represents the distance of the k-th sub-load data set obtained by clustering. represents the cluster center of the k-th sub-load data set.
[0085] In one embodiment, in step 103, the load data sets corresponding to each group of feature weights are clustered into n sub-load data sets according to the K-medoids algorithm. The features extracted from the load curve are clustered using the K-medoids algorithm to obtain the features of k cluster centers, and the load curves corresponding to the features are the corresponding cluster centers. In the embodiment of the present invention, instead of directly clustering the load data, the features are clustered, which can effectively reduce the dimension. At the same time, since the k-medoids is used, the load curves corresponding to the features can also be found.
[0086] Step 105: Determine the n sub-load data sets corresponding to the feature weights with the largest distance gain as the clustering result of this round of clustering.
[0087] The larger the distance gain indicates that the clustering result is more accurate and ideal, and it is the best set of splitting features. In step 105, that is, select the n sub-load data sets corresponding to the feature weights with the largest distance gain from h clustering results as the clustering result of this round of clustering.
[0088] Step 106: Determine whether the clustering result of this round of clustering meets the clustering stop condition.
[0089] In step 106, if the judgment result is no, it means that the clustering result obtained through this round of clustering is not ideal enough, then return to step 101 for the next round of clustering; if the judgment result is yes, stop clustering and generate the clustering result, that is, use the multiple load data sets obtained through this round of clustering as the clustering result of the load data sets to be clustered.
[0090] In one embodiment, the clustering stop condition includes at least one of the following: the number of clustering times is greater than the number threshold; at least one of the n sub-load data sets corresponding to the feature weights with the largest distance gain contains a number of data samples less than the number threshold; the distance gain is less than the gain threshold. The final number of clustering categories can be automatically determined through the clustering stop condition.
[0091] Among them, the above-mentioned number threshold, number threshold, and gain threshold can be set according to the actual situation.
[0092] In the embodiment of the present invention, by traversing the clustering effects under different feature weights, a set of optimal feature weights for each round of clustering can be selected, and the best set of clustering features can be determined, so as to well adapt to the clustering of different types of load data, judge the influence of different category information obtained by clustering on the final clustering result, make the clustering model interpretable, and have good clustering effect.
[0093] In one embodiment, the first distance and the second distance are determined by the silhouette distance, that is, the silhouette distance is used as an effect measurement index for different numbers of clusters. The silhouette coefficient formula for a single data sample j is:
[0094]
[0095] Among them, S i represents the silhouette distance, also known as the silhouette coefficient; a j represents the average distance of the j-th data sample to other points within the same category; b j represents the average distance of the j-th data sample to the data sample in the nearest different category. The silhouette coefficient of the entire load data set is the average value of the silhouette coefficients of all samples.
[0096] See Figure 7 , the figure shows a schematic diagram of the silhouette coefficients when n is 2 to 11 respectively. The silhouette coefficient characterizes the clustering effect. In the figure, the abscissa "Number of clusters" represents n, and the ordinate "silhouette score" represents the silhouette coefficient. The larger this value of the silhouette coefficient, the better the clustering effect. It can be seen from the figure that when n is selected from 2 to 5, a better clustering effect can be obtained.
[0097] In Figure 1 On the basis of the clustering method of the load data shown, Figure 2 is a flowchart of another clustering method for load data provided by an exemplary embodiment of the present invention. In the figure, taking n = 2, that is, each round of clustering clusters the load data set into 2 sub-load data sets as an example, the clustering process of the load data is described. See Figure 2 , this clustering method includes the following steps:
[0098] Step 201: Obtain a load data set, and perform preprocessing, normalization processing, and feature extraction on the load data set.
[0099] Among them, the features of the data samples may include, but are not limited to, at least one of the maximum value, minimum value, mean value, integral value, standard deviation, variance, skewness, kurtosis, etc. of the load.
[0100] Step 202: Calculate a first distance characterizing the clustering effect of the load data set.
[0101] Before the first round of clustering, there is only one load data set, and one first distance is calculated. The first distance is determined according to the distances between each data sample in the load data set to be clustered and the clustering center M1 of the load data set to be clustered. For the specific calculation process, refer to Step 101.
[0102] Step 203: Set h groups of feature weights.
[0103] Each group of feature weights is different from each other.
[0104] Step 204: For each group of feature weights, cluster the load data set into 2 sub-load data sets, and calculate the distance gain corresponding to each group of feature weights.
[0105] See Figure 2 , after the first round of clustering, the load data set a is divided into 2 sub-load data sets of the load data set a, namely the load data set b and the load data set c. For the h groups of feature weights, h Figure 2 As shown in the results, correspondingly h distance gains are obtained. One needs to be selected from the h results as the basis for the next round of clustering, then execute Step 205.
[0106] For each group of feature weights, the calculation formula of the distance gain may be expressed as follows, but is not limited to:
[0107] D g = D 1 - D 2 ;
[0108]
[0109] Among them, respectively represent the 2 sub-load data sets obtained by clustering.
[0110] Step 205: Determine the 2 sub-load data sets corresponding to the feature weights with the largest distance gain as the clustering result of this round of clustering.
[0111] See Figure 3, the two sub-load datasets corresponding to the feature weights with the largest distance gain in the figure are the load dataset b and the load dataset c respectively, and the clustering results of the first round of clustering are generated based on them.
[0112] Step 206: Determine whether the clustering result of this round of clustering meets the clustering stop condition.
[0113] In step 206, if the judgment result is yes, stop clustering and generate the clustering result. If the judgment result is no, it means that the clustering result obtained through this round of clustering is not ideal enough, then return to step 201 for the second round of clustering. The load datasets for the second round of clustering are the leaf nodes of the first round of clustering. If both of the two load datasets do not meet the requirements, then perform steps 201-206 on these two load datasets respectively; if only one of the two load datasets does not meet the requirements, then perform steps 201-206 on this load dataset until the clustering stop condition is reached.
[0114] See Figure 3 , for the second round of clustering, perform steps 201-206 on the load dataset b and the load dataset c respectively, that is, cluster the load dataset b and the load dataset c respectively. The clustering of the load dataset b and the load dataset c is independent of each other and can be executed in parallel to improve the clustering efficiency. After the second round of clustering, the two sub-load datasets corresponding to the feature weights with the largest distance gain of the load dataset b are the load dataset d and the load dataset e respectively, and the load dataset d and the load dataset e are used as the leaf nodes of the load dataset b; the two sub-load datasets corresponding to the feature weights with the largest distance gain of the load dataset c are the load dataset f and the load dataset g respectively, and the load dataset f and the load dataset g are used as the leaf nodes of the load dataset c; the load dataset d, the load dataset e and the load dataset f meet the clustering stop condition, so there is no need to perform the third round of distance calculation. The load dataset g still does not meet the clustering stop condition and needs to perform the third round of clustering. Perform steps 201-206 on the load dataset g. After the third round of clustering, the two sub-load datasets corresponding to the feature weights with the largest distance gain of the load dataset g are the load dataset h and the load dataset i respectively, and the load dataset h and the load dataset i are used as the leaf nodes of the load dataset g. At this time, the clustering stop condition is met, so stop clustering, and cluster the load dataset a into 5 load datasets, namely the load dataset d, the load dataset e, the load dataset f, the load dataset h and the load dataset i.
[0115] In the embodiments of the present invention, by traversing the clustering effects under different feature weights, a set of optimal feature weights for each round of clustering is selected. Meanwhile, only the clustering stop condition needs to be set. Once the load data set obtained by clustering triggers the clustering stop condition, it will not continue to grow downward. Finally, all the leaf nodes are the required load category information, and all the leaf nodes are used as the typical features of the load curve obtained by clustering.
[0116] Taking the data of a certain enterprise in 2021 for analysis, which includes a total of 32,406 pieces of data, and the data is collected every 15 minutes. The original data selects the active power to represent the load. Before clustering, the original data is preprocessed to remove null values and outliers. Taking the time in days as the unit, the relevant features of the load curve are extracted. Here, the maximum value, minimum value, mean value, integral value, standard deviation, variance, skewness, and kurtosis of the active power are selected as the feature representatives. When performing the first clustering operation, a set of weights for the features is selected. Figure 4a For the distance results shown, it is selected to set the weights of the maximum value, minimum value, mean value, and integral system to be larger, indicating that more attention is paid to the overall load quantity. Figure 4b The figure shows the clustering result diagram of selecting another set of feature weights. The weights of skewness and kurtosis are set to be larger, indicating that more attention is paid to the morphological distribution of the load, such as whether it is symmetric, whether it is concentrated, etc.
[0117] It can be seen that different weights correspond to different features that the clustering focuses on. Then, the weight corresponding to the best distance gain is selected as the classification feature point of this layer.
[0118] See Figure 4c , the figure shows the clustering result diagram of obtaining six results. It can be seen from the figure that the clustering effect is good. In the figure, the first type of load curve all occurs in June, July, August, and September, and all appears from Monday to Saturday, with the highest load, presumably due to the high electricity load caused by air conditioners in summer; the second / third type of load curve has the highest proportion and is distributed throughout the year. The peak electricity consumption is from 9 am to 6 pm, and the load trough is during the one-hour lunch break; the fourth type of load curve is concentrated in February; the fifth type of load curve mostly occurs during Saturdays and Sundays; the sixth type of load all occurs during the Spring Festival or National Day holidays, with a relatively low overall load and a relatively uniform fluctuation.
[0119] Corresponding to the foregoing embodiments of the clustering method for load data, the present invention also provides an embodiment of a clustering device for load data.
[0120] Figure 5 The module schematic diagram of a clustering device for load data provided by an exemplary embodiment of the present invention. The clustering device includes:
[0121] A calculation module 51 for calculating a first distance characterizing the clustering effect of each load data set; the load data set includes a plurality of data samples; each data sample characterizes the load data of a load curve.
[0122] An extraction module 52 for extracting features from the data samples included in each load data set and setting multiple groups of feature weights.
[0123] A clustering module 53 for clustering the load data sets corresponding to each group of feature weights into n sub-load data sets respectively and calculating the distance gain corresponding to each group of feature weights; where n≥2; the distance gain characterizes the deviation between the first distance and the second distance, and the second distance characterizes the clustering effect of the n sub-load data sets.
[0124] A determination module 54 for determining the n sub-load data sets corresponding to the feature weights with the largest distance gain as the clustering result of this round of clustering.
[0125] A judgment module 55 for judging whether this round of clustering meets the clustering stop condition; in the case where the judgment result is no, the calculation module is called; in the case where the judgment result is yes, the generation module is called.
[0126] The generation module 56 for generating a clustering result.
[0127] Optionally, the first distance is determined according to the distances of the respective data samples included in the load data set to the cluster center.
[0128] The second distance is determined according to the distances of the n sub-load data sets, and the distances of each sub-load data set are determined according to the distances of the respective data samples included in the sub-load data set to the cluster center.
[0129] Optionally, the clustering stop condition includes at least one of the following:
[0130] The number of clustering times is greater than the number threshold.
[0131] There is at least one sub-load data set among the n sub-load data sets corresponding to the feature weights with the largest distance gain that contains a number of data samples less than the number threshold.
[0132] The distance gain is less than the gain threshold.
[0133] Optionally, the clustering module is specifically configured to:
[0134] Cluster the load data sets corresponding to each group of feature weights into n sub-load data sets according to the K-medoids algorithm.
[0135] Optionally, it further includes:
[0136] A preprocessing module for preprocessing data samples included in each load data set; and / or normalizing the data samples included in each load data set.
[0137] For the apparatus embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The apparatus embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0138] Figure 6 FIG. is a schematic structural diagram of an electronic device shown in an exemplary embodiment of the present invention, showing a block diagram of an exemplary electronic device 60 suitable for implementing the embodiments of the present invention. Figure 6 The shown electronic device 60 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0139] As Figure 6 shown, the electronic device 60 may be presented in the form of a general-purpose computing device, for example, it may be a server device. The components of the electronic device 60 may include but are not limited to: at least one of the above-mentioned processors 61, at least one of the above-mentioned memories 62, and a bus 63 connecting different system components (including the memory 62 and the processor 61).
[0140] The bus 63 includes a data bus, an address bus, and a control bus.
[0141] The memory 62 may include volatile memory, such as a random access memory (RAM) 621 and / or a cache memory 622, and may further include a read-only memory (ROM) 623.
[0142] The memory 62 may further include a program tool 625 (or utility) having a set (at least one) of program modules 624. Such program modules 624 include but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0143] The processor 61 executes various functional applications and data processing by running computer programs stored in the memory 62, such as the methods provided in any of the above embodiments.
[0144] The electronic device 60 can also communicate with one or more external devices 64 (such as a keyboard, a pointing device, etc.). Such communication can be carried out through the input / output (I / O) interface 65. Moreover, the model-generated electronic device 60 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 66. As shown in the figure, the network adapter 66 communicates with other modules of the model-generated electronic device 60 through the bus 63. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the model-generated electronic device 60, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (redundant array of independent disks) systems, tape drives, and data backup storage systems, etc.
[0145] It should be noted that, although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present invention, the features and functions of two or more of the above-described units / modules can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.
[0146] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method provided by any one of the above embodiments is implemented.
[0147] Among them, the more specific forms that the readable storage medium can adopt can include but are not limited to: portable disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories, optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0148] In a possible implementation manner, the embodiments of the present invention can also be implemented in the form of a program product, which includes program code, and when the program product runs on a terminal device, the program code is used to enable the terminal device to execute the method provided by any one of the above embodiments.
[0149] Among them, the program code for executing the present invention can be written in any combination of one or more programming languages, and the program code can be executed entirely on the user device, partially on the user device, executed as an independent software package, partially on the user device and partially on a remote device, or executed entirely on a remote device.
[0150] Although the specific embodiments of the present invention have been described above, those skilled in the art should understand that this is only an example, and the protection scope of the present invention is defined by the appended claims. Without departing from the principle and essence of the present invention, those skilled in the art can make various changes or modifications to these embodiments, but these changes and modifications all fall within the protection scope of the present invention.
Claims
1. A clustering method for load data, characterized in that, it includes: Calculating a first distance characterizing the clustering effect of each load data set; The load data set contains multiple data samples; Each data sample represents the load data of a load curve; Performing feature extraction on the data samples included in each load data set and setting multiple groups of feature weights; Respectively clustering the load data sets corresponding to each group of feature weights into n sub-load data sets, and calculating the distance gain corresponding to each group of feature weights; where n≥2; the distance gain characterizes the deviation between the first distance and the second distance, and the second distance characterizes the clustering effect of the n sub-load data sets; The first distance is determined according to the distances from the respective data samples included in the load data set to the clustering center; the second distance is determined according to the distances of the n sub-load data sets, and the distances of each sub-load data set are determined according to the distances from the respective data samples included in the sub-load data set to the clustering center; Determining the n sub-load data sets corresponding to the feature weight with the largest distance gain as the clustering result of this round of clustering; Judging whether the clustering result of this round of clustering meets the clustering stop condition; In the case where the judgment result is no, returning to the step of calculating the first distance of each load data set; in the case where the judgment result is yes, generating a clustering result.
2. The clustering method for load data according to claim 1, characterized in that, The clustering stop condition includes at least one of the following: The number of clustering times is greater than the number threshold; There is at least one sub-load data set among the n sub-load data sets corresponding to the feature weight with the largest distance gain that contains a number of data samples less than the number threshold; The distance gain is less than the gain threshold.
3. The clustering method for load data according to claim 1, characterized in that, Clustering the load data sets corresponding to each group of feature weights into n sub-load data sets includes: Clustering the load data sets corresponding to each group of feature weights into n sub-load data sets according to the K-medoids algorithm.
4. The clustering method for load data according to claim 1, characterized in that, Before performing feature extraction on the data samples included in each load data set, it includes: Performing preprocessing on the data samples included in each load data set; and / or, performing normalization processing on the data samples included in each load data set.
5. A clustering device for load data, characterized in that, it includes: A calculation module for calculating a first distance characterizing the clustering effect of each load data set; The load data set contains multiple data samples; Each data sample represents the load data of a load curve; An extraction module for performing feature extraction on the data samples included in each load data set and setting multiple groups of feature weights; A clustering module for respectively clustering the load data sets corresponding to each group of feature weights into n sub-load data sets and calculating the distance gain corresponding to each group of feature weights; where n≥2; the distance gain characterizes the deviation between the first distance and the second distance, and the second distance characterizes the clustering effect of the n sub-load data sets; The first distance is determined according to the distances of the respective data samples included in the load data set to the cluster center; the second distance is determined according to the distances of the n sub-load data sets, and the distances of the respective sub-load data sets are determined according to the distances of the respective data samples included in the sub-load data set to the cluster center; A determination module, configured to determine the n sub-load data sets corresponding to the feature weights with the largest distance gain as the clustering result of the current round of clustering; A judgment module, configured to judge whether the clustering result of the current round of clustering meets the clustering stop condition; in the case that the judgment result is negative, call the calculation module; in the case that the judgment result is positive, call the generation module; The generation module is configured to generate a clustering result.
6. The clustering device for load data according to claim 5, wherein, The clustering stop condition includes at least one of the following: The number of clustering times is greater than the number threshold; There is at least one sub-load data set among the n sub-load data sets corresponding to the feature weights with the largest distance gain, and the number of data samples included in the sub-load data set is less than the number threshold; The distance gain is less than the gain threshold.
7. The clustering device for load data according to claim 5, wherein, The clustering module is specifically configured to: Cluster the load data sets corresponding to each group of feature weights into n sub-load data sets according to the K-medoids algorithm.
8. The clustering device for load data according to claim 5, wherein, It further includes: A preprocessing module, configured to preprocess the data samples included in each load data set; and / or, perform normalization processing on the data samples included in each load data set.
9. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the computer program, it implements the clustering method for load data according to any one of claims 1 to 4.
10. A computer-readable storage medium, on which a computer program is stored, wherein, When the computer program is executed by a processor, it implements the clustering method for load data according to any one of claims 1 to 4.