User clustering method, apparatus, device, storage medium, and program product
By adjusting the Gaussian mixture model and using the Bayesian information criterion to determine the clustering termination condition, the problem of difficulty in determining the number of clusters caused by the diversity of user numbers was solved, achieving more efficient and accurate user clustering.
Patent Information
- Application Number
- CN202311057898.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-21
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2043-08-21
AI Technical Summary
Existing user clustering methods suffer from poor clustering results due to the large number of users and the diversity of electricity load curves, making it difficult and unsuitable to predetermine the number of clusters.
By adjusting the initial Gaussian mixture model based on historical load curves, the target Gaussian mixture model is determined, and the Bayesian information criterion is used to determine the clustering termination condition, avoiding the pre-setting of the number of clusters. User clustering is performed using the Gaussian mixture model and the Bayesian information criterion.
It improves the accuracy and efficiency of user clustering, ensures the representativeness and accuracy of clustering results, and avoids errors caused by an inappropriate number of clusters.
Smart Images

Figure CN117171600B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart grid big data technology, and in particular to a user clustering method, apparatus, device, storage medium and program product. Background Technology
[0002] In power systems, clustering users based on their electricity load curves is crucial for electricity load analysis and optimization. Currently, clustering users requires processing each user's electricity load curve using statistical analysis or clustering algorithms to determine the number of clusters. In other words, related technologies necessitate pre-determining the number of clusters when clustering users.
[0003] However, due to the massive number of users and the diversity of electricity load curves, the types and number of user electricity load curves are usually unknown. Predetermining the number of clusters is extremely difficult, and even a predetermined number may not be appropriate, resulting in poor user clustering performance in related technologies. Therefore, there is an urgent need to provide a user clustering method that can improve the accuracy of user clustering. Summary of the Invention
[0004] Therefore, it is necessary to provide a user clustering method, apparatus, device, storage medium, and program product that can improve the accuracy of user clustering in response to the above-mentioned technical problems.
[0005] Firstly, this application provides a user clustering method. The method includes:
[0006] Based on the historical load curves of the users to be clustered within a historical time period, determine the sub-probability values of each probability distribution function in the initial Gaussian mixture model corresponding to the historical load curves. Then, adjust the probability distribution functions contained in the initial Gaussian mixture model according to the sub-probability values of each probability distribution function corresponding to the historical load curves to obtain the target Gaussian mixture model.
[0007] Based on the sub-probability values of the probability distribution functions in the target Gaussian mixture model corresponding to the historical load curves of the users to be clustered, the target load curves corresponding to the users to be clustered are determined.
[0008] Based on the target load curves corresponding to the users to be clustered, the users to be clustered are clustered, and based on the Bayesian information criterion values corresponding to each type of user after clustering, it is determined whether each type of user after clustering has met the clustering termination condition.
[0009] If so, then obtain the clustering results for the users to be clustered.
[0010] In one embodiment, based on the historical load curves of the users to be clustered within a historical time period, the sub-probability values of the historical load curves belonging to each probability distribution function in the initial Gaussian mixture model are determined, including:
[0011] Based on the historical load curves of the users to be clustered within the historical time period, determine the probability density value of the historical load curve belonging to the initial Gaussian mixture model.
[0012] Based on the probability density values corresponding to the historical load curves of the users to be clustered, and the weights of each probability distribution function included in the initial Gaussian mixture model, the sub-probability values of the historical load curves belonging to each probability distribution function in the initial Gaussian mixture model are determined.
[0013] In one embodiment, the probability distribution functions contained in the initial Gaussian mixture model are adjusted according to the sub-probability values of each probability distribution function corresponding to the historical load curve to obtain the target Gaussian mixture model, including:
[0014] Based on the sub-probability values of each probability distribution function corresponding to the historical load curves of the users to be clustered, determine the total risk value corresponding to each probability distribution function for the historical load curves of the users to be clustered.
[0015] Determine whether the total risk value is less than the risk threshold;
[0016] If so, the initial Gaussian mixture model corresponding to the total risk value will be used as the target Gaussian mixture model.
[0017] If not, the probability distribution functions contained in the initial Gaussian mixture model are adjusted, and the operation of determining the sub-probability values of each probability distribution function in the initial Gaussian mixture model is returned based on the historical load curves corresponding to the users to be clustered in the historical time period.
[0018] In one embodiment, the historical load curve includes at least two sub-curve segments; the target load curve corresponding to the user to be clustered is determined based on the sub-probability values of the historical load curves corresponding to the users in the target Gaussian mixture model, including:
[0019] Based on the sub-probability values of each sub-curve segment of the user to be clustered belonging to each probability distribution function in the target Gaussian mixture model, determine the typical load curve corresponding to each probability distribution function from each sub-curve segment of the user to be clustered;
[0020] Based on the sub-probability values of each sub-curve segment of the user to be clustered belonging to each probability distribution function in the target Gaussian mixture model, determine the typical load curve corresponding to each sub-curve segment of the user to be clustered.
[0021] Based on the typical load curves corresponding to each sub-curve segment of the user to be clustered, the target load curve of the user to be clustered is determined from the typical load curves corresponding to each probability distribution function.
[0022] In one embodiment, after determining whether all clustered users have met the clustering termination condition, the method further includes:
[0023] Users who have not met the clustering termination criteria are treated as users to be clustered, and the operation of clustering users to be clustered is performed.
[0024] In one embodiment, based on the Bayesian information criterion values corresponding to each type of user after clustering, it is determined whether all types of users after clustering have met the clustering termination condition, including:
[0025] Determine whether there is no target Bayesian information criterion value less than the historical minimum Bayesian information criterion value among the Bayesian information criterion values of each type of user after this round of clustering; where the historical minimum Bayesian information criterion value is the minimum Bayesian information criterion value among the Bayesian information criterion values of each type of user obtained from previous rounds of clustering.
[0026] If so, then it is determined that all users in each cluster have met the clustering termination condition;
[0027] If not, then it is determined that there are users in the class corresponding to the target Bayesian information criterion value that is less than the historical minimum Bayesian information criterion value, and the clustering termination condition has not been met.
[0028] In one embodiment, the method further includes:
[0029] Obtain the unit load curves for each time period of the users to be clustered within the historical time period;
[0030] The load curves of each unit corresponding to the user to be clustered are subjected to load reduction processing to obtain the historical load curves of the user to be clustered within the historical time period.
[0031] Secondly, this application also provides a user clustering apparatus. The apparatus includes:
[0032] The target model determination module is used to determine the sub-probability values of the historical load curves corresponding to the users to be clustered within the historical time period, and to adjust the probability distribution functions contained in the initial Gaussian mixture model according to the sub-probability values of the probability distribution functions corresponding to the historical load curves, so as to obtain the target Gaussian mixture model.
[0033] The target curve determination module is used to determine the target load curve corresponding to the user to be clustered based on the sub-probability values of the probability distribution functions in the target Gaussian mixture model corresponding to the historical load curve of the user to be clustered.
[0034] The clustering termination judgment module is used to cluster users according to the target load curves corresponding to the users to be clustered, and to determine whether all types of users after clustering have met the clustering termination condition based on the Bayesian information criterion values corresponding to each type of user after clustering.
[0035] The clustering result retrieval module is used to retrieve the clustering results of the user to be clustered if the clustering is specified.
[0036] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0037] Based on the historical load curves of the users to be clustered within a historical time period, determine the sub-probability values of each probability distribution function in the initial Gaussian mixture model corresponding to the historical load curves. Then, adjust the probability distribution functions contained in the initial Gaussian mixture model according to the sub-probability values of each probability distribution function corresponding to the historical load curves to obtain the target Gaussian mixture model.
[0038] Based on the sub-probability values of the probability distribution functions in the target Gaussian mixture model corresponding to the historical load curves of the users to be clustered, the target load curves corresponding to the users to be clustered are determined.
[0039] Based on the target load curves corresponding to the users to be clustered, the users to be clustered are clustered, and based on the Bayesian information criterion values corresponding to each type of user after clustering, it is determined whether each type of user after clustering has met the clustering termination condition.
[0040] If so, then obtain the clustering results for the users to be clustered.
[0041] Fourthly, this application also provides a computer-readable storage medium. This computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0042] Based on the historical load curves of the users to be clustered within a historical time period, determine the sub-probability values of each probability distribution function in the initial Gaussian mixture model corresponding to the historical load curves. Then, adjust the probability distribution functions contained in the initial Gaussian mixture model according to the sub-probability values of each probability distribution function corresponding to the historical load curves to obtain the target Gaussian mixture model.
[0043] Based on the sub-probability values of the probability distribution functions in the target Gaussian mixture model corresponding to the historical load curves of the users to be clustered, the target load curves corresponding to the users to be clustered are determined.
[0044] Based on the target load curves corresponding to the users to be clustered, the users to be clustered are clustered, and based on the Bayesian information criterion values corresponding to each type of user after clustering, it is determined whether each type of user after clustering has met the clustering termination condition.
[0045] If so, then obtain the clustering results for the users to be clustered.
[0046] Fifthly, this application also provides a computer program product. This computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0047] Based on the historical load curves of the users to be clustered within a historical time period, determine the sub-probability values of each probability distribution function in the initial Gaussian mixture model corresponding to the historical load curves. Then, adjust the probability distribution functions contained in the initial Gaussian mixture model according to the sub-probability values of each probability distribution function corresponding to the historical load curves to obtain the target Gaussian mixture model.
[0048] Based on the sub-probability values of the probability distribution functions in the target Gaussian mixture model corresponding to the historical load curves of the users to be clustered, the target load curves corresponding to the users to be clustered are determined.
[0049] Based on the target load curves corresponding to the users to be clustered, the users to be clustered are clustered, and based on the Bayesian information criterion values corresponding to each type of user after clustering, it is determined whether each type of user after clustering has met the clustering termination condition.
[0050] If so, then obtain the clustering results for the users to be clustered.
[0051] The aforementioned user clustering method, apparatus, device, storage medium, and program product determine the target Gaussian mixture model by identifying the sub-probability values of the historical load curves corresponding to the probability distribution functions in the initial Gaussian mixture model based on the historical load curves of the users to be clustered within a historical time period. Since the probability distribution functions in the target Gaussian mixture model are determined based on the sub-probability values of the historical load curves belonging to the probability distribution functions in the initial Gaussian mixture model, adding the step of determining the target Gaussian mixture model makes the target load curves corresponding to the users to be clustered more representative. Furthermore, since each user to be clustered corresponds to only one target load curve, rather than all historical load curves, the efficiency of clustering users to be clustered can be improved. In addition, during the clustering process, there is no need to pre-set the number of clusters; instead, the clustering termination condition is determined by calculating the BIC corresponding to each type of user. When the BIC corresponding to each type of user reaches the termination condition, the clustering result (including the number of clusters and each type of user) can be determined, resulting in better and more accurate clustering of users to be clustered. Attached Figure Description
[0052] Figure 1 This is an application environment diagram of a user clustering method provided in this embodiment;
[0053] Figure 2 This is a flowchart illustrating the first user clustering method provided in this embodiment;
[0054] Figure 3 A flowchart illustrating a method for determining a historical load curve provided in this embodiment;
[0055] Figure 4 This embodiment provides a flowchart for determining the target load curve corresponding to a user to be clustered.
[0056] Figure 5 This is a flowchart illustrating the second user clustering method provided in this embodiment;
[0057] Figure 6 This is a structural block diagram of the first user clustering device provided in this embodiment;
[0058] Figure 7 This is a structural block diagram of the second type of user clustering device provided in this embodiment;
[0059] Figure 8 This is a structural block diagram of the third type of user clustering device provided in this embodiment;
[0060] Figure 9 This is a structural block diagram of the fourth user clustering device provided in this embodiment;
[0061] Figure 10 This is a structural block diagram of the fifth user clustering device provided in this embodiment;
[0062] Figure 11 This is a structural block diagram of the sixth user clustering device provided in this embodiment;
[0063] Figure 12 This is a structural block diagram of the seventh user clustering device provided in this embodiment;
[0064] Figure 13 This is an internal structural diagram of a computer device provided in this embodiment. Detailed Implementation
[0065] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0066] The user clustering method provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, in one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows. Figure 1 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores relevant data for user clustering. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a user clustering method.
[0067] In one embodiment, such as Figure 2 As shown, a user clustering method is provided, which can be applied to... Figure 1 Taking a computer device as an example, the explanation includes the following steps:
[0068] S201. Based on the historical load curves of the users to be clustered within the historical time period, determine the sub-probability values of the probability distribution functions in the initial Gaussian mixture model that the historical load curves belong to. Based on the sub-probability values of the probability distribution functions corresponding to the historical load curves, adjust the probability distribution functions contained in the initial Gaussian mixture model to obtain the target Gaussian mixture model.
[0069] In this study, the users to be clustered can be any users requiring cluster analysis, and the number of users to be clustered is not unique. Taking electricity users under a certain power grid as an example, the purpose of clustering these users is to classify the power grid users, thereby analyzing their load patterns and behavioral characteristics. The historical time period can be a pre-set time cycle for analyzing the load patterns and behavioral characteristics of electricity users, such as a month or a quarter. The historical load curves can be the electricity load curves corresponding to each user in the historical time period, or they can be obtained by processing the electricity load curves corresponding to each user in the historical time period. The historical load curves can correspond to one day or one time period. The initial Gaussian mixture model can be pre-set, containing a preset number of probability distribution functions. The target Gaussian mixture model can be a Gaussian mixture model in which all the probability distribution functions are determined.
[0070] It should be noted that the number and type of probability distribution functions included in the initial Gaussian mixture model are uncertain; it can be understood that the initial Gaussian mixture model contains all types of probability distribution functions. This step is to determine the specific probability distribution functions included in the initial Gaussian mixture model, and once determined, this model will be used as the target Gaussian mixture model.
[0071] Optionally, in this embodiment, the method for determining the target Gaussian mixture model can be as follows: Obtain the historical load curves corresponding to the users to be clustered within a historical time period. For each historical load curve, compare its curve distribution with various probability distribution functions to determine the similarity between the historical load curve and each probability distribution function. Use this similarity as the sub-probability value of the historical load curve belonging to that type of probability distribution function. For example, taking 100 types of probability distribution functions as an example, for each historical load curve, the above steps can determine the sub-probability values of the historical load curve belonging to each of the 100 types of probability distribution functions, i.e., determine 100 probability distribution values. Further, for each type of probability distribution function, calculate the sum of the sub-probability values of all historical load curves belonging to that type of probability distribution function. Retain probability distribution functions whose sum of sub-probability values is greater than a preset value, and discard the remaining probability distribution functions, thereby completing the adjustment of the probability distribution functions included in the initial Gaussian mixture model. The initial Gaussian mixture model containing the retained probability distribution functions is used as the target Gaussian mixture model.
[0072] In addition, this embodiment also provides another method for determining the target Gaussian mixture model. The method for determining the sub-probability values of the historical load curve belonging to each probability distribution function in the initial Gaussian mixture model can be as follows: based on the historical load curves of the users to be clustered within the historical time period, determine the probability density value of the historical load curve belonging to each probability distribution function in the initial Gaussian mixture model; based on the probability density value corresponding to the historical load curves of the users to be clustered and the weights of each probability distribution function included in the initial Gaussian mixture model, determine the sub-probability values of the historical load curve belonging to each probability distribution function in the initial Gaussian mixture model.
[0073] The probability density value represents the probability that a historical load curve belongs to the initial Gaussian mixture model. The weights of each probability distribution function represent the proportion of each probability distribution function in the initial Gaussian mixture model. It can be understood that in this embodiment, each historical load curve corresponds to a probability density value indicating that it belongs to the initial Gaussian mixture model.
[0074] In this embodiment, for each historical load curve, the probability density value belonging to the initial Gaussian mixture model can be predetermined. Then, based on the probability density value corresponding to the historical load curve and the weights of the probability distribution functions included in the initial Gaussian mixture model, the sub-probability values of the historical load curve belonging to each probability distribution function in the initial Gaussian mixture model are determined. For example, the probability density value corresponding to the historical load curve and the weights of the probability distribution functions included in the initial Gaussian mixture model can be substituted into the predetermined sub-probability value determination formula for determination. The sub-probability value determination formula can be as shown in formula (1) below:
[0075]
[0076] In the formula, P' n Let λ be the nth historical load curve; k is the number of probability distribution functions included in the initial Gaussian mixture model; λ is the number of historical load curves. k It is the probability distribution function f k The weights of (·); ψ represents all parameters included in the initial Gaussian mixture model; θ k f is the k-th probability distribution function k The parameter of (·); λ k f k (P' n ;θ k ) represents the historical load curve P′ n Belongs to the k-th probability distribution function f k The weighted probability (prior probability) of (·); f(P') n ;ψ) represents the historical load curve P' n These are the probability density values belonging to the initial Gaussian mixture model. Since the sum of the sub-probabilities of the probability distribution function included in the initial Gaussian mixture model is 1, therefore...
[0077] It should be noted that in the process of determining the sub-probability values of the probability distribution functions in the initial Gaussian mixture model, the weights of the probability distribution functions in the initial Gaussian mixture model can be predetermined or determined by the expectation-maximization algorithm.
[0078] The prior probability of the historical load curve belonging to each probability distribution function in the initial Gaussian mixture model can be determined using the above formula (1). Furthermore, by processing the prior probability using Bayes' theorem, the sub-probability values of the historical load curve belonging to each probability distribution function in the initial Gaussian mixture model can be determined. For example, this can be determined using the following formula (2):
[0079]
[0080] In the formula, f(P′)n ;k) represents the historical load curve P' n The k-th probability distribution function f in the initial Gaussian mixture model k The subprobability value of (·); λ k ,f(P' n ;θ k ) represents the historical load curve P' n Belongs to the k-th probability distribution function f k The weighted probability (prior probability) of (·); f(P′) n ;ψ) represents the historical load curve P' n The probability density value belongs to the initial Gaussian mixture model; ω is the evidence factor.
[0081] It should be noted that prior probability and posterior probability (i.e., sub-probability values in this embodiment) are two important concepts in Bayes' theorem. Prior probability is an initial estimate of the probability of an event occurring before considering any new evidence. Posterior probability, on the other hand, is a revised estimate of the probability of an event occurring after considering new evidence; in other words, posterior probability is more accurate.
[0082] Using the above formulas (1) and (2), the sub-probability values of the historical load curve belonging to each probability distribution function in the initial Gaussian mixture model can be determined. Furthermore, the sub-probability values determined in the above manner are more accurate, providing a basis for the subsequent process of determining the target Gaussian mixture model.
[0083] Furthermore, the target Gaussian mixture model is determined based on the sub-probability values of the historical load curves belonging to each probability distribution function in the initial Gaussian mixture model. For example, the total risk value corresponding to the historical load curves of the users to be clustered under each probability distribution function can be determined based on the sub-probability values of the historical load curves corresponding to each probability distribution function. It is then determined whether the total risk value is less than a risk threshold. If so, the initial Gaussian mixture model corresponding to the total risk value is used as the target Gaussian mixture model. If not, the probability distribution functions included in the initial Gaussian mixture model are adjusted, and the operation of determining the sub-probability values of the historical load curves belonging to each probability distribution function in the initial Gaussian mixture model based on the adjusted initial Gaussian mixture model is returned.
[0084] The total risk value can be the sum of risk values used to determine whether the historical load curve belongs to various probability distribution functions. The risk threshold can be a pre-determined value used to judge whether the total risk value meets the requirements.
[0085] Specifically, in this embodiment, the risk value of the historical load curve belonging to various probability distribution functions can be the value obtained by subtracting a sub-probability value. For a given historical load curve, given the known risk values of each probability distribution function corresponding to the historical load curve, the sum of all risk values is taken as the total risk value of the historical load curve under each probability distribution function. Further, the sum of the total risk values of all historical load curves under each probability distribution function is taken as the total risk value of the historical load curve of the user to be clustered under each probability distribution function. It is determined whether the total risk value is less than a risk threshold. If so, the initial Gaussian mixture model corresponding to the total risk value is taken as the target Gaussian mixture model; that is, the current initial Gaussian mixture model is taken as the target Gaussian mixture model. If not, the probability distribution functions contained in the initial Gaussian mixture model are adjusted. This adjustment includes adjusting the number and type of probability distribution functions contained in the initial Gaussian mixture model. Based on the adjusted initial Gaussian mixture model, the process returns to perform the operation of determining the sub-probability values of the historical load curves corresponding to the historical load curves of the users to be clustered within the historical time period, until the target Gaussian mixture model is determined. The specific operation process of determining the sub-probability values of the historical load curves corresponding to the probability distribution functions in the initial Gaussian mixture model has been explained in detail in the above embodiments and will not be repeated here.
[0086] In the above embodiments, by comparing the total risk value and risk threshold corresponding to the historical load curves of users to be clustered under each probability distribution function, it can be determined whether to adjust the initial Gaussian mixture model to the target Gaussian mixture model, which makes the process of determining the target Gaussian mixture model simpler and more convenient.
[0087] S202, determine the target load curve corresponding to the user to be clustered based on the sub-probability values of the probability distribution functions in the target Gaussian mixture model corresponding to the historical load curve of the user to be clustered.
[0088] The target load curve can be a historical load curve that represents the electricity load of the users to be clustered within a historical time period.
[0089] Optionally, in this embodiment, for a historical load curve of a user to be clustered, the sub-probability values of the historical load curve belonging to each probability distribution function in the target Gaussian mixture model can be statistically analyzed, and the probability distribution function with the largest sub-probability value can be used as the target probability distribution function corresponding to the historical load curve. Further, the target probability distribution functions corresponding to all historical load curves of a user to be clustered are determined, and the probability distribution function with the highest frequency of occurrence is used as the target load curve of the user to be clustered. Additionally, to make the determined target load curve more realistic, one of the historical load curves corresponding to the probability distribution function with the highest frequency of occurrence can also be used as the target load curve corresponding to the user to be clustered.
[0090] S203. Based on the target load curves corresponding to the users to be clustered, cluster the users to be clustered, and based on the Bayesian information criterion values corresponding to each type of user after clustering, determine whether each type of user after clustering has met the clustering termination condition.
[0091] The Bayesian Information Criterion (BIC) is a model selection metric that evaluates and compares different models given a dataset. In this embodiment, the clustering termination condition can be determined based on the magnitude of each BIC value.
[0092] It should be noted that the clustering of users can be performed using X-Means clustering algorithm, K-Means clustering algorithm, hierarchical clustering algorithm, density clustering algorithm, or mean shift clustering algorithm.
[0093] Taking the clustering method of using the X-Means clustering algorithm as an example, in this embodiment, users to be clustered can be clustered, and the BIC corresponding to the user class obtained after each clustering can be determined. If the BIC corresponding to the user class obtained after this clustering is less than the preset BIC threshold, then it is determined that the clustering of all user classes has reached the clustering termination condition.
[0094] Alternatively, another implementation method involves clustering the users to be clustered and determining whether any of the Bayesian information criterion values for each user group after this round of clustering contain a target Bayesian information criterion value less than the historical minimum Bayesian information criterion value. If so, it is determined that all user groups after clustering have met the clustering termination condition; otherwise, it is determined that users in the group corresponding to a target Bayesian information criterion value less than the historical minimum Bayesian information criterion value have not met the clustering termination condition. Here, the historical minimum Bayesian information criterion value is the smallest Bayesian information criterion value among the Bayesian information criterion values for each user group obtained from previous rounds of clustering.
[0095] For example, an initial number of clusters can be predetermined. Users to be clustered are initially clustered according to this initial number, resulting in the initial number of user classes. The BIC (Balanced Indicator) for each user class is then calculated, and the minimum BIC is determined. Further, the users to be clustered within each user class are clustered again, and the BIC for each user class is calculated separately. The minimum BIC for each user class obtained after this clustering is compared with the minimum BIC obtained after the previous clustering. It is determined whether the BIC for each user class after this clustering is less than the minimum BIC obtained after the previous clustering. If so, it is determined that all users in each cluster have met the clustering termination condition, and operation S204 is executed. If not, users in the class corresponding to the target Bayesian information criterion value that is less than the historical minimum Bayesian information criterion value have not met the clustering termination condition. These users who have not met the clustering termination condition are then treated as users to be clustered, and the operation of clustering these users is returned.
[0096] Specifically, in this embodiment, if any of the BIC values for each user category obtained in this clustering is less than the minimum BIC value obtained from previous clustering, then the user category corresponding to the BIC value less than the minimum BIC value obtained from previous clustering is designated as the user to be clustered, and the clustering operation for the user to be clustered is returned. This process continues until the BIC value for each user category after this round of clustering is no longer less than the minimum BIC value obtained from previous clustering, at which point operation S204 is executed. The clustering operation for the user to be clustered has been described in detail in the above embodiments and will not be repeated here.
[0097] In the above embodiments, the clustering termination condition is determined based on the relationship between the BIC of each type of user after clustering and the historical minimum BIC, making the process of determining the termination of clustering more rigorous.
[0098] S204, obtain the clustering results for the users to be clustered.
[0099] Specifically, in this embodiment, if the BIC of each type of user after this clustering is not less than the minimum BIC obtained after the previous clustering, then the previous clustering result is used as the clustering result of the user to be clustered.
[0100] In the above embodiments, based on the historical load curves corresponding to the users to be clustered within a historical time period, the sub-probability values of the historical load curves belonging to each probability distribution function in the initial Gaussian mixture model are determined, thereby determining the target Gaussian mixture model. Since the probability distribution functions included in the target Gaussian mixture model are all determined based on the sub-probability values of the historical load curves belonging to each probability distribution function in the initial Gaussian mixture model, adding the step of determining the target Gaussian mixture model makes the target load curves corresponding to the users to be clustered more representative. Furthermore, since each user to be clustered corresponds to only one target load curve, rather than all historical load curves, the efficiency of clustering the users to be clustered can be improved. In addition, during the clustering process, the number of clusters is not preset in advance; instead, the clustering termination condition is determined by calculating the BIC corresponding to each type of user. When the BIC corresponding to each type of user reaches the termination condition, the clustering result (including the number of clusters and each type of user) can be determined, resulting in better and more accurate clustering of the users to be clustered.
[0101] Furthermore, to improve the clustering effect on users to be clustered, one embodiment provides a method for determining the historical load curves of users to be clustered within a historical time period. Specifically, such as... Figure 3 As shown, it includes the following steps:
[0102] S301, Obtain the unit load curves for each unit time period of the user to be clustered within the historical time period.
[0103] The unit time period can be a month, a week, a day, or an hour. The unit load curve can be the load curve corresponding to the users to be clustered within the unit time period.
[0104] For example, taking a historical period of 30 days and each unit period as 1 day. Specifically, in this embodiment, it may be to obtain the unit load curve of the user to be clustered for each day within the historical 30 days.
[0105] S302, perform load reduction processing on the load curves of each unit corresponding to the user to be clustered, and obtain the historical load curves of the user to be clustered within the historical time period.
[0106] Among them, load reduction can be carried out by segmenting the load curve of each unit, with the aim of reducing the volatility of each unit load curve and making the determined historical load curve easier to analyze and process.
[0107] Specifically, in this embodiment, the unit load curves can be segmented first, and then the segmented unit load curves can be averaged to obtain the historical load curves corresponding to the historical time period.
[0108] For example, taking a unit time period of 1 day, a total length of the unit load curve of T, and a window width of w (i.e., the number of sub-segments contained in each segment) to divide the unit load curve into T / w segments, the segmented unit load curve can be represented by the following formula (3):
[0109] P n =[P n1 P n2 P n3 P n4 , ..., P nT / w (3)
[0110] In the formula, n is the identifier of the unit load curve, and P n This is the nth unit load curve; P n1 This is the first segment of the nth unit load curve; P nT / w Let w be the T / w segment of the nth unit load curve; where w is an integer and divisible by T.
[0111] Furthermore, the average approximation is performed on each segment of the segmented unit load curve using the following formula (4):
[0112]
[0113]
[0114] ...
[0115]
[0116] In the formula, This is the average approximate result of the first segment of the nth unit load curve; P ni This refers to the i-th sub-segment within the first segment of the n-th unit load curve; This is the average approximate result of the T / w segment of the nth unit load curve.
[0117] The average approximate processing results of each segment in the unit load curve after segment processing are integrated to obtain the historical load curve, as shown in the following formula (5):
[0118] P n =[P′ n1 , P′ n2 , P′ n3 , P′ n4 , ..., P′ n,Tw (5)
[0119] In the formula, P′ n1 This is the average approximate processing result of the first segment of the nth unit load curve; This is the average approximate result of the T / w segment of the nth unit load curve.
[0120] In the above embodiments, by performing load reduction processing on the unit load curves corresponding to each unit time period of the users to be clustered in the historical time period, the historical load curves corresponding to the users to be clustered in the historical time period are determined, reducing the volatility of the historical load curves and facilitating subsequent processing.
[0121] Furthermore, in one embodiment, the historical load curve includes at least two sub-curve segments. To make the process of determining the target load curve for users to be clustered more rigorous, and thus make the determined target load curve more representative, in one embodiment, such as... Figure 4 As shown, the process of determining the target load curve corresponding to the users to be clustered is described in detail, including:
[0122] S401, based on the sub-probability values of each sub-curve segment of the user to be clustered belonging to each probability distribution function in the target Gaussian mixture model, determine the typical load curve corresponding to each probability distribution function from each sub-curve segment of the user to be clustered.
[0123] The sub-curve segment can be a historical load curve consisting of a unit time period or a historical load curve composed of multiple unit time periods. For example, taking a historical time period of 30 days and a unit time period of 1, the sub-curve segment can be the historical load curve corresponding to 1 day within the historical time period or the historical load curve corresponding to multiple days.
[0124] In this embodiment, for each sub-curve segment of the user to be clustered, the sub-probability value belonging to each probability distribution function in the target Gaussian mixture model is determined. The specific determination method is the same as that used to determine the sub-probability values of the historical load curve belonging to each probability distribution function in the target Gaussian mixture model, and will not be described in detail here.
[0125] S402, based on the sub-probability values of each sub-curve segment of the user to be clustered belonging to each probability distribution function in the target Gaussian mixture model, determine the typical load curve corresponding to each sub-curve segment of the user to be clustered.
[0126] Among them, a typical load curve can be a sub-curve segment that can be used to characterize a class of probability distribution functions.
[0127] Specifically, in this embodiment, for a probability distribution function, each sub-curve segment belonging to the probability distribution function can be determined, and then the sub-probability value of each sub-curve segment belonging to the probability distribution function can be determined. The sub-curve segment with the largest probability value is taken as the typical load curve of the probability distribution function.
[0128] S403. Based on the typical load curves corresponding to each sub-curve segment of the user to be clustered, determine the target load curve of the user to be clustered from the typical load curves corresponding to each probability distribution function.
[0129] The target load curve can be a typical load curve that can be used to characterize the load characteristics of a user to be clustered.
[0130] Specifically, in this embodiment, for a user to be clustered, after determining the typical load curve corresponding to the sub-curve segment of all its historical load curves, the typical load curve that appears most frequently among the typical load curves corresponding to the user to be clustered is determined as the target load curve for the user to be clustered.
[0131] The above embodiments specifically illustrate the method for determining the target load curve of users to be clustered. First, the typical load curves corresponding to each sub-curve segment of the user to be clustered are determined. Then, based on the frequency of occurrence of each typical load curve corresponding to the user to be clustered, the target load curve of the user to be clustered is determined, making the process of determining the target load curve of the user to be clustered more rigorous and complete.
[0132] To facilitate understanding by those skilled in the art, the above user clustering method will be described in detail, such as... Figure 5 As shown, the method may include:
[0133] S501, obtain the unit load curves for each unit time period of the user to be clustered within the historical time period.
[0134] S502, perform load reduction processing on the load curves of each unit corresponding to the user to be clustered, and obtain the historical load curves of the user to be clustered within the historical time period.
[0135] S503, based on the historical load curves of the users to be clustered within the historical time period, determine the probability density value of the historical load curve belonging to the initial Gaussian mixture model.
[0136] S504. Based on the probability density values corresponding to the historical load curves of the users to be clustered, and the weights of each probability distribution function included in the initial Gaussian mixture model, determine the sub-probability values of the historical load curves belonging to each probability distribution function in the initial Gaussian mixture model.
[0137] S505: Based on the sub-probability values of each probability distribution function corresponding to the historical load curves of the users to be clustered, determine the total risk value corresponding to each probability distribution function for the historical load curves of the users to be clustered, and determine whether the total risk value is less than the risk threshold. If yes, proceed to S506; otherwise, proceed to S507.
[0138] S506, the initial Gaussian mixture model corresponding to the total risk value is used as the target Gaussian mixture model.
[0139] S507 adjusts the probability distribution function contained in the initial Gaussian mixture model and returns to execute the operation of S504 based on the adjusted initial Gaussian mixture model.
[0140] S508. Based on the sub-probability values of each sub-curve segment of the user to be clustered belonging to each probability distribution function in the target Gaussian mixture model, determine the typical load curve corresponding to each probability distribution function from each sub-curve segment of the user to be clustered.
[0141] S509. Based on the sub-probability values of each sub-curve segment of the user to be clustered belonging to each probability distribution function in the target Gaussian mixture model, determine the typical load curve corresponding to each sub-curve segment of the user to be clustered.
[0142] S510, based on the typical load curves corresponding to each sub-curve segment of the user to be clustered, determine the target load curve of the user to be clustered from the typical load curves corresponding to each probability distribution function.
[0143] S511: Based on the target load curves corresponding to the users to be clustered, cluster the users to be clustered, and determine whether there is no target Bayesian information criterion value less than the historical minimum Bayesian information criterion value among the Bayesian information criterion values of each type of user after this round of clustering. If yes, execute operation S512; otherwise, execute operation S513.
[0144] Among them, the historical minimum Bayesian information criterion value is the minimum Bayesian information criterion value among the Bayesian information criterion values of various types of users obtained from historical rounds of clustering.
[0145] S512, determine that all users in each cluster have met the clustering termination condition, and obtain the clustering results of the users to be clustered.
[0146] S513, determine that there are users in the class corresponding to the target Bayesian information criterion value that is less than the historical minimum Bayesian information criterion value who have not met the clustering termination condition, and treat each class of users who have not met the clustering termination condition as users to be clustered, and return to execute the operation of S511.
[0147] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0148] Based on the same inventive concept, this application also provides a user clustering apparatus for implementing the user clustering method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more user clustering apparatus embodiments provided below can be found in the limitations of the user clustering method described above, and will not be repeated here.
[0149] In one embodiment, such as Figure 6 As shown, a user clustering device 1 is provided, including: a target model determination module 10, a target curve determination module 11, a clustering end judgment module 12, and a clustering result acquisition module 13, wherein:
[0150] The target model determination module 10 is used to determine the sub-probability values of the historical load curves corresponding to the historical load curves of the users to be clustered within the historical time period, and to adjust the probability distribution functions contained in the initial Gaussian mixture model according to the sub-probability values of the probability distribution functions corresponding to the historical load curves, so as to obtain the target Gaussian mixture model.
[0151] The target curve determination module 11 is used to determine the target load curve corresponding to the user to be clustered based on the sub-probability values of the probability distribution functions in the target Gaussian mixture model corresponding to the historical load curve of the user to be clustered.
[0152] The clustering termination judgment module 12 is used to cluster users according to the target load curves corresponding to the users to be clustered, and to determine whether all types of users after clustering have met the clustering termination condition based on the Bayesian information criterion values corresponding to each type of user after clustering.
[0153] Clustering result acquisition module 13 is used to acquire the clustering results of the user to be clustered if the clustering result is obtained.
[0154] In one embodiment, such as Figure 7As shown, the target model determination module 10 includes a density value determination unit 100 and a sub-probability value determination unit 101. Wherein:
[0155] The density value determination unit 100 is used to determine the probability density value of the historical load curve belonging to the initial Gaussian mixture model based on the historical load curve of the user to be clustered within the historical time period.
[0156] The sub-probability value determination unit 101 is used to determine the sub-probability value of the historical load curve belonging to each probability distribution function in the initial Gaussian mixture model based on the probability density value corresponding to the historical load curve of the user to be clustered and the weight of each probability distribution function contained in the initial Gaussian mixture model.
[0157] In one embodiment, such as Figure 8 As shown, the target model determination module 10 also includes a risk value determination unit 102, a risk judgment unit 103, a target model determination unit 104, and an initial model adjustment unit 105.
[0158] in:
[0159] The risk value determination unit 102 is used to determine the total risk value corresponding to the historical load curve of the user to be clustered under each probability distribution function based on the sub-probability values of each probability distribution function corresponding to the historical load curve of the user to be clustered.
[0160] Risk assessment unit 103 is used to determine whether the total risk value is less than the risk threshold.
[0161] The target model determination unit 104 is used to take the initial Gaussian mixture model corresponding to the total risk value as the target Gaussian mixture model if the condition is met.
[0162] The initial model adjustment unit 105 is used to adjust the probability distribution function contained in the initial Gaussian mixture model if no, and return to perform the operation of determining the sub-probability value of each probability distribution function in the initial Gaussian mixture model based on the historical load curve corresponding to the user to be clustered in the historical time period.
[0163] In one embodiment, such as Figure 9 As shown, the target curve determination module 11 includes a first curve determination unit 110, a second curve determination unit 111, and a target curve determination unit 112. Wherein:
[0164] The first curve determination unit 110 is used to determine the typical load curve corresponding to each probability distribution function from each sub-curve segment of the user to be clustered, based on the sub-probability values of each probability distribution function in the target Gaussian mixture model.
[0165] The second curve determination unit 111 is used to determine the typical load curve corresponding to each sub-curve segment of the user to be clustered based on the sub-probability values of each probability distribution function in the target Gaussian mixture model.
[0166] The target curve determination unit 112 is used to determine the target load curve of the user to be clustered from the typical load curves corresponding to each probability distribution function, based on the typical load curves corresponding to each sub-curve segment of the user to be clustered.
[0167] In one embodiment, such as Figure 10 As shown, the user clustering device 1 also includes a user clustering module 14, which is used to treat each type of user that has not met the clustering termination condition as a user to be clustered, and return to perform the clustering operation on the user to be clustered.
[0168] In one embodiment, such as Figure 11 As shown, the clustering end determination module 12 includes a clustering end determination unit 120, a clustering end determination unit 121, and a clustering not-end determination unit 122. Wherein:
[0169] The clustering termination judgment unit 120 is used to determine whether there is no target Bayesian information criterion value less than the historical minimum Bayesian information criterion value among the Bayesian information criterion values of all types of users after this round of clustering; wherein, the historical minimum Bayesian information criterion value is the minimum Bayesian information criterion value among the Bayesian information criterion values of all types of users obtained from previous rounds of clustering.
[0170] Clustering termination determination unit 121 is used to determine that if the clustering conditions are met, then all types of users after clustering have met the clustering termination conditions.
[0171] Clustering not terminated determination unit 122 is used to determine, if not, that there are users in the class corresponding to the target Bayesian information criterion value that is less than the historical minimum Bayesian information criterion value, and that the clustering termination condition has not been met.
[0172] In one embodiment, such as Figure 12 As shown, the user clustering device 1 further includes a load curve determination module 15, comprising a load curve acquisition unit 150 and a load curve determination unit 151. Wherein:
[0173] The load curve acquisition unit 150 is used to acquire the unit load curves corresponding to each unit time period of the users to be clustered within the historical time period.
[0174] The load curve determination unit 151 is used to perform load reduction processing on the load curves of each unit corresponding to the users to be clustered, so as to obtain the historical load curves of the users to be clustered within the historical time period.
[0175] Each module in the aforementioned user clustering device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0176] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 13 As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a user clustering method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.
[0177] Those skilled in the art will understand that Figure 13 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0178] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0179] Based on the historical load curves of the users to be clustered within a historical time period, determine the sub-probability values of each probability distribution function in the initial Gaussian mixture model corresponding to the historical load curves. Then, adjust the probability distribution functions contained in the initial Gaussian mixture model according to the sub-probability values of each probability distribution function corresponding to the historical load curves to obtain the target Gaussian mixture model.
[0180] Based on the sub-probability values of the probability distribution functions in the target Gaussian mixture model corresponding to the historical load curves of the users to be clustered, the target load curves corresponding to the users to be clustered are determined.
[0181] Based on the target load curves corresponding to the users to be clustered, the users to be clustered are clustered, and based on the Bayesian information criterion values corresponding to each type of user after clustering, it is determined whether each type of user after clustering has met the clustering termination condition.
[0182] If so, then obtain the clustering results for the users to be clustered.
[0183] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0184] Based on the historical load curves of the users to be clustered within a historical time period, determine the sub-probability values of each probability distribution function in the initial Gaussian mixture model corresponding to the historical load curves. Then, adjust the probability distribution functions contained in the initial Gaussian mixture model according to the sub-probability values of each probability distribution function corresponding to the historical load curves to obtain the target Gaussian mixture model.
[0185] Based on the sub-probability values of the probability distribution functions in the target Gaussian mixture model corresponding to the historical load curves of the users to be clustered, the target load curves corresponding to the users to be clustered are determined.
[0186] Based on the target load curves corresponding to the users to be clustered, the users to be clustered are clustered, and based on the Bayesian information criterion values corresponding to each type of user after clustering, it is determined whether each type of user after clustering has met the clustering termination condition.
[0187] If so, then obtain the clustering results for the users to be clustered.
[0188] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:
[0189] Based on the historical load curves of the users to be clustered within a historical time period, determine the sub-probability values of each probability distribution function in the initial Gaussian mixture model corresponding to the historical load curves. Then, adjust the probability distribution functions contained in the initial Gaussian mixture model according to the sub-probability values of each probability distribution function corresponding to the historical load curves to obtain the target Gaussian mixture model.
[0190] Based on the sub-probability values of the probability distribution functions in the target Gaussian mixture model corresponding to the historical load curves of the users to be clustered, the target load curves corresponding to the users to be clustered are determined.
[0191] Based on the target load curves corresponding to the users to be clustered, the users to be clustered are clustered, and based on the Bayesian information criterion values corresponding to each type of user after clustering, it is determined whether each type of user after clustering has met the clustering termination condition.
[0192] If so, then obtain the clustering results for the users to be clustered.
[0193] It should be noted that the user information (including but not limited to user historical load curve information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0194] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0195] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0196] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A user clustering method, characterized in that, The method includes: Based on the historical load curves of the users to be clustered within a historical time period, the sub-probability values of the probability distribution functions in the initial Gaussian mixture model are determined. Based on the sub-probability values of the probability distribution functions corresponding to the historical load curves, the probability distribution functions contained in the initial Gaussian mixture model are adjusted to obtain the target Gaussian mixture model. The target load curve corresponding to the user to be clustered is determined based on the sub-probability values of each probability distribution function in the target Gaussian mixture model corresponding to the historical load curve of the user to be clustered. Based on the target load curves corresponding to the users to be clustered, the users to be clustered are clustered, and based on the Bayesian information criterion values corresponding to each type of user after clustering, it is determined whether each type of user after clustering has met the clustering termination condition. If so, obtain the clustering results for the users to be clustered; The step of adjusting the probability distribution functions contained in the initial Gaussian mixture model according to the sub-probability values of the probability distribution functions corresponding to the historical load curves to obtain the target Gaussian mixture model includes: determining the total risk value corresponding to the historical load curves of the users to be clustered under each probability distribution function according to the sub-probability values of the probability distribution functions corresponding to the historical load curves of the users to be clustered; determining whether the total risk value is less than a risk threshold; if so, then using the initial Gaussian mixture model corresponding to the total risk value as the target Gaussian mixture model; if not, then adjusting the probability distribution functions contained in the initial Gaussian mixture model, and returning to the operation of determining the historical load curves of the users to be clustered within the historical time period as belonging to the sub-probability values of the probability distribution functions in the initial Gaussian mixture model based on the adjusted initial Gaussian mixture model; Among them, the risk value of the historical load curve belonging to each probability distribution function is the value obtained by subtracting the sub-probability value. The sum of all risk values is taken as the total risk value corresponding to the historical load curve under each probability distribution function. The sum of the total risk values corresponding to all historical load curves under each probability distribution function is taken as the total risk value corresponding to the historical load curve of the user to be clustered under each probability distribution function.
2. The method according to claim 1, characterized in that, The step of determining the sub-probability values of the historical load curves belonging to various probability distribution functions in the initial Gaussian mixture model based on the historical load curves of the users to be clustered within a historical time period includes: Based on the historical load curves of the users to be clustered within a historical time period, determine the probability density value of the historical load curves belonging to the initial Gaussian mixture model. Based on the probability density value corresponding to the historical load curve of the user to be clustered, and the weights of each probability distribution function included in the initial Gaussian mixture model, the sub-probability value of the historical load curve belonging to each probability distribution function in the initial Gaussian mixture model is determined.
3. The method according to claim 1, characterized in that, The historical load curve contains at least two sub-curve segments; determining the target load curve corresponding to the user to be clustered based on the sub-probability values of the historical load curve corresponding to the user to be clustered belonging to each probability distribution function in the target Gaussian mixture model includes: Based on the sub-probability values of each sub-curve segment of the user to be clustered belonging to each probability distribution function in the target Gaussian mixture model, determine the typical load curve corresponding to each probability distribution function from each sub-curve segment of the user to be clustered; Based on the sub-probability values of each sub-curve segment of the user to be clustered belonging to each probability distribution function in the target Gaussian mixture model, determine the typical load curve corresponding to each sub-curve segment of the user to be clustered; Based on the typical load curves corresponding to each sub-curve segment of the user to be clustered, the target load curve of the user to be clustered is determined from the typical load curves corresponding to each probability distribution function.
4. The method according to claim 1, characterized in that, After determining whether all clustered users meet the clustering termination condition, the process further includes: Users who have not met the clustering termination criteria are designated as users to be clustered, and the clustering operation is performed on these users.
5. The method according to claim 4, characterized in that, The step of determining whether all users in each cluster have met the clustering termination condition based on the Bayesian information criterion values corresponding to each cluster includes: Determine whether there is no target Bayesian information criterion value less than the historical minimum Bayesian information criterion value among the Bayesian information criterion values of each type of user after this round of clustering; wherein, the historical minimum Bayesian information criterion value is the minimum Bayesian information criterion value among the Bayesian information criterion values of each type of user obtained from previous rounds of clustering; If so, then it is determined that all users in each cluster have met the clustering termination condition; If not, then it is determined that there are users in the class corresponding to the target Bayesian information criterion value that is less than the historical minimum Bayesian information criterion value, and the clustering termination condition has not been met.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: Obtain the unit load curves for each time period of the users to be clustered within the historical time period; The load curves of each unit corresponding to the users to be clustered are subjected to load reduction processing to obtain the historical load curves of the users to be clustered within the historical time period.
7. A user clustering device, characterized in that, The device includes: The target model determination module is used to determine the sub-probability values of the historical load curves corresponding to the users to be clustered within a historical time period, and to adjust the probability distribution functions contained in the initial Gaussian mixture model according to the sub-probability values of the probability distribution functions corresponding to the historical load curves, so as to obtain the target Gaussian mixture model. The target curve determination module is used to determine the target load curve corresponding to the user to be clustered based on the sub-probability values of the probability distribution functions in the target Gaussian mixture model corresponding to the historical load curve of the user to be clustered. The clustering termination judgment module is used to cluster the users to be clustered according to the target load curve corresponding to the users to be clustered, and to determine whether all types of users after clustering have met the clustering termination condition according to the Bayesian information criterion value corresponding to each type of user after clustering. The clustering result acquisition module is used to acquire the clustering result for the user to be clustered if the clustering result is positive. Specifically, the target model determination module is used to: determine the total risk value corresponding to the historical load curve of the user to be clustered under each probability distribution function based on the sub-probability values of each probability distribution function corresponding to the historical load curve of the user to be clustered; determine whether the total risk value is less than a risk threshold; if so, use the initial Gaussian mixture model corresponding to the total risk value as the target Gaussian mixture model; if not, adjust the probability distribution functions contained in the initial Gaussian mixture model, and return to execute the operation of determining the sub-probability values of each probability distribution function in the initial Gaussian mixture model based on the historical load curve of the user to be clustered within the historical time period; wherein, the risk value of the historical load curve belonging to each probability distribution function is the value obtained by 1 minus the sub-probability value, the sum of all risk values is used as the total risk value corresponding to the historical load curve under each probability distribution function, and the sum of the total risk values corresponding to all historical load curves under each probability distribution function is used as the total risk value corresponding to the historical load curve of the user to be clustered under each probability distribution function.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Intelligent equipment control method and device, storage medium and electronic equipment
CN114694650A
GMM clustering-based comprehensive energy system typical day generation method
CN115221728A