Clustering method of low voltage area user electricity consumption data based on mean shift clustering
Through the method based on mean drift clustering, the power consumption data of users in low-voltage station areas is clustered, which solves the problem that the value of power big data information is not utilized, and improves the data analysis quality and classification robustness.
Patent Information
- Application Number
- CN202210336294.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-31
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-03-31
AI Technical Summary
During the operation of the power system, the real information value of power big data is not fully utilized, resulting in the waste of power big data resources.
The low-voltage drift clustering method is used to calculate the density estimation function, gradient function and Gaussian kernel function of the user's electricity data, calculate the mean drift vector, update the clustering center point, and complete the clustering of electricity data.
It effectively improves the quality of electricity consumption data analysis, is suitable for scenarios with uncertain user types, does not require training of prior knowledge, and improves the robustness of classification.
Smart Images

Figure CN114662608B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of unsupervised data clustering of electricity consumption data, and in particular relates to a method for clustering electricity consumption data of users in a low-voltage area based on mean shift clustering. Background Art
[0002] With the popularization and application of smart meters, the interaction between power supply companies and users is becoming more and more frequent, and the power consumption data on the user side is growing rapidly, which has promoted the development of power big data on the user side. Classification and analysis of power user power consumption data can more accurately understand the power consumption behavior of different types of power users, and provide a basic basis for power companies to provide high-quality and precise services. However, in the operation of the power system, the real information value of power big data has not been fully utilized, resulting in a waste of power big data resources. Summary of the invention
[0003] The purpose of the present invention is to provide a method for clustering electricity consumption data of low-voltage substation users based on mean shift clustering, aiming to solve the problem that the real information value of power big data is not fully utilized in the operation process of the power system, thus causing a waste of power big data resources.
[0004] To achieve the above object, the technical solution adopted by the present invention is: to provide a method for clustering power consumption data of low-voltage area users based on mean shift clustering, comprising the following steps:
[0005] S1: Given an m-dimensional electricity consumption dataset X∈R m×n , calculate the density estimation function of user electricity consumption data xi (i = 1, 2, ..., n);
[0006] S2: According to the extreme value theorem of the function, the local density maximum point is located at the gradient zero point of the density function. The gradient function is derived from the above formula to find the maximum point of the electricity consumption data density.
[0007] S3: Considering the influence of noise data in electricity consumption data, the Gaussian kernel function is introduced to achieve high-dimensional separability of user electricity consumption data and improve the robustness of classification;
[0008] S4: Calculate the mean shift vector M h , the size and direction of the center point offset are calculated according to the drift vector, so as to determine whether the center point iteration is completed and calculate the new center point.
[0009] S5: Update the candidate point of the cluster center point to the mean value of all electricity consumption data points within the search radius, find the dense area in the electricity consumption data X, locate the center points of different types of electricity consumption data, and complete the clustering of electricity consumption data.
[0010] As another embodiment of the present application, the step S1 includes the following contents:
[0011] The density estimation function of user electricity consumption data is:
[0012]
[0013] In the formula, C k is a constant, k(x) is the kernel function, h is the kernel width, and x is the center point of the initial data set.
[0014] As another embodiment of the present application, step S2 includes the following contents:
[0015] The gradient function of the density estimation function of user electricity consumption data is:
[0016]
[0017] As another embodiment of the present application, step S3 includes the following contents:
[0018] Define the Gaussian kernel function:
[0019]
[0020] Where x is the user's electricity consumption data vector, and h is the search radial kernel width.
[0021] As another embodiment of the present application, step S4 includes the following contents:
[0022] According to the gradient function (2) of the power consumption data density estimation function, substitute it into the Gaussian kernel function (3) to obtain the drift vector Mh(x):
[0023]
[0024]
[0025] As another embodiment of the present application, in step S1, obtaining a set of power consumption data points of users in a low-voltage area includes:
[0026] Obtaining original electricity consumption data points of the user to be detected;
[0027] Normalizing the values of the original power consumption data points to obtain normalized power consumption data points;
[0028] Based on the normalized power consumption data points, the power consumption data point set is generated.
[0029] As another embodiment of the present application, in step S1, in the set of electricity consumption data points, electricity consumption data of low-voltage substation users are obtained according to a preset electricity consumption cycle, and the electricity consumption cycle is day, month, quarter, half a year, or year.
[0030] As another embodiment of the present application, in step 5, the search radius is used to classify the user stations according to the hierarchical principle of the power grid company.
[0031] As another embodiment of the present application, in step 5, the user electricity consumption data of different typical substations are classified using an improved K-Means clustering algorithm, the user's electricity consumption behavior characteristics are analyzed and compared with the low-voltage electricity consumption characteristics, and according to the obtained electricity metering characteristics, the mortgaged users are assigned to one of the five typical substations: high-rise residential areas, old residential areas, urban-rural fringe areas, isolated small settlements and rural areas.
[0032] As another embodiment of the present application, after step S5, it also includes: analyzing the electricity consumption data of different categories of low-voltage substation users obtained after cluster analysis; the adopted analysis system includes an influencing factor determination module, a standardization module, a cluster analysis module and a feature analysis module connected in sequence, the influencing factor determination module is used to determine the low-voltage substation user electricity load characteristic index; the standardization module is used to standardize the low-voltage substation user electricity consumption information according to the determined user electricity load characteristic index; the cluster analysis module is used to perform cluster analysis on the standardized low-voltage substation user electricity consumption information; the feature analysis module is used to analyze the electricity consumption data of different categories of low-voltage substation users obtained after cluster analysis.
[0033] The present invention provides a low-voltage area user electricity consumption data clustering method based on mean shift clustering, which has the following beneficial effects compared with the prior art:
[0034] (1) In the low-voltage power consumption information collection scenario, the user types are diverse, and the power consumption scenarios and behaviors are complex, which effectively improves the quality foundation of power consumption data analysis.
[0035] (2) The traditional k-means clustering algorithm requires a clear number of clusters k as input. The present invention is suitable for scenarios where the user types in the low-voltage area are uncertain and the number of types k cannot be directly determined.
[0036] (3) The electricity consumption data clustering method proposed in the present invention is an unsupervised clustering method and does not require training based on prior knowledge. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0038] Figure 1 It is a flow chart of a method for clustering power consumption data of low-voltage area users based on mean shift clustering provided by an embodiment of the present invention;
[0039] Figure 2 It is a schematic diagram of a power consumption data matrix of a low-voltage area user power consumption data clustering method based on mean shift clustering provided by an embodiment of the present invention;
[0040] Figure 3 It is a relative error result of missing data repair of a low-voltage area user electricity consumption data clustering method based on mean shift clustering provided by an embodiment of the present invention;
[0041] Figure 4 It is a root mean square error result of missing data repair of a low-voltage area user electricity consumption data clustering method based on mean shift clustering provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0042] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0043] See also Figure 1 Now, a method for clustering power consumption data of low-voltage area users based on mean shift clustering provided by the present invention is described. The method for clustering power consumption data of low-voltage area users based on mean shift clustering includes the following steps:
[0044] S1: Given an m-dimensional electricity consumption dataset X∈R m×n , calculate the density estimation function of user electricity consumption data xi (i = 1, 2, ..., n);
[0045] S2: According to the extreme value theorem of the function, the local density maximum point is located at the gradient zero point of the density function. The gradient function is derived from the above formula to find the maximum point of the electricity consumption data density.
[0046] S3: Considering the influence of noise data in electricity consumption data, the Gaussian kernel function is introduced to achieve high-dimensional separability of user electricity consumption data and improve the robustness of classification;
[0047] S4: Calculate the mean shift vector M h , the size and direction of the center point offset are calculated according to the drift vector, so as to determine whether the center point iteration is completed and calculate the new center point;
[0048] S5: Update the candidate point of the cluster center point to the mean value of all electricity consumption data points within the search radius, find the dense area in the electricity consumption data X, locate the center points of different types of electricity consumption data, and complete the clustering of electricity consumption data.
[0049] The present invention provides a low-voltage area user electricity consumption data clustering method based on mean shift clustering, which has the following characteristics compared with the prior art:
[0050] (1) In the low-voltage power consumption information collection scenario, the user types are diverse, and the power consumption scenarios and behaviors are complex, which effectively improves the quality foundation of power consumption data analysis.
[0051] (2) The traditional k-means clustering algorithm requires a clear number of clusters k as input. The present invention is suitable for scenarios where the user types in the low-voltage area are uncertain and the number of types k cannot be directly determined.
[0052] (3) The electricity consumption data clustering method proposed in the present invention is an unsupervised clustering method and does not require training based on prior knowledge.
[0053] in, Figure 2 It is a schematic diagram of finding the density center point of data and completing data clustering according to the mean shift vector in the mean shift clustering process of electricity consumption data of the present invention.
[0054] In order to verify the effectiveness of the low-voltage area user power consumption data clustering method based on mean shift clustering of the present invention, the actual low-voltage area user power consumption data is clustered and the effect is analyzed. Figure 3 The image shown is an actual image of electricity consumption data of users in the low-voltage area. The types of electricity consumption data are diverse and the behaviors are complex.
[0055] Figure 4 The singular value distribution characteristics of the original electricity consumption data, the random data matrix and the clustered data matrix are shown in FIG. 1 . The clustering result obtained by the clustering method of the present invention classifies the data of similar electricity consumption behaviors into one category, and the singular value distribution characteristics of the matrix are significantly changed.
[0056] As a specific implementation of a method for clustering power consumption data of low-voltage area users based on mean shift clustering provided by the present invention, step S1 includes the following contents:
[0057] The density estimation function of user electricity consumption data is:
[0058]
[0059] In the formula, C k is a constant, k(x) is the kernel function, h is the kernel width, and x is the center point of the initial data.
[0060] As a specific implementation of a method for clustering power consumption data of low-voltage area users based on mean shift clustering provided by the present invention, step S2 includes the following contents:
[0061] The gradient function of the density estimation function of user electricity consumption data is:
[0062]
[0063] As a specific implementation of a method for clustering power consumption data of low-voltage area users based on mean shift clustering provided by the present invention, step S3 includes the following contents:
[0064] Define the Gaussian kernel function:
[0065]
[0066] Where x is the user's electricity consumption data vector, and h is the search radial kernel width.
[0067] As a specific implementation of a method for clustering power consumption data of low-voltage area users based on mean shift clustering provided by the present invention, step S4 includes the following contents:
[0068] According to the gradient function (2) of the power consumption data density estimation function, substitute it into the Gaussian kernel function (3) to obtain the drift vector Mh(x):
[0069]
[0070]
[0071] As a specific implementation of a method for clustering power consumption data of low-voltage area users based on mean shift clustering provided by the present invention, in step S1, obtaining a set of power consumption data points includes:
[0072] Obtaining original electricity consumption data points of the user to be detected;
[0073] Normalizing the values of the original power consumption data points to obtain normalized power consumption data points;
[0074] The set of electricity usage data points is generated based on the normalized electricity usage data points.
[0075] The original electricity consumption data points may refer to the user's daily electricity consumption data, monthly electricity consumption data, quarterly electricity consumption data, semi-annual electricity consumption data and annual electricity consumption data.
[0076] In a specific implementation, the process of obtaining a set of electricity consumption data points by the computer device specifically includes: the computer device can obtain the user's daily electricity consumption data and normalize the data until the data volume reaches a certain scale, and then obtain the normalized electricity consumption data points. Finally, the computer device generates a set of electricity consumption data points based on the normalized electricity consumption data points.
[0077] In practical applications, the electricity consumption data point set may be data in a two-dimensional format generated monthly.
[0078] In some embodiments, in step 5, the search radius is used to classify the user stations according to the hierarchical principles of the power grid companies.
[0079] In some embodiments, in step 5, the user electricity consumption data of different typical substations are classified using an improved K-Means clustering algorithm, the user's electricity consumption behavior characteristics are analyzed and compared with the low-voltage electricity consumption characteristics, and according to the obtained electricity metering characteristics, the mortgage users are assigned to one of the five typical substations: high-rise residential areas, old residential areas, urban-rural fringe areas, isolated small settlements and rural areas.
[0080] Since the K-means algorithm has the disadvantage that the selection of the initial value of the number of clusters affects the clustering effect, this influencing factor is considered and the K-means algorithm is optimized.
[0081] The KL index is used to determine the optimal K value, as shown in formula (4). By calculating the evaluation criterion function, the number of clusters corresponding to the maximum value is taken as the optimal number of clusters.
[0082] k=argmax[KL(h)] (4)
[0083] Where p is the data dimension; h is the number of clusters; W h is the sum of squares of intra-class distances when the number of clusters is h; DIEF is the change in intra-class distances when the number of clusters changes from h-1 to h for clustering p-dimensional data.
[0084] In some embodiments, after step S5, it also includes: analyzing the electricity consumption data of different categories of low-voltage substation users obtained after cluster analysis, and formulating load management measures according to the analysis results; the adopted analysis system includes an influencing factor determination module, a standardization module, a cluster analysis module and a feature analysis module connected in sequence, the influencing factor determination module is used to determine the low-voltage substation user electricity load characteristic index; the standardization module is used to standardize the low-voltage substation user electricity consumption information according to the determined user electricity load characteristic index; the cluster analysis module is used to perform cluster analysis on the standardized low-voltage substation user electricity consumption information; the feature analysis module is used to analyze the different categories of low-voltage substation user electricity consumption data obtained after cluster analysis.
[0085] Optionally, the low-voltage substation user power load characteristic indicators determined by the influencing factor determination module specifically include: typical daily maximum / minimum load, daily average load, daily load rate, daily minimum load rate, daily peak-to-valley difference rate, monthly load rate, annual average monthly load rate, seasonal imbalance coefficient, maximum load utilization hours, annual load rate, daily load curve, and annual load curve.
[0086] Optionally, the influencing factor determination module is also used to determine climate factors and time factors.
[0087] Optionally, the standardization module includes a user power load characteristic index collection module, a parameter input module, an index weight calculation module and a weighted calculation module which are connected in sequence. The user power load characteristic index collection module is used to determine the user power load characteristic index; the parameter input module is used to substitute the user power load characteristic index into the standard power characteristic parameter matrix; the index weight calculation module is used to calculate the corresponding different characteristic quantities for the user standard power characteristic quantity matrix, and calculate the index weight of the characteristic quantity by using the entropy weight method; the weighted calculation module is used to perform weighted calculation on the standard power characteristic quantity matrix according to the weights of the power characteristic quantities to be evaluated obtained by calculation, so as to obtain a weighted matrix that can reflect the user's comprehensive power consumption information.
[0088] Through reasonable demand response strategies, we provide users with personalized and differentiated power supply services, guide users to change their electricity usage behavior, reduce peak power load, and achieve the effect of peak shaving and valley filling, thereby improving energy utilization efficiency, reducing users' electricity costs, and improving the stability of power system operation.
[0089] The analysis system provided by the embodiment of the present invention needs to be operated with the help of terminal computing devices such as desktop computers, notebooks, PDAs and cloud servers. Those skilled in the art will understand that the controller includes, but is not limited to, a processor and a memory. Those skilled in the art will understand that the controller may include more or fewer components than shown in the figure, or combine certain components, or different components, for example, the controller may also include input and output devices, network access devices, buses, etc.
[0090] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0091] The memory may be an internal storage unit of the controller, such as a hard disk or memory of the controller. The memory may also be an external storage device of the controller, such as a plug-in hard disk equipped on the controller, a smart media card (SMC), a secure digital (SD) card, a flash card (Flash Card), a random access memory (RAM), or a non-volatile memory (NVM), such as at least one disk storage.
[0092] Furthermore, the memory may include both an internal storage unit of the controller and an external storage device. The memory is used to store computer programs and other programs and data required by the controller. The memory may also be used to temporarily store data that has been output or is to be output.
[0093] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0094] The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk (SSD)).
[0095] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for clustering power consumption data of low-voltage area users based on mean shift clustering, characterized in that: The steps include: S1: Given m Electricity consumption dataset in dimensional space X ∈R m × n , calculate user electricity consumption data xi(i=1,2,…,n) Density estimation function of ; S2: According to the extreme value theorem of the function, the local density maximum point is located at the gradient zero point of the density function. The gradient function is derived from the above formula to find the maximum point of the electricity consumption data density. S3: Considering the influence of noise data in electricity consumption data, the Gaussian kernel function is introduced to achieve high-dimensional separability of user electricity consumption data and improve the robustness of classification; S4: Calculate the mean shift vector M h , the size and direction of the center point offset are calculated according to the drift vector, so as to determine whether the center point iteration is completed and calculate the new center point; S5: Update the candidate point of the cluster center to the mean of all electricity consumption data points within the search radius to find the electricity consumption data X The dense areas in the grid are used to locate the center points of different types of electricity consumption data and complete the clustering of electricity consumption data; The step S1 comprises the following contents: The density estimation function of user electricity consumption data is: (1) In the formula, C is a constant, k(x) is the kernel function, h is the core width, x The data center point for initial setup; The step S2 comprises the following contents: The gradient function of the density estimation function of user electricity consumption data is: (2); The step S3 comprises the following contents: Define the Gaussian kernel function: (3) In the formula, x is the user's electricity consumption data vector, h To search for radial kernel width; The step S4 comprises the following contents: According to the gradient function (2) of the power consumption data density estimation function, the Gaussian kernel function (3) is substituted to obtain the drift vector Mh (x) : 。 2. A method for clustering power consumption data of low-voltage area users based on mean shift clustering as claimed in claim 1, characterized in that: In the step S1, obtaining a set of power consumption data points of users in a low-voltage area includes: Obtaining original electricity consumption data points of the user to be detected; Normalizing the values of the original power consumption data points to obtain normalized power consumption data points; The set of electricity usage data points is generated based on the normalized electricity usage data points.
3. A method for clustering low-voltage area user electricity consumption data based on mean shift clustering as claimed in claim 1, characterized in that: In the step S1, in the set of electricity consumption data points, the electricity consumption data of low-voltage area users is obtained according to a preset electricity consumption cycle, and the electricity consumption cycle is daily, monthly, quarterly, half-yearly, and annual.
4. A method for clustering low-voltage area user electricity consumption data based on mean shift clustering as claimed in claim 1, characterized in that: In step 5, the search radius is divided into user areas according to the hierarchical principle of the power grid company.
5. A method for clustering low-voltage area user electricity consumption data based on mean shift clustering as claimed in claim 1, characterized in that: In step 5, the user electricity consumption data of different typical substations are classified using an improved K-Means clustering algorithm to analyze the user's electricity consumption behavior characteristics and compare them with the low-voltage electricity consumption characteristics. According to the obtained electricity metering characteristics, the mortgaged users are classified into one of the five typical substations: high-rise residential areas, old residential areas, urban-rural fringe areas, isolated small settlements and rural areas.
6. A method for clustering low-voltage area user electricity consumption data based on mean shift clustering as claimed in claim 1, characterized in that: After step S5, it also includes: analyzing the electricity consumption data of different categories of low-voltage area users obtained after cluster analysis; the adopted analysis system includes an influencing factor determination module, a standardization module, a cluster analysis module and a feature analysis module connected in sequence, the influencing factor determination module is used to determine the low-voltage area user electricity load characteristic index; the standardization module is used to standardize the low-voltage area user electricity consumption information according to the determined user electricity load characteristic index; the cluster analysis module is used to perform cluster analysis on the standardized low-voltage area user electricity consumption information; the feature analysis module is used to analyze the different categories of low-voltage area user electricity consumption data obtained after cluster analysis.
Citation Information
Patent Citations
Large industrial user subdivision method based on clustering algorithm
CN110852370A
Analysis method and analysis device for residence vacancy state, and electronic equipment
CN113094448A