Power consumption data processing method for multiple user types

By constructing heterogeneous power graphs and applying graph convolution and clustering algorithms, the problem of difficulty in dealing with multi-user type electricity consumption data in the prior art is solved, and multi-scale analysis of user electricity consumption behavior and accurate identification of group behavior patterns are achieved.

CN119359344BActive Publication Date: 2025-06-03STATE GRID ZHEJIANG ELECTRIC POWER CO LTD NINGBO YINZHOU DISTRICT POWER SUPPLY CO
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411899278.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-06-03
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process multi-user type electricity consumption data, ignores the differences and correlations of electricity consumption patterns between different user types, and lacks the ability to mine user electricity consumption behavior patterns from a multi-scale perspective in time and space.

Method used

By constructing heterogeneous power graphs, applying heterogeneous graph convolution and spatiotemporal convolution models, multi-scale spatiotemporal dependence patterns are extracted, and combined with fuzzy C-mean clustering and community discovery algorithms, typical electricity consumption behavior patterns within the user group are identified.

Benefits of technology

The comprehensive processing of multi-user type electricity consumption data is realized, the comprehensiveness and accuracy of mining of the correlation of electricity consumption behavior is improved, and the user groups with similar electricity consumption behavior patterns can be accurately divided, and typical electricity consumption habit combinations and significant correlation characteristics are found within the group.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119359344B_ABST
    Figure CN119359344B_ABST
Patent Text Reader

Abstract

The present invention provides a power consumption data processing method for multiple user types, which relates to the technical field of power grids. The method includes constructing a heterogeneous power grid of a target area based on multiple node types, node attributes, multiple edge types, and edge attributes; applying heterogeneous graph convolution on the heterogeneous power grid, introducing a time dimension on the basis of heterogeneous graph convolution to construct a spatio-temporal heterogeneous graph convolution model, applying one-dimensional convolution in the time dimension to extract multi-scale spatio-temporal dependence patterns, and obtaining spatio-temporal heterogeneous features corresponding to the heterogeneous power grid; using the fuzzy C-means clustering algorithm to cluster users for the spatio-temporal heterogeneous features, obtaining a membership matrix of users for each cluster by optimizing the objective function, determining the optimal number of clusters according to the clustering validity index, and obtaining the user clustering result; by converting the power consumption load curve into a symbol sequence and applying a community discovery algorithm to identify the behavior clusters within the user group, obtaining typical power consumption behavior patterns.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power grids, and particularly to a method for processing power consumption data for multiple user types. Background Art

[0002] With the rapid development of smart grids and energy Internet, a large amount of multi-source heterogeneous power consumption data is continuously generated and accumulated. The power consumption behaviors of different user types are significantly different. Mining the valuable information contained in the power data of multiple users is of great significance for improving power supply service quality, optimizing energy resource allocation, and innovating power-added services.

[0003] Traditional power consumption data processing methods mainly target single user types, ignoring the differences and correlations in power consumption patterns among different user types. At the same time, existing methods often use single data sources and single time scales for modeling and analysis, lacking the ability to mine the regularities of user power consumption behaviors from the perspective of spatio-temporal multi-scales. In addition, the large amount of multi-source heterogeneous power data brings challenges in storage, calculation, analysis, etc. to existing processing methods. Summary of the Invention

[0004] Embodiments of the present invention provide a method for processing power consumption data for multiple user types, which can solve the problems in the prior art.

[0005] In the first aspect of the embodiments of the present invention,

[0006] A method for processing power consumption data for multiple user types is provided, including:

[0007] Constructing a heterogeneous power grid of a target area based on multiple node types, node attributes, multiple edge types, and edge attributes, where the node types include user nodes, transformer nodes, and power supply area nodes, the node attributes include user types, power consumption, and load curves, the edge types include user-transformer connection edges and transformer-power supply area connection edges, and the edge attributes include power transmission capacity and distance;

[0008] Applying heterogeneous graph convolution on the heterogeneous power grid, introducing a time dimension on the basis of heterogeneous graph convolution, constructing a spatio-temporal heterogeneous graph convolution model, regarding each time step as a graph snapshot, constructing a heterogeneous power grid corresponding to the time step, applying the spatio-temporal heterogeneous graph convolution model on the heterogeneous power grid of each time step, and extracting multi-scale spatio-temporal dependence patterns by applying one-dimensional convolution in the time dimension while applying graph convolution on the heterogeneous power grids of different time steps, to obtain spatio-temporal heterogeneous features corresponding to the heterogeneous power grid;

[0009] The fuzzy C - means clustering algorithm is used to cluster users based on the spatio - temporal heterogeneous features. The membership degree matrix of users for each cluster is obtained by optimizing the objective function. The optimal number of clusters is determined according to the clustering validity index, and the user clustering result is obtained. Association rule mining is applied to the data within each user group. Based on the results of association rule mining, the community discovery algorithm is used to identify the behavior clusters within the user group, and typical electricity consumption behavior patterns are obtained.

[0010] In an alternative embodiment,

[0011] The obtained spatio - temporal heterogeneous features corresponding to the heterogeneous power graph include:

[0012] ;

[0013] Among them, represents the spatio - temporal heterogeneous feature of the l+1 -th layer network at the t moment. σ represents the activation function. P, K respectively represent the size of the convolutional kernel in the time dimension and the size of the convolutional kernel of the heterogeneous graph convolution. represents the spatio - temporal convolutional kernel parameter. represents t the adjacency matrix of the heterogeneous power graph at the represents the heterogeneous graph convolution operation. represents the spatio - temporal heterogeneous feature of the l - th layer network at the t moment.

[0014] In an alternative embodiment,

[0015] The fuzzy C - means clustering algorithm is used to cluster users based on the spatio - temporal heterogeneous features. The membership degree matrix of users for each cluster obtained by optimizing the objective function includes:

[0016] Initialize the membership degree matrix, and randomly assign initial membership degree values that satisfy the sum of membership degrees equal to 1 for each data point. Among them, the spatio - temporal heterogeneous features corresponding to the heterogeneous power graph include the spatio - temporal heterogeneous feature vectors of each node in the heterogeneous power graph, and each data point corresponds to the spatio - temporal heterogeneous feature vector of a user node.

[0017] According to the current membership degree matrix, calculate the center vector of each cluster. The center vector is the weighted average of the membership degrees of all data points, and update the cluster center according to the center vector.

[0018] Update the membership degree matrix based on the distance between the data point and the cluster center, and realize the inverse ratio relationship between the membership degree value and the distance by optimizing and minimizing the objective function.

[0019] Among them, the objective function is the sum of the squares of the distances between the data points and their membership - weighted cluster centers.

[0020] Iteratively execute the calculation of the cluster centers and the update of the membership matrix until the change in the objective function is less than a preset threshold or the maximum number of iterations is reached, and output the optimized membership matrix to achieve the soft assignment of each cluster by the user.

[0021] In an alternative embodiment,

[0022] Determine the optimal number of clusters according to the cluster validity index, and the user grouping result includes:

[0023] Set the search range of the candidate number of clusters, and perform fuzzy C-means clustering for each candidate number of clusters to obtain the corresponding membership matrix and cluster centers, where the candidate number of clusters is randomly set and the value is in the cluster set between the minimum number of clusters and the maximum number of clusters;

[0024] Calculate the cluster validity index for each candidate number of clusters, and the cluster validity index includes the partition coefficient, the partition entropy, and the Xie-Beni index;

[0025] Among them, the partition coefficient represents the mean of the sum of squares of the elements in the membership matrix; the partition entropy represents the mean of the product of the elements in the membership matrix and their logarithms; the Xie-Beni index is the ratio of the cluster compactness to the cluster separation degree, where the cluster compactness represents the sum of squares of the distances between the data points and their membership-weighted cluster centers, and the cluster separation degree represents the square of the minimum distance between different cluster centers;

[0026] Select the candidate number of clusters with the largest partition coefficient or the smallest partition entropy or the smallest Xie-Beni index as the optimal number of clusters;

[0027] Output the membership matrix and cluster centers under the optimal number of clusters, and determine the belonging label of each data point according to the membership matrix under the optimal number of clusters to obtain the user grouping result.

[0028] In an alternative embodiment,

[0029] Apply association rule mining to the data within each user group, and based on the association rule mining results, apply a community discovery algorithm to identify the behavior clusters within the user group, and the typical electricity consumption behavior patterns obtained include:

[0030] Perform Z-score normalization on the electricity load curve of each user in the user group, equally divide the normalized sequence into several intervals, calculate the mean value of each interval, and map the mean value of each interval to the corresponding symbol according to the predefined symbol alphabet to obtain the electricity load symbol sequence;

[0031] Apply the FP-growth algorithm to mine the frequent patterns of electricity consumption behavior. By scanning the internal dataset of the user group, calculate the occurrence frequency of each electricity consumption load symbol and sort them in descending order of frequency. Create an FP-tree to insert the symbol sequences in descending order of frequency, and recursively construct the conditional pattern base and conditional FP-tree to generate frequent patterns. Among them, the internal dataset of the user group includes the electricity consumption load symbol sequences of all users;

[0032] Generate the association rules of electricity consumption behavior based on the mined frequent patterns. Filter the frequent item sets from the set of frequent patterns according to the minimum support threshold, generate all non-empty subsets of each frequent item set, calculate the confidence of the association rules between the non-empty subsets and the frequent item sets, output the association rules that meet the requirements according to the minimum confidence threshold, and use the Lift index to evaluate the significance value of the association rules. Define the electricity consumption behavior corresponding to the association rules with a significance value greater than the preset significance threshold as the typical electricity consumption behavior pattern. Among them, the set of frequent patterns includes the frequent patterns of each electricity consumption load symbol.

[0033] In an alternative embodiment,

[0034] The method further includes:

[0035] Based on the frequent patterns and association rules, apply the community discovery algorithm to identify the associated behavior clusters within the user group. Dynamically adjust the modularity by optimizing the modularity objective function until the community attributes of all nodes no longer change, realizing community division. The process of dynamically adjusting the modularity is as follows:

[0036] ;

[0037] Among them, △Q represents the change in modularity, represents the sum of the weights of the edges within the community, k i,in represents the node i The sum of the weights of the edges connecting to other nodes in the community, m represents the fuzzification parameter, represents the sum of the weights of the edges connecting the nodes inside and outside the community, k i represents the node i degree.

[0038] In an alternative embodiment,

[0039] The method further includes:

[0040] Statistically count the set of frequent patterns within each community, and define the frequent patterns with an occurrence frequency exceeding the preset threshold as the candidate typical electricity consumption behavior patterns of the community;

[0041] Calculate the similarity between candidate typical patterns, merge similar patterns by setting a similarity threshold, and obtain a set of deduplicated typical electricity consumption behavior patterns;

[0042] For each typical electricity consumption behavior pattern, traverse the set of association rules in the community, filter out the high-confidence association rules with this pattern as the antecedent, use them as the constraint rules for this typical pattern, and use the consequent of the rules as the supplementary attributes of the typical pattern;

[0043] Filter out the duplicate electricity consumption behaviors already included in the set of typical electricity consumption behavior patterns in the supplementary attributes, and update the set of typical electricity consumption behavior patterns;

[0044] For the set of typical electricity consumption behavior patterns of each community, count the support degrees of each typical electricity consumption behavior pattern, and use the top K patterns with the highest support degrees as the main electricity consumption behavior preferences of this community, and the patterns with support degrees in the range from K to M as the secondary electricity consumption behavior preferences of this community, forming a set of typical electricity consumption behavior patterns that reflect the group behavior preferences and habits, where K is a preset main preference threshold and M is a preset secondary preference threshold.

[0045] This application constructs a heterogeneous power graph containing multiple node types, attributes, multiple edge types, and attributes, fully characterizing multi-dimensional heterogeneous information such as the topological structure of the power system, user characteristics, and power supply equipment parameters, providing a complete data basis for subsequent feature learning and association mining. By constructing a heterogeneous power graph and applying deep learning models such as graph convolution, the topological structure information, spatio-temporal dependence relationships, and other intrinsic features of multi-source heterogeneous power data are fully utilized, improving the comprehensiveness and accuracy of the mining of electricity consumption behavior correlations.

[0046] By introducing fuzzy C-means clustering and adaptive clustering evaluation, soft clustering of users is realized based on spatio-temporal heterogeneous features, overcoming the limitations of traditional hard clustering methods, and accurately dividing user groups with similar electricity consumption behavior patterns. Through data mining algorithms such as frequent pattern mining, association rule learning, and community discovery, typical electricity consumption habit combinations and significant association features within the group are discovered from multiple perspectives, and a structured knowledge base of electricity consumption behavior patterns is formed, laying a foundation for subsequent applications such as personalized services and precision marketing. Brief Description of the Drawings

[0047] Figure 1 It is a schematic flowchart of the method for processing electricity consumption data for multiple user types according to the embodiment of the present invention. Detailed Embodiment

[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0049] The following uses specific embodiments to elaborate on the technical solutions of the present invention in detail. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0050] Figure 1 The following is a schematic flowchart of the method for processing electricity consumption data for multiple user types according to the embodiments of the present invention. As Figure 1 shown, the method includes:

[0051] S101. Construct a heterogeneous power grid of the target area based on multiple node types, node attributes, multiple edge types, and edge attributes. Among them, the node types include user nodes, transformer nodes, and power supply area nodes; the node attributes include user types, electricity consumption, and load curves; the edge types include user-transformer connection edges and transformer-power supply area connection edges; and the edge attributes include power transmission capacity and distance.

[0052] Exemplarily, the detailed steps for constructing a heterogeneous power grid of the target area based on multiple node types, node attributes, multiple edge types, and edge attributes are as follows:

[0053] Collect multi-source heterogeneous data such as users, transformers, and power supply areas in the target area, including user type information, historical electricity consumption data, load curve data, transformer type parameters, connection relationships with users and power supply areas, geographical locations of power supply areas, jurisdiction scopes, and other data. Perform preprocessing operations such as cleaning, deduplication, and normalization on the collected heterogeneous data, handle missing values and outliers, and convert the data into a format suitable for graph model representation.

[0054] Based on the preprocessed data, construct the nodes of the heterogeneous power grid. Create node objects of three types: user nodes, transformer nodes, and power supply area nodes. The node attributes include: User node attributes: user type, electricity consumption, load curve, etc.; Transformer node attributes: transformer type, capacity, belonging power supply area, etc.; Power supply area node attributes: area name, jurisdiction scope, total regional electricity consumption, etc.

[0055] Construct the edges of the heterogeneous power grid according to the connection relationships between nodes. Create two types of edge objects: user-transformer connection edges and transformer-power supply area connection edges. The edge attributes include: User-transformer connection edge attributes: power transmission capacity, power transmission distance, etc.; Transformer-power supply area connection edge attributes: power supply area to which the transformer belongs, number of transformers within the power supply area, etc.

[0056] Combine all node objects and edge objects to form a complete heterogeneous power grid. Use graph data structures such as adjacency matrices and adjacency lists to store the information of nodes and edges, and at the same time record the attribute information of different types of nodes and edges. Perform data verification on the constructed heterogeneous power grid to check the integrity and consistency of the attributes of nodes and edges, and ensure the correctness and usability of the graph model.

[0057] Persistently store the constructed heterogeneous power grid. It can be stored using a graph database (such as Neo4j, JanusGraph) or a custom graph data format (such as JSON, XML) for subsequent graph data analysis and applications. Through the above steps, a heterogeneous power grid that comprehensively describes the power system topology, user characteristics, and device attributes of the target area can be constructed based on multi-source heterogeneous data, laying a foundation for subsequent graph data mining and applications.

[0058] S102. Apply heterogeneous graph convolution on the heterogeneous power grid. Introduce the time dimension on the basis of heterogeneous graph convolution to construct a spatio-temporal heterogeneous graph convolution model. Consider each time step as a graph snapshot, construct the heterogeneous power grid corresponding to the time step, and apply the spatio-temporal heterogeneous graph convolution model on the heterogeneous power grid at each time step. By applying graph convolution on the heterogeneous power grids at different time steps and applying one-dimensional convolution in the time dimension, extract multi-scale spatio-temporal dependence patterns to obtain the spatio-temporal heterogeneous features corresponding to the heterogeneous power grid.

[0059] In an alternative embodiment,

[0060] Obtaining the spatio-temporal heterogeneous features corresponding to the heterogeneous power grid includes:

[0061] ;

[0062] Where, represents the spatio-temporal heterogeneous feature of the l+1 -th layer network at time t , σ represents the activation function, P, K respectively represent the size of the convolutional kernel in the time dimension and the size of the convolutional kernel of the heterogeneous graph convolution, represents the spatio-temporal convolutional kernel parameters, represents t the adjacency matrix of the heterogeneous power grid at time represents the heterogeneous graph convolution operation, Represents the spatio-temporal heterogeneous features of the l-th layer of the network at time t.

[0063] Exemplarily, the heterogeneous graph convolution design: For different types of nodes and edges in the heterogeneous power graph, independent graph convolution kernel functions are designed. For each type of node, a corresponding node convolution kernel function is designed to aggregate the features of the node itself and the features transmitted by its neighbor nodes; for each type of edge, a corresponding edge convolution kernel function is designed to converge the relationship features between the connected nodes.

[0064] Combine the heterogeneous graph convolution kernel functions (node convolution kernel function and edge convolution kernel function) into a complete heterogeneous graph convolution layer. For each node, apply the corresponding type of node convolution kernel function to aggregate the node features and neighbor information; for each edge between nodes, apply the corresponding type of edge convolution kernel function to converge the connection relationship features; add the results of node convolution and edge convolution to obtain the output features of the heterogeneous graph convolution layer.

[0065] Based on the heterogeneous graph convolution, introduce the time dimension to construct a spatio-temporal heterogeneous graph convolution model. Consider each time step as a graph snapshot and construct the corresponding heterogeneous power graph for the time step; apply the heterogeneous graph convolution layer on the heterogeneous power graph at each time step to extract the spatial dependency relationship features; apply one-dimensional convolution in the time dimension to extract the temporal evolution patterns between different time steps.

[0066] Construction of the spatio-temporal heterogeneous graph convolution model: Combine the heterogeneous graph convolution layer and the time dimension convolution to construct a complete spatio-temporal heterogeneous graph convolution model. For the heterogeneous power graph at each time step, apply the heterogeneous graph convolution layer to extract the spatial dependency features; stack the outputs of the heterogeneous graph convolution at different time steps in the time dimension to form time series data; apply one-dimensional convolution on the time series data to extract the evolution patterns between time steps; add the output of the one-dimensional convolution to the output of the heterogeneous graph convolution to obtain the final output features of the spatio-temporal heterogeneous graph convolution model.

[0067] By setting different sizes of convolution kernels and stacking multiple layers of spatio-temporal heterogeneous graph convolution models, multi-scale spatio-temporal dependency patterns are extracted. In the spatial dimension, by setting different sizes of heterogeneous graph convolution kernels, the features of nodes and edges within different hop ranges are extracted; in the time dimension, by setting different sizes of one-dimensional convolution kernels, the evolution patterns at different time scales are extracted; by stacking multiple spatio-temporal heterogeneous graph convolution layers, the receptive field is gradually expanded to extract more global and high-level spatio-temporal features.

[0068] Take the output of the last layer of the spatio-temporal heterogeneous graph convolution model as the spatio-temporal heterogeneous feature representation of the heterogeneous power graph. For each node, output the corresponding spatio-temporal heterogeneous feature vector; for different types of nodes, the dimensions of their spatio-temporal heterogeneous feature vectors can be different to adapt to the feature representation requirements of different types of nodes.

[0069] Through the above steps, the spatio-temporal heterogeneous graph convolutional model can be applied to the heterogeneous power grid graph to extract both spatial dependence relationships and temporal evolution patterns simultaneously, obtaining the spatio-temporal heterogeneous feature representations of each node in the heterogeneous power grid graph, providing more comprehensive and accurate feature inputs for subsequent tasks such as user clustering and behavior analysis.

[0070] S103. Use the fuzzy C-means clustering algorithm to cluster users based on the spatio-temporal heterogeneous features, obtain the membership degree matrix of users for each cluster by optimizing the objective function, determine the optimal number of clusters according to the clustering validity index, and obtain the user clustering result; apply association rule mining to the data within each user group, and based on the results of association rule mining, apply the community discovery algorithm to identify the behavior clusters within the user group, obtaining typical electricity consumption behavior patterns.

[0071] Exemplarily, use the fuzzy C-means clustering algorithm to cluster users based on the spatio-temporal heterogeneous features. Take the spatio-temporal heterogeneous feature vector of each user node as the input data for clustering; initialize the cluster centers and the membership degree matrix, set the range of the number of clusters and the fuzzy factor; iteratively optimize the objective function to update the cluster centers and the membership degree matrix until the convergence condition is met; the objective function is to minimize the weighted sum of the squares of the distances between user nodes and cluster centers, where the weight is the fuzzy factor power of the membership degree.

[0072] Determine the optimal number of clusters according to the clustering validity index. For each candidate number of clusters, calculate the corresponding membership degree matrix and cluster centers; calculate common clustering validity indicators such as the partition coefficient, partition entropy, Xie-Beni index, etc.; comprehensively consider multiple indicators and select the cluster number with the best score as the final number of user clusters. Output of the user clustering result: According to the membership degree matrix corresponding to the optimal number of clusters, obtain the user clustering result. For each user node, select the cluster with the largest membership degree as its belonging group; output the group label of each user node and the list of user nodes included in each group. The user clustering result includes multiple user groups obtained after clustering.

[0073] Apply association rule mining to the electricity load curve data within each user group. Convert the electricity load curve of each user into a symbol sequence, for example, divide it into three intervals of high, medium, and low according to the load value and represent them with different symbols; construct a transaction database for the symbol sequences of all users within the group, where the symbol sequence of each user is a transaction; apply frequent pattern mining algorithms (such as Apriori, FP-growth, etc.) to extract frequent symbol combination patterns; generate association rules based on the frequent patterns, calculate indicators such as the support degree and confidence degree of the rules, and filter out strong association rules.

[0074] Apply community discovery algorithms to identify behavioral clusters within user groups. Based on the results of association rule mining, construct a user behavior association graph with users as nodes and edges representing the behavioral similarity between users; apply community discovery algorithms (such as Louvain, Infomap, etc.) to identify closely related user subgroups on the behavior association graph; output the list of users included in each community, as well as the frequent behavior patterns and association rules within the community.

[0075] Extraction of typical electricity consumption behavior patterns: Integrate frequent patterns, association rules, and community structure to extract typical electricity consumption behavior patterns. For each community, count the frequent symbol sequence patterns within it, and take the pattern with the highest occurrence frequency as the typical electricity consumption behavior; for each typical behavior pattern, mine other behavior patterns strongly associated with it to form an association pattern set for the typical behavior; compare the typical behavior patterns of different communities to extract common patterns and characteristic differences, and characterize the group behavior preferences.

[0076] Through the above steps, the spatio-temporal heterogeneous features can be fully utilized for accurate user grouping, and typical electricity consumption behavior patterns can be mined within the group to identify the behavior characteristics and association rules of different user subgroups, providing decision support for subsequent applications such as personalized services and precision marketing.

[0077] In an optional implementation manner,

[0078] Use the fuzzy C-means clustering algorithm to group users based on spatio-temporal heterogeneous features, and obtain the membership matrix of users for each cluster by optimizing the objective function, including:

[0079] Initialize the membership matrix, and randomly assign initial membership values that satisfy the sum of memberships equal to 1 for each data point. Among them, the spatio-temporal heterogeneous features corresponding to the heterogeneous power graph include the spatio-temporal heterogeneous feature vectors of each node in the heterogeneous power graph, and each data point corresponds to the spatio-temporal heterogeneous feature vector of a user node;

[0080] According to the current membership matrix, calculate the center vector of each cluster. The center vector is the membership-weighted average value of all data points, and update the cluster center according to the center vector;

[0081] Update the membership matrix based on the distance between the data point and the cluster center, and realize the inverse relationship between the membership value and the distance by optimizing and minimizing the objective function;

[0082] Among them, the objective function is the sum of the squares of the distances between the data points and their membership-weighted cluster centers;

[0083] Iteratively execute the calculation of the cluster center and the update of the membership matrix until the change in the objective function is less than the preset threshold or the maximum number of iterations is reached, and output the optimized membership matrix to achieve the soft assignment of users to each cluster.

[0084] Exemplarily, the detailed steps of using the fuzzy C-means clustering algorithm to perform user clustering on spatio-temporal heterogeneous features are as follows:

[0085] Take the spatio-temporal heterogeneous feature vectors of each user node as the input data points of the clustering algorithm.

[0086] Set the number of clusters C and the dimension of the membership matrix U to N×C, where N is the number of user nodes; randomly generate the initial value of the membership matrix U. For each user node, randomly assign C membership values and normalize to ensure that the sum of the memberships in each row is 1. According to the current membership matrix U, calculate the center vector of each cluster; for the i-th cluster, its center vector is the membership-weighted average of all data points, and the weight is the m-th power of the membership of the data point belonging to the i-th class, where m is the fuzzy factor (usually taken as 2); update the center vectors of all C clusters.

[0087] Calculate the distances between each data point and all cluster centers, usually using Euclidean distance or Mahalanobis distance; update the membership matrix U based on the distances, and realize the inverse relationship between the membership values and the distances by optimizing and minimizing the objective function; the objective function is the sum of the squares of the distances between all data points and their membership-weighted cluster centers, and the weight of the cluster center is the m-th power of the corresponding membership; use the Lagrange multiplier method to solve the minimum value of the objective function to obtain the update formula of the membership matrix U. Alternately calculate the cluster centers and update the membership matrix; calculate the value of the objective function after each iteration, and stop the iteration when the change is less than the preset threshold or the maximum number of iterations is reached; output the final membership matrix U and the cluster centers.

[0088] According to the optimized membership matrix U, perform soft clustering on each user node; for each user node, its memberships belonging to each cluster are represented by the corresponding row vector in the matrix U; the cluster with the largest membership is used as the main cluster of the user, and other clusters are used as the secondary clusters of the user to achieve the soft division of users into clusters. Calculate common clustering validity indicators, such as partition coefficient, partition entropy, Xie-Beni index, etc.; compare the indicator values under different numbers of clusters, and select the number of clusters with the best comprehensive performance as the final number of user clusters.

[0089] Through the above steps, the fuzzy C-means clustering algorithm can be used to perform soft clustering on user spatio-temporal heterogeneous features, obtain the memberships of each user to different clusters, and realize the fuzzy division of user behavior characteristics. Compared with traditional hard clustering methods, fuzzy C-means clustering can better describe the diversity and complexity of user behavior, and provide more refined and comprehensive user clustering information for subsequent user profiling and personalized services.

[0090] In an alternative embodiment,

[0091] Determine the optimal number of clusters according to the clustering validity index, and the user clustering result includes:

[0092] Set the search range of candidate cluster numbers, perform fuzzy C-means clustering on each candidate cluster number, and obtain the corresponding membership matrix and cluster centers. Among them, the candidate cluster numbers are randomly set, and the values are in the cluster set between the minimum number of clusters and the maximum number of clusters;

[0093] Calculate the clustering validity index for each candidate cluster number. The clustering validity index includes the partition coefficient, partition entropy, and Xie-Beni index;

[0094] Among them, the partition coefficient represents the mean of the sum of squares of the elements in the membership matrix; the partition entropy represents the mean of the product of the elements in the membership matrix and their logarithms; the Xie-Beni index is the ratio of the clustering compactness to the clustering separation degree, where the clustering compactness represents the sum of squares of the distances between the data points and their membership-weighted cluster centers, and the clustering separation degree represents the square of the minimum distance between different cluster centers;

[0095] Select the candidate cluster number with the largest partition coefficient or the smallest partition entropy or the smallest Xie-Beni index as the optimal number of clusters;

[0096] Output the membership matrix and cluster centers under the optimal number of clusters, determine the belonging label of each data point according to the membership matrix under the optimal number of clusters, and obtain the user clustering result.

[0097] Exemplarily, the detailed steps to determine the optimal number of clusters according to the clustering validity index and obtain the user clustering result are as follows: Set the candidate cluster number range: According to the number of user nodes and domain knowledge, set the search range [Cmin, Cmax] of the candidate cluster number; Determine the step size of the cluster number search to obtain the discrete value set of the candidate cluster number.

[0098] Perform fuzzy C-means clustering: For each candidate cluster number C, perform the fuzzy C-means clustering algorithm; Input the spatio-temporal heterogeneous features of the user nodes, set the fuzzy factor m and the maximum number of iterations; Initialize the membership matrix, iteratively calculate the cluster centers and update the membership matrix until convergence or the maximum number of iterations is reached; Obtain the optimal membership matrix and cluster centers under the current cluster number C.

[0099] For each candidate number of clusters C, based on the corresponding membership matrix and cluster centers, calculate the following cluster validity indices: Partition Coefficient (PC): The mean of the sum of squares of all elements in the membership matrix, with a value range of [1 / C, 1], and a larger value indicates a better clustering effect; Partition Entropy (PE): The mean of the product of elements in the membership matrix and their logarithms, with a value range of [0, logC], and a smaller value indicates a better clustering effect; Xie-Beni index (XB): The ratio of cluster compactness to cluster separation, where the compactness is the sum of squares of the distances between all data points and their membership-weighted cluster centers, and the separation is the square of the minimum distance between different cluster centers. A smaller XB value indicates a better clustering effect.

[0100] For each candidate number of clusters C, comprehensively consider multiple cluster validity indices, and select the number of clusters with the optimal comprehensive performance as the optimal number of clusters C*; Common selection strategies include: the largest partition coefficient, the smallest partition entropy, the smallest Xie-Beni index, etc.; If the optimal number of clusters for multiple indices is inconsistent, the final optimal number of clusters can be determined by weighted average or voting.

[0101] For the optimal number of clusters C*, output the corresponding optimal membership matrix U* and cluster centers V*; For each user node, find the cluster with the largest membership in the optimal membership matrix U*, and use it as the belonging label of the user; Divide the user nodes with the same belonging label into the same user group to obtain the final user grouping result.

[0102] For the user grouping result under the optimal number of clusters, calculate common external evaluation indices such as Rand index, mutual information, etc.; Compare the user grouping result with the known user attribute labels to evaluate the degree of agreement between the grouping result and the business understanding; Compare the user grouping results under different candidate numbers of clusters to analyze the rationality and stability of the selection of the optimal number of clusters.

[0103] Through the above steps, based on fuzzy C-means clustering, multiple cluster validity indices can be introduced to evaluate the clustering quality under different numbers of clusters, and the number of clusters with the optimal comprehensive performance can be selected as the final number of user groups. This adaptive method for selecting the number of clusters can automatically determine the optimal user grouping granularity according to data characteristics and business requirements, improving the accuracy and practicality of user grouping. At the same time, through the quality evaluation of the final user grouping result, the effectiveness of the grouping method can be further verified, providing a reliable basis for user profiling and precision marketing for subsequent user groups.

[0104] In an alternative embodiment,

[0105] Apply association rule mining to the internal data of each user group. Based on the results of the association rule mining, apply community discovery algorithms to identify the behavior clusters within the user group, and obtain typical electricity consumption behavior patterns, including:

[0106] Perform Z-score normalization on the electricity consumption load curves of each user within the user group, equally divide the normalized sequence into several intervals, calculate the mean value of each interval, and map the mean value of each interval to the corresponding symbol according to the predefined symbol alphabet to obtain the electricity consumption load symbol sequence;

[0107] Apply the FP-growth algorithm to mine the frequent patterns of electricity consumption behavior. By scanning the internal data set of the user group, calculate the occurrence frequency of each electricity consumption load symbol and sort them in descending order of frequency. Create an FP-tree to insert the symbol sequence with descending frequency, and recursively construct the conditional pattern base and conditional FP-tree to generate frequent patterns. Among them, the internal data set of the user group includes the electricity consumption load symbol sequences of all users;

[0108] Generate electricity consumption behavior association rules based on the mined frequent patterns. Screen the frequent item sets from the set of frequent patterns according to the minimum support threshold, generate all non-empty subsets of each frequent item set, calculate the confidence of the association rules between the non-empty subsets and the frequent item sets, output the association rules that meet the requirements according to the minimum confidence threshold, and use the Lift index to evaluate the significance value of the association rules. Define the electricity consumption behavior corresponding to the association rules with a significance value greater than the preset significance threshold as a typical electricity consumption behavior pattern. Among them, the set of frequent patterns includes the frequent patterns of each electricity consumption load symbol.

[0109] Exemplarily, apply association rule mining to the internal data of each user group. By converting the electricity consumption load curve into a symbol sequence, apply community discovery algorithms to identify the behavior clusters within the user group. The detailed steps to obtain typical electricity consumption behavior patterns are as follows:

[0110] Perform Z-score normalization on the electricity consumption load curves of each user within the group, shift the curve mean to 0, and scale the variance to 1; equally divide the normalized load sequence into several intervals (such as 10 intervals), and calculate the mean value of the load values within each interval; predefined a symbol alphabet (such as A-J), and map each interval to the corresponding symbol according to the ascending order of the interval mean; symbolize the load sequence of each user to obtain the electricity consumption load symbol sequence.

[0111] The power consumption load symbol sequences of all users within the group are formed into a data set, and the power consumption symbol sequence of each user is a transaction; the FP-growth algorithm is applied to mine frequent patterns: scan the data set, calculate the occurrence frequency of each symbol, and sort the symbol table in descending order of frequency; create an FP-tree, and insert the symbol sequences in each transaction into the FP-tree according to the symbol table sorted in descending order of frequency; recursively construct the conditional pattern base and the conditional FP-tree: construct the conditional pattern base for each symbol, that is, the set of paths with this symbol as the suffix; construct the conditional FP-tree for each conditional pattern base, and mine the frequent patterns in the conditional FP-tree; combine the frequent patterns of each symbol to generate a global set of frequent patterns.

[0112] According to the preset minimum support threshold, frequent itemsets are screened out from the set of frequent patterns; for each frequent itemset, all its non-empty subsets are generated; for each non-empty subset and the original frequent itemset, the confidence of its association rule is calculated, and the confidence is the ratio of the subset support to the frequent itemset support; according to the preset minimum confidence threshold, the association rules that meet the requirements are output; for each association rule, its Lift index is calculated, and the Lift index is the ratio of the association rule confidence to the support of the consequent symbol; according to the preset Lift significance threshold, the association rules with Lift values greater than the threshold are selected as significant association rules.

[0113] For each significant association rule, its antecedent symbol sequence is defined as a typical power consumption behavior pattern; all typical behavior patterns are clustered, and similar behavior patterns are grouped into one category, and each category represents a typical power consumption behavior cluster; for each behavior cluster, the number of users it covers and the load curve characteristics are statistically analyzed to characterize the user group attributes of this type of behavior pattern; the typical power consumption behavior clusters and the corresponding association rules of each user group are output as the description of the behavior characteristics of this group.

[0114] For each typical power consumption behavior cluster, the mean and variance of its load curve are extracted, and a load curve example is drawn; for each association rule, a distribution diagram of its antecedent symbol sequence on the time axis is drawn to show the time characteristics of this behavior pattern; a comparative analysis is carried out on the behavior clusters and association rules of different groups to reveal the differences and similarities in power consumption behavior between groups.

[0115] Through the above steps, the electricity consumption load curve within the user group can be transformed into a symbol sequence, and by using frequent pattern mining and association rule generation algorithms, typical electricity consumption behavior patterns and behavior clusters within the group can be discovered. By characterizing the load characteristics and user attributes of each behavior cluster, the electricity consumption habits and preferences of different user groups can be insighted, providing decision-making support for the precise marketing and personalized services of power companies. At the same time, by visually displaying typical behavior patterns and association rules, the interpretability and operability of the analysis results can be enhanced, facilitating business personnel to understand and apply the discoveries.

[0116] In an alternative embodiment,

[0117] the method further includes:

[0118] Based on frequent patterns and association rules, a community discovery algorithm is applied to identify behavior clusters associated within the user group, and the modularity is dynamically adjusted by optimizing the modularity objective function until the community attributes of all nodes no longer change, achieving community partitioning. Among them, the process of dynamically adjusting the modularity is as follows:

[0119] ;

[0120] where, △Q represents the change in modularity, represents the sum of the weights of the edges within the community, k i,in represents the node i and the sum of the weights of the edges connecting it to other nodes in the community, m represents the fuzzification parameter, represents the sum of the weights of the edges connecting the nodes inside and outside the community, k i represents the node i 's degree.

[0121] Exemplarily, through frequent pattern mining and association rule generation, typical electricity consumption behavior patterns and the association relationships between behaviors within the user group are discovered; by applying the community discovery algorithm, the behavior pattern nodes are divided into several behavior communities according to the association strength. The behavior patterns within each community are closely associated, while the associations between communities are relatively weak; the community structure reveals the hierarchical and organizational forms of the behavior associations within the user group, helping to understand the internal logic and influence mechanism of user behavior.

[0122] Traditional frequent pattern mining and association rule generation can only discover pairwise behavioral associations, while community discovery algorithms can automatically identify higher-order associations among multiple behavioral patterns, forming tightly associated behavioral clusters; each behavioral community corresponds to a subgroup of users, and users within the subgroup exhibit consistency in this behavioral cluster, while different subgroups exhibit differences in different behavioral clusters; automatically identifying behavioral clusters and user subgroups can discover the division of labor and role differentiation within the group, mine the diversity and heterogeneity of users, and provide precise segmentation for personalized services.

[0123] Modularity is an important indicator for measuring the quality of community division, representing the degree to which the association degree of nodes within a community exceeds that of a random network; by optimizing the modularity objective function and dynamically adjusting the community division, the actual association degree of nodes within the community is maximized, while the actual association degree between communities is minimized; the optimized community division can maintain the cohesion of behavioral clusters and the separation between behavioral clusters to the greatest extent, making the community discovery results more accurate and reliable.

[0124] In an alternative implementation,

[0125] The method further includes:

[0126] Statistically analyze the set of frequent patterns within each community, and define the frequent patterns with occurrence frequencies exceeding a preset threshold as the candidate typical electricity consumption behavior patterns of the community;

[0127] Calculate the similarity between candidate typical patterns, and merge similar patterns by setting a similarity threshold to obtain a deduplicated set of typical electricity consumption behavior patterns;

[0128] For each typical electricity consumption behavior pattern, traverse the set of association rules within the community, and filter out the high-confidence association rules with this pattern as the antecedent as the constraint rules for this typical pattern, and use the consequent of the rule as the supplementary attribute of the typical pattern;

[0129] Filter out the duplicate electricity consumption behaviors already included in the set of typical electricity consumption behavior patterns in the supplementary attributes, and update the set of typical electricity consumption behavior patterns;

[0130] For the set of typical electricity consumption behavior patterns of each community, statistically analyze the support degree of each typical electricity consumption behavior pattern, and take the top K patterns with the highest support degree as the main electricity consumption behavior preferences of the community, and the patterns with support degrees in the range from K to M as the secondary electricity consumption behavior preferences of the community, forming a set of typical electricity consumption behavior patterns that reflect the group's behavior preferences and habits, where K is a preset main preference threshold and M is a preset secondary preference threshold.

[0131] Exemplarily, the frequently co-occurring electricity consumption behaviors within each community are defined as typical electricity consumption patterns, and significant association rules are used to constrain and supplement the typical patterns. The detailed steps to finally obtain a set of typical electricity consumption behavior patterns reflecting the group's behavior preferences and habits are as follows:

[0132] For the behavior community corresponding to each user subgroup, count the set of frequent patterns within it; set a preset frequency threshold, and define the frequent patterns with frequencies exceeding the threshold as the candidate typical electricity consumption behavior patterns for this community; the candidate typical patterns reflect the high-frequency co-occurring behavior combinations of users within this community and represent the main electricity consumption preferences of this subgroup.

[0133] Calculate the similarity between candidate typical patterns, and indicators such as the Jaccard similarity coefficient can be used; set a similarity threshold, and merge the candidate typical patterns with similarities exceeding the threshold; the merged set of typical patterns removes highly similar duplicate patterns, improving the diversity and representativeness of the typical patterns.

[0134] Traverse the set of association rules within each community, and filter out the high-confidence association rules with typical electricity consumption behavior patterns as the antecedent; set a confidence threshold, and define the association rules with confidences exceeding the threshold as the constraint rules for the typical patterns; use the consequent of the constraint rules as the supplementary attributes of the typical patterns, enriching the semantic connotation and context association of the typical patterns.

[0135] Deduplicate the set of supplementary attributes for each typical pattern, filtering out duplicate electricity consumption behaviors that are already included in the typical pattern; the updated set of typical electricity consumption behavior patterns includes the original patterns and supplementary attributes and has no duplicate behaviors; the deduplicated set of typical patterns is more concise and comprehensive and can accurately express the behavior preferences and habits of the community.

[0136] For the set of typical electricity consumption behavior patterns of each community, count the support degrees of each typical electricity consumption behavior pattern within the community; set a main preference threshold K and a secondary preference threshold M, and define the top K typical patterns with the highest support degrees as the main electricity consumption behavior preferences of this community; define the typical patterns with support degrees between the top K and M as the secondary electricity consumption behavior preferences of this community; the main preferences reflect the most common and significant behavior patterns within the community, and the secondary preferences reflect the relatively less significant but still somewhat common behavior patterns within the community.

[0137] Exemplarily, "rule consequent" is an important concept in association rules, referring to the "B" part in the association rule "if A then B", that is, the subsequent event or attribute that will necessarily occur when the rule antecedent A appears. The rule consequent represents the inference result or prediction target of the association rule, reflecting the causal association or correlation between the antecedent and the consequent. In association rule mining, the support, confidence, and lift of the rule consequent are important indicators for evaluating the significance and effectiveness of the rule.

[0138] The following gives specific examples of rule consequents:

[0139] Association rule of electricity consumption behavior: "If a user has a high electricity load in the afternoon (14:00 - 16:00) in summer, then it is very likely that the user also has a high electricity load at night (20:00 - 22:00)".

[0140] Rule consequent: The user has a high electricity load at night (20:00 - 22:00).

[0141] Explanation of meaning: When the rule antecedent "the user has a high electricity load in the afternoon in summer" holds, it is inferred that the user also has a high electricity load at night, reflecting the correlation between the user's electricity consumption behaviors at different times.

[0142] Integrate the main preferences and secondary preferences of all communities to form a set of typical electricity consumption behavior patterns that reflect the behavior preferences and habits of the entire user group; for each typical pattern, count its support and association rules in different communities and the entire group to characterize the universality and relevance of the pattern; cluster and classify the set of typical patterns to identify different types of behavior preferences and habits, such as energy-saving type, peak type, stable type, etc.; the set of typical electricity consumption behavior patterns provides a multi-level and multi-granularity analysis perspective for characterizing the group's electricity consumption behavior, and can comprehensively reveal the behavior characteristics and internal laws of the group.

[0143] For each typical electricity consumption behavior pattern, extract its load curve samples at different times and in different environments, and draw a family of load curves; for the constraint rules and supplementary attributes of each typical pattern, generate an association network diagram to show the association structure of the typical pattern; for the main preferences and secondary preferences of the group, draw a bar chart of support degree rankings and a pie chart of community distributions to visually show the universality and distribution characteristics of the preferences; the visual display of typical patterns can concretize abstract behavior patterns into intuitive graphic images, facilitating business personnel to understand and apply the analysis results.

[0144] Through the above steps, community discovery can be combined with frequent pattern mining and association rule generation to extract typical electricity consumption behavior patterns from the frequent co-occurrence behaviors within the community, and use significant association rules to constrain and supplement the typical patterns. Finally, a set of typical electricity consumption behavior patterns reflecting the group behavior preferences and habits can be obtained. This comprehensive analysis method can fully exploit the behavior pattern information contained in the community structure, improve the representativeness and richness of the typical patterns, and provide accurate user insights for power companies to carry out targeted marketing and personalized services. At the same time, through the clustering classification and visual display of the typical patterns, the interpretability and operability of the analysis results can be enhanced, and the effective transformation of data mining results into business applications can be promoted.

[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for processing electricity consumption data for multiple user types, characterized in that: include: A heterogeneous power graph of the target area is constructed based on multiple node types, node attributes, and multiple edge types and edge attributes. The node types include user nodes, transformer nodes, and power supply area nodes. The node attributes include user type, power consumption, and load curve. The edge types include user-transformer connection edges and transformer-power supply area connection edges. The edge attributes include power transmission capacity and distance. Applying heterogeneous graph convolution on the heterogeneous power graph, introducing the time dimension on the basis of the heterogeneous graph convolution, constructing a spatiotemporal heterogeneous graph convolution model, regarding each time step as a graph snapshot, constructing a heterogeneous power graph corresponding to the time step, applying the spatiotemporal heterogeneous graph convolution model on the heterogeneous power graph at each time step, applying one-dimensional convolution on the time dimension while applying graph convolution on the heterogeneous power graphs at different time steps, extracting multi-scale spatiotemporal dependency patterns, and obtaining spatiotemporal heterogeneous features corresponding to the heterogeneous power graph; The fuzzy C-means clustering algorithm is used to group users based on the spatiotemporal heterogeneous features. The membership matrix of users to each cluster is obtained by optimizing the objective function. The optimal number of clusters is determined according to the clustering effectiveness index to obtain the user grouping results. Association rule mining is applied to the internal data of each user group. Based on the association rule mining results, a community discovery algorithm is used to identify the behavior clusters within the user group to obtain a typical electricity consumption behavior pattern. Initialize the membership matrix, and randomly assign an initial membership value to each data point so that the sum of the memberships is 1, wherein the spatiotemporal heterogeneous features corresponding to the heterogeneous power graph include the spatiotemporal heterogeneous feature vectors of each node in the heterogeneous power graph, and each data point corresponds to the spatiotemporal heterogeneous feature vector of a user node; According to the current membership matrix, the center vector of each cluster is calculated, the center vector is the weighted average of the membership of all data points, and the cluster center is updated according to the center vector; The membership matrix is ​​updated based on the distance between the data point and the cluster center, and the inverse relationship between the membership value and the distance is achieved by optimizing and minimizing the objective function; Among them, the objective function is the sum of the squares of the distances between the data points and their membership-weighted cluster centers; The cluster center calculation and membership matrix update are iteratively performed until the change of the objective function is less than the preset threshold or the maximum number of iterations is reached, and the optimized membership matrix is ​​output to achieve the soft assignment of users to each cluster.

2. The method according to claim 1, characterized in that The spatiotemporal heterogeneous features corresponding to the heterogeneous power graph include: ; in, Indicates l+1 Layer network in t The spatiotemporal heterogeneous features of the time, σ represents the activation function, P.K They represent the size of the convolution kernel in the time dimension and the size of the convolution kernel of the heterogeneous graph convolution respectively. represents the spatiotemporal convolution kernel parameters, express t The adjacency matrix of the heterogeneous power graph at time, represents the heterogeneous graph convolution operation, Indicates l The spatiotemporal heterogeneous characteristics of the layer network at time t.

3. The method according to claim 1, characterized in that The optimal number of clusters is determined based on the clustering effectiveness index, and the user grouping results include: Setting a search range for candidate cluster numbers, performing fuzzy C-means clustering on each candidate cluster number, and obtaining a corresponding membership matrix and cluster center, wherein the candidate cluster number is randomly set and the value is a cluster set between the minimum cluster number and the maximum cluster number; Calculate the clustering effectiveness index of each candidate cluster number, wherein the clustering effectiveness index includes partition coefficient, partition entropy and Xie-Beni index; Among them, the partition coefficient represents the mean of the sum of squares of elements in the membership matrix; the partition entropy represents the mean of the elements in the membership matrix multiplied by their logarithms; the Xie-Beni index is the ratio of cluster compactness to cluster separation, where cluster compactness represents the sum of squares of distances between data points and their membership-weighted cluster centers, and cluster separation represents the square of the minimum distance between different cluster centers; Select the candidate cluster number with the largest partition coefficient, the smallest partition entropy, or the smallest Xie-Beni index as the optimal cluster number; The membership matrix and cluster centers under the optimal number of clusters are output, and the belonging label of each data point is determined according to the membership matrix under the optimal number of clusters to obtain the user grouping result.

4. The method according to claim 1, characterized in that: Association rule mining is applied to the internal data of each user group. Based on the association rule mining results, the community discovery algorithm is used to identify the behavior clusters within the user group. The typical electricity consumption behavior patterns include: Perform Z-score normalization on the power load curve of each user in the user group, divide the normalized sequence into several intervals, calculate the mean of each interval, and map the mean of each interval to a corresponding symbol according to a predefined symbol alphabet to obtain a power load symbol sequence; Apply the FP-growth algorithm to mine frequent patterns of electricity consumption behavior, calculate the frequency of occurrence of each power load symbol by scanning the internal data set of the user group and arrange them in descending order of frequency, create an FP-tree to insert the symbol sequence in descending order of frequency, recursively construct the conditional pattern base and the conditional FP-tree to generate frequent patterns, wherein the internal data set of the user group includes the power load symbol sequence of all users; Based on the mined frequent patterns, electricity consumption behavior association rules are generated, frequent item sets are filtered from the frequent pattern set according to the minimum support threshold, all non-empty subsets of each frequent item set are generated, the confidence of the association rules between the non-empty subsets and the frequent item sets is calculated, and the association rules that meet the requirements are output according to the minimum confidence threshold. The significance value of the association rules is evaluated using the Lift indicator, and the electricity consumption behavior corresponding to the association rules whose significance value is greater than the preset significance threshold is defined as a typical electricity consumption behavior pattern, wherein the frequent pattern set includes the frequent pattern of each electricity load symbol.

5. The method according to claim 4, characterized in that The method further comprises: Based on frequent patterns and association rules, the community discovery algorithm is used to identify behavioral clusters associated within user groups. The modularity is dynamically adjusted by optimizing the modularity objective function until the community attributes of all nodes no longer change, thus achieving community division. The process of dynamically adjusting the modularity is as follows: ; in, △Q represents the change in modularity, represents the weight and the weight of the edges within the community, k i,in Representation Node i The sum of the weights of the edges connecting to other nodes in the community, m represents the fuzzification parameter, represents the sum of the weights of the edges connecting nodes inside and outside the community, k i Representation Node i The degree.

6. The method according to claim 5, characterized in that The method further comprises: Count the frequent pattern sets within each community, and define the frequent patterns whose occurrence frequency exceeds the preset threshold as the candidate typical electricity consumption behavior patterns of the community; Calculate the similarity between candidate typical patterns, merge similar patterns by setting a similarity threshold, and obtain a typical electricity consumption behavior pattern set after deduplication; For each typical electricity consumption behavior pattern, traverse the association rule set in the community, filter out high-confidence association rules with the pattern as the antecedent, and use the rule consequence as the supplementary attribute of the typical pattern; Filter out repeated power consumption behaviors in the supplementary attributes that have been included in the typical power consumption behavior pattern set, and update the typical power consumption behavior pattern set; For each community's set of typical electricity usage behavior patterns, the support of each typical electricity usage behavior pattern is counted, and the top K patterns with the highest support are taken as the community's main electricity usage behavior preferences, and the patterns with support ranging from K to M are taken as the community's secondary electricity usage behavior preferences, forming a set of typical electricity usage behavior patterns that reflect group behavior preferences and habits, where K is the preset main preference threshold and M is the preset secondary preference threshold.

Citation Information

Patent Citations

  • Data processing method and device based on artificial intelligence, equipment and medium

    CN116701706A

  • Data real-time efficient fusion processing method based on FCM clustering algorithm model

    CN117112871A

  • Customer relationship management method and system for electricity marketing

    CN118941298A