A method for profiling electricity consumption behavior based on big data multi-model analysis

By constructing electricity consumption vectors and performing similarity calculations and factor analysis, user categories are precisely segmented, solving the problem that user profiles in existing technologies fail to meet personalized needs, and enabling precise electricity packages and promotional strategies.

CN118469604BActive Publication Date: 2026-04-03GUANGZHOU POWER SUPPLY BUREAU GUANGDONG POWER GRID CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-15
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing methods for profiling electricity consumption behavior fail to fully consider the personalized needs of different user groups, making it impossible for power supply companies to accurately match and differentiate services when formulating electricity packages and promotional strategies.

Method used

By collecting historical load data to construct electricity consumption vectors, cosine similarity calculation and mean clustering are performed. Combined with multi-threshold segmentation and factor analysis, multiple user categories are divided, and separation necessity values ​​are calculated to determine the user's electricity consumption behavior profile.

Benefits of technology

This enables more refined user classification, meets the personalized needs of different user groups, and enhances the accuracy of power supply companies in formulating personalized electricity packages and promotion strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118469604B_ABST
    Figure CN118469604B_ABST
Patent Text Reader

Abstract

This invention relates to the field of data processing, and more specifically, to a method for profiling electricity consumption behavior based on big data multi-model analysis. The method includes: collecting historical load data to construct user electricity consumption vectors; calculating the cosine similarity of any two electricity consumption vectors as a first similarity, performing mean clustering to obtain multiple first categories; calculating the cluster centers of the first categories to obtain first and second electricity consumption vectors, and calculating the cosine similarity of the two electricity consumption vectors to obtain a second similarity, performing multi-threshold segmentation to divide into multiple second categories, and using factor analysis to obtain the information content of each unique factor; calculating the separation necessity value, separation requirement value, and separation value for each second category; and determining the user's electricity consumption behavior profile based on the separation value. This invention finely classifies users, obtains personalized needs for different user profiles, and achieves the goal of accurate recommendations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention generally relates to the field of data processing. More specifically, this invention relates to a method for profiling electricity consumption behavior based on big data multi-model analysis. Background Technology

[0002] Big data multi-modeling typically refers to using multiple different types of data models or processing methods when dealing with large-scale data to better meet diverse data needs. The goal of this approach is to fully leverage the advantages of various data processing tools and technologies to handle different types and properties of data.

[0003] User personas are abstract representations or descriptions of users based on their characteristics, behaviors, and preferences. In fields such as digital marketing, market research, and user experience design, user personas are widely used to understand target user groups in order to better meet their needs.

[0004] User profiling targets a group of people. It summarizes and analyzes the common characteristics of the target user group so that these characteristics can be taken into account when planning related activities and developing marketing strategies. For electricity consumption behavior, existing user profiling often only clusters users' electricity consumption data, resulting in a rather generic portrayal of user electricity consumption behavior and failing to fully consider the personalized needs of different user groups. This may prevent power companies from achieving truly accurate matching and differentiated services when developing electricity packages and promotional strategies. Summary of the Invention

[0005] To address one or more of the aforementioned technical problems, this invention proposes to cluster electricity consumption data while simultaneously updating the cluster categories based on their commonalities and characteristics. This updates the user profiles obtained from different cluster categories, enabling them to meet the personalized needs of different user groups and achieve accurate recommendations. To this end, this invention provides solutions in the following aspects.

[0006] A method for profiling electricity consumption behavior based on big data multi-model analysis includes: collecting historical load data, wherein the historical load data includes current, voltage, and power factor, and constructing user electricity consumption vectors; calculating the cosine similarity between any two electricity consumption vectors as a first similarity, performing mean clustering to obtain multiple first categories; obtaining the cluster center of the first category based on the sum of the cosine similarities of all electricity consumption vectors in the first category, and labeling it as a first electricity consumption vector; labeling all electricity consumption vectors in each first category other than the first electricity consumption vector as second electricity consumption vectors, calculating the cosine similarity between each second electricity consumption vector and the first electricity consumption vector to obtain a second similarity; performing multi-threshold segmentation based on the second similarity to obtain multiple second categories; applying factor analysis to the second categories to obtain the information content of each unique factor, and calculating the separation necessity value of each second category based on the information content of the unique factor; calculating the separation necessity value of the second category based on the separation necessity value; dividing multiple user categories based on the separation value, and determining the user's electricity consumption behavior profile based on the user categories.

[0007] In one embodiment, multi-threshold segmentation is performed based on the second similarity to obtain multiple second categories, including:

[0008] A histogram is constructed based on the second similarity, and the peak value in the histogram is counted. The second similarity is then segmented using multiple thresholds based on the peak value to obtain multiple second categories. Each second category corresponds to a peak value, and the second category with the largest peak value is marked as the target category.

[0009] By adopting the above technical solution, multiple second categories are divided by calculating the similarity of electricity consumption of the cluster centers of the first category, thereby further refining the user classification and meeting the personalized needs of users.

[0010] In one embodiment, factor analysis is used on the second category to obtain the information content of each of the individual factors, including:

[0011] Construct an electricity consumption matrix for each electricity consumption vector in the second category, and apply factor analysis to the electricity consumption matrix to obtain the common factor vector of the electricity consumption matrix and the unique factor vector of each electricity consumption vector, wherein the common factor and unique factor of the target category are denoted as the target category common factor and the target category unique factor, respectively.

[0012] Calculate the cosine similarity between any two common factors, mark the common factors corresponding to the largest cosine similarity as a common factor pair, mark the two electricity consumption matrices corresponding to the common factor pair as an electricity consumption matrix pair, and perform factor analysis on the electricity consumption matrix pair to obtain the unique factors of the electricity consumption matrix pair.

[0013] Based on the single factor and common factor pair, calculate the cosine similarity of any common factor in the single factor and common factor pair respectively, and perform normalization processing to obtain the information content of each single factor;

[0014] Calculate the cosine similarity between the common factors of the second category and the common factors of the target category to obtain the separation requirement value.

[0015] By adopting the above technical solution, each histogram of second similarity corresponds to a peak, which is denoted as the target category. This can best represent the category. The smaller the similarity between the unique factor of each second category and the unique factor of the target category, the greater the information content of the unique factor in the second category.

[0016] In one embodiment, the separation necessity value satisfies the following relationship:

[0017] g = k * (1 - s)

[0018] Where g represents the separation necessity value of the second category, k represents the information content of the unique factor of the second category, and s represents the similarity between the unique factor corresponding to the second category and the unique factor of the target category.

[0019] By adopting the above technical solution, the information content of a single factor can be represented by the proportion of the single factor. The greater the information content of a single factor, the smaller the similarity between the single factor and the single factor of the target category, and the greater the necessity for the second category to be separated from the first category.

[0020] In one embodiment, the separability value satisfies the following relationship:

[0021]

[0022] Where p represents the separability value of the second category, f represents the separability requirement value, and g represents the separability necessity value of the second category.

[0023] By adopting the above technical solution, the greater the separation necessity value, the more necessary it is to separate the second category and make it an independent category, thereby enhancing the personalization of the user group.

[0024] In one embodiment, multiple user categories are divided based on dissociation values, and a user's electricity consumption behavior profile is determined based on the user categories, including:

[0025] If the separation value of the second category is less than a preset threshold, the user is classified into the first category, and a user's electricity consumption behavior profile is determined based on the user category.

[0026] The present invention has the following effects:

[0027] 1. This invention uses similarity clustering of users' electricity consumption vectors to perform more refined and personalized classification of users in the original classification results, resulting in more personalized classification results. By calculating the common features and personalized features of user categories, a more refined user classification is completed, so that the resulting user profile can meet the personalized needs of different user groups and achieve the purpose of accurate recommendation.

[0028] 2. This invention categorizes users into multiple user classes by dividing them into electricity consumption vectors, thereby enhancing the personalization of user categories. This helps power supply companies develop personalized electricity packages and promotional strategies, truly achieving precise matching and differentiated services. Attached Figure Description

[0029] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of the invention are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein:

[0030] Figure 1 This is a flowchart of steps S1-S8 in a method for profiling electricity consumption behavior based on big data multi-model analysis according to an embodiment of the present invention.

[0031] Figure 2 This is a flowchart of steps S60-S64 in a method for profiling electricity consumption behavior based on big data multi-model analysis according to an embodiment of the present invention. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0034] Reference Figure 1 A method for profiling electricity consumption behavior based on big data multi-model analysis includes steps S1-S8, as follows:

[0035] S1: Collect historical load data, including current, voltage, and power factor, to construct the user's electricity consumption vector.

[0036] For example, a user's electricity consumption vector is constructed based on load data. The vector contains data such as current, voltage, and power factor, and is used to classify users according to their electricity consumption vectors.

[0037] S2: Calculate the cosine similarity of any two power vectors as the first similarity, perform mean clustering, and obtain multiple first categories.

[0038] For example, using (1-cosine similarity) as the distance parameter in clustering, multiple first categories are obtained through the k-means clustering algorithm. The similarity of power consumption vectors in the same first category is relatively high, while the similarity of power consumption vectors in different first categories is relatively low.

[0039] S3: Based on the sum of the cosine similarities of all electricity consumption vectors in the first category, obtain the cluster center of the first category and mark it as the first electricity consumption vector.

[0040] For example, the electricity consumption vector corresponding to the maximum sum value is used as the cluster center of the first category, thus obtaining the cluster center of each first category.

[0041] S4: Label all power consumption vectors in each first category except for the first power consumption vector as second power consumption vectors, calculate the cosine similarity between each second power consumption vector and the first power consumption vector, and obtain the second similarity.

[0042] For example, there are multiple second similarities in each second category.

[0043] S5: Perform multi-threshold segmentation based on the second similarity to obtain multiple second categories.

[0044] A histogram is constructed based on the second similarity. The peak value in the histogram is counted. The second similarity is segmented using multiple thresholds based on the peak value to obtain multiple second categories. Each second category corresponds to a peak value. The second category with the largest peak value is marked as the target category. Among them, the peak value in the histogram is the second category with the most identical second similarity values.

[0045] For example, by calculating the similarity between the second electricity consumption vector in each second category and the electricity consumption vector of the cluster center of the corresponding first category, multiple second categories are formed. The electricity consumption vector in a second category has the highest similarity to the electricity consumption vector of the cluster center of the first category to which the second category belongs (this category refers to the category in the initial classification result), and can best represent that category.

[0046] The smaller the difference between each secondary category and the target category, the greater the separation degree required for the secondary category to separate from its parent primary category and become an independent category. In other words, the greater the separation degree of the secondary category, the greater the personalization of the user vector of the secondary category, and the greater the separation degree.

[0047] Further explanation: In the first category A, there are two second categories a1 and a2. The similarity between a1 and the target category is 0.9, and the similarity between a2 and the target category is 0.8. Since the separation degree of a1 is greater than that of a2, a1 should be separated from the first category and become an independent second category.

[0048] The smaller the similarity between the unique factor in each second category and the unique factor in the target category, the greater the information content of the unique factor in that second category, and the greater the necessity for separation. This is used to enhance the personalization of user groups and facilitate subsequent accurate recommendations.

[0049] S6: Use factor analysis on the second category to obtain the information content of each individual factor. Based on the information content of the individual factors, calculate the separation necessity value for each second category, referring to... Figure 2 This includes steps S60-S64:

[0050] S60: Construct the electricity consumption matrix for each electricity consumption vector in the second category, and use factor analysis on the electricity consumption matrix to obtain the common factor vector of the electricity consumption matrix and the unique factor vector of each electricity consumption vector. The common factor and unique factor of the target category are denoted as the target category common factor and the target category unique factor, respectively.

[0051] For example, each electricity consumption vector in the second category is taken as a row of a matrix to obtain an electricity consumption matrix. Through factor analysis, we find that there is only one common factor vector, which represents the common feature of the second category to which the row vectors in the electricity consumption matrix belong. The unique factor vector corresponds to one electricity consumption vector and represents the independent feature of the electricity consumption vector, that is, the feature that is different from other electricity consumption vectors.

[0052] Further explanation: The power and voltage changes of a certain electricity consumption data point differ significantly from the power and voltage changes in other electricity consumption vectors. This feature is highly specific and therefore serves as an independent feature of that electricity consumption vector.

[0053] Factor analysis is a method for analyzing multiple input data to obtain the common features of all input data and the unique features of each input data. The mathematical model of factor analysis is X = A * F + ε, where X has a size of [M,1], A has a size of [M,n], F has a size of [N,1], and ε has a size of [M,1]. Inputting multiple [M,1] data points results in an [M,n] matrix, which is the factor loading matrix. F is a common factor vector with only one value, representing the common features of all input data. ε is a unique feature vector, with each input data point corresponding to a unique feature vector.

[0054] S61: Calculate the cosine similarity between any two common factors, mark the common factors corresponding to the largest cosine similarity as a common factor pair, mark the two electricity consumption matrices corresponding to the common factor pair as electricity consumption matrix pairs, perform factor analysis on the electricity consumption matrix pairs to obtain the unique factors of the electricity consumption matrix pairs.

[0055] For example, taking any two common factors a and b, the matrix corresponding to a is denoted as a0, and the matrix corresponding to b is denoted as b0. By superimposing matrices a0 and b0, a composite matrix is ​​obtained, which is the power matrix pair, denoted as ab. Further, for example, if the size of matrix a0 is 20*40 and the size of matrix b0 is 30*40, then the size of the resulting matrix ab is 50*40. Through factor analysis, the unique factor c of the composite matrix ab is obtained.

[0056] The greater the similarity between the independent factors in a composite matrix and the independent factors in the first two matrices, the more independent features that matrix contains. Therefore, these features are more readily reflected in the composite matrix, preserving more independent feature information. In other words, the independent factors of that matrix have a higher similarity to the independent factors of the composite matrix.

[0057] S62: Based on the single factor and common factor pairs, calculate the cosine similarity of any common factor in the single factor and common factor pairs respectively, and perform normalization to obtain the information content of each single factor;

[0058] For example, calculate the cosine similarity between the unique factor c and the common factor a of the composite matrix, denoted as ca, and calculate the cosine similarity between the unique factor c and the common factor b of the composite matrix, denoted as cb. Normalize ca, and obtain the normalized ca value by normalizing the sum of ca and cb. Similarly, the normalized cb value can be obtained. Use the normalized ca value as the information content of the unique factor in matrix a0, and use the normalized cb value as the information content in matrix b0. Since the common factors are similar, the proportion of unique factors can represent the information content of unique factors. The greater the information content of a unique factor, the smaller the similarity between the unique factor and the unique factor of the target category, and the greater the necessity for separation.

[0059] S63: Calculate the separation necessity value for each second category to satisfy the following relationship:

[0060] g = k * (1 - s)

[0061] Where g represents the separation necessity value of the second category, k represents the information content of the unique factor of the second category, and s represents the similarity between the unique factor corresponding to the second category and the unique factor of the target category;

[0062] For example, the greater the information content of a single factor, the smaller the similarity between the single factor and the single factor of the target category, and the greater the necessity for the second category to be separated from the first category.

[0063] S64: Calculate the cosine similarity between the common factors of the second category and the common factors of the target category to obtain the separation requirement value.

[0064] For example, the greater the similarity, the closer the category is to the standard common factor, and the less likely it is to be separated. Therefore, the separation requirement is greater. Here, similarity is used as the separation requirement.

[0065] S7: Calculate the separation value for the second category based on the separation necessity value and the separation requirement value.

[0066] The separability values ​​satisfy the following relationship:

[0067]

[0068] Where p represents the separability value of the second category, f represents the separability requirement value, and g represents the separability necessity value of the second category.

[0069] For example, the greater the necessity for separation, the less demand for separation, and the greater the separation of the second category, that is, the more it needs to be separated to increase the diversity of the user group.

[0070] S8: Divide users into multiple categories based on their separation values, and determine the user's electricity consumption behavior profile based on the user category.

[0071] If the separation value of the second category is greater than the preset threshold, it is classified into the first category, and electricity packages and promotion strategies are formulated based on the user category.

[0072] For example, with a preset threshold of 0.7, the separability of each second category in each first category can be calculated. Second categories with a separability greater than 0.7 are classified as first categories, resulting in two user categories. Electricity packages and promotion strategies can then be formulated based on the user categories.

[0073] In the description of this specification, "multiple" or "several" means at least two, such as two, three or more, unless otherwise explicitly specified.

[0074] While this specification has shown and described numerous embodiments of the invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of this invention.

Claims

1. A method for profiling electricity consumption behavior based on big data multi-model analysis, characterized in that, include: Collect historical load data, including current, voltage, and power factor, to construct the user's electricity consumption vector; Calculate the cosine similarity between any two electricity consumption vectors as the first similarity, perform mean clustering, and obtain multiple first categories; The cluster center of the first category is obtained based on the sum of the cosine similarities of all electricity consumption vectors in the first category, and is marked as the first electricity consumption vector. All power consumption vectors in each of the first categories, excluding the first power consumption vector, are labeled as second power consumption vectors. The cosine similarity between each second power consumption vector and the first power consumption vector is calculated to obtain the second similarity. Multiple threshold segmentation is performed based on the second similarity to obtain multiple second categories; Factor analysis was used on the second category to obtain the information content of each individual factor, and the separation necessity value of each second category was calculated based on the information content of the individual factors. Based on the separation necessity value, calculate the separation value for the second category; Multiple user categories are divided based on the separability value, and user electricity consumption behavior profiles are determined based on these categories. Factor analysis is applied to the second category to obtain the information content of each unique factor, including: constructing an electricity consumption matrix for each electricity consumption vector in the second category; applying factor analysis to the electricity consumption matrix to obtain the common factor vector and the unique factor vector for each electricity consumption vector, wherein the common factor and unique factor of the target category are respectively denoted as the target category common factor and the target category unique factor; calculating the cosine similarity between any two common factors; marking the common factor corresponding to the largest cosine similarity as a common factor pair; marking the two electricity consumption matrices corresponding to the common factor pair as an electricity consumption matrix pair; applying factor analysis to the electricity consumption matrix pair to obtain the unique factors of the electricity consumption matrix pair; calculating the cosine similarity of any one common factor in the unique factor and the common factor pair, and performing normalization processing to obtain the information content of each unique factor; calculating the cosine similarity between the common factor of the second category and the common factor of the target category to obtain the separability value. The separation necessity value satisfies the following relationship: g = k * (1 - s); where g represents the separation necessity value of the second category, k represents the information content of the unique factor of the second category, and s represents the similarity between the unique factor corresponding to the second category and the unique factor of the target category. The separation value satisfies the following relationship: Where p represents the separability value of the second category, f represents the separability requirement value, and g represents the separability necessity value of the second category.

2. The method for profiling electricity consumption behavior based on big data multi-model analysis according to claim 1, characterized in that, Based on the second similarity, multi-threshold segmentation is performed to obtain multiple second categories, including: A histogram is constructed based on the second similarity, and the peak value in the histogram is counted. The second similarity is then segmented using multiple thresholds based on the peak value to obtain multiple second categories. Each second category corresponds to a peak value, and the second category with the largest peak value is marked as the target category.

3. The method for profiling electricity consumption behavior based on big data multi-model analysis according to claim 1, characterized in that, Users are categorized into multiple user groups based on their separation scores, and their electricity consumption behavior profiles are determined based on these user groups, including: If the separation value of the second category is less than a preset threshold, the user is classified into the first category, and a user's electricity consumption behavior profile is determined based on the user category.

Citation Information

Patent Citations

  • Power customer portraying method based on intersection-parallel ratio density clustering

    CN113935410A

  • Method and device for establishing power consumer portrait, and electronic equipment

    CN115130811A