Predictive segmentation of energy consumers

By generating consumer segments through machine learning and decision tree technology, the effectiveness of energy consumer segmentation in existing technologies is insufficient, improving the efficiency of marketing and communication and enabling the identification of consumers with high engagement potential.

CN115082102BActive Publication Date: 2026-02-13SIRUI ARTIFICIAL INTELLIGENCE CO
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210473492.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2015-12-18
Filing Date
2016-12-15
Publication Date
2026-02-13
Estimated Expiration
2036-12-15

AI Technical Summary

Technical Problem

Existing market segmentation techniques lack effectiveness in energy consumer analysis, failing to accurately identify consumer groups with high engagement potential, leading to inefficient marketing and communication strategies.

Method used

By employing machine learning techniques, combined with consumer data and demographics, an initial set of patterns is generated through decision trees and pattern recognition algorithms. After refinement, the patterns are assigned to the optimal consumer segments to ensure the effectiveness and predictive power of the segments.

Benefits of technology

It achieved twice the average level of consumer segmentation engagement, and improved the effectiveness of marketing and communication by describing consumer behavior through a small number of intuitive rules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115082102B_ABST
    Figure CN115082102B_ABST
Patent Text Reader

Abstract

A computer system receives consumer records listing consumer attributes and consumer adoption status, such as whether the consumer has participated in a particular energy efficiency program. An initial set of patterns is identified in the consumer records, such as according to a decision tree. The initial set is pruned to obtain a set of patterns that satisfy minimum support and validity, and maximum overlap requirements. The patterns are assigned to segments according to an optimization algorithm that seeks to maximize the minimum validity of each segment, where validity indicates the number of consumers with a positive adoption status that conform to the patterns of each segment. The optimization algorithm can be a bisection algorithm that evaluates a linear fractional integer program (LFIP-F) to iteratively approximate an optimal distribution of patterns.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional of Chinese Patent Application No. 201680074515.3, filed December 15, 2016, entitled "Predictive Segmentation of Energy Consumers," which corresponds to PCT Application No. PCT / US2016 / 067002, filed December 15, 2016, both of which are incorporated by reference herein in their entirety. TECHNICAL FIELD

[0002] This application claims the benefit of U.S. Provisional Application Serial No. 62 / 269,793, filed December 18, 2015, entitled "Predictive Segmentation of Energy Consumers," the entire contents of which are incorporated herein by reference. TECHNICAL FIELD

[0004] The present invention relates to a computer algorithm for analyzing energy consumers. BACKGROUND

[0005] In recent years, energy utility companies have become increasingly interested in improving their relationships with consumer groups that have traditionally been disconnected from the power suppliers. In the past, both energy utility companies and their consumers understood the role of the power company to be "to keep the lights on." However, current technological trends and shifts in consumer attitudes, driven in particular by the growth of consumer-facing internet companies that are good at understanding and anticipating the preferences of their consumers, have led utilities to increase their interest in engaging with their consumers.

[0006] The combination of these trends has brought an increase in data availability (high-granularity consumption data collected through sensing infrastructure such as smart meters, as well as other "meta-data" about the consumers themselves) and computational methods for processing such data (e.g., Li and Yang (2015), Liu and Nielsen (2015)). As a result, energy utility companies are increasingly relying on analytics techniques that can provide them with ways to increase consumer satisfaction and engagement, as well as participation in environmentally friendly programs within their consumer base. Consumer segmentation is a technique for understanding consumers and identifying ways to act on that understanding that is a cornerstone of the marketing toolkit for both large and small organizations. It is widely used in marketing (for a comprehensive review, see Association (2014)), online advertising (e.g., Yan (2009) et al.) or e-tail (e.g., Bhatnagar and Ghose (2004)), to name a few applications.

[0007] As utilities strive to build more personalized and modern relationships with their customers, they are keen to use segmentation as a means to tailor their communications about efficiency measures and other programs to improve engagement and outreach. Most market segmentation techniques used in practice focus on the application of a fixed set of rules. For example, customers living in large families and with children are assigned to a "high consumption" category, while those who subscribe to an environmentalist magazine are placed in a "green advocate" group. Often, these rules come from anecdotal or hearsay experience, behavioral research, or small-scale psychological experiments and are treated as "common knowledge" in practice. As a crystallization of aggregated domain knowledge, such segmentation strategies are certainly valuable and should inform theory and practice.

[0008] The methods described herein provide improved methods for segmenting energy consumers. BRIEF DESCRIPTION OF DRAWINGS

[0009] In order that the advantages of the application will be readily understood, a more particular description of the application briefly described above will be rendered by reference to specific embodiments illustrated in the appended drawings. Understanding that these drawings depict only typical embodiments of the application and are not therefore to be considered to be limiting of its scope, the application will be described and explained with additional specificity and detail by reference to the accompanying drawings, in which:

[0010] Figure 1 is a schematic block diagram of components for implementing predictive segmentation of consumers according to embodiments of the application;

[0011] Figure 2 is a schematic block diagram of a computing device;

[0012] Figure 3 is a process flow diagram of a method for performing predictive segmentation according to embodiments of the application;

[0013] Figure 4 is a schematic diagram illustrating a decision tree extracted from consumer data;

[0014] Figure 5 is a plot of pattern validity versus number of rules;

[0015] Figure 6 is a chart showing predictive variables for explaining engagement of energy consumers;

[0016] Figure 7 is a feasibility matrix for patterns assigned to consumer segments;

[0017] Figure 8 is a plot showing distribution of inter-pattern overlap;

[0018] Figure 9 is an example of patterns associated with two segments;

[0019] Figure 10 is a plot of lower and upper bounds on segment effectiveness as a function of number of iterations;

[0020] Figure 11 illustrates pattern-to-segment assignment matrix;

[0021] Figure 12 illustrates overlap of segments;

[0022] Figure 13 is another schematic plot illustrating subdivision overlap;

[0023] Figure 14 is a list of subdivisions and corresponding patterns according to a segmentation algorithm;

[0024] Figure 15 is a plot showing sensitivity analysis of segmentation algorithm for and π ;

[0025] Figure 16 is a plot of segment effectiveness as a function of ;

[0026] Figure 17 is a scatter plot of segment effectiveness as a function of subdivision complexity; DETAILED DESCRIPTION

[0027] The present application introduces a predictive segmentation technique for identifying subgroups in a large population that are homogeneous with respect to certain patterns in consumer attributes and predictive with respect to desired outcomes. The incentive setting is creating a highly interpretable and intuitive segmentation and targeting process for the energy utility's customers that is also, in some sense, optimal. In this setting, the energy utility wants to design a small number of message types to send to a properly selected set of customers that are most likely to respond to these different types of messages. The approach presented here uses standard machine learning techniques to extract basic predictive patterns using consumption, demographic, and program participation data. The approach next defines feasible potential ways of assigning patterns to a small number of segmentation types that are described by expert guidelines and assumptions about customer characteristics that can be derived from current behavioral research. The algorithm then identifies the optimal way of assigning patterns to segmentation types that is both feasible and maximizes predictive power. The approach is implemented on a large-scale data set from a large U.S. energy company and obtains a customer segmentation whose likelihood of participation is more than twice as high as the average population and can be described by a small number of simple and intuitive rules.

[0028] 1. Operating Environment and Overview

[0029] refer to Figure 1 The methods disclosed herein can be implemented using the illustrated operating environment 100. A server system 102 or other type of computer system can host or access the database 104. The server system 102 can also be replaced by a desktop computer, laptop computer, or even a mobile device with sufficient computing power. The database 104 may include consumer records 106 for multiple consumers. The methods disclosed herein are described in relation to energy consumers. Therefore, each consumer record 106 may include data from a single household or consumer account, and thus may include data from multiple cohabiting individuals.

[0030] Consumer record 106 may include information such as identifiers 108a of one or more consumers in the form of names, account numbers, or other unique identifiers. Consumer record 106 may include the consumer's address 108b and demographic information 108c (such as age, income, gender, occupation, education level, and any other information that may characterize the consumer) of one or more individuals associated with consumer record 106.

[0031] When the methods disclosed herein are applied to energy consumers, consumer record 106 may further include usage data 108d, such as the number of kilowatt-hours used annually, monthly, or daily. Usage data 108d may include daily, monthly, or quarterly usage patterns obtained from the analysis of energy consumption data. In other applications, usage data 108d may include the use of another service or the purchase of a specific item or supply.

[0032] Consumer records 106 may include any other data 108e that is available to the consumer and that may help identify patterns describing consumer types and consumer behavior.

[0033] The methods disclosed herein are used to analyze data to identify patterns in consumer data (demographics, usage, and others) that predict consumers will take specific actions. In the case of energy consumers, this might include participating in energy efficiency programs or taking other actions to reduce consumption or otherwise reduce the consumer's environmental impact. Therefore, consumer record 106 may further include one or more adoption states 108f indicating whether the consumer has opted to participate in a particular program. For example, adoption state 108f could be 1 if the consumer opts to participate, and 0 otherwise. In other embodiments, adoption state could be one of a series of values ​​indicating the degree of conformity with program guidelines or the amount spent on a particular objective.

[0034] The database 104 can further store segments 110 having a plurality of patterns 112 assigned thereto. Each segment 110 has an effectiveness 114 that is a measure of the number of consumer records that have a positive adoption status and match one of the patterns 112 assigned to the segment 110.

[0035] The segments 110 can be defined by an analysis module 116 that implements the methods disclosed below. In particular, the analysis module 116 can include a pattern generation module 118a. The pattern generation module 118a identifies a series of attributes that co-occur in consumer records. As described below, a pattern can be described in terms of thresholds for the values of various attributes in each consumer record. As also described below, patterns can be generated using decision trees or other pattern recognition algorithms.

[0036] The analysis module 116 can further include a pattern pruning module 118b. As described below, the pattern pruning module 118b can prune patterns that do not meet minimum support, effectiveness, or non-overlapping criteria. The analysis module can include a segmentation module 118c. The segmentation module 118c assigns patterns 112 that remain effective after passing through the pattern pruning module 118b to segments 110 such that a set of segments 110 is obtained in which the minimum effectiveness 114 of the segments has been increased by an algorithm that assigns patterns among the segments.

[0037] Figure 2 is a block diagram illustrating an example computing device 200. The computing device 200 can be used to perform various procedures, such as those discussed herein. The server system 102 can have some or all of the properties of the computing device 200.

[0038] The computing device 200 includes one or more processors 202, one or more memory devices 204, one or more interfaces 206, one or more mass storage devices 208, input / output (I / O) device(s) 210, and a display device 230, all of which are coupled to a bus 212. The processor(s) 202 include one or more processors or controllers that execute instructions stored in the memory device(s) 204 and / or mass storage device(s) 208. The processor(s) 202 can also include various types of computer-readable media, such as cache memory.

[0039] The memory device(s) 204 include various computer-readable media, such as volatile memory (e.g., random access memory (RAM) 214) and / or non-volatile memory (e.g., read-only memory (ROM) 216). The memory device(s) 204 can also include rewritable ROM, such as flash memory.

[0040] The mass storage device(s) 208 include various computer-readable media, such as magnetic tapes, magnetic disks, optical disks, solid-state memory (e.g., Flash memory), and so forth. As indicated, a specific mass storage device is a hard disk drive 224. Various drives can also be included in the mass storage device(s) 208 to enable reading from and / or writing to various computer- readable media. The mass storage device(s) 208 include removable media 226 and / or non-removable media. Figure 2

[0041] The input / output (I / O) device(s) 210 include various devices that allow data and / or other information to be input to or retrieved from the computing device 200. Example input / output (I / O) device(s) 210 include cursor control devices, keyboards, keypads, microphones, monitors or other display devices, speakers, network interface cards, modems, lenses, CCDs or other image capture devices, and the like.

[0042] The display device 230 includes any type of device capable of displaying information to one or more users of the computing device 200. Examples of display devices 230 include monitors, display terminals, video projection devices, and the like.

[0043] The interface(s) 206 include various interfaces that allow the computing device 200 to interact with various systems, devices, or computing environments. Example interface(s) 206 include any number of different network interfaces 220, such as interfaces to local area networks (LANs), wide area networks (WANs), wireless networks, and the Internet. Other interfaces include a user interface 218 and a peripheral device interface 222. The interface(s) 206 can also include one or more peripheral interfaces, such as interfaces for pointing devices (mice, track pads, etc.), keyboards, and the like.

[0044] ​Bus 212 allows processor(s) 202, memory device(s) 204, interface(s) 206, mass storage device(s) 208, input / output (I / O) device(s) 210, and display device(s) 230 to communicate with one another, and with other devices or components coupled to bus 212. Bus 212 represents one or more of several types of bus structures, such as a system bus, a PCI bus, an IEEE 1394 bus, a USB bus, and the like.

[0045] For illustration purposes, programs and other executable program components are shown herein as discrete blocks, although it is understood that such programs and components can reside at various times in different storage components of computing device 200, and are executed by processor(s) 202. Alternatively, the systems and procedures described herein can be implemented in hardware, or a combination of hardware, software, and / or firmware. For example, one or more application specific integrated circuits (ASICs) can be programmed to carry out one or more of the systems and procedures described herein.

[0046] With reference to Figure 3 Server system 102 can perform the illustrated method 300. Method 300 can include receiving 302 consumer data. This can include receiving data over a period of time as data about consumers is collected. The received data can include some or all of the data described above as included in consumer records 106.

[0047] Method 300 can further include determining 304 consumer adoption status. Adoption status 108f can be included in the received consumer records, or the adoption status 108f can be received as part of a subsequent procedure of providing offers to consumers and receiving responses. In either case, data is provided to server system 102, either manually or automatically, that indicates the adoption status of each consumer. In some embodiments, method 300 can be performed only for consumers to whom an offer was received.

[0048] Method 300 can further include generating 306 an initial pattern set. For example, generating 306 an initial pattern set can include traversing a decision tree as is known in the art, where each node of the decision tree is an attribute value or attribute value range corresponding to attributes 108b-108e of consumer records 106. Figure 5 An example decision tree is illustrated, and the generation of an initial pattern set is described in more detail in Section 3.2, “Extracting Predictive Patterns from Data,” and Section 5.2, “Predictive Patterns Extracted from Data,” below.

[0049] The method 300 can further include trimming 308 the initial set of patterns. This can include removing patterns that do not have sufficient support, e.g., a number of consumer records 106 that match the pattern is insufficient; patterns that do not have sufficient effectiveness, e.g., a number of consumer records 106 that match the pattern and have a positive adoption status is insufficient; and patterns that have a higher than a threshold percentage of matching consumer records that also match another pattern. A more detailed explanation of the trimming process is described in Sections 3.2, "Extracting Predictive Patterns from Data," and 5.2, "Predictive Patterns Extracted from Data," below.

[0050] The method 300 can further include assigning 310 patterns to segments according to an algorithm that iteratively approaches a maximum of a minimum effectiveness of the segments, where effectiveness is a measure of how many consumer records that match a pattern assigned to each segment have a positive adoption status. This can include executing an optimization algorithm, as described in Section 3.1, "Maximizing Minimum Effectiveness," below.

[0051] The segments can then be further processed 312. In particular, the segments can be used for targeted marketing: advertisements can be made and delivered only to consumers that match a pattern of a segment in order to improve the effectiveness of the advertisements. The segments can also be used to visualize consumer behavior or for any other business goal.

[0052] The algorithms implemented by the method 300 are described in more detail in Sections 2 through 3 below. Section 4 includes a summary of existing methods, and Section 5 describes experimental results using actual consumer data.

[0053] Please note that the following description is with respect to an optimization algorithm that seeks to maximize the minimum effectiveness of the segments. Thus, references to "maximum," "optimal," "optimization," "minimum," and "minimal" should not be understood to refer to an absolute or actual maximum, optimal, or minimum value, but rather a determined maximum, optimal, or minimum value subject to the limitations of the disclosed algorithm and subject to a finite number of iterations of execution of the disclosed algorithm.

[0054] In particular, "maximizing" a value, "maximization" of a value, and "maximum" of a value should be understood to refer to an increase in the value compared to a previous iteration of the disclosed algorithm or without execution of the disclosed algorithm, unless in the case of a closed set of values and a maximum value in the closed set can be positively determined.

[0055] "Minimizing" a value, "minimization" of a value, and "minimum" of a value should be understood to refer to a decrease in the value compared to a previous iteration of the disclosed algorithm or without execution of the disclosed algorithm, unless in the case of a closed set of values and a minimum value in the closed set can be positively determined.

[0056] "Optimization" should be understood to mean finding a value that is closer to the absolute best value than would be found without the disclosed algorithm, and should not be understood to mean actually finding the absolute best value. Similarly, "optimal" values should be understood to be approximate best values, where "approximate" refers to limitations in the accuracy of representing and performing mathematical operations on numbers, limitations that the disclosed algorithm can theoretically achieve, and limitations in the number of iterations that can actually be performed.

[0057] 2. Predictive Segmentation

[0058] A transparent and useful segmentation strategy should achieve the following objectives:

[0059] 1. Internalize existing, valuable domain knowledge and best practices so that practitioners can easily relate to and adopt them;

[0060] 2. Be understandable and intuitive to non-technical planning managers at energy utility companies, and useful for crafting marketing communications;

[0061] 3. Provide some guarantee of optimality in terms of effectiveness, i.e., be highly discriminative for the purpose of identifying sub-populations whose members are more likely to take action than consumers randomly selected from the entire population.

[0062] First, energy utility companies have rich expertise and practical experience that allow them to make assumptions about certain high-level consumer types that they would like to identify from their customers. For example, most experienced planning managers would agree that "green advocates" consumers would respond to other types of communications (that emphasize environmental impact) than those who are more "cost conscious" (who might respond to arguments about saving on expenses).

[0063] Second, the method can start from existing domain knowledge that associates certain variables with each given segment (e.g., "green advocates" can be defined by their income, family type, and education level), and identify simple logical rules that involve those variables that lead to the most effective segmentation strategies. Such intuitive segments should allow for crafting appropriate messaging strategies. For example, consumers in the "green advocates" group would receive information that emphasizes energy conservation and environmental aspects, while consumers in the "high consumption" category would be told about ways they can reduce their large bills.

[0064] The challenge then becomes (as suggested in point 3 above) to develop an algorithmic segmentation method that internalizes the requirements of points 1 and 2, while ensuring that the resulting segments have useful properties, and that the best possible segmentation given the imposed structure has been achieved. The desired outcome is to maximize the impact of marketing communications about energy efficiency program participation, i.e., to target those consumers who are more likely to participate. Since both custom communications and managed activities are expensive, there is a real incentive to create information for a small number of segments, and to have those segments include consumers who are likely to take action.

[0065] 2.1. Problem Setup

[0066] A population of N consumers A service is provided by an operator (an energy utility company); for each consumer, the utility company observes a number M of characteristics that include both consumption and consumer characteristics, such as socio-demographic and physical building attributes Thus, all of the consumers' characteristic data is stored in a matrix For each consumer i, the utility company also observes whether he has participated in any program in the past year, and only if consumer i has participated is this encoded as a binary variable y: y i = 1.

[0067] The utility company wishes to use the data (X, y) to identify K segments within the population that are "homogeneous" with respect to the attributes X, with the goal of informing, simplifying, and improving the effectiveness of targeted communications for demand-side efficiency program participation. Based on existing market research, the utility company can have certain assumptions about which "types of consumers" it serves. Assume that this existing knowledge is in the form of:

[0068] "Green advocates" have relatively high income or at least a university degree.

[0069] "Home improvers" are homeowners or have a large equity share in their home.

[0070] ...

[0071] The data (X, y) can then be used to concretize these assumptions by extracting a set of V patterns from the matrix that are both descriptive (they refer to consumer characteristics that exhibit these patterns) and predictive (because consumers that fall into a certain pattern are more likely to participate than consumers chosen at random from the entire population). Thus, a pattern can be defined as a logical expression of the form:

[0072]

[0073] ​where P is a base rule (a logical statement). Thus, a pattern is defined as a succession of conjunctions. Alternatively, a pattern can be referred to as a set of consumers that follow the logical definition of the pattern. We consider a base rule to have the form:

[0074] r j (x) := x j ≤ t j or r j (x) := x j ≥ t j (2)

[0075] Thus, a base rule is defined by its referenced variable x j (the jth variable in x), the trend (“≥” or “≤”), and the threshold t j learned from the data. We consider a rule P j (·) to be consistent with a hypothesis if both the variable and the trend defining the rule match the hypothesis. Likewise, we define a pattern P to be δ-consistent with a hypothesis if the pattern contains at least δ > 1 rules that are consistent with the hypothesis.

[0076] It is useful to define a coverage matrix C that summarizes the degree to which an item i is covered by a pattern m:

[0077]

[0078] The effectiveness of a pattern P can be computed as the (empirical) engagement probability of the consumers covered by the pattern:

[0079]

[0080] With the above setup, we define a set of K segments such that each pattern in each segment is δ-consistent with the hypothesis that defines the segment. With we define a (known) consistency matrix that describes the allowable relationships between segments and patterns:

[0081]

[0082] Finally, a refinement is a set of segments as if they were individual segments

[0083]

[0084] 2.2. Effective Refinements

[0085] Here, we consider a segmentation strategy to be valid if it differentiates between the segments in terms of the participation rate. That is, a good strategy (for K segments) would identify those segments in the population that participate with probabilities q k , k = 1,..., K that are significantly different (smaller or larger) from the overall rate q observed in the whole population. For example, if the segmentation consists of K = 2 groups A and B, and if all consumers in A participate while none of the consumers in B participate (so q A = 1 and q B = 0), then the segmentation is perfectly valid. A perfectly invalid segmentation is one in which the consumers in A participate at the same rate as the consumers in B (so q A = q B ). Of course, one can always segment the consumers into two segments by letting all those consumers who have already participated in the incentive program enter one of the segments; however, the challenge is to identify patterns in the consumer characteristics X that lead to interpretable, intuitive segment definitions that also allow to predict the participation behavior.

[0086] The validity of each segment can be computed in a similar way as the validity of a pattern, as the (empirical) participation probability of the consumers in the segment:

[0087]

[0088] Thus, if |q k -q| » 0, the segment is a good representative of the participation, where

[0089]

[0090] is the participation rate in the whole population. The problem we want to solve is to assign to each segment at least π and at most patterns, such that the resulting segments have the desired validity properties, e.g.:

[0091] • maximize the minimum validity:

[0092]

[0093] • ensure a proper balance of the validity across segments:

[0094] max θ1q(S1)+... + θ K q(S K ) (9)

[0095] where θ is a given weight vector.

[0096] To this end, define decision variables such that

[0097]

[0098] Thus, segment k is defined as

[0099]

[0100] Thus, the problem reduces to finding the values of z mk that maximize one of the objectives (8-9) and satisfy the following feasibility constraints (F0):

[0101] including only patterns in allowed segments

[0102] limiting the number of patterns in each segment

[0103] patterns can belong to only one segment

[0104] selecting or not selecting a pattern

[0105] For a given segment, there can be many feasible patterns, e.g., |{P m |b mk > 0}| > 1; moreover, patterns can overlap (that is, they define non-disjoint sets of consumers, ). Therefore, segments in S can also overlap if they contain exactly those patterns that overlap in the consumers they describe. This adds additional complexity to the optimization problem of properly formulating the solution to (8-9) and satisfying the constraints (F0).

[0106] If patterns do not overlap, segment validity can be written as:

[0107]

[0108] where

[0109] a≡C T y

[0110] and

[0111] d≡C T 1

[0112] However, since pattern overlap can be substantial, the above expression counts the patterns Consumers in multiple patterns. One simplification we use to solve this problem is to relax the definition of the coverage matrix C, considering that consumers covered by n different patterns can be considered to have a fractional coverage of 1 / n for each pattern. This translates to a modified coverage matrix.

[0113]

[0114] Therefore, the modified coverage matrix assigns weights to each consumer i that indicate the score coverage of a single mode (giving equal importance to each mode). For simplicity, we will still refer to the modified matrix as C.

[0115] 3. Calculate predictive segmentation

[0116] The design of the algorithm used to compute predictive segmentation will be determined by the specific form of the objective function (all constraints are simple linear constraints). Here, we focus on cases where the objective is to assign permissible patterns to segments, such as maximizing minimum effectiveness across K segments—see Equation (8). This is a natural requirement for program managers who want to ensure the minimum effectiveness of their objective communication strategy.

[0117] 3.1. Improve minimum effectiveness

[0118] The above formula in equation (12) uses K vectors z k The K vectors z k Encode the decision variables for each segment. To express the objective and constraints in a more familiar affine form using a single vector of decision variables, we can use the following notation:

[0119]

[0120]

[0121] Where z, v k and Therefore, effectiveness can be expressed as

[0122]

[0123] And the feasibility condition in F0 is:

[0124] z≤vec(B)

[0125]

[0126] z mk ∈{0,1} (F)

[0127] Under the max-min objective case (8), the optimization can attempt to increase the lower bound of effectiveness across segments as much as possible. This results in a relatively uniform distribution of For example, this case can be desirable when taking action on each of the segments. In this case, the optimization problem can be expressed as:

[0128]

[0129] subject to z < vec(B)

[0130]

[0131] z mk ∈ {0, 1} (LFIP)

[0132] Problem (LFIP) is a generalized (max-min) linear fractional integer program with linear constraints. This class of problems has been extensively studied in the literature (see, e.g., Horst and Pardalos (1995), Feng et al. (2011), Schaible and Shi (2004) for a review). Following Boyd and Vandenberghe (2004), we present an equivalent formulation of (LFIP) as a linear integer programming feasibility problem (LFIP-F):

[0133]

[0134] subject to (A - λD)z > 0

[0135] z - vec(B) < 0

[0136]

[0137] z mk ∈ {0, 1} (LFIP-F)

[0138] where A is a matrix with rows D is a matrix with rows k = 1,..., K. For a given value of λ, the above feasibility problem (LFIP-F) can be solved using standard mixed-integer programming packages. Although the initial consumer characteristic data can be quite large (here N ~ 1M consumers), the number of patterns is expected to be much smaller (M ~ 1,000), as is the number of segments (here K = 5). As such, standard packages can provide excellent out-of-the-box performance. In turn, an iterative bisection algorithm 1 (see Table 1 below) can be used to efficiently find the maximum λ*≡ max zλ, the algorithm solves a feasibility problem (LFIP-F) at each step. Starting from a large interval [lo, uo] that ensures the optimal λ* (here [0, 1]), the algorithm progressively narrows the interval [l, u] ensuring at each step that λ* ∈ [u, b]. This is outlined in the following Lemma 1, which builds on Patel et al. (2013).

[0139] Algorithm 1, bisection algorithm

[0140]

[0141] Lemma 1. The output of Algorithm 1 is the optimal z* corresponding to the maximum value λ* of λ within a tolerance ∈ and log2(∈0 / ∈) iterations.

[0142] To prove Lemma 1, we must show that Algorithm 1 will find the unique value λ * for which λ can take its maximum feasible value. To this end, define the feasible set

[0143]

[0144] With this notation we have

[0145] λ * = sup{λ ∈ Λ}

[0146] and the optimal pattern of z* assignments corresponds to λ * . By definition, the optimal λ is the (upper) turning point between the feasible set Λ and the infeasible set

[0147]

[0148] So for a tolerance parameter ∈ > 0 (small) the following conditions must be satisfied:

[0149]

[0150]

[0151] To prove that Algorithm 1 will find the optimal λ * , we need to show that it satisfies the above conditions. We only focus on the terms containing λ in our analysis.

[0152] To prove the first condition, we take

[0153] λ ∈ Λ,

[0154] and we must prove that

[0155] λ - ∈ ∈ Λ.

[0156] The fact that λ ∈ Λ implies

[0157]

[0158] Thus, for λ+∈ and the same Z λ We have

[0159] (A-(λ-∈)D)z λ =(A-λD)z λ +∈Dz λ ≥0.

[0160] Because ∈>0 and D and z λ Both have only non-zero entries, so the second item above is positive.

[0161] To prove the second condition, determine the value. We then want to show that for ∈>0 fact Implicit st(A-λD)z≥D; therefore, we must have (A-λD)z<0. Assign the value to λ. This makes z... λ It is the decision variable vector corresponding to λ, which produces the maximum value of (A-λD)z and satisfies all other conditions defining the feasibility set Λ.

[0162] Then

[0163]

[0164] From the infeasibility of λ, we further have (A-λD)|z λ <0. Thus, for λ+∈, the determining vector z is taken. λ+∈ It produces the maximum value of (A-(λ+∈)D)z, but from the above we have z λ+∈ (A-λD)z λ ≥(A-λD)z. Thus, for z λ+∈ We have

[0165] (A-(λ+∈)D)z λ+∈ =(A-λD)z λ+∈ -∈Dz λ+∈

[0166] ≤(A-λD)z λ -∈Dz λ+∈

[0167] <0. (19)

[0168] Then, from (A-(λ+∈)D)Z λ+∈ <0 We conclude Therefore, Algorithm 1 will always find the maximum feasible λ corresponding to the optimal assignment vector z*.* Furthermore, since each step of the algorithm halves the search interval [l, u], it requires at most

[0169]

[0170] The step completes with the completion condition |u - l| < ε. As it is apparent, the optimization algorithm approaches the optimal solution, so that the optimal solution is within the search box [l, u] < ε that becomes smaller with each iteration as described above.

[0171] 3.2. Extracting predictive patterns from data

[0172] Given a set of observations coded as a feature matrix X and a binary response (engagement) vector y, we want to extract efficient (q » q0) patterns P. To this end we employ the following methodology:

[0173] 1. Use an ensemble method such as Random Forests or AdaBoost Hastie et al. (2009) with classification trees as base learners to generate a number of decision trees of different depths (here we generate trees up to 5 levels). This step allows us to build an initial pattern list P0obtained by traversing the decision trees to each leaf. Depending on the level of the tree used as base classifier in the boosting set, these rules can take different forms of complexity from a single assertion (tree of depth 1 or decision stump) to the conjunction of multiple base rules.

[0174] 2. Trim the pattern list P0to eliminate those rules that do not correspond to some "quality" set of criteria. For this purpose, we should consider a pattern P to be "efficient" if it satisfies the following two criteria:

[0175] • Minimum support: |P| > η, i.e. the number of consumers matching the pattern must be greater than η, so that η + 1 is the minimum number of people matching each pattern. Here we make η = 500.

[0176] • Minimum efficiency: q(P) > ζq0. Here we make ζ = 2.

[0177] 3. Further remove patterns that overlap more than v% (here v = 70%, however v values between 60% and 75% can also be used) with other patterns and have lower efficiency q. For example, for a pattern P1with matching consumers C1and efficiency q1and a pattern P2with matching consumers C2and efficiency q2less than q1, if more than v% of the consumers C2are included in C1, then trim pattern P2as it has lower efficiency.

[0178] This process results in a trimmed set of patterns P = {P1, P2,..., Pn}.​

[0179] 4. Literature review

[0180] Consumer profiling for energy programs has recently received attention in seemingly disparate literatures in engineering and computer science, operations management, and marketing. This work contributes to a broader discussion in these fields by providing a simple and transparent method to create interpretable segments built on the existing domain knowledge of energy utility operations and marketing departments. Engineering research on demand side management has recently been motivated by detailed consumer data, including fine-grained consumption data and social demographic information. This data is often focused in several major areas: i) using data from smart meters or from custom instrumentation for an entire household to describe consumption patterns of a user group with the goal of informing programs such as time-of-use pricing or smart thermostat control, Kwac et al. (2013), Albert and Rajagopal (2015); ii) collecting experimental data for entire households and individual appliances to reconstruct individual end-use from aggregate signals, Carrie Armel et al. (2013), Kolter and Jaakkola (2012); and iii) studying the average effect of different external factors, especially weather, on energy usage, Houde et al. (2012), Kavousian et al. (2013), Kavousian et al. (2015).

[0181] Recent literature on energy analytics has focused on the extension of traditional demand management practices by utility companies using aggregate demand curves to characterize consumption patterns (load profiling) for the purpose of providing information for planning. Albert and Rajagopal (2014) proposed a disaggregation strategy for the cost imposed on the grid by consumers' consumption behavior as a way to target consumer groups that contribute the most to demand fluctuations. A popular research topic is the heterogeneity in typical daily load profiles (this usually requires the use of off-the-shelf unsupervised algorithms such as KMeans to cluster daily customer consumption load shapes) that can be subsequently used for interventions such as differential pricing or incentives to reduce energy consumption. This approach is adopted in, for example, Flath et al. (2012), Rasanen and Kolehmainen (2009), Figueiredo et al. (2005), Smith et al. (2012), Tsekouras et al. (2007), Espinoza et al. (2005). Other variations of segmenting load profiles generate models based on the first learning of consumption, and clustering the resulting models have been discussed in, for example, Albeit and Rajagopal (2013), Alzate et al. (2009). However, this series of research is descriptive in nature, as there is usually no explicit use case provided for the identified load patterns - and there are few projects in utility companies that can incorporate such information at present.

[0182] On the other hand, the operations management and marketing literature has seen an increasing interest in energy applications over the last few years. This can be influenced by the fact that in many utility companies, the department that is traditionally concerned with the allocation, registration, and targeting of consumers for efficiency programs is the operations or marketing department.

[0183] 5. Experimental setup

[0184] 5.1. Consumer characteristic data

[0185] The data we use in this application comes from a large energy company in the Northeastern United States and contains about 100 socio-demographic and building characteristics, as well as monthly energy consumption readings for N = 957,150 consumers over two years. After standard data cleaning procedures, 43 variables of interest were selected that had at least 80% valid entries across the entire population. Out of these, 19 variables were categorical and 24 were numerical. Converting the categorical variables into binary dummy variables, we obtained a final dataset of P = 304 variables. Overall, 48,310 consumers (equivalent to a fraction q0= 4.9%) enrolled in any energy efficiency program in the two years prior to the data collection.

[0186] Table 1. Example categorical consumer characteristics

[0187]

[0188] Table 2. Example numerical societal demographic and building-related consumer characteristics

[0189]

[0190] Table 1 describes several categorical variables of interest. The vast majority of consumers (~80%) own their own homes, with only ~16% renting. The education level reflects the society as a whole, with one quarter of consumers having a college degree and one quarter having a graduate degree, while half have a high school diploma or less. The "green consciousness" variable summarizes the results of third-party analysis considering factors such as magazine subscriptions, community involvement, political leanings, membership in various organizations, and so on, to arrive at a level of interest in environmental issues.

[0191] Table 2 summarizes several more numerical variables of interest. The average birth year is 1957, which indicates a baby boom demographic. The average family in the sample lives in a large house (6 rooms) with over 12 years of land tenure.

[0192] 5.2. Predictive patterns extracted from the data

[0193] As described in Section 3 above, predictive rules were extracted from the data. After pruning, the list of predictive patterns (with a minimum validity of 2 x q0~ 0.10 and a minimum support of η = 500) contained Mo = 2,965 patterns, each with a maximum of 5 base rules (1,852 patterns had 5 base rules, 963 patterns had 4 base rules, 143 patterns had 3 base rules, and 7 patterns had 2 base rules). Figure 4 An example decision tree of height 3 extracted from the data is illustrated. The highlighted pattern is the path in the decision tree starting from the tree root with a validity (proportion of positive samples) of 8%. Figure 5 The distribution of pattern validity q(R) for the different complexity patterns (2-5 base rules) of the Mo = 2,965 patterns extracted from the data is illustrated. As expected, the distribution exhibits an exponential behavior with many lower-validity patterns and fewer high-validity patterns.

[0194] Figure 6The top 20 most important variables predicting enrollment are listed. These include measures of home ownership (loan-to-value ratio, available equity), home and household size, and household income, among others. This suggests that enrollment depends on awareness of financial guarantees and capabilities related to home improvements. The current analysis considers enrollment in any energy efficiency program; an analysis of specific programs targeting more specific types of customers would likely produce more nuanced distinctions in important variables, such as insulation rebates versus high-efficiency appliances.

[0195] 5.3. Associating patterns with segments

[0196] Segments are defined using the results of existing behavioral research and extensive interactions with energy utility companies providing data. Utilities want to identify customers falling into a small number of segments they have defined based on their own internal expertise and research, as well as independent third-party behavioral and market research, such as Frankel et al. (2013). As described in Section 2 above, the purpose of segments is twofold: i) to produce a small number of marketing communications, such as standardized emails, with appropriate information and framing for each segment, and ii) to identify customers likely to enroll in energy efficiency programs corresponding to each segment.

[0197] Based on this existing technology, the utility company believes that customers fall into K = 5 segments: "Green Advocates," "High Consumers," "Home Improvers," "Cost Conscious," and "Culture Drivers." Table 3 summarizes the segment meanings encoding this hypothesis. Given these segment definitions, the latent patterns P from P0 are associated with different segments by ensuring that each pattern P is consistent with the hypothesis about the respective segment meanings (see Section 2). That is, for a given segment S, the rules contain δ basic rules P matching both on variable j and on trend (greater or less than the threshold learned from the data) j ∈ R. The resulting set of patterns P contains M = 219 patterns. Not all customers are covered by the reduced set of patterns with P = 219 (64% of the original sample), but 89% of the enrolled customers are included in the reduced set.

[0198] Table 3. Segment definitions and associated patterns extracted from the data

[0199]

[0200] Table 3 also lists the number of patterns obtained for each segment and their coverage (number of customers in the patterns). Although the number of patterns is smaller than the initial ~3000, selecting a small enough number that approximates maximum effectiveness is still a formidable task. Figure 7The diagram shows the correlation matrix B between encoding patterns and the feasibility of segment allocation. Some patterns may belong to multiple segments, such as... Figure 8 The diagram illustrates the distribution of the number of patterns covering users. Most users are covered by a few patterns; however, a small number of users fall into more than 50 patterns simultaneously.

[0201] Figure 9 The text shows two examples of rules extracted from the data and assigned to the "high consumption" and "cost-conscious" segments. The patterns assigned to "high consumption" contain at least δ=1 basic rules, which involve conditions where consumption exceeds a given threshold.

[0202] 6. Results

[0203] exist In the case of |ul| < ∈ = 10, Algorithm 1 is used to obtain the approximate optimal feasible assignment from pattern to segment Z. -14 At this point, the algorithm narrows the search region from [0, 1] (width ∈ 0) until convergence after 14 iterations. Therefore, the assignment of the pattern to the segment Z is in ∈ = 10. -14 The solution is close to the optimal one. Figure 10 The diagram illustrates the binary search process, which demonstrates how to solve iteratively. and π The feasibility problem of 1 (LFIP-F) is to find the maximum lower bound λ for the piecewise validity. Figure 11 The text shows the target and π The optimal assignment matrix Z is obtained when =1. The horizontal axis is arranged according to... Figure 10 The algorithm uses arbitrary ID number permutations of the same format to represent the allowed allocation matrix B. It selects a small number of patterns that have the best efficiency properties and satisfy the constraints in (F0).

[0204] The optimal solution comprises 10 patterns across 5 segments. Table 4 summarizes the validity and size of the segments. The final effective numbers are all greater than 2×q0, with consumers assigned to one segment (“Culture Drivers”) signing up at almost three times the rate of the total group.

[0205] Table 4. Validity and Size of Segmentation

[0206]

[0207] Figure 12An example of overlap between segments is shown. This overlap is inductive because the patterns that make up the segments themselves may and do overlap among the consumers they encompass. However, segment overlap is actually a natural concept because consumers may possess certain characteristics that allow them to be categorized into one segment (such as "cost-conscious"), while sharing other characteristics with consumers in different segments (such as "home improvement enthusiasts"). Segmentation techniques clearly explain this situation. As a network diagram... Figure 13 The diagram presents a more detailed view of the segmented overlap. Each segment is represented as a node of a size proportional to the number of consumers in that segment; the weights of the connections between segments indicate the paired overlap of the segments. When constraints are... (Left image) changed to (Right figure) When the subdivision structure changes as more patterns are used to construct some segments.

[0208] Please note that "segmentation" is a construct defined by program managers to aid in the creation and management of communications that differentiate consumers to some extent, while maintaining low operating costs and complexity. It reveals some heterogeneity but also does not allow for completely customized interventions on an individual basis. Imposing that every consumer belongs to a segment would only impose unrealistic assumptions, which this approach avoids.

[0209] Figure 14 Including definitions and The optimal pattern assignment list corresponds to the pattern of the segment. Table 3 contains assumptions about the meaning of each segment, such as a threshold t. j (Precisely defining the meaning of "high" and "low") and adding specific information such as basic rules to enrich the definition. For example, one category of "housing improvers" with a high participation rate in energy efficiency programs is South Asians who earn over $75,000 annually and have housing interests exceeding $306,870. Similarly, one category of "green advocates" consists of households with an annual income exceeding $75,000, whose income level is at least two and a half times the national average, who have children, and who do not live in multi-family homes. Patterns within each segment can then be used to design marketing communications targeting that segment, incorporating elements that consumers in that segment are considered likely to respond to. Furthermore, the exclusivity of the patterns (in terms of thresholds learned from data) allows targeting those consumers most likely to enroll.

[0210] It is evident from the above discussion that the obtained subdivision structure is highly dependent on the nature of the constraints, especially on... and π The value of . In order to study this dependency, we focus on the value of . and grid Run Algorithm 1. Objective The optimal value, and the number of modes selected for subdivision in Figure 15 The diagram in the middle shows... When π The best result is obtained when = 1 (therefore the algorithm will not force more patterns into the segments than necessary). For medium to large values ​​and π The low value of (1-3) indicates a good result (λ). * (≈12%). Therefore, these figures provide guidelines on how to balance model complexity and segmentation effectiveness.

[0211] For a given π Value, observed target λ * and the single segment validity value q k k = 1, ..., K varies The changes in this can then be used to adjust parameters for the complexity of the resulting subdivisions, which can be designed to adapt to the desired effectiveness values ​​of the individual segments of interest. This is in π Value = 2 Figure 16 The diagram shows, for example, the focus on "culture drivers," which has... The subdivision is preferred. Note that for all values ​​of k, q k Significantly greater than λ * .

[0212] Finally, the dependence of individual segment effectiveness on segment complexity (the total number of patterns selected across segments) is as follows: Figure 17 The diagram illustrates this. It highlights the best possible efficiency value achievable for a fixed, given segment complexity value. For example, if an efficiency planning manager wants to select a total number of patterns between 20 and 25, he can expect the optimal efficiency of the "culture-driven" segment to always be greater than that of the "cost-conscious" segment. Within the scope, “Family Improvers,” “Green Advocates,” and “Culture Drivers” all have an effectiveness value of around 11%.

[0213] 7. Conclusion

[0214] This application presents a method for programmatically constructing interpretable, predictive segments of energy consumers. The predictive segmentation problem is formulated based on first extracting predictive patterns (conjuncts) from the data, and then optimally assigning the patterns to segments. These segments are defined using existing behavior and marketing research of the energy utility. The optimal assignment is formulated as a generalized (max-min) linear fractional integer program with linear constraints. To solve this program, an efficient bisection algorithm is used. The method is used to identify optimal predictive segments in a population of ~1M electricity consumers of a large US energy utility. An optimal subset of consumers is identified that are consistent with the utility's general assumptions about the types of consumers it serves, and that enroll at least double the enrollment rate of the total population, at ~5%. These segments represent the consumers for which the utility can make appropriate information and target more effectively and economically.

[0215] In the above disclosure, reference has been made to drawings which form a part hereof, and in which are shown by way of illustration specific implementations in which the disclosure can be practiced. It is understood that other implementations can be utilized and structural changes can be made without departing from the scope of the present disclosure. References in the specification to "one implementation", "an implementation", "an example implementation", etc. indicate that the described implementation can include a particular feature, structure, or characteristic, but every implementation can not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same implementation. Furthermore, when a particular feature, structure, or characteristic is described in connection with an implementation, it is submitted that it is within the knowledge of those skilled in the art to effect such feature, structure, or characteristic in connection with other implementations whether or not explicitly described.

[0216] Implementations of the systems, devices, and methods disclosed herein can include or utilize special-purpose or general-purpose computers including computer hardware, such as, for example, one or more processors and system memory as discussed herein. Implementations within the scope of the present disclosure can also include physical and other computer-readable media for carrying or storing computer- executable instructions and / or data structures. Such computer-readable media can be any available media that can be accessed by a general-purpose or special-purpose computer. Computer-readable media that store computer- executable instructions are computer storage media (devices). Computer-readable media that carry computer-executable instructions are transmission media. Thus, by way of example, and not limitation, implementations of the disclosure can comprise at least two distinctly different kinds of computer-readable media: computer storage media (devices) and transmission media.

[0217] Computer storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid state drives ("SSDs") (e.g., based on RAM), Flash memory, phase- change memory ("PCM"), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer.

[0218] Implementations of the devices, systems and methods disclosed herein can communicate over a computer network. A "network" is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a computer, the computer properly views the connection as a transmission medium. Transmissions media can include a network and / or data links which can be used to carry desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer. Combinations of the above should also be included within the scope of computer-readable media.

[0219] Computer-executable instructions comprise, for example, instructions and data which, when executed at a processor, cause a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. The computer executable instructions can be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.

[0220] Those skilled in the art will realize that the subject matter of the present disclosure can be practiced with many computer system configurations, including embedded vehicle computers, personal computers, desktop computers, laptop computers, message processors, hand-held devices, multi-processor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, tablets, pagers, routers, switches, various storage devices, and the like. The present disclosure can also be practiced in distributed system environments where local and remote computer systems, which are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network, both perform tasks. In a distributed system environment, program modules can be located in both local and remote memory storage devices.

[0221] Furthermore, the functionality described herein can be performed, in suitable cases, in one or more of hardware, software, firmware, digital components, or analog components. For example, one or more application specific integrated circuits (ASICs) can be programmed to perform one or more of the systems and processes described herein. Certain terminology is used throughout the description and claims to refer to particular system components. As one skilled in the art will appreciate, components can be referred to by different names in different contexts. In this document, no special significance is to be placed upon whether a component can be called a regulator, a controller, a system, or by some other nomenclature, unless expressly so indicated.

[0222] It should be noted that the sensor implementations discussed above can include computer hardware, software, firmware, or any combination thereof to perform at least a portion of their functions. For example, a sensor can include computer code configured to be executed in one or more processors, and can include hardware logic / circuitry controlled by the computer code. These example devices are provided herein for purposes of illustration, and are not intended to be limiting. Embodiments of the present disclosure can be implemented in further types of devices, as would be known to persons skilled in the relevant art.

[0223] At least some embodiments of the present disclosure have been described in terms of a computer program product, comprising such logic stored on any computer useable medium (e.g., in the form of software). Such software when executed in one or more data processing devices causes the devices to operate as described herein.

[0224] While various embodiments of the present disclosure have been described above, it should be understood that they have been presented by way of example only, and not limitation. It will be apparent to persons skilled in the relevant computer and software arts that a number of modifications, additions, substitutions, and changes can be made to the embodiments described without departing from the spirit and scope of the disclosure. Thus, the breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined in accordance with the following claims and their equivalents. The foregoing description has been presented for the purpose of illustration and description. It is not intended to be exhaustive or to limit the disclosure to the precise form disclosed. Many modifications and variations are possible in light of the above teaching. Further, it should be noted that any or all combinations of the above-described alternative implementations can be used in any combination desired to form additional hybrid implementations of the disclosure.

[0225] The present invention provides, among other things, the following embodiments:

[0226] 1. A method comprising:

[0227] receiving, by a computer system, a plurality of consumer records, each consumer record comprising attributes of a consumer and an adoption status;

[0228] identifying, by the computer system, a plurality of patterns, each pattern defining a group of consumer records of the plurality of consumer records according to the attributes of the plurality of consumer records; and

[0229] distributing, by the computer system, the plurality of patterns among a plurality of segments such that each pattern of the plurality of patterns is assigned to one segment of the plurality of segments, the distribution of the plurality of patterns among the plurality of segments being performed according to an algorithm that iteratively increases a minimum effectiveness of the plurality of segments, the effectiveness of each segment being a function of the adoption status of consumer records of the plurality of consumer records that match a pattern of the plurality of patterns assigned to the each segment.

[0230] 2. The method of embodiment 1, wherein the algorithm comprises solving the equation LFIP-F.

[0231] 3. The method of embodiment 2, wherein the algorithm comprises solving a linear integer programming feasibility problem according to binary search algorithm 1.

[0232] 4. The method of embodiment 1, wherein for each attribute of the attributes of the each pattern, each pattern of the plurality of patterns comprises a threshold value for the each attribute.

[0233] 5. The method of embodiment 1, wherein identifying the plurality of patterns comprises processing the plurality of consumer records according to a decision tree.

[0234] 6. The method of embodiment 1, wherein identifying the plurality of patterns comprises:

[0235] processing the plurality of consumer records according to a decision tree to obtain an initial pattern set; and

[0236] refining the initial pattern set to obtain the plurality of patterns.

[0237] 7. The method of embodiment 6, wherein refining the initial pattern set to obtain the plurality of patterns comprises:

[0238] removing patterns having a support that is below a minimum support, the support of each pattern of the pattern set indicating a number of consumer records that match the each pattern.

[0239] 8. The method of embodiment 7, wherein refining the initial pattern set to obtain the plurality of patterns comprises:

[0240] removing patterns having an effectiveness below a minimum effectiveness, the effectiveness of each pattern in the set of patterns indicating a proportion of consumer records matching the each pattern that have a positive adoption status.

[0241] 9. The method of embodiment 8, wherein the minimum effectiveness is a multiple of a proportion of all consumer records in the plurality of consumer records that have a positive adoption status, the multiple being greater than 1.

[0242] 10. The method of embodiment 9, wherein the multiple is at least 2.

[0243] 11. The method of embodiment 6, wherein refining the initial set of patterns to obtain the plurality of patterns comprises removing from the initial set of patterns those patterns having a set of the plurality of consumer records matching the each pattern such that more than a maximum percentage of the set of consumer records also match another pattern in the initial set of patterns.

[0244] 12. The method of embodiment 11, wherein the maximum percentage is between 60% and 75%.

[0245] 13. The method of embodiment 1, wherein the effectiveness of each segment is also a function of an inclusion matrix assigned to each pattern of the each segment, wherein the inclusion matrix of each pattern has a value of 1 / n for each consumer record in the plurality of consumer records matching the each pattern, wherein n is a number of patterns in the plurality of patterns matching the each consumer record.

[0246] 14. A system comprising one or more processing devices and one or more memory devices operably coupled to the one or more processing devices, the one or more memory devices storing executable code causing the one or more processors to perform operations of:

[0247] receiving a plurality of consumer records, each consumer record including attributes of a consumer and an adoption status;

[0248] identifying a plurality of patterns, each pattern defining a grouping of consumer records of the plurality of consumer records according to the attributes of the plurality of consumer records; and

[0249] allocating the plurality of patterns among the plurality of segments such that each pattern of the plurality of patterns is allocated to one segment of the plurality of segments such that the plurality of patterns is allocated among the plurality of segments according to an algorithm that seeks to maximize a minimum effectiveness of the plurality of segments, the effectiveness of each segment being a function of the adoption status of consumer records of the plurality of consumer records that match a pattern of the plurality of patterns allocated to the each segment.

[0250] 15. The system of embodiment 14, wherein the algorithm comprises solving a linear integer programming feasibility problem according to equation LFIP-F using a binary search algorithm 1.

[0251] 16. The system of embodiment 14, wherein each pattern of the plurality of patterns comprises a threshold for each attribute of the attributes of the each pattern.

[0252] 17. The system of embodiment 14, wherein the executable code further causes the one or more processes to identify the plurality of patterns by:

[0253] processing the plurality of consumer records according to a decision tree to obtain an initial pattern set; and

[0254] refining the initial pattern set to obtain the plurality of patterns.

[0255] 18. The system of embodiment 17, wherein the executable code further causes the one or more processes to refine the initial pattern set to obtain the plurality of patterns by:

[0256] removing patterns having a support that is below a minimum support, the support of each pattern of the pattern set indicating a number of consumer records that match the each pattern;

[0257] removing patterns having an effectiveness that is below a minimum effectiveness, the effectiveness of each pattern of the pattern set indicating a proportion of consumer records that match the each pattern that have a positive adoption status; and

[0258] removing patterns of the initial pattern set that have a set of the plurality of consumer records that match the each pattern such that more than a maximum percentage of the set of consumer records also match another pattern of the initial pattern set.

[0259] 19. The system of embodiment 18, wherein the multiple is at least 2.

[0260] 20. The system of embodiment 14, wherein the executable code further causes the one or more processes to calculate the validity of each segment as a function of an inclusion matrix assigned to each pattern of the each segment, wherein the inclusion matrix of each pattern has a value of 1 / n for each consumer record of the plurality of consumer records that matches the each pattern, where n is the number of patterns of the plurality of patterns that match the each consumer record.

[0261] The application can take other forms, all without departing from its spirit or essential characteristic. The described embodiments are to be considered in all respects only as illustrative and not restrictive. The scope of the application is, therefore, indicated by the appended claims rather than by the description preceding them. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.

Claims

1. A method for consumer segmentation, comprising: receiving, by a computer system, a plurality of consumer records, each consumer record including attributes of a consumer and an adoption status, wherein the adoption status indicates whether the consumer opted in to a particular program or indicates a degree of alignment with a program; processing, by the computer system, the plurality of consumer records according to a decision tree to obtain an initial set of patterns, wherein each node of the decision tree defines an attribute value or an attribute value range corresponding to an attribute of the plurality of consumer records, each pattern defines a grouping of consumer records of the plurality of consumer records according to the attributes of the plurality of consumer records, and each pattern is a combination of rules that respective consumer records of consumers to be included within the pattern must satisfy; refining, by the computer system, the initial set of patterns to obtain a plurality of patterns by removing invalid or duplicate patterns from the initial set of patterns; and allocating, by the computer system, the plurality of patterns among a plurality of segments such that each pattern of the plurality of patterns is allocated to one segment of the plurality of segments, the allocation of the plurality of patterns among the plurality of segments being performed according to an algorithm that iteratively increases a minimum effectiveness of the plurality of segments, wherein the effectiveness is a measure of a number of consumer records that match patterns allocated to each segment and have a positive adoption status, the effectiveness of each segment being a function of the adoption status of consumer records of the plurality of consumer records that match a pattern of the plurality of patterns allocated to the each segment.

2. The method of claim 1, wherein the algorithm comprises solving an equation LFIP-F.

3. The method of claim 2, wherein the algorithm comprises solving a linear integer programming feasibility problem according to a binary search algorithm.

4. The method of claim 1, wherein for each attribute of the attributes of the each pattern, each pattern of the plurality of patterns includes a threshold value of the each attribute.

5. The method of claim 1, wherein refining the initial set of patterns to obtain the plurality of patterns comprises: removing patterns having a support that is below a minimum support, the support of each pattern of the set of patterns indicating a number of consumer records that match the each pattern.

6. The method of claim 5, wherein refining the initial set of patterns to obtain the plurality of patterns comprises: removing patterns having an effectiveness that is below a minimum effectiveness, the effectiveness of each pattern of the set of patterns indicating a proportion of consumer records that match the each pattern that have a positive adoption status.

7. The method of claim 6, wherein the minimum effectiveness is a multiple of a proportion of all consumer records of the plurality of consumer records that have a positive adoption status, the multiple being greater than one. ​ 8. A system for consumer segmentation comprising one or more processing devices and one or more memory devices operably coupled to the one or more processing devices, the one or more memory devices storing executable code that causes the one or more processors to perform operations of: receiving a plurality of consumer records, each consumer record comprising attributes of a consumer and an adoption status, wherein the adoption status indicates whether the consumer opted in to a particular program or indicates a degree of alignment with a program; processing the plurality of consumer records according to a decision tree to obtain an initial set of patterns, wherein each node of the decision tree defines an attribute value or an attribute value range corresponding to an attribute of the plurality of consumer records, each pattern defines a grouping of consumer records of the plurality of consumer records according to the attributes of the plurality of consumer records, and each pattern is a combination of rules that respective consumer records of consumers to be included within the pattern must satisfy; refining the initial set of patterns by removing invalid or duplicate patterns from the initial set of patterns to obtain a plurality of patterns; and allocating the plurality of patterns among a plurality of segments such that each pattern of the plurality of patterns is allocated to one segment of the plurality of segments such that the plurality of patterns is allocated among the plurality of segments according to an algorithm that seeks to maximize a minimum effectiveness of the plurality of segments, wherein the effectiveness is a measure of a number of consumer records that match patterns allocated to each segment and have a positive adoption status, the effectiveness of each segment is a function of the adoption status of consumer records of the plurality of consumer records that match a pattern of the plurality of patterns allocated to the each segment.

Citation Information

Patent Citations

  • TIC: Customization of electronic content based on user interpretation of online reports, with hierarchical models of consumer attributes for targeting content in privacy-preserving manner

    CN1316078A

  • Method for constructing segmentation-based predictive models

    US20030176931A1

  • Customer energy consumption segmentation using time-series data

    US20150161233A1