Generation of scaled-up real-time aggregation for inclusion within one or more modified fields within a subset of the created data

The data processing system addresses the challenge of segmenting and aggregating data across multiple sources by allowing for the creation of data subsets, modification of attributes, and real-time aggregation, thereby enhancing data management and analysis.

JP7692349B2Active Publication Date: 2025-06-13AB INITIO TECHNOLOGY LLC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021506521
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-10-18
Filing Date
2019-08-05
Publication Date
2025-06-13
Estimated Expiration
2039-08-05

AI Technical Summary

Technical Problem

Existing database management systems lack an efficient method for segmenting data records across multiple data sources while allowing for real-time aggregation and modification of data attributes.

Method used

A data processing system that enables segmentation by creating subsets of data from multiple data sources, modifying attributes of fields within these subsets, and displaying representations of the modified fields, allowing for the identification and segmentation of data records based on selected criteria.

Benefits of technology

Enables efficient segmentation and real-time aggregation of data records across multiple data sources, improving data management and analysis capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007692349000001
    Figure 0007692349000001
  • Figure 0007692349000002
    Figure 0007692349000002
  • Figure 0007692349000003
    Figure 0007692349000003
Patent Text Reader

Abstract

1. A data processing system for creating a subset of data from a plurality of data sources, the data processing system including: a memory for storing a plurality of data sources represented in an editor interface; a data structure modification module for selecting a plurality of data sources represented in the editor interface and generating a subset of data contained within the plurality of data sources; a memory for storing selected data structures included within the subset, at least one of the stored data structures including one or more modified attributes of one or more respective fields; a rendering module for displaying a representation of the stored data structures in the editor interface; and a segmentation module for segmenting a plurality of received data records.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Technical Field This application relates to the aggregation of data in a networked database environment. This application also relates to the segmentation of data in a networked database environment.

Background Art

[0002] Background In a database management system, a primary data source is a database that can be in a disk or a remote server. A data source for a computer program can be a file, a data sheet, a spreadsheet, an XML file, or even hard-coded data within a program.

Summary of the Invention

Means for Solving the Problems

[0003] Summary In one general aspect, a method implemented by a data processing system is described for displaying an editor interface that enables segmentation of data records by creating a subset of data from multiple data sources, modifying one or more attributes of one or more fields of each of the subsets, and displaying one or more representations of the one or more modified fields. The method includes selecting a plurality of data sources represented within the editor interface, creating a subset of data included within the plurality of data sources for each of the data sources by selecting one or more data structures, each data structure including one or more fields, from that data source, modifying one or more attributes of one or more of the respective fields within at least one of the selected data structures, storing in memory the selected data structures included within the subset, wherein at least one of the stored data structures includes one or more modified attributes of one or more of the respective fields, displaying in the editor interface a representation of the stored data structures, wherein at least one of the representations is of one or more modified attributes of one or more of the respective fields, each representation including one or more selectable portions, the selectable portions representing fields of the data structures, displaying, receiving through the editor interface selection data that designates selection of one or more of the selectable portions, and segmenting a plurality of received data records by identifying which of the received data records have one or more fields corresponding to the one or more fields represented within the selected one or more selectable portions.

[0004] In aspect 2 according to aspect 1, the data structure includes a key field representing a key for the data structure, the record is associated with the value of the key, and this method includes selecting a plurality of fields from a plurality of selected data structures, when executed, for a specified value of the key, selecting the values of the respective selected fields, for the specified value of the key, combining the selected values and storing in memory executable instructions for outputting the combined values.

[0005] In aspect 3 according to any one of aspects 1 to 2, the representation is a first representation, and this method further includes displaying a second representation of the executable instructions within an editor interface.

[0006] In aspect 4 according to any one of aspects 1 to 3, this method further includes receiving, through the editor interface, additional selection data that further specifies that one or more criteria are applied to the output combined values of one or more given fields represented by one or more selectable portions selected through the editor interface by specifying a selection of the second representation.

[0007] In aspect 5 according to any one of aspects 1 to 4, this method further includes displaying a user interface having one or more first controls for selecting a data structure and one or more second controls for modifying one or more fields.

[0008] In aspect 6 according to any one of aspects 1 to 5, the method further includes receiving, through an editor interface, additional data that specifies one or more criteria to be applied to one or more given fields represented by one or more selectable portions selected through the editor interface, and segmenting includes identifying which of the received data records have one or more values for one or more fields that correspond to one or more fields represented within the selected one or more selectable portions and satisfy one or more criteria, thereby segmenting the plurality of received data records.

[0009] In aspect 7 according to any one of aspects 1 to 6, the data structure includes one or more records, and each record has one or more values for a particular field.

[0010] In aspect 8 according to any one of aspects 1 to 7, at least one of the data sources includes an unselected data structure.

[0011] In general aspect 9, a data processing system for displaying an editor interface that enables segmentation is described by creating a subset of data from a plurality of data sources, modifying one or more attributes of one or more fields of each of the subsets, and displaying one or more representations of one or more modified fields. The data processing system includes one or more processing devices and selecting a plurality of data sources represented within the editor interface, for each of the plurality of data sources, selecting one or more data structures from that data source, each data structure including one or more fields, and generating a subset of data included within the data sources by modifying one or more attributes of one or more fields within at least one selected data structure, storing in memory the selected data structures included within the subset, wherein at least one of the stored data structures includes one or more modified attributes of one or more fields, displaying a representation of the stored data structures within the editor interface, wherein at least one of the representations is of one or more modified attributes of one or more fields, each representation including one or more selectable portions, the selectable portions representing fields of the data structure, displaying, receiving through the editor interface selection data that designates a selection of one or more of the selectable portions, and identifying which of the received data records have one or more fields corresponding to one or more fields represented within the selected one or more selectable portions, and storing one or more machine-readable hardware storage devices storing instructions executable by one or more processing devices to perform operations including segmenting a plurality of received data records.

[0012] In aspect 10 according to aspect 9, the data structure includes a key field representing a key for the data structure, the record is associated with the value of the key, and one or more operations include selecting a plurality of fields from a plurality of selected data structures, when executed, selecting the values of the respective selected fields for a specified value of the key, combining the selected values for the specified value of the key, and storing in memory executable instructions for outputting the combined values.

[0013] In aspect 11 according to any one of aspects 9 to 10, the representation is a first representation, and one or more operations further include displaying a second representation of the executable instructions within an editor interface.

[0014] In aspect 12 according to any one of aspects 9 to 11, one or more operations further include receiving, through the editor interface, additional selection data that specifies a selection of the second representation and applies one or more criteria to those output combined values of one or more given fields represented by one or more selectable portions selected through the editor interface.

[0015] In aspect 13 according to any one of aspects 9 to 12, one or more operations further include displaying a user interface having one or more first controls for selecting a data structure and one or more second controls for modifying one or more fields.

[0016] In aspect 14 according to any one of aspects 9 to 13, one or more operations further include receiving, through an editor interface, additional data that specifies one or more criteria to be applied to one or more given fields represented by one or more selectable portions selected through the editor interface, and segmenting includes identifying which of the received data records correspond to one or more fields represented within the one or more selected selectable portions and have one or more values that meet one or more criteria, thereby segmenting a plurality of received data records.

[0017] In aspect 15 according to any one of aspects 9 to 14, the data structure includes one or more records, and each record has one or more values for a particular field.

[0018] In aspect 16 according to any one of aspects 9 to 15, at least one of the data sources includes an unselected data structure.

[0019] In an overall aspect 17, an editor interface that enables segmentation is displayed by creating a subset of data from a plurality of data sources, modifying one or more attributes of one or more fields of each of the subsets, and displaying one or more representations of one or more modified fields. One or more machine-readable hardware storage devices are described. The one or more machine-readable hardware storage devices are configured to select a plurality of data sources represented within the editor interface; for each of the plurality of data sources, select one or more data structures from that data source, each data structure including one or more fields, and generate a subset of data included within the plurality of data sources by modifying one or more attributes of one or more fields within at least one selected data structure; store in memory the selected data structures included within the subset, wherein at least one of the stored data structures includes one or more modified attributes of one or more fields; display a representation of the stored data structures within the editor interface, wherein at least one of the representations is of one or more modified attributes of one or more fields, and each representation includes one or more selectable portions, the selectable portions representing fields of the data structure; receive selection data specifying selection of one or more of the selectable portions through the editor interface; and identify which of the received data records have one or more fields corresponding to one or more fields represented within the selected one or more selectable portions, to execute operations including segmenting a plurality of received data records. Store instructions executable by one or more processing devices for performing the operations.

[0020] In aspect 18 according to aspect 17, the data structure includes a key field representing a key for the data structure, the record is related to the value of the key, and one or more operations include selecting a plurality of fields from a plurality of selected data structures, when executed, for a specified value of the key, selecting the values of each selected field, for the specified value of the key, combining the selected values and storing in memory executable instructions for outputting the combined values.

[0021] In aspect 19 according to any one of aspects 17 to 18, the representation is a first representation, and one or more operations further include displaying a second representation of the executable instructions within an editor interface.

[0022] In aspect 20 according to any one of aspects 17 to 19, one or more operations further include receiving, through the editor interface, additional selection data that specifies the selection of the second representation and further specifies that one or more criteria are applied to those output combined values of one or more given fields represented by one or more selectable portions selected through the editor interface.

[0023] In aspect 21 according to any one of aspects 17 to 20, one or more operations further include displaying a user interface having one or more first controls for selecting a data structure and one or more second controls for modifying one or more fields.

[0024] In aspect 22 according to any one of aspects 17 to 21, one or more operations further include receiving, through an editor interface, additional data that specifies one or more criteria to be applied to one or more given fields represented by one or more selectable portions selected through the editor interface, and segmenting includes identifying which of the received data records correspond to one or more fields represented within the one or more selected selectable portions and have one or more values that meet one or more criteria, thereby segmenting a plurality of received data records.

[0025] In aspect 23 according to any one of aspects 17 to 22, the data structure includes one or more records, and each record has one or more values for a particular field.

[0026] In aspect 24 according to any one of aspects 17 to 23, at least one of the data sources includes an unselected data structure.

[0027] In general aspect 25, a method is described that is executed by a data processing system to generate a near real-time aggregation. The method includes intermittently receiving data records from one or more data sources, for a given received data record, identifying at least a first field and a second field within the given data record, detecting a first value within the first field and a second value within the second field, and generating a composite key according to the first value of the first field and the second value of the second field, accessing from memory aggregated data related to at least the first field or the second field, generating a data record having a field for storing the composite key and one or more fields for storing items of the aggregated data respectively, thereby generating a composite key value, where the composite key value represents a near real-time aggregation of data related to at least the first field or the second field, generating, and recording the occurrence of the given data record by storing the composite key value in memory.

[0028] In aspect 26 according to aspect 25, generating the composite key includes concatenating the first value with the second value.

[0029] In aspect 27 according to any one of aspects 25 - 26, the method further includes hashing the composite key and storing the hashed composite key in a hash table together with the composite value.

[0030] In aspect 28 according to any one of aspects 25 - 27, the composite value is the aggregated data.

[0031] In aspect 29 according to any one of aspects 25 to 28, the method comprises, for a given record, detecting the values of each field included in the given record and generating a plurality of unique combinations of at least two detected values, wherein each unique combination is a composite key, generating, for each composite key, identifying one or more fields in the given record, wherein for the one or more fields, the composite key includes one or more respective values of the one or more fields, accessing from memory aggregated data related to at least one of the one or more identified fields, and generating a data record having a field storing the composite key and a field storing the aggregated data, thereby generating a composite key value and storing the composite key value in memory.

[0032] In aspect 30 according to any one of aspects 25 to 29, generating a plurality of unique combinations includes generating all unique combinations of the detected values of the fields in the given record.

[0033] Aspect 31 according to any one of aspects 25 to 30, the method further comprising receiving a request for aggregation of specified values over a period of time, selecting from memory composite key values that store the occurrence of the specified values, and extracting the requested aggregation from the composite key values.

[0034] In aspect 32 according to any one of aspects 25 to 31, the method further comprises aggregating one or more items of aggregated data with the values of fields in a given data record, generating an approximately real-time aggregated value of the fields based on the aggregation, and storing the approximately real-time aggregated value in the composite key value.

[0035] In aspect 33 according to any one of aspects 25 - 32, the method further includes receiving a request for aggregation related to one or more specified values, generating a composite key from the one or more specified values, hashing the composite key, retrieving a composite value stored with the hashed composite key from a hash table stored in memory, and extracting an item of the requested aggregated data from the composite value.

[0036] In general aspect 34, a data processing system for generating near real - time aggregations, includes one or more processing devices, intermittently receiving data records from one or more data sources, for a given received data record, identifying at least a first field and a second field within the given data record, detecting a first value within the first field and a second value within the second field and generating a composite key according to the first value of the first field and the second value of the second key, accessing from memory aggregated data related to at least the first field or the second field, generating a data record having a field for storing the composite key and one or more fields for storing respectively an item of the aggregated data, thereby generating a composite key value, the composite key value representing a near real - time aggregation of data related to at least the first field or the second field, generating and storing the composite key value in memory, and storing one or more machine - readable hardware storage devices storing instructions executable by the one or more processing devices to perform operations including recording the occurrence of a given data record. A data processing system for generating near real - time aggregations is described.

[0037] In aspect 35 according to aspect 34, generating the composite key includes concatenating the first value with the second value.

[0038] In aspect 36 according to any one of aspects 34 to 35, one or more operations further include hashing the composite key and storing the hashed composite key in a hash table together with the composite value.

[0039] In aspect 37 according to any one of aspects 34 to 36, the composite value is aggregated data.

[0040] In aspect 38 according to any one of aspects 34 to 37, one or more operations are to detect the values of each field included in a given record for the given record and generate a plurality of unique combinations of at least two detected values, wherein each unique combination is a composite key, to generate, for each composite key, one or more fields in the given record, for the one or more fields, the composite key includes one or more respective values of those one or more fields, to identify one or more fields, to access from memory the aggregated data related to at least one of the one or more identified fields, to generate a data record having a field for storing the composite key and a field for storing the aggregated data, thereby generating a composite key value and storing the composite key value in memory.

[0041] In aspect 39 according to any one of aspects 34 to 38, generating a plurality of unique combinations includes generating all unique combinations of the detected values of the fields in the given record.

[0042] In aspect 40 according to any one of aspects 34 to 39, one or more operations further include receiving a request for aggregation of specified values over a period of time, selecting from memory a composite key value that stores the occurrence of the specified values, and extracting the requested aggregation from the composite key value.

[0043] In aspect 41 according to any one of aspects 34 to 40, one or more operations further include aggregating one or more items of aggregated data with the values of fields in a given data record, generating a substantially real-time aggregated value of the field based on the aggregation, and storing the substantially real-time aggregated value within the composite key value.

[0044] In aspect 42 according to any one of aspects 34 to 41, one or more operations further include receiving a request for aggregation related to one or more specified values, generating a composite key from the one or more specified values, hashing the composite key, requesting the composite value stored with the hashed composite key from a hash table stored in memory, and extracting the item of the requested aggregated data from the composite value.

[0045] In general aspect 43 according to any one of the described aspects 1 to 42, one or more machine-readable hardware storage devices for generating substantially real-time aggregation, the one or more machine-readable hardware storage devices intermittently receive data records from one or more data sources, for a given received data record, identify at least a first field and a second field within the given data record, detect a first value within the first field and a second value within the second field and generate a composite key according to the first value of the first field and the second value of the second key, access from memory the aggregated data related to at least the first field or the second field, generate a data record having fields for storing the composite key and one or more fields for storing items of aggregated data respectively, thereby generating a composite key value, the composite key value representing a substantially real-time aggregation of data related to at least the first field or the second field, generating and storing the composite key value in memory, and storing instructions executable by one or more processing devices for performing operations including recording the occurrence of a given data record.

[0046] In aspect 44 according to any one of aspects 1 to 44, generating a composite key includes concatenating a first value with a second value.

[0047] In aspect 45 according to any one of aspects 1 to 44, the one or more operations further include hashing the composite key and storing the hashed composite key in a hash table together with the composite value.

[0048] In aspect 46 according to any one of aspects 1 to 45, the composite value is aggregated data.

[0049] In aspect 47 according to any one of aspects 1 to 46, the one or more operations are, for a given record, detecting the values of each field included in the given record and generating a plurality of unique combinations of at least two detected values, each unique combination being a composite key, generating, for each composite key, one or more fields in the given record, for the one or more fields, the composite key including one or more respective values of the one or more fields, identifying one or more fields, accessing from memory aggregated data related to at least one of the one or more identified fields, generating a data record having a field for storing the composite key value and a field for storing the aggregated data, thereby generating a composite key value and storing the composite key value in memory.

[0050] In aspect 48 according to any one of aspects 1 to 47, generating a plurality of unique combinations includes generating all unique combinations of the detected values of the fields in the given record.

[0051] In aspect 49 according to any one of aspects 1 to 48, the one or more operations further include receiving a request for aggregation of specified values over a period of time, selecting from memory a composite key value storing the occurrence of the specified values, and extracting the requested aggregation from the composite key value.

[0052] In aspect 50 according to any one of aspects 1 to 49, one or more operations further include aggregating one or more items of aggregated data with the values of fields in a given data record, generating a substantially real-time aggregated value of the field based on the aggregation, and storing the substantially real-time aggregated value in the composite key value.

[0053] In aspect 51 according to any one of aspects 1 to 50, one or more operations further include receiving a request for aggregation related to one or more specified values, generating a composite key from the one or more specified values, hashing the composite key, retrieving a composite value stored with the hashed composite key from a hash table stored in memory, and extracting an item of the requested aggregated data from the composite value.

[0054] In aspect 52 according to any one of aspects 1 to 50, a data processing system for displaying an editor interface that enables segmentation of data records by creating a subset of data from a plurality of data sources, modifying one or more attributes of one or more fields of the subset, and displaying one or more representations of the one or more modified fields, the system comprising: a memory storing a plurality of data sources represented within the editor interface; a data source selection module that selects a plurality of data sources represented within the editor interface and, for each of the plurality of data sources, selects one or more data structures from that data source, each data structure including one or more fields, and generates a subset of data included within the plurality of data sources by modifying one or more attributes of one or more fields within at least one selected data structure; a memory storing the selected data structures included within the subset, at least one of the stored data structures including one or more modified attributes of one or more fields; a rendering module that displays a representation of the stored data structure within the editor interface, at least one of the representations being of one or more modified attributes of one or more fields, each representation including one or more selectable portions that represent fields of the data structure, and the rendering module receiving selection data that designates selection of one or more of the selectable portions through the editor interface; and a segmentation module that segments a plurality of received data records by identifying which of the received data records and which fields have one or more fields corresponding to the one or more fields represented within the selected one or more selectable portions.

[0055] Other features and advantages of the present invention will become apparent from the following description and claims.

Brief Description of the Drawings

[0056]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 6

Figure 7

Figure 8A

Figure 8B

Figure 8C

Figure 8D

Figure 8E

Figure 8F

Figure 8G

Figure 8H

Figure 8I

Figure 8J

Figure 8K

Figure 8L

Figure 8M

Figure 8N

Figure 8O

Figure 8P

Figure 8Q

Figure 8R

Figure 8S

Figure 8T

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Best Mode for Carrying Out the Invention

[0057] Detailed Description Referring to FIG. 1, a networked system 10 for modifying a data structure (e.g., a table) is shown. Specifically, the networked system 10 enables obtaining a data structure from a plurality of data sources and modifying one or more fields (or other attributes or attributes of fields) of those data structures from the data sources. Generally, a field includes, for example, a designated portion of a data record for storing data and / or rows within a relational database. Generally, an attribute includes, for example, characteristics such as a data format. The networked system 10 includes data sources 12a - 12c. The networked system 10 includes, for example, an execution system 14 that specifies which of those data structures to provide to a client device for accessing the data structure and that modifies those data structures. The execution system 14 includes a memory 16 (e.g., including volatile memory, non-volatile memory, etc.) for receiving and storing data structures from the data sources 12a - 12c.

[0058] In one example, the memory 16 stores references to each of the data sources 12a - 12c. The memory 16 receives data structures included within the data sources 12a - 12c from those respective data sources. The memory 16 stores the data structures (and data included within the data structures, such as records in a table) in relation to references to the data sources that transmitted the data structures to the execution system 14. The execution system 14 also includes, for example, a data structure selection and data structure modification module 18 (hereinafter "module 18") for selecting one or more data sources that are data providers, selecting one or more data structures within those selected one or more data sources, and modifying one or more data structures within the selected one or more data sources (e.g., by modifying field names). In one example where the data structure is a table, columns within the table are referred to as fields and rows within the table are referred to as records. The module 18 also enables enhancement of data structures and / or fields within data structures, for example, by enabling generation of new data structures that include combinations of two or more fields from various data structures. The execution system 14 includes, for example, a rendering module 20 for rendering a visual representation of the modified data structure within a user interface 24 (displayed on a client device).

[0059] To specify instructions for segmentation, the user selects one or more fields or portions of a data structure via the user interface 24. Generally, segmentation involves the process of defining and subdividing a set of data records to only those data records that meet one or more specified criteria. The client device that renders the user interface 24 transmits data specifying the selection of one or more fields or portions of the data structure to the rendering module 20. The rendering module 20 transmits this data (specifying the selection) to the segmentation module 22, and the segmentation module 22 performs segmentation of various data records stored in the memory 16 or other data repositories to create an output data set 26.

[0060] Referring to FIG. 2, the networked system 10 has an execution system 14 configured to modify data structures A - D (e.g., tables). In this specific example, data structures A and C are modified by a data structure modification module 18. As shown in the graphical user interface 18b, data is input into fields titled "Card Purchase Visa" and "Customer Engagement" from the data structure modification module 18, and modified data structures A and B are illustrated that include field names of "ID", "Transaction Amount" and "ID" and "Period", respectively. The data structure modification module 18 also creates modified field data 18a to send to a rendering module 20, and the rendering module 20 renders representations 21 of the modified structures A and C at a user interface 24 (FIG. 1) along with instructions for joining return data records based on a joining ID, e.g., the value of an ID field included within a return data record. Generally, the modified field data includes data specifying one or more modifications to fields of a data structure, such as modifications to the name of a field or a column or row within the data structure.

[0061] The execution system 14 sends the modified field data 18a (joined by ID (transaction amount > $5000) & (customer engagement < 6 months)) to the segmentation module 22, and the segmentation module 22 creates a query (query (transaction amount > $5000) & (customer engagement < 6 months)) based on the segmentation logic 22a for accessing a data source, for example 12d. The data source 12d returns two records 13a and 13b, each containing content such as "ID: f423543 VISA: 7349.00" and "ID: f423543 Customer Engagement: 2 months" as shown. The returned record 22b is sent back to the segmentation module 22 and given to the logic module 25 for specifying the joining of the returned records (by joined ID) 18a "ID: f423543 VISA: $7349.00 Customer Engagement: 2 months".

[0062] Next, referring to FIG. 3, an execution system 14 configured to create a composite key by a composite key module 30 is shown. The data structure modification module 18 sends the modified field data to the composite key module 30. The composite key module 30 also sends the concatenation 18a of the return record "ID: f423543 VISA: $7349.00 Customer Engagement: 2 months", which is the concatenation of the composite key 31a and the composite value 31b, to create the composite key value 31 stored in the data store 12e. Generally, a composite key includes a key generated from one or more values of another field in a data record. Generally, a composite value includes some aggregation or other data related to one or more of the values from which the composite key is generated. In this example, the composite value includes an aggregation of "5550.32, 345.24, 12.01, 23", each of which represents the current amount of the current transaction, the average transaction amount related to that specific ID over a specified period (e.g., the past 30 days), the minimum transaction amount that occurred during that period, and the count of the number of transactions that occurred during that period, respectively. As will be described below, the composite key module 30 and the rendering module 20, the segmentation module 22, and the logic module 25 that operate on the outputs from each other are also illustrated.

[0063] Referring to FIG. 4, the structure 32 of the composite key and the associated composite values derived from the values of the data record fields is shown. In this example, the system 10 receives a data record 32a that includes four fields: a subscriber ID (SubID) field, an event type field, a date field, and a length field (which specifies the length of the voice event). In this example, the value of the SubID field is "43054421". The value of the event type field is "voice". The value of the date field is "4 / 3 / 2018". The value of the length field is "4.34 minutes". From the values of the first three fields within the data record 32a, the system 10 generates several keys, for example, one key for each possible combination of fields. In some examples, if the data record has "n" fields, the number of different combinations of fields is 2 n as follows. As shown in Table 33a, in this example, the system 10 generates seven distinct keys from the fields within the data record 32. The seven distinct keys are as follows. Key 1: SubID Key 2: SubID.Event Type Key 3: SubID.Date Key 4: SubID.Event Type.Date Key 5: Event Type Key 6: Date Key 7: Event Type.Date

[0064] For each of the composite keys, the system generates a composite value that includes one or more specified values. For the "SubID" key (i.e., Key 1 in Table 33a), the composite value (represented by composite value 1 in Table 33a) is the average number ("Average") of events received over a specified period (e.g., 5 days) for the subscriber represented by the SubID and the count ("Count") of the number of events received over the specified period for that subscriber. That is, for the "SubID" key, as shown in Table 33a, the composite value is "Average,Count".

[0065] For the key of "SubID.Event Type" (i.e., key 2 in Table 33a), the composite value (indicated by composite value 2 in Table 33a) is the average number ("Average") of events of the event type specified in the key received over a specified period (e.g., 5 days) for the subscriber represented by SubID, the shortest time ("Min") of the events of the specified event type, the longest time ("Max") of the events of the specified event type, and the count ("Count") of the number of events of the event type specified in the key received over a specified period for that subscriber. That is, for the key of "SubID.Event Type", as shown in Table 33a, the composite value is "Average,Min,Max,Count".

[0066] For the key of "SubID.Date" (i.e., key 3 in Table 33a), the composite value (indicated by composite value 3 in Table 33a) is the count ("Count") of the number of events received on the date specified by the date field for the subscriber specified by the SubID field. That is, for the key of "SubID.Date", as shown in Table 33a, the composite value is "Count".

[0067] For the key of "SubID.Event Type.Date" (i.e., key 4 in Table 33a), the composite value (indicated by composite value 4 in Table 33a) is the shortest time ("Min") of the events of the specified event type for the subscriber specified within the SubID field and on the specified date within the date field, the longest time ("Max") of the events of the specified event type for the subscriber specified within the SubID field and on the specified date within the date field, and the count ("Count") of the number of events of the event type specified in the key for the subscriber specified within the SubID field and on the specified date within the date field. That is, for the key of "SubID.Event Type.Date", as shown in Table 33a, the composite value is "Min,Max,Count".

[0068] For the key of "Event Type" (i.e., key 5 in Table 33a), the composite value (indicated by the composite value 5 in Table 33a) is the average number ("Average") of events of the event type specified in the key received over a specified period (e.g., 5 days), the shortest time ("Min") of events of the specified event type over the specified period, the longest time ("Max") of events of the specified event type over the specified period, and the count ("Count") of the number of events of the event type specified in the key over the specified period. That is, for the key of "Event Type", as shown in Table 33a, the composite value is "Average, Min, Max, Count".

[0069] For the key of "Date" (i.e., key 6 in Table 33a), the composite value (indicated by the composite value 6 in Table 33a) is the count ("Count") of the number of events received on the date specified by the date field. That is, for the key of "Date", as shown in Table 33a, the composite value is "Count".

[0070] For the key of "Event Type.Date" (i.e., key 7 in Table 33a), the composite value (indicated by the composite value 7 in Table 33a) is the shortest time ("Min") of events of the specified event type on the specified date in the date field, the longest time ("Max") of events of the specified event type on the specified date in the date field, and the count ("Count") of the number of events of the event type specified in the key on the specified date in the date field. That is, for the key of "Event Type.Date", as shown in Table 33a, the composite value is "Min, Max, Count".

[0071] Table 33b shows the actual keys 1 to 7 and the composite values 1 to 7 which are the related composite values respectively. In this example, the values of keys 1 to 7 are generated from the values of fields within data record 33a. The system generates the composite values by updating the previously calculated composite values and / or by accessing specified data from the persistent memory 58 (Figure 6). For example, for key 4 (i.e., SubID.Event Type.Date Key), the memory 56 (Figure 6) may already store an entry for that key. The stored entry could be as follows: Key 4: 43054421.Voice.4 / 3 / 2018, Composite Value 4: Shortest 0.9, Longest 8.09, 2. In this example, when record 32a is received, system 10 identifies that it already stores the composite value for the key 43054421.Voice.4 / 3 / 2018. Therefore, system 10 updates the composite value according to the length of the voice event (i.e., 4.34 minutes) specified within data record 268a. Based on this update, the system determines a new composite value of "Shortest 1.2, Longest 8.09, 3" as shown in Figure 4.

[0072] In other examples, there may be cases where the system still fails to identify the composite value for key 4. In this example, the system accesses memory 56 (Figure 6) and / or persistent memory 58 (Figure 6) for the data record of the subscriber (i.e., SubID: 43054421) referred to within data record 32a. From the accessed data records, the system determines which data records refer to voice events for the specified date, i.e., 4 / 3 / 2018. From the data records referring to voice events for the specified date, the system determines the shortest time of the voice events that occurred for the specified subscriber on the specified date (e.g., from the "Length" field of each record), the longest time of the voice events (e.g., from the "Length" field of each record), and the count of the number of voice events. From the determined values, the system determines the composite value and stores it in relation to the key of "43054421.Voice.4 / 3 / 2018", i.e., within the hash table 268d for key 4.

[0073] In this example, a memory (not shown) stores a hash table 33c having the respective hashed key values 35a to 35g of keys 1 to 7. In this example, the system generates a hashed key value by applying a hashing algorithm to a composite key. The hash table 33c also stores composite values 36a to 36g respectively corresponding to the composite values 1 to 7 within the table 33c. Generally, corresponding or correspondence refers to matching or having a similar threshold. In this example, each record is stored independently by storing a composite key and a related composite value.

[0074] In these examples, the system pre-computes data values for various combinations of keys. For example, for the key "43054421.Voice.4 / 3 / 2018", the system pre-computes the minimum, maximum, and count values for the voice events of the specified subscriber that occurred on 4 / 3 / 2018. By pre-computing these values, the system reduces (or eliminates) the runtime latency with respect to determining real-time aggregations and other real-time values. For example, at runtime, the system needs to determine the number of voice events that occurred on 4 / 3 / 2018 for the subscriber represented by SubID 43054421. The system can determine this real-time aggregation by querying various data repositories and warehouses for data records that include a SubID field with the value 43054421. Then, from all the returned data records, the system can parse the data records to determine a subset of the data records that have the value "Voice" for the event type field and the value "4 / 3 / 2018" for the date field. Then, the system 10 can count the number of records returned within the subset to determine the count. However, this querying and processing causes latency associated with the system performing the querying and parsing. To reduce or eliminate this latency, the system pre-computes the aggregations (or other values such as minimum and maximum values) and stores those aggregations in relation to the hash value of the composite key. Thus, to look up the count of the number of voice calls made by a particular subscriber on a particular date, the system generates the appropriate key such as SubID.Voice.Date or 43054421.Voice.4 / 3 / 2018. The system hashes the value of the key and uses the hashed key value to access the composite value within the hash table 33c. By doing so, the system eliminates or reduces the latency associated with having to calculate the aggregation in real time.

[0075] Another advantage of storing composite values in relation to composite keys is that when adding a new field to a data record, a new composite key with a value for that new field can be generated and the occurrence of values within the associated composite values can be easily tracked by keeping track of the count (or other aggregation) within the associated composite values. For example, a "new customer" field is added to data record 32a. In this example, if a customer has applied for telephone company service within the past six months, then that customer is a new customer. In this case, the new customer field has a value of yes. Otherwise, the new customer field has a value of no. In this example, the system tracks the occurrence of new customers who made a voice call on a specified date of 4 / 3 / 2018 and the system generates a new key of new customer.event type.date with a value of yes.voice.4 / 3 / 2018. The composite value for this new key is "count". Then, when a new record is received, the system generates or updates the composite value according to the number of records that reference the voice event of new customers received on 4 / 3 / 2018. The advantage of generating and storing the composite value in relation to the composite key is that when adding new fields to a data record, there is no need to add columns to the table to track the values of those new fields. Rather, the new values of the fields can be tracked by generating new keys, and the generation of new keys simply requires adding new rows to the table rather than changing the table's structure by adding new columns.

[0076] Next, referring to FIG. 5A, the execution system 14 accesses the data source 12a, and the data source 12a returns two records 18a, 18b each containing the respective content of "ID: 53054423, Transaction Amount: $5550.32, Date: 4 / 3 / 2018, Card Type: Visa" and "ID: 53054423, Customer Engagement: 2 months, Date: 4 / 3 / 2018". The returned records 18a, 18b are sent to the composite key module 30 to create a composite key 25a for record 18a and a composite key 25b for record 18b. The composite key values 25a, 25b are stored in the data store 12f, and the composite key module 30 sends the composite key values 25a, 25b to the rendering module 20 that renders the representation shown in FIG. 5B. In this example, the composite key value 25a includes the composite key of "53054423.VISA" and the composite value of "5550.32,345.24,12.01,23" representing the total purchase amount of the current transaction, the average purchase amount over the specified number of days (e.g., the past 30 days), the minimum purchase amount over the specified number of days, and the count of the number of purchases made over the specified number of days. The composite key value 25b includes the composite key of "53054423.Customer Engagement" and the composite value of "2,1 / 1 / 2018,9.00,1045,104" representing the period for which a particular user represented by the ID field is a customer, the date the user became a customer, the minimum purchase amount of the customer, the maximum purchase amount of the customer, and the average purchase amount of the customer, respectively. As will be described in more detail below with respect to FIG. 5B, the composite key values 25a, 25b are sent to the rendering module 20 so that the rendering module can determine what aggregations can be displayed as part of the segmentation template. As will be described in more detail below, in this example, depending on which tables are selected and which composite key values are associated with or available for those tables, the aggregations (e.g., minimum, count, etc.) included in the composite key values 25a, 25b become available for use in the definitions in the segmentation.

[0077] The rendering module 20 receives the segmentation logic 23c (specified within the graphical user interface shown in FIG. 5B) and transmits the segmentation logic 23c to the segmentation module 22. In this example, the segmentation logic is as follows: "(Combined by ID (Count(Visa transaction amount > $5000) > 2) and (Customer engagement < 6 months and total transaction count > 100))". In response, the segmentation module 22 creates the composite key queries 23a (53054423.VISA) and 23b (53054423.Customer engagement) and transmits the composite key queries 23a, 23b to the data repository 12f. In response, the data repository 12f searches for (e.g., within a table) composite key values that have composite keys matching the composite keys specified in the queries 23a, 23b. In this example, the composite key query 23a includes the composite key of "53054423.VISA". Based on this composite key, the data repository 12f obtains the composite key value 25a that has the composite key of "53054423.VISA" and thus matches the composite key specified in the composite key query 23a. In this example, the composite key query 23b includes the composite key of "53054423.Customer engagement". Based on this composite key, the data repository 12f obtains the composite key value 25b that has the composite key of "53054423.Customer engagement" and thus matches the composite key specified in the composite key query 23b. The data repository 12f returns the composite key values 25a, 25b to the segmentation module 22 as the return record 22c. In response, the segmentation module 22 transmits the return record 22c to the logic module 25 for further processing. In this example, the logic module 25 also receives the segmentation logic 23c (e.g., from the segmentation module 22) and implements the combination logic to combine the return record 22c to create the aggregated or combined record 27.

[0078] Referring to FIG. 5B, the graphical user interface 40 is a modified form of the graphical user interface 21 (FIG. 2). In this example, the graphical user interface 40 includes components 40a to 40e. In this example, component 40a is specified to use the "Customer Engagement" table to segment customers by including only customers who have been customers for more than six months. In this example, certain aggregations (e.g., counts) are associated with the customer engagement table. In some examples, each composite value includes the same type of aggregation. Therefore, each table can be associated with the same type of aggregation. In other examples, a table can be associated with only a specific type of aggregation. In this example, when a composite key query is sent to the data repository and a composite key value that does not include the aggregations required for segmentation is returned, the execution system simply discards the composite key value to be retained). In this example, component 40e is specified to use the Visa card purchase table for segmentation and to include only customers who have made more than three card purchases, as specified by component 40d, in the segment. Component 40c is specified to combine the records returned by executing the logic specified within components 40a to 40b, 40d, and 40e.

[0079] Generation of Scaled-Up Real-Time (or Near Real-Time) Aggregations In some examples, when the system 10 receives a large amount of data, the system 10 aggregates the data within the fields in real-time (or near real-time) upon receipt of the data so that the system 10 can perform aggregations without significant latency, and further scales up to aggregate the data. In some examples, these aggregations are used when generating data that is accessed or retrieved when the system performs segmentation.

[0080] Next, referring to FIG. 6, the networked system 10 (FIG. 1) also includes a system 50 that generates real-time aggregations. In some examples, system 50 is the execution system 14 in FIG. 1. In this example, system 50 receives data records 52a-52c from data source 12a and data records 52d-52f from data source 12b. Each of data records 52a-52f includes one or more fields, such as a key field (i.e., a subscriber identifier (SubID) field) having a value that uniquely identifies a user. Each of data records 52a-52f may also include a communication type ("communication type") field for storing a value identifying the type of communication (i.e., voice, SMS, or data).

[0081] The networked system 10 further includes an execution system 14 (FIG. 1) and a storage area including a memory 56 (e.g., shared memory, semiconductor memory, low-persistence memory, etc.) and a persistent memory 58. The memories 56 and 58 may form the memory 16 of FIG. 1. Generally, the memory 56 includes a memory that can be accessed by the system 50 while reducing latency as compared to, for example, the latency when acquiring data or data records from the persistent memory 58. Since the memory 56 is not a disk memory (e.g., data records are not stored on a disk when stored in the memory 56), the latency of the memory 56 is low. In some examples, the memory 56 includes a memory cache, sometimes referred to as a cache store or a RAM cache, which is a portion of the memory made of fast static RAM (SRAM) rather than slow dynamic RAM (DRAM) used for the main memory, e.g., the persistent memory 58. In this example, the memory 56 stores only recent data (or records of the occurrence of recent data), and the data is "recent" if it has been received within a threshold period (e.g., less than 14 days). When the data becomes older than this threshold period, the system 50 or the memory 56 transmits the data to the persistent memory 58 for more persistent storage. Since the system 50 most frequently accesses more recent data, memory caching on the memory 56 is effective. That is, the data stored in the memory 56 is data frequently used by the system 50. By keeping this information in SRAM or the memory 56 as much as possible, the system 50 avoids accessing the slow DRAM or the persistent memory 58.

[0082] In this example, the system 50 stores records of events rather than the received data or the data records themselves. Generally, an event includes the occurrence of a particular value in a particular field. In this example, the system 50 designates that each possible value (i.e., voice, SMS, or data) in the "communication type" field is an event. The memory 56 stores data records 60 that store records of individual detected events (and the subscriber ID associated with that event). Specifically, the data record 60 includes columns 60a - 60c. Column 60a stores data indicating the occurrence of a "voice event" (detection of the "voice" value in the "communication type" field). Column 60b stores data indicating the occurrence of a "data event" (detection of the "data" value in the "communication type" field). Column 60c stores data indicating the occurrence of an "SMS event" (detection of the "SMS" value in the "communication type" field).

[0083] Specifically, the system 50 receives the data record 52a and detects the occurrence of a voice event within the data record 52a. Therefore, the system 50 inserts the value of the subscriber ID into column 60a of the data record 60. The system 50 receives the data record 52b and detects the occurrence of an SMS event within the data record 52b. Therefore, the system 50 inserts the subscriber ID specified within the SubID field in the data record 52b into column 60c.

[0084] System 50 receives data record 52c and detects the occurrence of an audio event within data record 52c. To that end, System 50 inserts the subscriber ID specified within the SubID field in data record 52c into column 60a. System 50 receives data record 52d and detects the occurrence of an audio event within data record 52d. To that end, System 50 inserts the value of the subscriber ID specified within the SubID field in data record 52d into column 60a. System 50 receives data record 52e and detects the occurrence of an audio event within data record 52e. To that end, System 50 inserts the value of the subscriber ID specified within the SubID field in data record 52e into column 60a. System 50 receives data record 52f and detects the occurrence of a data event within data record 52f. To that end, System 50 inserts the value of the subscriber ID specified within the SubID field in data record 52f into column 60b.

[0085] In this example, data record 60 is a data record with increased flexibility, as it can be modified to track the occurrence of other types of events (such as video conference events) by adding another column to data record 60. Therefore, data record 60 can be modified during execution to track the aggregation of new events. This provides an improvement in flexibility beyond simply storing the received data record itself, as the event can be tracked by generating a new composite key for the new event, as will be explained in more detail below. Additionally, querying data record 60 results in a reduction in latency when executing the query compared to the latency when executing a query against individual data records. For example, system 50 can query data record 60 for subscribers involved in voice communication. In this example, system 50 generates a query of "communication type = voice". Based on this query, memory 56 returns the values within column 60a that simply search for the values of the subscriber IDs contained within column 60a. System 50 returns the results of this query more quickly (compared to the speed required to search individual data records 52a - f to identify data records that satisfy the query) because system 50 (or memory 56) only needs to identify the columns that match or satisfy the query, rather than exhaustively searching the data records to identify the record that stores the value that satisfies the query. In some examples, after the expiration of a threshold time, the data within data record 60 is transferred to persistent memory 58 and stored in one of data records 62a...62n.

[0086] In some examples, since each of columns 60a - 60c represents an aggregation of a particular type of event, the data contained within those columns are each referred to as in - memory aggregations. Generally, in - memory aggregations (e.g., counts, averages, etc.) include aggregations of data stored in memory 56. In other examples, system 50 may perform operations on the data contained within record 60 to generate in - memory aggregations. For example, system 50 can query memory 56 for the count of records where "communication type = voice" and "SubID = 53054423". In this example, since column 60a indicates that a subscriber with "SubID = 53054423" had two voice communications, memory 56 will return a value of "2". In this example, memory 56 generates an in - memory aggregation for the query, and the in - memory aggregation has a value of 2. Memory 56 (or system 50) stores the value of the in - memory aggregation in a shared variable. In this example, upon receiving a query for "count of communication type = voice and SubID = 53054423", memory 56 generates a shared variable with the name "count of communication type = voice and SubID = 53054423" and sets the value of the shared variable to be "2". In this example, the shared variable stores the value of the in - memory aggregation. As described above, these in - memory aggregations are retrieved from memory 56 more quickly (compared to the speed of retrieving these in - memory aggregations from persistent memory 58 or constructing these in - memory aggregations by exhaustively searching individual data records 52a - f).

[0087] When the data in data record 60 is moved to disk (i.e., to persistent memory 58), the values in columns 60a - 60c are not stored in memory but are on - disk aggregations, including records of occurrences stored, for example, on disk. In some examples, after the elapse of a threshold time, the value of the shared variable is also moved to persistent memory 58.

[0088] Rather than storing the data records themselves, by recording the occurrence of events, system 50 can scale up and make aggregation decisions such that there is no increase (or only a minimal increase) in latency when the number of records represented within data record 60 increases. This is because system 50 only needs to identify the relevant fields (or relevant cells within a column) within data record 60, rather than painstakingly parsing each individual record 52a - f and identifying its content. Since the number of fields does not increase as the number of occurrences within a record increases, identifying the relevant fields within a data record is a scalable process. Therefore, identifying these real - time aggregations is scalable and does not introduce latency as the number of data records to be processed increases.

[0089] As will be described in more detail below, in a modified form, memory 56 stores a hash table, and within the hash table, hash values of composite keys are stored in relation to composite values. Generally, a composite key includes a key assembled from (or including) a plurality of distinct values. Generally, a composite value includes a concatenation or assembly of a plurality of distinct values.

[0090] Referring to FIG. 7, the selection and modification functions executed by the system 10 (e.g., by module 18 of FIG. 1) are shown. In this example, the system causes the rendering of various user interfaces where one or more data sources are selected, one of more data structures are selected, and one or more fields of those data structures are modified. In this example, data sources 72a - 72d (where data sources 72a - 72c respectively correspond to data sources 12a - 12c) are candidate data sources from which data is provided by a rendering module (e.g., rendering module 20 of FIG. 1). The rendering module 20 (FIG. 1) provides various graphical user interfaces for an end user (e.g., a business user) to view and access a curated subset of data. From the curated subset, the system generates instructions for performing various operations and actions, for example, based on data received by the graphical user interface. The subset of data is curated from a superset of data across data sources (e.g., data sources 72a - 72d) into a subset for a specified operation such as a segmentation operation. The rendering module also provides a graphical user interface for receiving instructions on how to curate the data. In this example, data sources 72a - 72d are the superset of data that is the source, e.g., the curation source, for generating the subset. The system selects data sources 72b, 72d as the original data sources for modifying various data structures (e.g., tables). In some examples, based on the selection of data sources 14, 18, the system identifies references for data sources 72a, 72d and searches in memory 16 (FIG. 1) to determine which tables are associated with those identified references. The system then retrieves those tables from data sources 72b, 72d or from memory 16 if memory 16 itself stores the tables.

[0091] In this example, data source 72a includes tables 73, 74, and 75. Data source 72d includes tables 78, 79, 80, and 81. From tables 73, 74, and 75, the system selects table 26 as the data structure to be modified (e.g., curated) by rendering the visual representation of table 26 by rendering module 20 (FIG. 1). From tables 78, 79, 80, and 81, the system selects table 32 as the data structure to be modified (e.g., curated) by rendering the visual representation of table 32 by rendering module 20 (FIG. 1). These selections are made, for example, in accordance with user instructions received by the user interface to select tables 26, 32. That is, not all of tables 73, 74, 75, 78, 79, 80, and 81 are modified and curated. Only the selected tables from the selected data sources are modified and provided within user interface 24 (FIG. 1) by rendering module 20.

[0092] View 75a of table 75 shows the content of table 75. Without limitation and for convenience, in this specification, view 75a and table 75 may be collectively referred to as "table 75". Table 75 includes a title portion 75g that specifies the title "Plan Statistics". In this example, table 75 includes columns 75b, 75c, 75d (also referred to herein as "fields 75b, 75c, 75d" respectively). The names of fields 75b, 75c, 75d are "sub_id", "min_usd", and "prc_pln_id" respectively. Table 75 also includes rows 75e, 75f. In one example, table 75 (or a visual representation of table 75) is rendered within the user interface in order to be able to modify and / or rename titles and / or fields, and also in order to be able to specify one of the fields as a key, for example when joining fields of various tables. Table 76 is a modified version of table 75. Table 76 is rendered on client device 85a based on receiving graphical user interface data from execution system 14, and the graphical user interface data specifies the content of table 75. In this modified version of table 75, the original title specified within portion 75g has been modified to the title "Plan Statistics" as specified within portion 76a. Additionally, each of fields 75b, 75c, 75d has been renamed to a more descriptive name (for example, a name that is more meaningful to business users). As shown within field 76b, in this example, the name of field 75b has been renamed to "Subscriber ID". Generally, a subscriber includes a user of the system and is identified by a key also called a subscriber ID. As shown by icon 76c, in this example, field 76b has been selected as the key. As specified by field 76e, the name of field 75c has been renamed to "Usage Fraction". As specified within field 76f, the name of field 75d has been renamed to "Rate Plan Name". Table 76 also includes rows 76g, 76h, and the content of each of them corresponds to rows 75e, 75f respectively.In this example, table 76 is presented to the end user by the system. To enable viewing and / or selection of data provided from data source 14, in this example, only table 76 is presented to the end user (tables 73, 74, 75 are not presented). Table 76 (or a visual representation of table 76 (not shown)) represents a curated version or curated configuration of the data from data source 14. In some examples, a curated version of table 75 (e.g., table 76) may include only a subset of the fields within table 75. For example, fields 75c or 75d may be removed and not included within table 76. In another example, the user may select rows within table 75 (or table 76) to be pivot rows, such as if they are multiple rows regarding a particular subscriber ID.

[0093] View 83 of table 81 shows the content of table 81. Without limitation and for convenience, in this specification, view 83 and table 81 may be collectively referred to as "table 81". Table 81 includes a title portion 83a that specifies the title of "Bndld_vc_data". In this example, table 81 includes columns 81b, 81c (also referred to herein as "fields 81b, 81c" respectively). The names of fields 81b, 81c are "sub_id" and "sub_bndledvd" respectively. Table 81 also includes rows 81d, 81e. In one example, table 81 is rendered within a user interface to enable modification and / or renaming of titles and / or fields and / or to enable designation of one of the fields as a key, for example, when combining fields of various tables. Table 83 is a modified version of table 81. Table 83 is rendered on client device 85b based on receipt of graphical user interface data from execution system 14, and the graphical user interface data specifies the content of table 81. In this modified version, the original title specified within portion 81a has been modified to the title of "Audio and Data Bundle" as specified within portion 83g. Additionally, each of fields 81b, 81c has been renamed to a more descriptive name (e.g., a name that is more meaningful to a business user). As shown within field 83a, in this example, the name of field 81b has been renamed to "Subscriber ID". As indicated by icon 83b, in this example, field 83a has been selected as the key.

[0094] As specified by field 83c, the name of field 81c has been renamed to "Voice and Data Bundle". Table 83 also includes rows 83d, 83e, and the content of each of them corresponds to rows 81d, 81e respectively. In this example, table 83 is presented to the end user by the system. In order to enable viewing and / or selection of data provided from data source 72d, in this example, only table 83 is presented to the end user (tables 78, 79, 80, 81 are not presented). Table 83 represents a curated version or curated configuration of data from data source 72d.

[0095] Referring to FIG. 8A, for example, a graphical user interface 90 is rendered by the system to enable selection of one or more data sources to modify their data. In this example, graphical user interface 90 is one of the graphical user interfaces rendered by rendering module 20 (FIG. 1). Graphical user interface 90 includes a menu portion 92 having a control 92a, and selection of control 92a causes visual representations 93a - 93d of available data sources to be displayed in portion 94 of graphical user interface 90. In this example, visual representations 93a - 93d represent data sources 72a - 72d (FIG. 7) respectively. As indicated by visual representations 95a, 95d alongside each of visual representations 93a and 93d, as the data source from which to select the data structure to be modified, the user selects data sources 72a, 72d (FIG. 7) or 12, 14 (FIG. 2) via graphical user interface 90. In this example, column 96 includes selectable controls (not shown), and selecting such a control causes the display of a visual representation such as one of visual representations 93a and 93d.

[0096] Referring to FIG. 8B, a developed state of a graphical user interface 90 is shown that enables selection of one or more data structures (e.g., tables) to be modified (from the selected data source). In this example, the graphical user interface 90 displays a menu portion 92 having a control 92a (which may correspond to the menu portion 92 of FIG. 8A), and selection of the control 92a causes the data structures contained within the selected data source 72a, 72d (FIG. 7) to be displayed in the graphical user interface 90 in portion 94 for each selected data source. For example, portion 94 displays a table 96 including columns 95a - 95j. Column 95a displays visual representations 96a - 96g having the name of the selected data source as selected, for example, in FIG. 3 or FIG. 7. The visual representations 96a - 96g represent the selected data source 72d according to, for example, the visual representation 95d (FIG. 8A) that designates data source 72d as the data source for which the table is selected for curation. The visual representations 96a - 96g represent the data source 72d according to, for example, the visual representation 58 (FIG. 8A) that designates data source 72d as the data source for which the table is selected for curation. Column 95b displays visual representations 97a - 97g, each of which represents the name of the table within the corresponding data source represented by the visual representations 96a - 96g. Column 95c displays the data type for each table represented within column 95b. Column 95d provides a control for entering an explanation for each table represented within column 95a. Column 95e displays selectable controls, and selection thereof designates the table for which data and / or data structures are selected for modification (e.g., by display of visual representations such as visual representations 95e', 95e''). In this example, visual representations 95e', 95e'' are displayed within column 95e to designate that the tables represented by visual representations 97c, 97f respectively are the tables selected for modification.

[0097] Table 96 also includes columns 95f - 95h that respectively specify the specific functions that can be selected and applied. Table 96 includes a "Last Update" column 95i that displays, for each row in Table 96, the table represented by that row, the user who last updated it, and data specifying when the update was made. Table 96 also includes an edit column 95j that represents controls (e.g., controls 95j', 95j'') for each row in Table 96. For a specific row specified by column 95e as being editable, selecting the control displayed within the edit column 95j for that row enables that table to be edited and modified. For example, selecting control 95j' enables the table represented by visual representation 97c to be modified. In this example, visual representation 97c represents Table 75 (Figure 7), and Table 75 is modified as previously described. Selecting control 95j'' enables the table represented by visual representation 97f to be modified. In this example, visual representation 97f represents Table 81 (Figure 7), and Table 81 is modified as previously described.

[0098] Referring to Figure 8C, a graphical user interface 150 is displayed (e.g., by the rendering module 20 of Figure 1) to enable user access, viewing, and generation of instructions from a modified data structure by enabling the user to generate instructions for performing, for example, segmentation. The graphical user interface 150 includes a menu portion 152 that includes controls, the selection of which enables the generation of various types of instructions. In this example, menu portion 152 includes control 154, the selection of which enables the specification of one or more fields of one or more of the selected data structures, such as 72a, 72d, that are the basis for selecting (and / or operating on) data records.

[0099] The graphical user interface 150 includes an editor interface 156 for segmenting data and specifying various other operations such as filtering, joining, etc. When a control 154 is selected, a component 158 is displayed within the editor interface 156. Generally, a component represents executable logic (or instructions) such as segmentation logic for performing various operations. The component receives inputs such as selected data or other input data (by the editor interface 156) and the system uses the received inputs when generating the executable logic. In one example, the system stores a pre-configured mapping between the executable logic and the components. Then, based on the input or as specified by the component, the executable logic (for that component) is updated or modified to include the input. The component 158 enables the specification of instructions for selecting specific tables such as 97a - 97g (e.g., curated tables). The component 158 includes an icon 160, and its selection enables the selection of a specific curated table.

[0100] Referring to FIG. 8D, a graphical user interface 170 is displayed. The graphical user interface 170 is an updated version of the graphical user interface 150 (FIG. 8C) that is updated, for example, after the selection of the icon 160. When the icon 160 is selected, an overlay portion 172 is rendered within the graphical user interface 170. The overlay portion 172 includes a visual representation 172a of (the modified and curated table 83 (FIG. 2), 97f (FIG. 8B)) and a visual representation 172b of (the modified and curated table 76 (FIG. 7)). In this example, the visual representation 172a is selected.

[0101] Referring to FIG. 8E, a graphical user interface 174 is displayed. The graphical user interface 174 is an updated version of the graphical user interface 170 (FIG. 8D). When the visual representation 172a is selected, an overlay 176 is displayed. The overlay 176 displays the contents of Table 38 and includes a selectable portion 176a, and the selection thereof enables a user to select a field (i.e., "audio and data bundle field") represented within the selectable portion 176a for inclusion within the executable logic represented by the component 158. The overlay 176 also includes a portion 176b representing a subscriber ID field. In this example, since it is necessary to associate the value within the field represented within the selectable portion 176a with the subscriber ID in order to assign the value to an appropriate subscriber, the selection of the selectable portion 176a automatically causes the selection of the portion 176b. In some examples, the executable logic represented by the component 158 is modified or updated according to the field selected by selecting the selectable portion 176a.

[0102] Referring to FIG. 8F, a graphical user interface 180 is displayed. The graphical user interface 180 is an updated version of the graphical user interface 174 (FIG. 8E). In this example, the component 158 is updated with parts 158a and 158b. Part 158a specifies a title for the component 158, and the title is based on the selection of a field represented within the selectable part 176a in FIG. 8E. Part 158b specifies that executable logic (represented by the component 158) is configured to select data records in which the value of the "voice and data bundle" field is equal to "1" from, for example, the data records received or stored by the system. In a more general expression, part 158b specifies the segmentation of the received or stored data records by identifying which of the received data records have one or more fields corresponding to one or more fields represented within the selected one or more selectable parts. Generally, segmentation includes assigning data records to a specified group, dividing data records, and / or excluding data records from a specified set and including data records in other sets. This enables the user to start the segmentation of data records "during execution" by means of an editor interface. Thus, the representation displayed by the editor interface provides a graphical shortcut for setting the conditions for the segmentation process, which results in a more efficient (in terms of time and resources) segmentation of the input data records compared to, for example, stepwise defining those conditions by means of condition type input and the like. Overall, the user is provided with more direct, less error-prone, and faster control over the segmentation of the input data records. In the example of FIG. 8F, when the selectable part 176a (FIG. 8E) is selected, part 158b is automatically populated with the string "voice and data bundle subscriber = ____". In this example, the graphical user interface 180 displays a prompt (not shown) prompting the user to fill in the value of "1" or "0" for the empty field "____" within the aforementioned string.In this example, the user selects a value of "1". In this example, the subscriber referred to within the aforementioned string is represented by the subscriber ID field represented within part 176b (Figure 7).

[0103] In this example, the user also selects control 154 to cause editor interface 156 to add component 182. Component 182 enables specifying instructions for selecting one or more fields from another specific table (e.g., a curated table). Component 182 includes icon 182a, and its selection enables selection of a specific curated table and / or fields from a specific curated table.

[0104] Referring to Figure 8G, graphical user interface 184 is displayed. Graphical user interface 184 is an updated version of graphical user interface 180 (Figure 8F) that is updated, for example, after the selection of icon 182a. When icon 182a is selected, overlay portion 186 is rendered within graphical user interface 184. Overlay portion 186 includes a visual representation 186a (of the revised and curated table 83 (Figure 7)) and a visual representation 186b (of the revised and curated table 76 (Figure 7)). In this example, visual representation 186b is selected.

[0105] Referring to FIG. 8H, a graphical user interface 188 is displayed. The graphical user interface 188 is an updated version of the graphical user interface 184 (FIG. 8G). When the visual representation 186b (FIG. 8G) is selected, an overlay 190 is displayed. The overlay 190 displays the content of the table 76 (FIG. 7) and includes selectable portions 190a, 190b, 190c, the selection of which enables the user to select a subscriber ID field, a usage fraction field, and a rate plan name field, respectively, for inclusion within the executable logic represented by the component 182. In this example, since it is necessary to associate the values within the fields represented within the selectable portions 190a, 190b with a subscriber ID in order to attribute the values to an appropriate subscriber (e.g., user), one or more selections of the selectable portions 190b, 190c automatically cause a selection of the selectable portion 190a. In some examples, the executable logic represented by the component 182 is modified or updated according to the fields selected by selecting the selectable portion 190b, which is the selectable portion selected in this example.

[0106] Referring to FIG. 8I, a graphical user interface 192 is displayed. The graphical user interface 192 is an updated version of the graphical user interface 188 (FIG. 8H). In this example, component 182 is updated with parts 182b and 182c. Part 182c designates a title for component 182, and the title is based on the selection of a field represented within selectable part 190b in FIG. 8H. Part 182b designates that executable logic (represented by component 182) is configured to select a data record including the "usage fraction" field specified within selectable part 190b (FIG. 8H) from, for example, data records received or stored by the system. In this example, the "usage fraction" field is a curated field and does not actually match the name of the field within the actual data record. Therefore, the system stores a copy of table 75 of FIG. 7 (including the actual names of the fields within the data record itself) and table 76 of FIG. 7 (including the curated names of the fields) as well as the mapping between each of the field names in table 75 and the field names in table 76 for the field names within table 76. Based on this mapping, the system searches for the actual field name regarding the curated field name. For example, the system uses this mapping to identify that the "usage fraction" field referred to within part 182b is actually the "min_used" field within the data record.

[0107] Referring to FIG. 8J, the graphical user interface 194 shows that the editor interface is updated by the joining component 200 after the selection of the control 198. In this example, the joining component 200 represents executable logic for joining two separate data streams or sets of data records. In this example, the joining component 200 represents the joining of the output of component 158 (which is a data record where the value of the "voice and data bundle" field is equal to "1") and the output of component 182 (which is a data record having a value within the "usage fraction by subscriber" field). In a modified form, the output of component 158, when the value is equal to "1", is the value of the "voice and data bundle" field (i.e., field 81c or field 82c of FIG. 7) and the associated value of the "subscriber ID" field (i.e., field 83a or field 81b of FIG. 7). In this modified form, the output of component 182 is the value of the "usage fraction" field (field 76c or field 75c of FIG. 7) and the associated value of the "subscriber ID" field (i.e., field 76b or field 75b of FIG. 7). The executable logic represented by the joining component 200 joins the outputs of components 158 and 182 for each subscriber ID. In this example, to specify that the outputs of components 158 and 182 are input into the joining component 200 to join the outputs of components 158 and 182, the connectors 201 and 203 can be selected from the menu portion 152, for example, by selecting the control 197.

[0108] In this example, by selecting the save control 196, the editor interface 156 displays the definition of a specific segment that can be saved (e.g., for future use), and the selection of the save control 196 prompts the user to enter a name for the segment so that the segment can be retrieved later by name.

[0109] Referring to FIG. 8K, a graphical user interface 206 is displayed to enable the creation of predefined data aggregations (also referred to as "entities"). Specifically, based on instructions received by, for example, the user interface 206, the system combines data from different tables to generate these entities, and the generated entities can be utilized (e.g., via the system) when causing segmentation. These entities group various data into specified categories.

[0110] The graphical user interface 206 includes a menu portion 207 having a subscriber control 207a, and selection of the subscriber control 207a causes the display of an entity portion 208. The entity portion 208 displays the defined entities and includes a generation control 205, and selection of the generation control 205 causes the display of a series of prompts and controls for the user to define a new entity. In this example, the entity portion displays a visual representation 209 of predefined entities that specify aggregations of data related to a handset or device associated with a particular key (e.g., representing a subscriber).

[0111] Referring to FIG. 8L, for example, when the generation control 205 (FIG. 8K) is selected, the graphical user interface 210 is displayed as an overlay on the graphical user interface 206 (FIG. 8K). The graphical user interface 210 enables the generation of new entities and includes a name portion 210a for entering information specifying the name of the entity being created, a prefix portion 210b for entering a prefix specifying the practical or database field name of the entity to be generated, and a description portion 210c for entering information specifying the description of the entity being created. The graphical user interface 210 also includes a field portion 212 for specifying fields 212a-212j included within the entity defined within the graphical user interface 210. In this example, the fields 212a-212j are fields selected from various data sources, such as the tables included within the data sources 12a-12c of FIG. 1. In this example, the field portion 212 displays both the original field name (e.g., from an uncurated table) within column 212k and the modified or curated field name within column 212l. In this example, the field 212h corresponds to the field 75d of FIG. 2, and the field 212m is a modified or curated version of the field 212h. For example, to show how the field portion 212 enables selection of fields by both the original field name of the field and the modified field name of the field, the field 212m corresponds to the field 76f of FIG. 7. The field portion 212 also includes a search control 212n for entering a search query or one or more search terms. The system uses the input of the search control 212n to search for field names (from tables within the data source) that match or correspond to the search criteria. When appropriate fields are selected, those fields are displayed within the field portion 212.

[0112] Referring to FIG. 8M, the graphical user interface 213 is a modified version of the graphical user interface 206 (FIG. 8K). In this example, the entity portion 208 is updated to display a visual representation 214 of an entity (hereinafter referred to as the "subscriber enrichment entity") created according to the specifications and selections made within the graphical user interface 210 (FIG. 8L). As will be described in more detail below, in this example, when defining segments, the user is provided with the entities represented within the visual representations 209, 214.

[0113] Referring to FIG. 8N, the graphical user interface 216 is displayed, and the graphical user interface 216 is the same graphical user interface as the graphical user interface 194 (FIG. 8J). For example, in order to access a predefined entity represented within the entity portion 208 (FIG. 15), the control 218 is selected from the menu portion 152.

[0114] Referring to FIG. 8O, the graphical user interface 220 is displayed, and the graphical user interface 220 is a modified version of the graphical user interface 216 (FIG. 8N), and an overlay portion 222 is rendered within the editor interface 156. The overlay portion 222 includes a visual representation 222a of a subscriber enrichment entity (defined within the graphical user interface 210 of FIG. 8L) and a visual representation 122b of a handset entity (represented by the visual representation 209 of FIG. 8K). For example, in this example, the visual representation 222a is selected to add subscribers who meet certain criteria specified by the subscriber enrichment entity to the segments defined within the editor interface 156.

[0115] Referring to FIG. 8P, a graphical user interface 224 is displayed. The graphical user interface 224 is a modified version of the graphical user interface 220 (FIG. 8O), and an entity component 226 is added to the editor interface 156. In this example, the entity component 226 represents executable logic that obtains at runtime a data record having specified fields (e.g., fields 212k to 212j in FIG. 8M). In a modified form, the entity component 226 represents executable logic that obtains at runtime the values of fields 212k to 212j (FIG. 8M) and the values of keys associated with those fields. In this example, profile data is output from the entity component 226 and is combined with the outputs of components 158 and 182 via a connector 228 and coupled to a combining component 200. By adding the entity component 226 to the editor interface 156, a segment 230 is defined to include only subscribers specified by the entity component 226, subscribers having a voice and data bundle plan (specified by component 158), and subscribers having a usage fraction (specified by component 183). In this example, the segment 230 is saved (for later use) by selecting a save control 196. In this example, selecting the save control 196 causes the executable logic represented by components 158, 182, 200, 226 and connectors 201, 203, 228 to be saved (e.g., in a data structure) for later retrieval, for example when further defining other segments. In this example, the execution system 14 (FIG. 1) can execute the segment 230 to segment a plurality of data records stored in the memory 16 (FIG. 1) such that only data records having fields and / or field values that satisfy the definition specified by, for example, the segment 230 are included. In this example, access to the segment 230 can be obtained by a segment control 232 that enables selection of the segment component.The segment component enables the use of segments that are already defined as sources for data selection when constructing segments.

[0116] In this example, segment 230 is an executable data flow graph executed by the segmentation module 22 of FIG. 1 (or the segmentation module 22 of FIG. 9) for performing segmentation. Generally, an executable data flow graph includes a directed data flow graph, where the vertices in the graph represent components (data files or processes), and the links or "edges" in the graph indicate the data flow between components. A system for performing such graph-based computations and executing data flow graphs is described in the past U.S. Patent No. 5,966,072 entitled "EXECUTING COMPUTATIONS EXPRESSED AS GRAPHS", which is incorporated herein by reference. By performing segmentation as a data flow graph executed on system 14 (FIG. 1), the system resources and memory 16 (FIG. 1) of system 14 are freed because the system does not directly execute functions such as sorting, filtering, and joining of data and data records in memory 16 (FIG. 1) or other data storage devices, but rather performs such functions (by executing the data flow graph). That is, compared to the processing power and speed of memory 16 when the operations required for segmentation are directly executed within memory 16, the processing power and speed of memory 16 (FIG. 1) are increased by executing the data flow graph to perform the segmentation function (rather than performing the segmentation function by directly operating on the data in memory).

[0117] In some examples, the systems described herein identify segments for various campaigns. Generally, a campaign is a definition of a selected offer for delivery to a specified user at a specified time, for example. For instance, if the executable logic of a dataflow graph specifies how to determine which offers to send to segments assigned to a campaign, the dataflow graph can define the campaign. A campaign may be assigned a type of offer that includes instructions to specify or limit membership in the campaign. For example, there are various types of offers including "none" type, "global" type, "voice" type, "data" type, "SMS" type, "package" type, "reload" type, etc. A campaign with a none offer type places no restrictions on subscribers. A campaign with a global offer type specifies that a subscriber can be included in that campaign and not in other campaigns simultaneously. A campaign with a voice offer type specifies that a subscriber must be included in only a single campaign of the voice offer type at a time. A campaign with a data offer type specifies that a subscriber must be included in only a single campaign of the data offer type at a time. A campaign with an SMS offer type specifies that a subscriber must be included in only a single campaign of the SMS offer type at a time. A campaign with a package offer type specifies that a subscriber must be included in only a single campaign of the package offer type at a time. A campaign with a reload offer type specifies that a subscriber must be included in only a single campaign of the reload offer type at a time. When configuring a campaign, a user can select whether to release a subscriber from a campaign offer when the campaign cycle ends.

[0118] This system also assigns to each campaign a theme that represents the campaign's goal (e.g., a data structure that stores data representing the theme). Campaign themes have priorities. As will be described below, the system uses this priority during the mediation of campaigns along with the contact policy.

[0119] At any given time, a subscriber may acquire the rights to one or more campaigns. It is important not to send an undesired large number of messages to the subscriber with too many offers, and to focus on the most important offers to the subscriber, which in some cases is regulated by the regulatory authority. The contact policy manages the frequency with which the system can communicate with the subscriber during a campaign. The contact policy sets a limit on the number of outbound offers that the system can transmit. When a subscriber has the rights to multiple campaigns, the system needs to extend those offers according to the contact policy and the relative priorities of the various campaigns in progress, such that when an offer can be extended to obtain a higher-priority campaign, the offer is not extended to obtain a lower-priority campaign. For example, a subscriber may be at a particular stage of a higher-priority campaign but begin to fall into a low-engagement segment. As will be described below, the system executes a campaign mediation order to ensure that a lower-priority offer is not assigned to a subscriber when a higher-priority offer should be assigned to the subscriber.

[0120] When the contact policy restricts the number of offers that can be sent to a target customer, the system executes a campaign mediation order to select the best offers based on the campaign theme and priority. For example, if the contact policy stipulates that only two offers can be sent to a customer per day and the customer has the rights to five offers, the system selects the top two offers based on the campaign theme and priority (e.g., the system selects two offers each related to one of the campaigns having the top two priorities relative to the priorities of the other campaigns).

[0121] For example, when configuring thousands of campaigns, the system may spend the limit of the contact policy it gives to high-priority campaigns on campaigns that are scheduled to be executed earlier in the day. To solve this problem, the system uses the priority of the campaign theme to secure communication slots for subscribers using the communication subsystem included in the system. Generally, the system stores a data structure or queue for each subscriber, inserts data into an entry in the data structure using data representing when a message is transmitted to the subscriber, and / or inserts data into an entry in the data structure using data that secures that entry for a specific campaign. When it is time to send a message, the system checks the contact policy and data structure for that subscriber so that the low-priority campaigns do not spend the limit of the contact policy required for that day.

[0122] Referring to FIG. 8Q, a graphical user interface 270 is displayed that enables modification of a subscriber enrichment entity defined within the graphical user interface 210 (FIG. 8L). In this example, for instance, a creation control 272 is selected to add one or more additional fields to the definition of the subscriber enrichment entity. In this example, when the creation control 272 is selected, one or more in-memory aggregations are added to the subscriber enrichment entity.

[0123] Referring to FIG. 8R, for example, an overlay 274 is shown as an overlay to the graphical user interface 270 (FIG. 8Q). In this example, the overlay 274 displays various types of in-memory aggregations 274a, 274b, 274c, 274d that can be added to include within the subscriber enrichment entity. In this example, one of the in-memory aggregations 274d (i.e., in-memory aggregation 274d’) is selected to be included within the subscriber enrichment entity. In this example, the in-memory aggregation 274d’ is an in-memory aggregation that specifies a voice usage count (e.g., the number of voice calls made by the user).

[0124] Referring to FIG. 8S, the graphical user interface 276 is an updated version of the graphical user interface 270 (FIG. 8Q), and the field portion 212 is updated to include a field 278 in accordance with the selection of the in-memory aggregation 274d’ of FIG. 21. The field 278 stores the value of the in-memory aggregation 274d’, and is thus a field for adding the in-memory aggregation 274d’ to the subscriber enrichment entity. In this example, the in-memory aggregation 274d’ is an in-memory aggregation of voice usage count that provides a count of the voice call volume per key.

[0125] Referring to FIG. 8T, the graphical user interface 280 is an updated version of the graphical user interface 224 (FIG. 8N), and for example, after the selection of the filter control 284, the filter component 282 is added to the editor interface 156. The filter component represents executable logic for filtering data against the flow of data according to specified rules. As specified by the connector 286, the profile data output from the entity component 226 is input into the filter component. In this example, the profile data output from the entity component 226 includes values for in-memory aggregation of voice usage counts (e.g., for each key value). The filter component 282 is configured to exclude data records from the data records included in the output profile data in which the value of the in-memory aggregation of the voice usage count is greater than or equal to the value of 2. In this example, only the received data records that specify that the subscriber has made less than two voice calls pass through the logic represented by the filter component 282 because those data records are the only received data records that specify that the subscriber has made less than two voice calls. The data records that pass through the criteria or logic represented by the filter component 282 are the filtered data, and they are input into the combining component 200 as specified by the connector 288. The content of the editor interface 156 defines a new segment, namely segment 290.

[0126] Referring to FIG. 9, the execution environment 300 uses a data source 303 (for receiving data records, for example, whose fields are curated or modified as described above) and a data source 302 (for receiving data records to be processed), and includes a system 314 for implementing collection-detection-action (CDA). Generally, CDA refers to the process of a system 314 that collects data records, processes those data records to detect which data records contain values that meet specified criteria, and executes one or more actions with respect to those data records. This environment also includes a pre-execution module 14' (similar to the execution module 14), which includes a data structure modification module 18 (FIG. 1) for selecting one or more data sources from which data is provided and for selecting one or more data structures within those one or more selected data sources to modify one or more data structures (for example, by modifying field names) as discussed above. As discussed in FIG. 1, the pre-execution module 14' also includes a rendering module 20 (FIG. 1) for rendering a visual representation of the modified data structure within a user interface (displayed on the client device 301).

[0127] Referring to FIG. 10, system 314 includes a collection module 316 that collects data records (e.g., data records 302a - c received from data source 302 and other data), converts the data within data records 302a - 302c into converted data 306, and distributes the converted data 306 to downstream applications including, for example, segmentation module 326, detection module 318, and operation module 320. The content of data record 302a includes "ID: 34213, Event Type: Voice, Date: 4 / 3 / 2018". The content of data record 302b is "ID: 34214, Event Type: SMS, Date: 4 / 3 / 2018". The content of data record 302c is "ID: 34215, Plant Type: Bundle Plan, Date: 4 / 3 / 2018". In this example, operation module 320 includes an interface to a third - party system and / or an external system. Specifically, collection module 316 is located at various sources or various positions, such as data source 302, and collects batch or real - time data and real - time data streams from various servers interconnected by a network, for example, real - time data arriving from various servers located at various positions and interconnected by a network. The storage device providing data source 302 can be local to system 314, for example, stored on a storage medium (e.g., hard drive) connected to the computer executing system 314, or can be remote to system 314, for example, hosted on a remote system that communicates with system 314 over a local area data network or a wide - area data network.

[0128] The collection module 316 records the records of each event occurring within the received data record in the memory 322. As previously explained, the collection module 316 records these occurrences within the data record or within a table. In this example, the collection module 316 records the occurrence of the received event in a hash table within the memory 322. For each received record, the collection module 316 generates (i) a composite key value by generating a composite key, which is later hashed to generate a hashed composite key as shown in column 346a, and (ii) a composite value as shown in column 346b. In this example, the composite value is generated from the data contained within the fields in the received data record that represent aggregations (e.g., count, minimum value, average value, maximum value, etc.) and / or data stored previously (e.g., stored in the memory 322 or the memory 324). In this example, an entry 346c within the hash table 346 is generated from the received record 302a. To generate the entry 346c, the collection module 316 hashes the value of the ID field in the record 302a ( "34213") to generate a hash value of "0111", which is stored in column 346a of the entry 346c. In this example, the collection module 316 includes the value of "voice" of the event field in the record 302a in the composite value column 346b of the entry 346c. The collection module 316 also includes other aggregations (e.g., obtained from the memory 322 or the persistent memory 324) in the composite value column 346b of the entry 346c for the ID of 34213. In this example, the count of the number of SMS messages used over a specified period (e.g., the past 30 days) ( "4") is represented by the obtained aggregated data. In this example, the hash table 346 also includes entries 346d, 346e having the illustrated values in columns 346a and 346b.In this example, entry 346d contains a composite value of "bundle plan, average 23 minutes", which for an ID with hash value "0010" indicates that the ID is related to a bundle plan (the value of which is obtained from a field of the received data record), and the average score of the subscriber having that ID is 23 minutes (the value of which is an aggregation obtained from memory 322 or persistent memory 324). Entry 346e contains a composite value of "50 minutes, average 2340 minutes", which for an ID with hash value "1101" indicates that the ID has a current event (e.g., a call) of length 50 minutes (the value of which is obtained from a field of the received data record), and the average score of the subscriber having that ID is 2340 minutes (the value of which is an aggregation obtained from memory 322 or persistent memory 324). As will be described in more detail below, the segmentation module 326 segments the data record by identifying which of the data records contain values of fields that meet various criteria of one or more segment definitions stored in memory 322. In this example, the segmentation module 326 includes segmentation logic 326a.

[0129] Next, referring also to FIG. 10, in this example, the detection module 318 executes segmentation logic 326a that may be stored in or included in a segmentation module 326 that may be the same as, for example, the segmentation module 22 of FIG. 1. Using the techniques described herein, the segmentation module 326 executes one or more segment definitions to identify subscribers (identified by key or subscriber ID) that meet one or more segment definitions, such as various criteria included within the segmentation logic 326a. For example, the segmentation module 326 executes a segment definition against one or more sets of data records stored, for example, in the memory 322 or the persistent memory 324 to identify a subset of data records that meet the various criteria of the segment definition. Generally, a segment definition includes a data structure that stores data representing the definition of the segment. In this example, the memory 322 stores records of event occurrences over a 14-day period. Records of event occurrences older than 14 days are stored in the persistent memory 324.

[0130] In some examples, the segment definition requires a real-time aggregation (or near real-time (e.g., live-time) aggregation) of events that occurred within the past 14 days. In this example, the segmentation module 326 generates appropriate queries 344a - 344c (regarding the segment definition) and sends those queries to the memory 322 to generate a real-time aggregation. In response, the memory 322 executes the appropriate ones of the queries 344a - 344c against a hash table 346. In this example, each of the composite values for each of the hashed key values within the hash table 346 is the result returned for the queries 344a - 344c. In this example, the memory returns a query result 347, and as will be described in more detail below, the query result 347 includes return entries 347a - 347c.

[0131] In this example, the segmentation module 326 generates a query 344a from the content of record 302a (and / or from data included within the transformed data 306 representing the content of record 302a). In this example, the segmentation module 326 detects that record 302a contains a value of voice for the event type field. Based on this field, the segmentation module 326 generates a query 344a that includes a composite key of "34213.voice", which is the concatenation of the values of the ID field and the event type field within the data record 302a. Since the first part of the composite key represents the value of the ID, when receiving the query 344a, the memory 322 is configured to hash the first part of the composite key. In this example, the hash value of "34213" is "0111", and the memory 322 returns entry 346c to the segmentation module 326 as the return entry 347a. In this example, when returning the value of entry 346c, the memory 322 unhashes the hashed key value, and thus the return entry 347a includes the ID value of "34213".

[0132] The segmentation module 326 generates a query 344b from the content of the record 302b (and / or from data included within the converted data 306 representing the content of the record 302b). In this example, the segmentation module 326 detects that the record 302b includes a value of SMS for the event field. Based on this field, the segmentation module 326 generates a query 344b that includes a composite key of "34214.SMS", which is the concatenation of the values of the ID field and the event type field within the data record 302b. Since the first part of the composite key represents the value of the ID, when receiving the query 344b, the memory 322 is configured to hash the first part of the composite key. In this example, the hash value of "34214" is "0010", and the memory 322 returns the entry 346d to the segmentation module 326 as the return entry 347b. In this example, when returning the value of the entry 346d, the memory 322 unhashes the hashed key value, and thus the return entry 347b includes the ID value of "34214".

[0133] The segmentation module 326 generates a query 344c from the content of the record 302c (and / or from data included within the transformed data 306 representing the content of the record 302c). In this example, the segmentation module 326 detects that the record 302c includes a value of a bundle plan for an event field. Based on this field, the segmentation module 326 generates a query 344c that includes a composite key of "34215. bundle plan", which is the concatenation of the values of the ID field and the event type field within the data record 302c. Since the first part of the composite key represents the value of the ID, when receiving the query 344c, the memory 322 is configured to hash the first part of the composite key. In this example, the hash value of "34215" is "1101", and the memory 322 returns the entry 346e to the segmentation module 326 as the return entry 347c. In this example, when returning the value of the entry 346e, the memory 322 unhashes the hashed key value, and thus the return entry 347c includes the ID value of "34215".

[0134] When query result 347 is received, the segmentation module 326 executes each of return entries 347a to 347c against segmentation logic 326a. In this example, return entry 347a passes through filter 1, and thus the user represented by the ID within return entry 347a is included in the segment. In this example, return entry 347b passes through filter 2, and thus the user represented by the ID within return entry 347b is included in the segment. In this example, return entry 347c passes through filter 3, and thus the user represented by the ID within return entry 347c is included in the segment. Based on each of entries 347a to 347c passing through at least one of the filters included within segmentation logic 326a, the segmentation module 326 executes the logic of "Combine: Output of filter 1, Output of filter 2, Output of filter 3" (specified within segmentation logic 326a) and generates a combined data record 327 that specifies the value of the ID field of the data records included within the segment defined by segmentation logic 326a. As will be described in more detail below, the segmentation module 326 transmits the combined data record 327 to the logic module 328 for further processing.

[0135] In some examples, the segmentation module 326 stores a shared variable representing the required real-time aggregation. The segmentation module 326 stores the values or entries returned from the memory 322 within the shared variable, and they become accessible by the various segment definitions that require the shared variable. In this example, since the system 314 does not need to obtain data from a disk (e.g., persistent memory 324), the aggregation returned from the memory 322 (e.g., as one or more items of a composite value included within a returned entry from the hash table 346) is a real-time aggregation (or near real-time aggregation). Rather, the aggregation can be obtained from on-system memory (e.g., memory 322). Additionally, the system 314 pre-computes the aggregation, for example, by recording the occurrence of events rather than storing the records themselves. By recording the occurrence of events, the system 314 can execute queries more quickly by simply identifying the columns having names that match (or correspond to) the types of data requested. Then, if the required aggregation is a count, for example, the system 314 can count the number of occurrences within the identified columns, and this process has a processing time that is shorter than the processing time required to comprehensively parse the data records and the fields within the data records to identify the data records having values that meet the criteria of the segment specification.

[0136] In another example, the segmentation module 326 requests a real-time aggregation for data records that occurred within the past 20 days. Therefore, the system 314 needs to query both the memory 322 and the persistent memory 324 to obtain the real-time aggregation. In this example, the system 314 queries the memory 322 as described above. The system 314 also queries the persistent memory 324 to obtain data that meets the criteria required for real-time aggregation (and / or to obtain the occurrence of data). Based on the required criteria, the persistent memory 324 retrieves appropriate and / or relevant data within the data record (e.g., from a hash table 324a that hashes the content of old records, e.g., records received before a specified date), and transmits the retrieved data to the detection module 318 (as data on the disk).

[0137] In this example, the data from the persistent memory 324 includes data that specifies the occurrence of events, and those occurrences occurred more than 14 days ago. In this example, to generate the requested real-time aggregation (e.g., the aggregation included in the composite value part of the composite key value), the segmentation module 326 aggregates the data from the persistent memory 324 with the data obtained from the memory 322. Specifically, the segment definition implemented by the segmentation module 326 specifies that only users who made more than 4 voice calls within the past 20 days are included in the segment. In this example, the executable logic implemented by the logical module 328 is executed on the data records of the users within the segment (e.g., data records that include a key representing the users within the segment). In this example, the system 314 queries the memory 322 about the count of voice calls related to a specific subscriber ID.

[0138] In this example, memory 322 returns the count of two voice calls made for the user represented by a specific subscriber ID. Next, system 314 queries persistent memory 324 for the count of voice calls (made within the past six days) and the associated specified subscriber ID. The data returned from persistent memory 324 (e.g., based on querying hash table 324a) specifies that one voice call has been made for the user associated with the specific subscriber ID. Therefore, the data on disk specifies the count of one voice call made for the user represented by the specific subscriber ID. Detection module 318 receives the data on disk and aggregates the count for the specific subscriber ID (contained within the data on disk) with the count for the specific subscriber ID obtained from memory 322 to identify that the user associated with the specific subscriber ID has made three calls within the past 20 days, thus meeting the criteria for segment definition and being included within the segment.

[0139] In this example, the segmentation module 326 transmits to the logic module 328 a data item that specifies the subscriber ID of a subscriber who meets the criteria of one or more segment definitions. In this example, each of the data items is a wide record that includes all composite values used by system 314. That is, the wide record is a wide record of all events or composite keys used by the system. In this example, as described above, the system stores all composite keys and associated composite values. However, the system includes only the composite keys that are actually being used (e.g., by segmentation module 326, logic module 328, or operation module 320) by system 314 within the wide record. In this example, to optimize the performance of the system and ensure no increase in latency with respect to data processing, the system pre-identifies which composite keys are being used and / or accessed and adds only those accessed composite keys (and associated composite values) to the wide record.

[0140] As described above, the logic module 328 stores and executes executable logic represented as a data flow graph. In this example, the logic module 328 executes a data flow graph for a particular segment of subscribers identified within the segmentation module 326. The data flow graph identifies, for example, various actions to be performed on various subscribers based on the attributes of the subscribers included within the segment.

[0141] When the detection module 318 detects subscribers (included within the segment) that meet various criteria for performing an action, the detection module 318 issues a trigger 332 (e.g., an instruction or a message) to a queue, the content of which is received and processed by the action module 320. Generally, the trigger 332 specifies one or more instructions for performing one or more actions.

[0142] The action module 320 executes the triggered action, such as sending a text message or an email, opening a work order ticket in a case management system, immediately disconnecting a service, providing a web service to a target system or device, transmitting packetized data along with one or more notifications, etc. In another example, the action module 320 generates the content of instructions and / or messages and transmits those instructions (and / or content) to a third-party system, and in response, the third-party system performs an action based on the instructions, such as transmitting a text message, providing a web service, etc. In one example, the action module 320 is configured to generate customized content for various recipients. In this example, the action module 320 is configured using rules or instructions that specify which customized content is to be sent or transmitted to which recipient.

[0143] Next, referring to FIG. 11, a process 350 for processing a data structure to modify an attribute of the data structure is shown. This process selects (352) one or more of a plurality of data sources represented within an editor interface, generates (354) a subset of data contained within the plurality of data sources, selects (356) one or more data structures from the selected data sources, each data structure including one or more fields that are modified by modifying (358) one or more attributes of one or more fields within the selected data structure. The modified data structure is stored (360) in a memory (e.g., memory 16 of FIG. 1, or memory 56 of FIG. 6, or memory 58 of FIG. 6), and the selected data structure is included within the subset. What is displayed within the editor interface is a representation of the stored data structure. This process receives (364) selection data that specifies a selection of one or more selectable portions through the editor interface, and process 350 segments (366) a plurality of received data records by identifying which of the received data records have one or more fields corresponding to one or more fields represented within the selected one or more selectable portions. Next, referring to FIG. 12, a process 380 for aggregating data records is shown. Process 380 intermittently receives (382) data records from one or more data sources, e.g., 72a - 72d (FIG. 7), and identifies (384) at least a first field and a second field within a given data record from the one or more data sources. Process 380 detects (386) the presence of a first value within the first field and a second value within the second field of the given record. Process 380 generates (388) a composite key value as described above. This process accesses (390) from memory 322 (FIG. 7) aggregated data related to at least the first field or the second field of the given record, generates (392) a composite key value, and records (394) the occurrence of the given data record.

[0144] Next, referring to FIG. 13, process 400 is executed, and the functional modules executed by each of data structure modification module 18, rendering module 20, segmentation module 22, composite key module 30, and logic module 25 are shown. Data structure modification module 18 receives a data structure (402), causes the rendering of the visualization of the data structure (404), and the execution system receives the modified data structure 418 resulting from modifying the data structure within data structure modification module 18 (406).

[0145] Rendering module 20 receives the modified data structure from data structure modification module 18 and causes the rendering of the segmentation template along with the modified data structure (410). The rendering module further receives segmentation logic (412) and transmits the segmentation logic to segmentation module 22 (414).

[0146] Segmentation module 22 receives the segmentation logic (420) and generates a composite key query that is transmitted to composite key module 30 (424) (422).

[0147] Composite key module 30 receives data records (e.g., from repository 12a of FIG. 1) (430), and a composite key is generated by generating a composite key value from the data records (434). Composite key module 30 stores the composite key value (e.g., within data repository 12e of FIG. 3) (436).

[0148] Next, referring to FIG. 14, a process 450 is shown that illustrates functional modules executed by each of the segmentation module 22, the composite key module 30, and the logic module 25. The composite key module 30 receives a composite key query (438), obtains a record having a composite key that meets the specified criteria (440), and transmits the obtained record to the segmentation module 22. The segmentation module 22 receives the obtained queried record from the composite key module 30 (426) and transmits the obtained record to the logic module 25 (428).

[0149] The logic module 25 receives the transmitted record from the segmentation module 22 (450), applies the logic specified by the segmentation logic such as the executable logic represented by the component 158 (452), and then outputs the data set resulting from the process (454).

[0150] The above techniques can be implemented using software for execution on a computer. For example, the software forms procedures by one or more computer programs executed on one or more programmed or programmable computer systems (which can be of various architectures such as distributed, client / server, grid, etc.), and such computer systems each include at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. The software can form one or more modules of a larger program that provides other services related to the design and configuration of charts and flowcharts, for example. The nodes, links, and elements of the chart can be implemented as other organized data conforming to a data structure stored in a computer-readable medium or a data model stored in a data repository.

[0151] The techniques described herein can be implemented by digital electronic circuitry, or by computer hardware, firmware, software, or combinations thereof. Apparatus can be implemented by a computer program product tangibly embodied in a machine-readable storage device (e.g., a non-transitory machine-readable storage device, a machine-readable hardware storage device, etc.) for execution by a programmable processor; the actions of the method can be performed by a programmable processor executing a program of instructions to perform functions by operating on input data to produce output. The described embodiments and other embodiments in the claims and the techniques described herein can be advantageously implemented by one or more computer programs executable on a programmable system including at least one programmable processor coupled to interact with a data storage system, at least one input device, and at least one output device. Each computer program can be implemented in a high-level procedural or object-oriented programming language, or in assembly or machine language as desired; in any case, the language can be a compiled or interpreted language.

[0152] Processors suitable for executing a computer program include, by way of example, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes, or is operatively coupled to, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, optical disks, or receives data therefrom, transfers data thereto, or does both. Computer-readable media for embodying computer program instructions and data include any form of non-volatile memory, including by way of example semiconductor memory devices, such as EPROM, EEPROM, flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks, and CD ROM disks and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special logic circuitry. Any of the foregoing can be supplemented by, or incorporated in, an ASIC (Application Specific Integrated Circuit).

[0153] To enable interaction with a user, embodiments can be implemented on a computer having a display device, such as an LCD (Liquid Crystal Display) monitor, for displaying information to the user, and a keyboard and a pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other kinds of devices can also be used to enable interaction with a user, for example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, tactile feedback, and the input received from the user can be received in any form including acoustic input, voice input, or tactile input.

[0154] Embodiments can be implemented, for example, by a computing system including a back-end component as a data server, by a computing system including a middleware component, such as an application server, by a computing system including a front-end component, such as a graphical user interface or a client computer having a web browser by which a user can interact with an implementation of the embodiment, or by a computing system including any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.

[0155] The system and method, or a part thereof, may use the "World Wide Web" (Web or WWW), which is a collection of servers on the Internet that utilize the Hypertext Transfer Protocol (HTTP). HTTP is a known application protocol that provides users with access to resources that can be various forms of information such as text, graphics, images, audio, video, Hypertext Markup Language (HTML), programs, etc. When a link is specified by a user, the client computer makes a TCP / IP request to the web server and receives information that can be another web page formatted according to HTML. The user can also access other pages on the same server or other servers by following instructions on the screen, by entering specific data, or by clicking on a selected icon. It should also be noted that in embodiments using web pages, any type of selection device known to those skilled in the art, such as checkboxes and dropdown boxes, can be used to enable the user to select options regarding a given component. Servers, including UNIX machines, run on various platforms, but other platforms such as Windows 2000 / 2003, Windows NT, Sun, Linux, Macintosh, etc. can also be used. Computer users can view information available on web servers or on the network by using browsing software such as Firefox, Netscape Navigator, Microsoft Internet Explorer, Mosaic browser, etc. The computing system can include clients and servers. The client and the server are generally separated from each other and typically interact via a communication network. The relationship between the client and the server is created by computer programs that are executed on their respective computers and have a client-server relationship with each other.

[0156] Other embodiments are within the scope and spirit of the specification and the claims. For example, due to the nature of software, the above functions can be implemented using software, hardware, firmware, hardwiring, or any combination thereof. The features implementing the functions can also be physically arranged at various locations, including being distributed, such that parts of the functions are implemented at various physical locations. Throughout this specification and the present application, the use of the word "a" is not used in a limiting sense and thus is not intended to exclude a plurality of meanings or the meaning of "one or more" for the word "a". Additionally, to the extent priority is claimed to a provisional patent application, the provisional patent application is to be understood to include, not by way of limitation, examples of how the techniques described herein can be implemented.

[0157] Some embodiments of the invention have been described. Nevertheless, those skilled in the art will understand that various modifications can be made without departing from the spirit and scope of the claims and the techniques described in this specification.

Claims

1. A data processing system for displaying an editor interface that enables segmentation of data records by generating a subset of data from a plurality of data sources, identifying one or more attributes of one or more respective fields of the subset, and displaying one or more representations of the one or more identified attributes, comprising: A memory for storing a plurality of data sources represented within the editor interface; Identifying a plurality of data sources represented within the editor interface, and for each of the data sources, identifying one or more first data structures, each first data structure including one or more fields, wherein one or more fields of at least one first data structure identify one or more first attributes; accessing a second data structure corresponding to the at least one first data structure from each of the data sources; and identifying one or more second attributes of one or more respective fields within the second data structure, wherein the one or more second attributes indicate a version of the one or more first attributes in which the one or more first attributes have been modified; thereby generating a subset of data including the first data structure and the second data structure included within the plurality of data sources; a data structure modification module; A memory for storing the second data structure corresponding to the first data structure included within the subset, wherein at least one of the second data structures includes one or more second attributes of one or more respective fields; a memory A rendering module that displays the representation of the second data structure within the editor interface, wherein at least one of the representations is of one or more second attributes of one or more respective fields of at least one second data structure, each representation includes one or more selectable portions, the selectable portions represent second attributes of fields of the second data structure, and the rendering module receives, through the editor interface, selection data that designates selection of one or more selectable portions that represent one or more second attributes representing a version of the one or more first attributes in which the one or more first attributes have been modified, and A segmentation module that segments the received plurality of data records by identifying which of the plurality of data records received from the plurality of data sources have one or more third attributes of one or more fields corresponding to the one or more second attributes represented within the selected one or more selectable portions A data processing system including the same. **Claim 2** A method, executed by a data processing system, for displaying an editor interface that enables segmentation of data records by generating a subset of data from a plurality of data sources, identifying one or more attributes of one or more respective fields of the subset, and displaying one or more representations of the one or more identified attributes, the method comprising: Identifying a plurality of data sources represented within the editor interface For each of the data sources, identifying, from the data source, one or more first data structures, each first data structure including one or more fields, wherein one or more fields of at least one first data structure specify one or more first attributes; accessing, for at least one first data structure from each of the data sources, a second data structure corresponding to the at least one first data structure; and specifying, for one or more fields of the one or more second data structures, one or more second attributes, wherein the one or more second attributes indicate a version of the one or more first attributes in which the one or more first attributes have been modified, thereby generating a subset of data including the first data structures and the second data structures included in the plurality of data sources. Storing, in memory, the second data structure corresponding to the first data structure included in the subset, wherein at least one of the second data structures includes one or more second attributes of one or more fields. Displaying, within the editor interface, a representation of the second data structure, wherein at least one of the representations is of the one or more second attributes of the one or more fields of at least one second data structure, and each representation includes one or more selectable portions, the selectable portions representing second attributes of fields of the second data structure. Receiving, through the editor interface, selection data specifying selection of one or more selectable portions representing the one or more second attributes indicating a version of the one or more first attributes in which the one or more first attributes have been modified; and Segmenting the received plurality of data records by identifying which of the plurality of data records received from the plurality of data sources have one or more third attributes of one or more fields corresponding to the one or more second attributes represented in the selected one or more selectable portions. A method comprising the above. Claim 3 The data structure includes a key field representing a key for the data structure, the data record is associated with the value of the key, the method, selecting a plurality of fields from a plurality of second data structures, selecting values of the respective selected fields for a specified value of the key, combining the selected values for the specified value of the key, and executing an executable data flow graph to segment the plurality of data records by outputting the combined values, The method according to claim 2, further comprising.

4. The representation is a first representation, the method, displaying the representation of the executable data flow graph within the editor interface as a second representation, The method according to claim 3, further comprising.

5. Receiving, through the editor interface, additional selection data that specifies a selection of the second representation and further specifies that one or more criteria are applied to the output combined values of the one or more given fields represented by the one or more selectable portions selected through the editor interface, the method according to claim 4.

6. Displaying a user interface having one or more first controls for selecting a data structure and one or more second controls for modifying the one or more fields, the method according to claim 2.

7. The data structure includes one or more data records, and each data record has one or more values for a specific field, the method according to claim 2.

8. At least one of the data sources includes an unselected data structure, the method according to claim 2.

9. The data structure includes a key field representing a key for the data structure, the data record is associated with the value of the key, the segmentation module, selects a plurality of fields from a plurality of second data structures, Executing an executable data flow graph to segment the plurality of data records by selecting the values of the respective selected fields for the specified value of the key, combining the selected values for the specified value of the key, and outputting the combined values The data processing system according to claim 1, configured as described above.

10. The representation is a first representation, The rendering module, Displays the representation of the executable data flow graph within the editor interface as a second representation The data processing system according to claim 1, configured as described above.

11. The rendering module, Receives, through the editor interface, additional selection data that specifies a selection of the second representation and further specifies that one or more criteria are applied to the output combined values of the one or more given fields represented by the one or more selectable portions selected through the editor interface The data processing system according to claim 1, configured as described above.

12. The rendering module, Displays a user interface having one or more first controls for selecting a data structure and one or more second controls for modifying the one or more fields The data processing system according to claim 1, configured as described above.

13. The data structure includes one or more data records, and each data record has one or more values for a specific field. The data processing system according to claim 1.

14. At least one of the data sources includes an unselected data structure. The data processing system according to claim 1.

15. A computer program for causing a data processing system to execute the method according to any one of claims 2 to 8

Citation Information

Patent Citations

  • Interactive data retrieval and extraction system for relational data base

    JP1994290221A

  • Database processing system

    JP1998097544A

  • Relational database management system and storage medium stored with same

    JP2001216307A

  • Real-time aggregation and scoring in an information handling system

    US20050125280A1

  • Fast aggregation of compressed data using full table scans

    US20050192941A1