Dynamic execution of parameterized applications that process keyed network data streams
A unified system for real-time data processing addresses latency issues by integrating data collection, detection, and action, using parameterized applications to handle large volumes of diverse data efficiently.
Patent Information
- Application Number
- JP2024068886
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-08-28
- Filing Date
- 2024-04-22
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2038-02-22
AI Technical Summary
Conventional methods for processing data records from heterogeneous sources introduce latency and delay due to multiple reformatting and non-integrated systems, preventing real-time or near-real-time data processing and analysis.
A single execution system for collecting, validating, and processing data streams in real-time or near-real-time, eliminating the need for separate systems by integrating data collection, detection, and action, and utilizing parameterized applications to process keyed data items with specified conditions and actions.
Enables immediate response to data records, reduces latency, increases bandwidth, and minimizes memory usage, allowing processing of large volumes of data from diverse sources in real-time.
Smart Images

Figure 0007770464000002 
Figure 0007770464000003 
Figure 0007770464000004
Abstract
Description
[Technical Field]
[0001] Technical Field The present invention relates to a network-aware computer-implemented method, computer system and computer-readable medium for collecting, validating, formatting and further processing (near-) real-time network data streams in parameterized applications performing certain operations, e.g., operations performed in a network, such as a telecommunications network.
[0002] Priority claims This application claims priority to U.S. Provisional Patent Application No. 62 / 462,498, filed February 23, 2017, and U.S. Patent Application Publication No. 15 / 688,587, filed August 28, 2017, the entire contents of each of which are incorporated herein by reference. [Background technology]
[0003] Background technology A typical method for processing data records received from heterogeneous sources involves collecting the data records with multiple heterogeneous systems, each configured to process a specific type of data. After processing the data with a system configured to process that type of data, that system then forwards the data to another downstream application for further processing. In conventional systems, the downstream application then performs additional formatting (and / or reformatting) on the received data records to convert them into a format acceptable to that application. As a result, this conventional method introduces latency and delays in the interaction between the system receiving and collecting the data and the downstream application processing the data. This conventional method also introduces increased latency because a single data record must be reformatted multiple times. For example, a data record is formatted (and / or reformatted) by each system that receives and processes it, resulting in increased accumulated latency. This non-integrated system framework (of heterogeneous systems, each with its own data formatting and processing) also introduces latency in processing the data because it cannot process and analyze the data in real time, or at least near real time. Consequently, tasks or actions that depend on the processed data are increasingly delayed or even not processed at all, and the execution that is performed becomes obsolete if the current situation has already changed again. Summary of the Invention [Means for solving the problem]
[0004] overview Unlike these common approaches, the methods, systems, and computer-readable media described herein simplify and accelerate data integration and data record management preparation. The described functionality allows for both batch and real-time collection, validation, formatting, and further processing of data streams (e.g., as data is received, e.g., in real time or near real time, without storing the collected data to disk). Furthermore, by providing a single execution system that performs collection, detection, and action tasks, the execution system eliminates the complexity associated with integrating data into one system for data collection and subsequently reintegrating the collected data into another system for detection and action. The functionality described herein can provide immediate response to data records or data items (e.g., as they are received), thereby enabling application results to be immediately visible. Because desired end actions typically depend on rapid processing of data, end actions such as operations in logistics or telecommunications greatly benefit from this rapid processing of large volumes of data coming from various disparate sources that produce data in various different formats. The system described herein can process over 2 billion data records or items per day from 50 million users. Unlike conventional methods, the system described herein provides increased bandwidth and reduced memory usage.
[0005] In this method for processing keyed data items, each associated with a key value, the keyed data items coming from a plurality of different data streams, the processing including collecting the keyed data items, determining satisfaction of one or more specified conditions for the performance of one or more actions based on the content of at least one keyed data item, and causing the performance of at least one of the one or more actions in response to the determination, the method includes the steps of: accessing first, second and third parameterized applications each including a first, second and third specification, where the first specification specifies one or more parameters defining one or more characteristics of the first parameterized application and one or more values for each of those one or more parameters; the second specification specifies one or more parameters defining one or more characteristics of the second parameterized application and one or more values for each of those one or more parameters, and further includes rules and respective conditions for those rules; and the third specification specifies one or more parameters defining one or more characteristics of the third parameterized application and one or more values for each of those one or more parameters; maintaining a state of the second specification associated with each value of the key according to the state of a particular value of the key specifying one or more portions of the second specification to be executed by the application; collecting data items from one or more data sources and one or more data streams, multiple data sources, or multiple data streams, where the format of the collected first data items differs from the format of the collected second data items, and the data items are associated with values of the keys; executing the first parameterized application according to one or more values of one or more parameters specified by the first specification to perform processing including: converting the first and second data items according to the first specification of the first parameterized application to obtain converted data items; and adding the converted data items to a queue;executing the second parameterized application with one or more values of one or more parameters specified by the second specification to process the transformed data items in the queue, together with processing of the transformed data items, including, for one or more transformed data items associated with a particular value of the key, identifying a current state of the second specification for the particular value of the key, identifying one or more rules in a portion of the second specification to be executed in the current state, executing the identified one or more rules, determining that at least one of the one or more changed data items satisfies one or more conditions of at least one of the one or more rules executed in the current state, creating a data structure that specifies the execution of one or more operations in response to said determination, transitioning the second specification from its current state to a next state for the particular value of the key, and sending the created data structure to a third parameterized application; and executing the third parameterized application with one or more values of one or more parameters specified by the third specification to perform an operation, including sending one or more instructions that cause the execution of at least one of the one or more operations based on at least one of the one or more operations specified in the data structure. A system may comprise one or more computers that perform specific tasks or operations by providing the system with software, firmware, hardware, or a combination thereof. Such software, hardware, etc., or a combination thereof, in operation, causes the system to perform those operations. One or more computer programs may comprise instructions that perform specific tasks or operations. These programs, when executed by a data processing device, cause the device to perform those operations. In certain circumstances, the methods, computer programs, and / or systems may include one or more of the following functions and / or operations:
[0006] The operations further include displaying one or more user interface elements for specifying one or more values for one or more parameters of each of the first, second, and third parameterized applications during operation of one or more user interfaces. Executing the first parameterized application includes executing the first parameterized application with one or more values specified by the one or more user interface elements for the one or more parameters of the first parameterized application. Executing the second parameterized application includes executing the second parameterized application with one or more values specified by the one or more user interface elements for the one or more parameters of the second parameterized application, where the specified one or more values are used as inputs by the rules to determine whether the one or more conditions are satisfied. Executing the third parameterized application includes executing the third parameterized application with one or more values specified by the one or more user interface elements for the one or more parameters of the third parameterized application. The data items include data records, and in this case, converting includes reformatting the data records according to a format specified by the first specification of the first parameterized application.
[0007] The operations further include augmenting the data record with data from a profile of a user associated with the data record based on execution of a second parameterized application. The augmentation is according to instructions specified by a second specification for the second parameterized application, which retrieves profile data for the user from memory and enters the retrieved profile data into one or more fields of the data record. The parameterized application includes an application for data processing, the application including one or more parameters configurable with one or more values. The operations further include executing a feedback loop (e.g., a synchronous or asynchronous feedback loop) to one or more third-party systems to request confirmation of execution of the one or more operations. The operations further include generating one or more key performance indicators (KPIs) related to specific values of the keys based on execution of the second parameterized application. The KPIs specify one or more values of data items associated with specific values of the keys. For example, the KPIs may include performance data, such as performance data specifying the performance of one or more portions or logical branches of a campaign or a set of predetermined logic.
[0008] The operations further include receiving data for a particular value of the key, the received data indicating feedback regarding at least one of the one or more operations. The operations also include updating a KPI related to the particular value of the key with the feedback data. The one or more operations include one or more of: sending a text message to an external device, sending an email to an external system, opening a work order ticket in a case management system, disconnecting a mobile phone connection, providing a web service to the target device, sending a data packet of one or more transformed data items including a notification, and executing a data processing application hosted on one or more external computers on the one or more transformed data items. The one or more instructions are sent via a network connection to cause execution of at least one of the one or more operations on the external device. The method further includes receiving a feedback message indicating whether at least one of the one or more operations (i) completed successfully or (ii) failed. The at least one of the one or more operations is considered to have failed if at least one of the one or more operations is not partially completed. In this case, the feedback message indicates which portions of the at least one of the one or more failed operations were not completed, and the feedback message indicates successful completion and / or failure result data of at least one of the one or more operations.
[0009] The operations further include modifying the one or more specified values of one or more parameters of the first, second, and / or third parameterized applications based on the result data. The operations also include re-executing the first, second, and / or third parameterized applications with the modified one or more specified values. The one or more instructions are transmitted over a network connection to cause performance of the one or more operations on the external device. The method further includes receiving a feedback message from the external device, the feedback message including result data of at least one performed operation of the one or more operations; comparing the result data with predetermined data regarding successful completion of execution of at least one of the one or more operations; and determining, based on the comparison, that execution of the at least one of the one or more operations has completed successfully or that execution of the at least one of the one or more operations has failed.
[0010] The operations further include modifying the one or more specified values of one or more parameters of the first, second, and / or third parameterized applications based on the results data, and re-executing the first, second, and / or third parameterized applications with the one or more specified values modified. The execution of the at least one of the one or more operations is determined to have completed successfully if the results data deviates from the predetermined data by less than a predetermined amount, and the execution of the at least one of the one or more operations is determined to have failed if the results data deviates from the predetermined data by at least a predetermined amount.
[0011] This operation further includes displaying, during operation of one or more user interfaces, one or more user interface elements specifying the predetermined data and the predetermined amount. This operation further includes outputting, during operation of one or more user interfaces via one or more display user interface elements, whether the at least one of the one or more operations (i) completed successfully or (ii) failed. The method of claim 15 further includes: outputting, during operation of one or more user interfaces via one or more display user interface elements, result data. This operation further includes receiving, via one or more display user interface elements, user-specified changes based on the result data to one or more specified values of one or more parameters of the first, second, and / or third parameterized applications, and changing the one or more specified values and re-executing the first, second, and / or third parameterized applications. The sending of the one or more instructions causing the execution of at least one of the one or more actions is performed automatically by a third parameterization application using as input an output specifying the execution of the one or more actions.
[0012] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. [Brief explanation of the drawings]
[0013] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1] FIG. 1 is a schematic diagram of a system that implements parameterized logic for processing keyed data. [Figure 2A] FIG. 2A is a diagram illustrating a typical data flow graph. [Figure 2B]FIG. 2B is a diagram showing the portions of the interface for customizing a data flow graph. [Figure 2C] FIG. 2C is a diagram showing the portions of the interface for customizing a data flow graph. [Figure 3A] FIG. 3A is a schematic diagram of a system for computing near real-time data record aggregations. [Figure 3B] FIG. 3B is a schematic diagram of a system for processing keyed data. [Figure 3C] FIG. 3C is a schematic diagram of a system for processing keyed data. [Figure 4] FIG. 4 is a schematic diagram of a system for processing keyed data. [Figure 5A] FIG. 5A is a schematic diagram of a system for processing keyed data. [Figure 5B] FIG. 5B is a schematic diagram of a system for processing keyed data. [Figure 6] FIG. 6 is an example of a graphic user interface for the configuration of parameterized logic. [Figure 7] FIG. 7 is an example of a graphic user interface for the configuration of parameterized logic. [Figure 8] FIG. 8 is an example of a graphic user interface for the configuration of parameterized logic. [Figure 9] FIG. 9 is an example of a graphic user interface for the configuration of parameterized logic. [Figure 10] FIG. 10 is a simplified flowchart. [Figure 11] FIG. 11 is a diagram of a flowchart instance. [Figure 12] FIG. 12 is a schematic diagram of near real-time logic execution with large records. [Figure 13] FIG. 13 is a diagram of an example process for processing keyed data with parameterized logic. DETAILED DESCRIPTION OF THE INVENTION
[0014] Detailed explanation Referring to FIG. 1, a system 100 is shown that collects data and data records from various sources, e.g., from various servers located in different locations and interconnected via a network, and integrates the data into modules for data detection and action execution. Generally, data items include data records, data records, or data indicative of events (e.g., data indicative of the occurrence of an action or a record containing data indicative of the occurrence of an action (e.g., the occurrence of a phone call or the duration of a phone call)). While the techniques described herein are primarily described with respect to data records, the techniques can also be used to process events. The system 100 includes a code management store 102, a development environment 104, data sources 106, and an execution system (also referred to as a runtime environment) 108. The execution system 108 includes a system that implements a collection-detection-action (CDA) environment for configuring and executing various applications and programs that perform the collection, detection, and action described above. These various applications and programs include, for example, a set of reusable graphs, plans, applications, etc. (e.g., to facilitate the development and simplify the maintenance of applications and programs). Generally, a graph (e.g., a data flow graph) includes vertices or nodes (components or data sets) connected by directed edges (indicating the flow of operational elements) between the vertices, as described, for example, in U.S. Patent Application Publication No. 2007 / 0011668, entitled "Managing Parameters for Graph-Based Applications," which is incorporated herein by reference. Generally, a plan includes an application that represents a process flow in which process steps, called tasks, are connected by flows that indicate execution order (e.g., dependencies).
[0015] Each of the methods described herein can be performed by system 100. The system includes a development environment 104 coupled to code management store 102. Development environment 104 is configured to create any of the applications described herein. Such applications are associated with a dataflow graph that implements graph-based computations performed on data flowing from one or more input datasets through a graph of processing graph components to one or more output datasets. The dataflow graph is specified by a data structure in code management store 102. The dataflow graph has multiple vertices or nodes that represent graph components specified by the data structure and connected by one or more edges. The edges are specified by the data structure and represent data flow between the graph components. A runtime environment 108 is also coupled to store 102 and is hosted on one or more computers. Runtime environment 108 also reads the stored data structure that specifies the dataflow graph and allocates and configures computing resources, such as processes, to perform computations on the graph components assigned to the dataflow graph by runtime environment 108. The runtime environment 108 also includes an execution module that schedules and controls the execution of assigned processes so that work according to the method is performed.
[0016] The execution system 108 includes a collection / integration module 114 (hereinafter collection module 114) that collects data records (and other data), transforms it, and distributes the transformed data to downstream applications (e.g., including the detection module 116 and the action module 118). In this example, the action module 118 includes interfaces to third-party and / or external systems. In particular, the collection module 114 collects data in batch or real time from various sources, such as data sources 106, or from various servers located at various locations and interconnected via a network, and collects real-time data streams 107, such as real-time data arriving from various servers located at various locations and interconnected via a network. The storage providing the data sources 106 may be locally connected to the system 100, e.g., on a storage medium (e.g., a hard drive) connected to the computer-operated execution system 108, or remotely connected to the execution system 108, e.g., hosted on a remote system that communicates with the execution system 108 via a local or wide-area data network.
[0017] The collection module 114 is configured to parse, validate, and augment data (e.g., data in slower batch data 109 and / or real-time data streams 107 received from data sources 106). For convenience, the terms “real-time” and “near real-time” will hereinafter be referred to collectively, without limitation, as “real-time” or “near real-time.” The collection module 114 stores the augmented data records in memory and also writes the augmented data records to disk to ensure archival and recovery of the records. Because the collection module 114 can handle any source, the collection module 114 allows for rapid and autonomous integration of data into the CDA environment (e.g., additional applications are required for the collection module 114 to successfully integrate), while addressing the complexities of arbitrarily large data volumes, low latency, and multiple data formats.
[0018] In one example, records are transferred from a remote system to system 108. In this example, thousands of records arrive each day. In this example, collection module 114 detects the transferred records in a local directory on system 108. The collection module 114 also transfers the records by adding them to a queue and copying the data to an archive (not shown). The collection module 114 also deletes the original records after parsing. In this example, the archive stores, and optionally compresses, a copy of the unchanged input data. The archive stores the data for a specified number of days. In this example, collection module 114 adds or transfers the collected data to a data queue. There is one data queue for each data feed. The collection module 114 also parses the records by reading them, removing duplicates, enriching them (e.g., by adding profile data to the records), sorting them by key (e.g., for further enrichment), and writing them to a queue. In this example, the collection module 114 removes duplicates by storing an ICFF (Indexed Compressed Flat File) archive that stores a copy of each record's hash for a specified number of days to be used to remove duplicates. In this example, the collection module 114 also performs maintenance by running a scheduled (e.g., daily) archive and a graph that cleans up (e.g., removes) old files from the ICFF archive.
[0019] In this example, the queue (to which the records are attached) includes a single centralized queue that is partitioned by key. This centralized queue has a standardized format that includes a data record type (indicating the source data feed name), standard fields (such as the key), and feed-specific fields. This queue includes parallel queues that can queue in parallel within a server or across servers and / or can process in parallel across one or more systems.
[0020] In this example, the collection module 114 checkpoints the data feed after receiving a specified number of records or files (e.g., a sequence of records) or after a specified number of seconds have passed. The collection module 114 performs this checkpointing by executing a graph with configurable parameters that pumps data through the system 108 with near-real-time latency as input files arrive. The execution system 108 also includes a detection module 116 that creates program logic to detect predetermined occurrences. Following collection and integration of the collected data by the collection module 114, the execution system 108 executes the detection module 116 on the collected and integrated data. In one example, the collection module 114 sends one or more (e.g., distinct) streams of collected data (e.g., already validated and formatted / reformatted) to the detection module 116 as the data is received, e.g., in real time, for further processing and / or execution of one or more predetermined rules on the collected data streams. As described in more detail below, the detection module 116 is uniquely flexible, allowing users to create both simple and highly sophisticated detection schemes (e.g., campaigns and / or sets of rules with various conditions that must be met prior to rule execution) based on multiple dynamic data record types, aggregation types, state definitions and transitions, composite functions, and timers. The detection module 116 also enables the detection of "synthetic" data (e.g., data records). Typically, synthetic data includes data indicating the absence of a condition or occurrence, e.g., detecting a user's failure to access a portal within a predetermined time period. After detecting a data record, the detection module 116 adds a command or message to a queue, the contents of which are received and processed by the action module 118.
[0021] The action module 118 performs the triggered action, such as sending a text message or email, opening a work order ticket in a case management system, immediately terminating a service, providing a web service to the targeted system or device, sending packetized data containing one or more notifications, etc. In another example, the action module 118 creates instructions and / or message content and sends these instructions (and / or content) to a third-party system. The third-party system then performs an action based on the instructions, such as sending a text message, providing a web service, etc. In one example, the action module 118 is configured to create customized content for various recipients. In this example, the action module 118 is configured with rules or instructions that specify which customized content to send or communicate to which recipients.
[0022] In the traditional model of data integration and data record discovery, one system performs data collection, batch collection, and stores the data in a data warehouse. Then, for data record discovery, another system searches and queries the batch data in the data warehouse and performs data record discovery on the warehouse data. This traditional model has several limitations, including, for example, an inability to support real-time data collection. Also, various types of inconsistent applications and / or engines have been devised to perform various types of data record discovery (e.g., for various segments). Unlike these traditional models, the CDA environment described herein prepares data for integration and data record management without requiring additional technology, for example, a separate system for data integration and yet another system for data record management. This functionality simplifies and accelerates end-to-end integration. Additionally, the CDA environment described herein is configured to process both batch and real-time data streams (e.g., in real time or near real time as the data is received). For example, the CDA environment can handle data in real time, without storing the collected data to disk, by the collection module 114 processing, validating, and / or formatting the data in near real time (e.g., as the data is received by the execution system 108). The collection module 114 is configured to validate and process streams of data as they are received, and then queue the processed data for further processing by the detection module 116, all without storing the data in a data warehouse for subsequent retrieval, which introduces latency. Also, by providing a single system to perform the tasks of collection, detection, and action, the execution system 108 eliminates the complexity involved in integrating data into one system for data collection and then reintegrating the collected data into another system for detection and action.Since the desired final task or action generally depends on the data being processed and may be increasingly delayed (or may not occur at all if the current situation has already changed again, making the execution obsolete), final tasks or actions, such as those in logistics or telecommunications, benefit greatly from this rapid processing of large amounts of data in a variety of different formats.
[0023] In this example, the code management store 102 is configured to communicate with the execution system 108 and stores a parameterized collection application 120, a parameterized detection application 122, and a parameterized action application 124. A parameterized application typically includes an electronic template or record preconfigured to perform a specified function or task for one or more parameters. Values for these parameters are passed to the parameterized application from one or more other applications and / or user input. After parameter values are specified or identified, the parameterized application (e.g., a parameterized template) represents a parameterized template specification (hereinafter, "specification"). For example, the parameterized application specifies state values, action values, transition values between states, etc. A specification typically represents executable logic and specifies values for parameters of the executable logic and various states of the executable logic based on states reached by execution of the executable logic on preceding data items. In some cases, the system includes action applications that are not parameterized.
[0024] In one example, a parameterized application is referred to as parameterized logic, e.g., when an electronic application or record is preconfigured with logical expressions (or other logic) that perform specified functions and operations with respect to one or more parameters.
[0025] In one example, a parameterized application includes a generic data flow graph (or other generic data processing program or application) with various parameters, the values of which are specified as inputs to the parameterized application. In this example, a parameterized collection application includes a parameterized application that performs data collection and integration. A parameterized detection application includes a parameterized application that performs data detection. A parameterized action application includes a parameterized application that performs action.
[0026] In one example, each of modules 114, 116, and 118 is realized by the execution of an instance of a parameterized application 120, 122, or 124, respectively. In this example, an instance of a parameterized application includes an execution of a parameterized application in which parameter values have been specified. Execution of collection module 114 is based on one or more parameterized collection applications 120, which specify how data is to be collected, formatted, validated, and integrated. By parameterizing input data streams, there is no need to create a new program describing the processing of each input data stream. Rather, execution system 108 maintains a program or application that processes input data. The program or application includes various parameters, the values of which can be set to configure the program or application to process a particular input data stream. In some examples, the program or application is a generic data flow graph, such as a parameterized application. Using interface module 126, a user customizes the generic data flow graph for a particular input data stream by specifying values for the parameters of the generic data flow graph. These specified values specify how collection module 114 processes that particular data stream. Rather than relying on the traditional technique of creating specialized programs to process each input data stream, parameterization of a generic data flow graph (or application) allows for reuse, which reduces memory requirements for the execution system 108 and also reduces data flow errors in data collection and integration, since only one generic (and error-free) mechanism is used and reused.
[0027] In this example, the parameterized collection application promotes reuse of a specified collection application because the parameters of the parameterized collection application facilitate changing values in the parameterized collection application, e.g., instead of a (non-parameterized) collection application whose code must be changed whenever there is a change related to the collection application's values. In this example, the execution system 108 compiles the parameterized collection application 120 into executable code. The collection module 114 is realized by executing this executable code.
[0028] Execution of the detection module 116 is based on one or more parameterized detection applications 122, which specify one or more predefined rules for detecting specified occurrences or lack thereof. In this example, the parameterized detection application 122 includes one or more parameters that specify one or more values (e.g., values used by the rules to perform detection). In this example, the parameterized detection application facilitates reuse of the specified detection application because the parameters of the parameterized detection application facilitate changing values in the parameterized detection application, e.g., instead of using a (non-parameterized) detection application whose code must be modified each time there is a change related to the detection application's values. Furthermore, parameterized applications in general, and the parameterized detection application 122 in particular, enable flexible on-the-fly detection between various segments (e.g., users) and between various types of data records. For example, one of the parameterized detection applications 122 can be configured to detect occurrences of a specified type of data record (e.g., by setting values for parameters in the parameterized detection application). The parameterized detection application can then be reused to detect occurrences of other types of data records (e.g., by specifying other values for the parameters in the parameterized detection application). This reuse of parameterized applications reduces errors in the execution of the detection module 116, for example, because error-free applications can be reused. These parameterized applications also promote flexibility in data record detection. In particular, deployment of conventional detection modules (e.g., detection engines) requires extensive modeling and configuration of various instructions and relationships (e.g., between the detection of various data record types and the rules that specify various actions in response to successful detection) before detection can occur.In contrast, the detection module 116 can be launched without such effort and rather can be configured on the fly, for example, by a user (e.g., via the user interface module 126) specifying values for parameters of the parameterized detection application 122. In this example, the execution system 108 compiles the parameterized detection application 122 into executable code. The detection module 116 is realized by the execution of this executable code.
[0029] Execution of the action module 118 is based on one or more parameterized action applications 124, which specify one or more predetermined actions or instructions to be performed based on, for example, commands or triggers received from the detection module 116. In this example, the parameterized action application 124 includes one or more parameters that specify one or more values (e.g., values used when a rule executes its action). In this example, the parameterized action application promotes reuse of the specified action application because the parameters in the parameterized action application make it easy to change values in the parameterized action application, e.g., instead of using a (non-parameterized) action application whose code must be changed whenever there is a change related to the action application's values. The parameterized action application 124 also enables flexible on-the-fly execution of actions between various segments (e.g., users) and between various types of data records. For example, one of the parameterized action applications 124 can be configured to perform a specified action (e.g., send a text message to alert a user) by setting the values of parameters in the parameterized action application. The parameterized action application can then be reused to send an action, e.g., an instruction to send a particular message, to a third party. A parameterized action application is reused by specifying or changing the values of the parameters of the parameterized action application. In this example, the execution system 108 compiles the parameterized action application 124 into executable code. The action module 118 is realized by executing this executable code.
[0030] The collection module 114, detection module 116, and action module 118 represent an improvement over existing technologies (e.g., those that include separate systems for collection, detection, and action) by providing benefits over conventional systems, such as increased flexibility, reduced processing time for real-time data streams, and reduced memory requirements. As previously described, the execution system 108 includes an integrated system that includes the collection module 114, the detection module 116, and the action module 118. By having these modules integrated into a single system, the execution system 108 is able to process data in real time (or near real time) because the collection module 114 processes the received data stream in memory and then, for example, adds the verified data to a queue for further processing rather than delegating the received data to data storage for subsequent retrieval. The integration of the collection module 114, the detection module 116, and the action module 118 into a single system also reduces memory requirements because the system no longer needs to store multiple different applications that collect, detect, and act on each data stream that contributes records. Rather, each parameterized application can be reused, for example, when different values of different parameters are specified.
[0031] The execution system 108 also changes the configuration of the parameterized application and sets the values of various parameters in the parameterized application, for example, based on user input specifying the values of the parameters or based on an executed action that results in such a change or setting. The user interface module 126 displays configuration information to the user and receives configuration actions from the user, for example, data specifying the values of the parameters and / or data specifying actions for configuring the parameterized application. In this example, the parameterized collection application 120, the parameterized detection application 122, and the parameterized action application 124 are each stored in the code management store 102. The code management store 102 is also accessible from the development environment 104. In this development environment, a developer can develop user interfaces that are stored in the code management store 102 and used by the user interface module 126 to display the user interfaces. The development environment 104 also enables a developer to develop parameterized applications, including, for example, the parameterized collection application 120, the parameterized detection application 122, and the parameterized action application 124. The performed operations resulting from the execution of these one or more applications allow the developer to determine whether these one or more applications performed correctly under one or more given parameter values, i.e., these one or more applications can be tested before they are made available to users so that users are not exposed to one or more applications that do not perform correctly.
[0032] Referring to FIG. 2A , dataflow graph 202 may include data sources 206 a, 206 b, components 208 a-c, 210, and a data sink 212. In this example, dataflow graph 202 is an example of a parameterized application. As will be described in more detail, dataflow graph 202 includes various parameters whose values are specified by user input. For example, if dataflow graph 202 is a parameterized collection application, each of data sources 206 a, 206 b, components 208 a-c, 210, and data sink 212 specifies how data is collected, processed, and integrated (e.g., into execution system 108) in near real-time as data streams are received. In another example, dataflow graph 202 is a parameterized detection application, each of data sources 206 a, 206 b, components 208 a-c, 210, and data sink 212 specifies various detection tasks and functions to be performed in detecting various data records. In yet another example where the data flow graph 202 is a parameterized operational application, each of the data sources 206a, 206b, components 208a-c, 210, and data sink 212 specifies various operational tasks and functions that are performed in response to triggers and / or instructions received from the detection module.
[0033] In this example, each of the sources, components, and sinks may be associated with a set of parameters 204a-g. A user may enter or otherwise specify values for these parameters using user interface module 126 (FIG. 1). The parameters of one source, component, or sink may be used to evaluate the parameters of a different source, component, or sink. Sources 206a, 206b are connected to input ports of components 208a, 208c. The output port of component 208a is connected to the input port of component 208b. The output port of component 210 is connected to data sink 212. The connections between the sources, components, and sinks define the data flow.
[0034] Some data sources, components, or sinks may have input parameters 204a-g that can define some of the graph's behavior. For example, a parameter may define the location of a data source or sink on a physical disk. Parameters may also define the behavior of a component. For example, a parameter may define how a classification component sorts input. In another example, a parameter may define how a data record is formatted or validated. In some configurations, the value of one parameter may depend on the value of another parameter. For example, source 206a may be stored in a file in a particular directory. Parameters 204a may include a parameter called "DIRECTORY" and another parameter called "FILENAME." In this case, the FILENAME parameter depends on the DIRECTORY parameter. (For example, DIRECTORY may be " / usr / local / " and FILENAME may be " / usr / local / input.dat"). Parameters may also depend on parameters of other components. For example, the physical location of sink 212 may depend on the physical location of source 206a. In this example, sink 212 includes a set of parameters 204g, including a FILENAME parameter that depends on the DIRECTORY parameter of source 206a (e.g., the FILENAME parameter in set 204g can be " / usr / local / output.dat", where the value " / usr / local / " comes from the DIRECTORY parameter in set 204a).
[0035] Within the user interface for the client, parameters in parameter groups 204a-204g can be combined and reorganized into various groups for interacting with the user that reflect business rather than technical considerations. The user interface for receiving parameter values based on user input can display various parameters according to their relationships in a flexible manner that is not necessarily constrained by the context of the development environment on the server.
[0036] See FIG. 2B. The user interface can be represented by relationships that display icons to indicate dependencies between parameters. In this example, the parameters are categorized into a first group of parameters represented by a first source icon 224 representing parameters of a first source dataset, a second source icon 226 representing parameters of a second source dataset, a sink icon 230 representing parameters of a sink dataset, and a transform icon 228 representing parameters of one or more components of the dataflow graph being constructed, showing the relationships between the source and sink datasets. This grouping of parameters can be based on a stored user interface specification 222. This specification defines how a user interacts with parameters from the dataflow graph within a user interface for a client and how user interface elements, such as icons 224, 226, 228, and 230, are interrelated and organized for presentation in the user interface. In some embodiments, this user interface specification is an XML document. The user interface specification also identifies dataflow graph components and may identify specific components that enable a user to perform certain functions, such as viewing sample data while constructing a graph, as described in more detail below.
[0037] In some cases, the user interface specification may include instructions regarding how to display parameters. Referring to Figures 2B and 2C, for example, the user interface specification 222 may define an interface 250 that is displayed to a user. Furthermore, the user interface specification 222 may indicate that, upon interaction with a source dataset icon 224, one parameter may be displayed in the user interface 250 as a user-writeable text box 252, while another parameter may be displayed in the user interface 250 as a drop-down list 254 containing pre-canned values, and yet another parameter may be displayed in the user interface as radio buttons 256. In this manner, the user interface specification provides flexibility in how parameters are presented to a user to customize the dataflow graph in a manner that may suit business and / or non-technical users.
[0038] In some cases, user interface specifications may mandate an order in which business users enter parameter values. As represented by the dotted lines, parameters for sink 230 may not be displayed to the user until the user meets certain conditions. For example, the user may be required to provide specific parameter values or fill in a set of parameters before the data sink parameters are displayed.
[0039] In some embodiments, the user interface specification may also include variables that define the characteristics of user interface elements (as opposed to parameters that define the characteristics of components in a dataflow graph). For example, these variables may be used to control the order in which user interface elements are used by a business user. A variable references at least one data value. In some examples, a variable references multiple data values, and each data value is defined as a property of the variable. Thus, a single variable may have multiple properties, each of which may be associated with a data value.
[0040] The user interface 250 defined by the user interface specification may be presented in a way that the user interface elements (e.g., text boxes 252, drop-down lists 254, radio buttons 256) do not directly correspond to parameters used to customize the dataflow graph. Instead, some of the user interface elements may correspond to configuration options appropriate for a user, e.g., a business user and / or a non-technical user who may not have knowledge of the parameters.
[0041] In these examples, the user interface 250 need not be associated with a particular element 224 of the dataflow graph. Furthermore, the user interface 250 may be associated with multiple dataflow graphs and other data processing and data storage structures.
[0042] For example, a user interface element may allow a user to change a configuration option that has business rather than technical meaning. This configuration option may be an option related to converting between types of currencies used in commercial transactions, or an option to update information about a particular category of product inventory, or another type of option that is not correlated to the configuration of a single parameter. The user interface specification 222 may be specified in a way that allows business / non-technical users to change the configuration options using terms they understand, and that allows parameter changes to be made using relationships and dependencies defined in the user interface specification 222.
[0043] The user interface specification 222 can specify how configuration options correspond to configuring parameters of dataflow graphs and other data elements configurable via the user interface 250. For example, interactions between a user and a user interface element can cause changes to parameters in multiple dataflow graph components and changes to data stored in a database, data file, metadata repository, or other type of data store. The user interface specification 222 can specify relationships between the user interface elements and data that changes in association with changes to the user interface elements during operation of the user interface 250.
[0044] User interface specification 222 may also define user interface elements based on data received from a database, a data file, a metadata repository, or another type of data store, or another type of data source, such as a web service. When displaying user interface 250, the received data is used to determine how to display the user interface elements. In one embodiment, during operation of user interface 250, data is received from an external source, such as a database, a data file, a metadata repository, or another type of data store, or another type of data source, such as a web service, and the data received from the external source is defined in user interface 222 as being associated with a parameter (e.g., updating the parameter to include the data received from the external source).
[0045] The user interface may also display component output data associated with at least one flow of data represented by an edge of the dataflow graph. See FIG. 2C. For example, data flows from one component 224 to another component 228. The flow of data between components may be viewed on the user interface 250. In one example, sample data (e.g., data retrieved for testing purposes, rather than for processing or transformation purposes) may be provided to one component 224 to determine how the component 224 handles the data.
[0046] 2C, user interface 250 (or multiple user interfaces) displays user interface elements for specifying values for parameters in each of parameterized collection application 120 (FIG. 1), parameterized detection application 122 (FIG. 1), and parameterized action application 124 (FIG. 1). In this example, user interface 250 displays user interface elements without regard to specifying which user interface elements correspond to which parameterized application. Rather, these user interface elements are presented to allow a user to specify, for example, how to collect and integrate data, how to perform detection, and how to perform action on a particular data stream.
[0047] Referring to FIG. 3A, environment 300 includes a collection and detection action (CDA) system 320 that collects data records, detects satisfaction of one or more predetermined conditions (specified in rules) in the data records, and performs appropriate actions on the detected data records. In this example, execution system 108 of FIG. 1 is shown as CDA system 320, and the data flow of the CDA system is shown. CDA system 320 receives data intermittently (e.g., periodically or continuously) from various data sources, e.g., various servers interconnected in a network. Because the data is received intermittently, the system collects the data into a single data stream (e.g., by adding the received data in batches to a queue) and combines the data into a single long record (e.g., by generating a long record that includes the queued data) in near real time (e.g., after 1 millisecond, 2 milliseconds, etc.). Data is collected from data sources in near real time, rather than being retrieved from a data warehouse (e.g., in batches). This collected data includes, for example, data records including data indicating the occurrence of an event or action (e.g., the occurrence of a phone call or the length of a phone call) or records including data indicating the occurrence of an event or action. By combining data from these various data sources, the long record includes various types of data records (e.g., short message service (SMS) data records, voice data records, etc.). CDA system 320 augments this long record with data record aggregations, data indicating the non-occurrence of an event, status data and various aspects such as customer data (e.g., customer profiles), account data, etc. Environment 300 generates long records of various types of data records (e.g., records that include and / or point to various sub-records) in near real time as the data records are received.
[0048] Typically, data collected from a data stream does not contain all the information required for processing by the CDA system, such as a user's name and profile information. In such cases, the data (i.e., data collected from the data stream) is augmented by combining profile data with data received in the real-time data stream and computing near-real-time aggregates. By combining profile data with data from the real-time data stream and computing near-real-time aggregates, the search and retrieval system creates key data records (e.g., data records including near-real-time received data associated with a key, profile data for that key, and near-real-time aggregates for that key) that meet the processing requirements of the search and retrieval system. Typically, the processing requirements include various tasks to be performed (and / or rules to be enforced) by the system and various data necessary to perform those tasks. This pre-computation or creation of data records including "all data records" or fields (and / or predetermined sets of fields) pre-filled with data corresponding to each data record in the data record can help avoid and reduce congestion at network bottlenecks, for example, when processing real-time data streams. This is because all the data required for processing is contained in a single record (e.g., a record of a group of records), thus eliminating or reducing data lookups, calculations, and database queries at each stage or step of data record processing or record collection. Also, by storing large amounts of augmentation data (e.g., profile data) in the CDA system's memory or cache index, the system can access the data more quickly because it generates pre-calculated records (of a group of records).
[0049] For example, the systems described herein may be configured to load augmentation / enhancement data into memory (or an indexed cache) when the system load is low relative to the load at other times. Because the system has the flexibility to pre-load the augmentation data when the system load is low, the system may enable load balancing by loading the augmentation data into memory when the load is low, rather than having to do it in real time when data records are being processed (which may be a period of increased load).
[0050] In one example, CDA system 320 processes over 2 billion data records per day for 50 million users and calculates aggregates for each data record type. In this example, CDA system 320 receives real-time data streams 340 (e.g., multiple separate data streams, each with its own unique format) from data sources 360. As used herein, real-time includes, but is not limited to, near-real-time and substantially real-time. In each of these cases, there may be a time lag between when the data is received or accessed and when the processing of that data actually occurs, but the data is still processed in real time as the data is received. From real-time data streams 340, CDA system 320 intermittently receives data including data records, also referred to as data items. The received data also includes data records of various types (e.g., various formats). In one example, one of the first real-time data streams includes data representing data records of a first type / format, and one of the second real-time data streams includes data representing data records of a second type / format. CDA system 320 includes a collection module 420 that collects various types of data records received in real-time data stream 340. Because collection module 420 operates on real-time data records rather than data extracted from the EDW, CDA system 320 can provide immediate response to data records (as they are received) and near-real-time aggregation of data records, which also provides immediate visibility of application results. Collection module 420 collects the data records into a single data stream and adds these data records together to a queue. In one example, collection module 420 collects data records by using a continuous flow to continuously process received data records.
[0051] As data records from the real-time data stream 340 continue to be intermittently received by the collection module 420, the collection module 420 detects (e.g., in a queue) two or more particular data records that share a common attribute, such as being included in a data record palette or being associated with a particular user attribute (e.g., a user identifier (ID), a user key, etc.). In one example, this common attribute is a converted value of a particular field (e.g., a user ID field) of the two or more particular data records, the two or more particular data records belong to a specified data record type, and / or the two or more particular data records are defined by the data record palette.
[0052] The collection module 420 creates a collection of data records that includes two or more specific data records detected. In this example, the collection module 420 creates a data record 460 that includes the collection of the detected data records. The collection module 420 also inserts augmentations 440 into the data record 460, e.g., a long record. Typically, the augmentations are data (previously received or calculated) stored in the data warehouse that are associated with the data record. For example, the data record may specify the number of SMS messages sent by a user and may also include the user ID of the user. In this example, the data warehouse 380 receives the data record 461 and stores data that includes (or is associated with) the same user ID. This stored data includes, for example, user profile data, including the user's most recent handset type. The collection module 420 appends or inserts customer profile data for the customer associated with the specific data record included in the data record 460 into the data record 460.
[0053] The collection module 420 may filter the received data, for example, so that only a portion of the received data is enriched and added to the data records 460. The collection module 420 may be configured to filter based on a key (associated with the record) and / or based on specified values of specified fields of the record. The collection module 420 may also correlate the received data and / or data records such that records associated with the same or similar keys are grouped together, for example, to enable complex data record processing (e.g., processing of records associated with a particular key that are staggered in time). In another example, the collection module 420 may correlate data records based on records having certain fields, certain values of certain fields, etc. In this example, values of the fields of the correlated records are inserted or added to a longer record.
[0054] The collection module 420 also calculates one or more aggregates (i.e., data record aggregates) of one or more data records contained in the data records 460. For a particular data record of a particular user (specified by a user ID contained in the data record), the collection module 420 retrieves batch data 400 of the particular data records of the particular user from the data warehouse 380. The batch data 400 includes historical aggregates for the particular data record, which are pre-computed aggregates of data record data from a previous period, e.g., from a start time to a particular time prior to the execution of the data record detection. Generally, the data record data includes data indicating particular properties, attributes, or characteristics of the data record (e.g., the amount of data usage of the data record). For example, the properties of the data record may include particular fields (contained in the data record), particular values of fields contained in the data record, particular user ID keys contained in or associated with the data record, the absence or value of particular fields in the data record, etc. Based on the data contained in the real-time data stream 340 for the particular data record for the particular user and based on the historical aggregation, the collection module 420 calculates combined data record data, e.g., a near real-time aggregation of the data record. The collection module 420 augments the data record 460 with the combined data record data for the at least one particular data record.
[0055] In one example, one data record in data records 460 is data usage for "someone" with user ID 5454hdrm. In this example, collection module 420 retrieves batch data 400 of "data usage" data records associated with user ID 5454hdrm from data warehouse 380. To calculate a near real-time aggregate of this data record for this particular user, collection module 420 aggregates batch data 400 with incremental data 410 to calculate a near real-time aggregate 430 of this data record.
[0056] In this example, incremental data 410 includes a portion of the data received from real-time data stream 340 for the data record type being aggregated for that particular user. The incremental data 410 spans the time from when the historical aggregation was last calculated to approximately the current time, e.g., the time the near real-time data stream was received. For example, batch data 400 may indicate that user "A" used 65 megabytes of data in the last month, and incremental data 410 may indicate that user "A" used 1 megabyte of data in the last five minutes. By aggregating batch data 400 with incremental data 410, collection module 420 calculates near real-time aggregation 430 for this particular data usage data record for customer "A." Collection module 420 inserts near real-time aggregation 430 into data record 460, e.g., as part of the record for this particular data record for this particular user. The collection module 420 also attaches an additional reference file (ALF) to the data record 460 with historical aggregations for that particular data record, for example, as specified by the batch file 400. The collection module 420 attaches the ALF with the historical aggregations to facilitate the use of the historical aggregations in the calculation of new near real-time aggregations as new data records are received, for example.
[0057] In this example, collection module 420 forwards data records 460 to detection module 480. Detection module 480 includes rules 500, which may include, for example, rules for implementing a variety of different applications for different types of entities. Detection module 480 includes a single module for implementing the various applications and performing aggregation.
[0058] The detection module 480 calculates one or more aggregates (i.e., data record aggregates) of one or more data records contained in the data records 460. For a particular data record of a particular user (designated by a user ID contained in the data record), the detection module 480 retrieves batch data 400 of the particular data record for the particular user from the data warehouse 380. The batch data 400 includes historical aggregates for the particular data record, which are pre-computed aggregates of data record data from a previous period, e.g., from a start time to a particular time prior to the execution of the data record detection. Generally, the data record data includes data indicating particular properties, attributes, or characteristics of the data record (e.g., the amount of data usage of the data record). For example, the properties of the data record may include particular fields (contained in the data record), particular values of fields contained in the data record, particular user ID keys contained in or associated with the data record, and particular fields or the absence of values for those particular fields in the data record. Based on the data contained in the real-time data stream 340 for the particular data record for the particular user and based on the historical aggregation, the detection module 480 calculates combined data record data, e.g., a near real-time aggregation of the data record. The detection module 480 augments the data record 460 with the combined data record data for the at least one particular data record.
[0059] In one example, one data record in data records 460 is a data usage for "someone" associated with user ID 5454hdrm. In this example, discovery module 480 retrieves batch data 400 of "data usage" data records associated with user ID 5454hdrm from data warehouse 380. To calculate a near real-time aggregation of this data record for this particular user, discovery module 480 aggregates batch data 400 with incremental data 410 to calculate a near real-time aggregation 430 of this data record.
[0060] In this example, incremental data 410 includes a portion of the data received from real-time data stream 340 for the data record type being aggregated for that particular user. The incremental data 410 spans the time from when the historical aggregation was last calculated to approximately the current time, e.g., the time the near real-time data stream was received. For example, batch data 400 may indicate that user "A" used 65 megabytes of data in the last month, and incremental data 410 may indicate that user "A" used 1 megabyte of data in the last five minutes. By aggregating batch data 400 with incremental data 410, detection module 480 calculates near real-time aggregation 430 for this particular data usage data record for customer "A." Detection module 480 may insert near real-time aggregation 430 into data record 460, e.g., as part of the record for this particular data record for this particular user. The detection module 480 also attaches an additional reference file (ALF) to the data record 460 with historical aggregations for that particular data record, for example, as specified by the batch file 400. The detection module 480 attaches the ALF with the historical aggregations to facilitate the use of the historical aggregations in the calculation of new near real-time aggregations as new data records are received, for example.
[0061] In this example, the CDA system 320 retrieves data representing one or more rules that define an application from a user's client device. For example, the user can use a data record palette to define the rules. The CDA system 320 generates the one or more rules that define the application based on the received data. The CDA system 320 passes the one or more rules to a process, e.g., the detection module 480, configured to implement the one or more rules. The detection module 480 implements the application based on the execution of the rules 500 against the data records 460. The detection module 480 also includes state transitions 530, which include data indicating, for example, a state to which the user has transitioned or progressed in the application. Based on the state transitions 530, the detection module 480 identifies actions to be performed in the application and / or choices to be made in the application. For example, based on a particular user's state in the application specified by the user's state transitions 530, the detection module 480 identifies which components of the application have already executed and which components of the application should be executed next according to the user's application state.
[0062] Data records 460 include various types of data records, such as SMS data records, voice data records, etc. Thus, rules 500 include rules related to conditions for a variety of different types of data records. Generally, rules include conditions that, when met, cause an action to be performed. In this example, one rule ("Rule 1") may have a condition that the user has sent 30 SMS messages within the last six months. When this condition is met, "Rule 1" specifies an action to provide the user with a $5 credit. Another rule ("Rule 2") may have a condition that the user has used less than 50 megabytes of data in the last month. When this condition is met, "Rule 2" may specify an action to provide the user with a usage discount to encourage increased data usage, for example. In this example, both "Rule 1" and "Rule 2" use different types of data records (i.e., SMS data records and voice data records, respectively). Because data record 460 is a single long record containing various data record types, detection module 480 can execute programs containing rules that depend on various types of data records. Additionally, detection module 480 is a single module that executes applications for multiple different applications because detection module 480 receives data records 460 that contain all data record types at all different operational levels. That is, rather than providing different modules that execute different applications on different data records (each of which contains a type of data appropriate for a respective application), detection module 480 is configured to execute multiple different applications on a single long record, i.e., data record 460.
[0063] Upon detecting a data record (or aggregation of data records) in data records 460 that satisfies at least one of the conditions in rule 500, detection module 480 adds action trigger 510 to queue 520 for initiation of one or more actions (e.g., those specified in the rule having the condition that is met). In one example, the action trigger includes data specifying the action to perform, the application for which the action is to be performed, and the user (e.g., the user for which the action is to be performed). Detection module 480 forwards queue 520 to action module 540 for execution of the action specified in action trigger 510. In this example, action module 540 is configured to perform various actions, such as issuing credit to a user account, sending a message, sending a discount message, etc.
[0064] See FIG. 3B. Diagram 541 illustrates the lifecycle of executable logic (e.g., of a campaign). In this example, the executable logic is part of the detection engine. In this example, the system does not start a campaign until a triggering event occurs. Typically, a triggering event includes an event that meets one or more specified conditions or attributes. The triggering event is based on the input event stream and the subscriber profile. The system defers calculation of the control group until the triggering event occurs, as described below. The system can also configure a campaign to end early, for example, if an acceptance event occurs. The system can chain campaigns so that another campaign can be launched after acceptance or expiration.
[0065] In this example, start node 542 represents the start of a campaign. In this example, start node 542 represents executable logic that specifies start and end dates for the campaign. The executable logic also specifies rules (part of a collection of rules, referred to herein as "decision rules") to be executed for the campaign. In this example, node 543 represents decision rules (e.g., executable logic) for detecting a triggering event, e.g., an event that indicates that a campaign should begin. In this example, the decision rules include logic used to trigger the campaign, e.g., by specifying which event will start the campaign. The decision rules also include logic that specifies offer content, campaign duration, and message priority. The decision rules also include logic to detect when the offer is accepted.
[0066] At the start of a campaign, the system executes a decision rule for every event received. Based on the event, the system decides whether to launch a campaign. Possible outcomes (of the execution of a decision rule on an event) are to ignore the event, end the campaign, ignore the event until the next day, or launch the campaign. As explained in more detail below, the decision rule also specifies what to do if the message cannot be sent. For example, the rule may specify to cancel the campaign, launch the campaign, or resume searching for triggers later, e.g., the next day.
[0067] After the campaign is launched, the system executes control group logic, represented by node 544, to determine whether an event (e.g., a data record) should be assigned to a control group that receives the offer or does not receive the offer. The control group assignment is just-in-time, e.g., because it occurs only after the campaign is launched. The system also calculates the control group using a match panel cube, which is described in more detail below.
[0068] Following assignment to a target group (e.g., a group that receives an offer) or a control group, the system performs message moderation, represented by node 545. In this example, message moderation refers to the process upon which the system relies when searching for relative message priority. In this example, the system is configured to honor new message limits in the subscriber profile. For example, the subscriber profile may specify a predetermined (e.g., maximum) amount of “general” messages per day that the subscriber can receive. The subscriber profile may also specify a predetermined (e.g., maximum) amount of “urgent” messages that the subscriber can receive on a particular day. In this example, each message has an urgency and a priority. Different types of urgency include normal, urgent, or unlimited. In this example, the system prioritizes messages so that urgent messages are sent before normal messages if there are too many messages to send at the same time. In this example, the decision rule specifies a message transmission time. If this time is in the future, the system postpones sending the message to allow higher priority messages an opportunity to be sent instead.
[0069] In one example, when a campaign is first launched, the system performs message moderation. If the first message cannot be sent (e.g., because a limit on the maximum amount of messages has been exceeded), the event launch can be canceled. In this example, for subsequent messages, there is no option to cancel the campaign if message moderation fails. In this example, the subscriber simply does not see the message.
[0070] Part of the reconciliation process is message (e.g., offer) priority lookup. In this example, message priority and urgency are specified in a lookup (reference) file. The search field keys include a subject field, a type field, and a priority key field. That is, each offer or message is pre-configured with a specified priority and urgency. In this example, based on the results of message reconciliation, a message can be sent, as represented by node 546.
[0071] Following the sending of the message, the system is configured to take a further waiting action, as represented by node 547 (e.g., by waiting until the next day or by retrieving (i.e., the offer fulfilled)). The system performs this further waiting action by executing an additional decision rule. In this example, a decision rule is executed for every event. The rule is configured to identify whether to end the campaign based on the event (e.g., based on detecting the occurrence of a target event). In this example, the target event is typically a notification that a fulfillment has occurred. However, any event or condition can be used to trigger a campaign phase end. In this example, the available options are to ignore the event, end the campaign, and start a new phase. Typically, a campaign phase refers to a distinct offer provided by the campaign. In this example, when a campaign phase ends, the decision rule can start a new phase. This decision rule is executed after the campaign phase has run for a specified number of days (e.g., N days), for example, if no fulfillment event has been detected. When a campaign phase ends, the decision rule starts a new phase. In this example, the available options are to end the campaign or start a new phase.
[0072] Following execution of the next rendezvous event, the system also performs one or more actions to execute a calendar rule, as represented by node 548. For example, a calendar rule may specify that the campaign is configured to send a reminder message every day after the initial message. This rule may also be used to suppress this message for some days. Based on the calendar rule, the system determines whether to resend the reminder message. If the system decides to resend the message, the system again performs message reconciliation, as represented by node 549. Based on the results of the message reconciliation, the system may send the message again, as represented by node 550. Typically, the executable logic represented by the nodes in this diagram can be configured by a business rules editor (e.g., as described in U.S. Pat. No. 8,069,129, the entire contents of which are incorporated herein by reference) or a flowchart user interface editor (e.g., as described in U.S. patent application Ser. No. 15 / 376,129, the entire contents of which are incorporated herein by reference).
[0073] See FIG. 3C. Networked environment 560 includes data sources 562, 570 and CDA system 577. In this example, CDA system 577 includes parameterized collection application 573, parameterized detection application 580, and parameterized action application 590. In this example, one of parameterized collection applications 573 includes, for example, dataflow graph 573a including nodes each representing one or more data processing operations. One of parameterized detection applications 580 includes dataflow graph 580a including nodes 582, 583, 584, and 585. In some examples, dataflow graph 580a includes a state diagram, for example, such that dataflow graph 580a specifies various executable logic to be executed in various states. In this example, the state diagram may be parameterized, for example, to allow input of values for various parameters in the state diagram. Also, a single key may be associated with multiple state diagrams (not shown), for example, to maintain the state of the key across various dataflow graphs. Each node represents one or more portions of executable logic that are executed in a particular state (of that executable logic). Accordingly, nodes 582, 583, 584, and 585 are hereinafter referred to as states 582, 583, 584, and 585, respectively. In this example, one of the parameterized behavior applications 590 includes a dataflow graph 590a.
[0074] In this example, parameterized collection application 573 includes specification 574 that specifies one or more parameters (e.g., parameters A and C) that define one or more characteristics of parameterized collection application 573 (or one of parameterized collection applications 573) and one or more respective values (e.g., values B and D) of those one or more parameters. For example, parameter A may be a parameter that specifies a data format, and parameter C may be a parameter that specifies a data source from which data items are collected. In this example, value B (of parameter A) specifies a data format to be applied to transform the collected data items. Value D (of parameter C) specifies a data source from which data items (e.g., to be transformed) are collected. In one example, the transformation includes correlation. In this example, data records are transformed by correlating data records associated with the same key and then adding a data record (e.g., a master record) that represents the correlated records to a queue. In this example, parameterized collection application 573 is also configured to perform enrichment, filtering, and formatting. In this example, the parameterized collection application 573 augments the events (e.g., data records) themselves, rather than augmenting data records with, for example, profile data (as would be done by the parameterized detection application 580). For example, an SMS message event is augmented by including data specifying the location of the cellular tower that relayed the SMS message. In this example, the parameterized collection application 573 augments the event with this data, for example, by retrieving this data from one or more internal or external data sources.
[0075] Parameterized detection application 580 includes specification 586, which specifies one or more parameters (e.g., parameters E and G) that define one or more characteristics of parameterized detection application 580 (or one of parameterized detection applications 580) and one or more respective values (e.g., values F and H) of those one or more parameters. For example, parameters E and G can be parameters included in a rule (e.g., a rule that detects specified events) executed by parameterized detection application 580, and values F and H are values of those parameters (e.g., values that specify various types of events to be detected). Specification 586 also includes rules 587 and their respective conditions. Specification 586 also includes state data 592, which specifies various states of the parameterized detection application, including, for example, data flow graph 580a. In this example, state data 592 specifies that parameterized detection application 580a has four states, i.e., states 1 through 4, corresponding to states 582 through 585, respectively. In this example, some of the rules 587 are executed in certain states according to data structures 593, 594, 595 (e.g., pointers). For example, data structure 593 specifies that rule 1 is executed in state 1. No rules are executed in state 2. Rule 2 is executed in state 3 according to data structure 594. Rule 3 is executed in state 4 according to data structure 595.
[0076] The parameterized detection application 580 stores profile data 575, including KPIs 576 and keyed state data 589. Typically, KPIs include data obtained through detection (e.g., data on various events contained in or represented in collected records). Thus, the values of KPIs are updated and changed on the fly and in real time. For example, the CDA system 577 defines KPIs that track data usage. In this example, when a new data record is received, the parameterized detection application 580 detects whether the new data record includes or specifies a data usage event. If the data record specifies a data usage event, the CDA system 577 updates the KPI (for the key associated with the data usage event) with data specifying the updated data usage. In this example, each KPI is a keyed KPI associated with a particular key. In one example, the KPIs are used in the execution of this rule to determine whether various conditions of the rule are satisfied, for example, by determining whether the content of the KPI for a particular key satisfies the condition of the rule to be executed in the current state of the key's specification 586. KPIs can also be calculated based on multiple events. For example, a KPI can be defined to indicate when a user has used a specified amount of data and a specified amount of voice usage. In this example, the KPI is based on two events: a data event and a voice usage event. KPIs themselves can also be aggregated, and aggregated KPIs are used in detecting events and / or detecting satisfaction of rule conditions. KPIs are stored as part of a customer profile. In some examples, KPIs are attributes defined by the user and / or a system administrator.
[0077] In this example, keyed state data 589 specifies a state of the parameterized detection application for each value of the key. In this example, each data item received by CDA system 577 is associated with a key value. For example, a key can be a unique identifier such as a subscriber identifier. CDA system 577 maintains a state of the parameterized detection application 580 for each key value. For example, the key value "349jds4" is associated with state 1 of the parameterized detection application 580. The key value "834edsf" is associated with state 3 of the parameterized detection application 580 at a first time point (T1) and state 4 at a second time point (T2).
[0078] In this example, parameterized action application 590 includes specification 597 that specifies parameters I, K and respective values J, L of these parameters that define one or more characteristics of parameterized action application 590. For example, parameter I may be a parameter that specifies that profile data is to be used in customizing a user's action. In this example, value J of parameter I specifies the profile data to use in this customization.
[0079] During operation, CDA system 577 executes one or more parameterized collection applications 573. This execution involves processing of data records with the one or more values of the one or more parameters specified by specification 574. In this example, this processing involves collection by parameterized collection application 573 of data items 566a, 566b...566n in data stream 564 from data source 562. In this example, each of data items 566a, 566b...566n is a keyed data item, e.g., a data item associated with a key. One or more parameterized collection applications 573 also collect data items 572a, 572b...572n (which are part of batch data 568) from data source 570. In this example, each of data items 572a, 572b...572n is a keyed data item. In this example, the format of data items 566a, 566b...566n is different from the format of data items 572a, 572b...572n.
[0080] One or more parameterized collection applications 573 transform data items 566a, 566b...566n and data items 572a, 572b...572n according to specification 574 to obtain transformed data items 579a...579f. In this example, each of transformed data items 579a...579f is converted to a data format suitable for parameterized detection application 580. One or more parameterized collection applications 573 add transformed data items 579a...579f to queue 579 and forward the added queue 579 to parameterized detection application 580.
[0081] CDA system 577 executes one or more parameterized detection applications 580. During this execution, the one or more values of the one or more parameters are specified by specification 586, and the transformed data items 579a...579f in queue 579 are processed as follows: CDA system 577 can augment the transformed data items with one or more portions of profile data 575 and / or KPIs 576. For example, CDA system 577 creates augmented transformed data item 598 by adding profile data (e.g., data usage, SMS usage, geolocation, and data plan data, etc.) to transformed data item 579f. In this example, augmented transformed data item 598 is associated with a particular value of key ("834edsf"). Parameterized detection application 580 identifies the current state of one or more parameterized detection applications 580 with respect to the particular value of key. In this example, at time T1, the current state of one or more parameterized detection applications 580 is state 3, as specified by keyed state data 589. Parameterized detection application 580 identifies one or more rules 587 within a portion of specification 586 as being to be executed in the current state. In this example, rule 2 is executed in state 3, as indicated by the dotted line around rule 2 in FIG. 3C. Parameterized detection application 580 executes the identified one or more rules (e.g., rule 2). Parameterized detection application 580 determines that at least one of the one or more transformed data items (e.g., augmented, transformed data item 598) satisfies one or more conditions of at least one of the one or more rules (e.g., rule 2) to be executed in the current state. In response to this determination, parameterized detection application 580 generates data structure 521 that specifies the execution of one or more actions (represented by action data 511). The parameterized detection application 580 also causes a transition from the current state of the specification 586 to the next state for a particular value of the key (eg, 834edsf).In this example, this transition is shown as a transition from time T1 to time T2, where the state of specification 586 transitions from state 3 to state 4 for key value 834edsf. Parameterized detection application 580 then forwards the generated data structure 521 to parameterized action application 590.
[0082] CDA system 577 executes parameterized action application 590, where the one or more values for the one or more parameters are specified by specification 597, and performs operations including: sending one or more instructions 591 to cause execution of at least one of the one or more actions based on at least one of the one or more actions specified in data structure 521;
[0083] As one transformation, parameterized collection application 573 accesses profile data (e.g., profile data 575) for each received data record and augments the data record with the accessed profile data. In this example, parameterized detection application 580 compares the content of the profile data with one or more rules and / or applications included in parameterized detection application 580 to detect the occurrence of one or more predetermined events. After detection, parameterized detection application 580 updates KPIs accordingly, for example, with the detected information and / or information specifying the detected events.
[0084] In yet another transformation, there are multiple queues between the collection application and the detection application. For example, there can be a priority queue for certain types of events, e.g., events whose processing cannot be delayed. Processing of other events can be delayed. These other events are assigned to another queue, e.g., a non-priority queue, and alarms are inserted into this queue to indicate that processing of these events will be delayed. Typically, an alarm is generated when the chart itself sends a special type of event (e.g., an alarm event) that arrives at a pre-calculated time in the future. Alarms are used whenever the chart logic wants to enter a queue. Having multiple queues allows the system to balance the event processing load and also reduces latency in processing events by processing events in the priority queue first and deferring events in the non-priority queue. In this example, the parameterized collection application is parameterized and configured with rules specifying alarm insertion rules and the various event types associated with the alarms.
[0085] Referring to FIG. 4, system 600 includes a collection unit 610 that receives data records 601-608, as indicated by 615, for example, in a batch manner and / or from a real-time data stream. In this example, collection unit 610 stores records 601-608 in memory, for example, in a data structure 620 (e.g., an index) in memory. In this example, data structure 620 is not a static data structure. Rather, data structure 620 is a dynamic data structure that is intermittently updated and changed as new records are received in data stream 619. Also, entries in data structure 620 are removed after the entries (e.g., logical rows) are processed, for example, by being assigned to data structure 630 or by being filtered from data structure 620. Data structure 620 includes logical rows 621-628. In this example, each of logical rows 621-628 corresponds to one of data records 601-608. Data structure 620 also includes logical columns 620a-620d, which correspond to fields and / or field values in data records 601-608, respectively. In this example, each logical row of data structure 620, such as logical rows 621-628, corresponds to a subset of relevant information extracted from a particular data record 601-608 by collection unit 610. Each logical column of data structure 620 conceptually defines a particular data attribute of the particular data record associated with the particular logical row. In an example where data structure 620 is an index, each of logical rows 621-628 is an indexed entry.
[0086] The collection unit 610 also stores in memory a data structure 617. In one example, the data structure 617 represents a static data structure stored in a data store or repository. In this example, the data structure 617 stores augmentation data, e.g., profile data. In this example, the augmentation data stored in the data structure 617 can be added to a received data record on the fly, e.g., by adding or appending the particular augmentation data to a logical row of the data structure 620. The data structure 617 also specifies the data types to be included in the data record 630, e.g., the data types to be dynamically added to a received record on the fly as the record is received. In this example, the data structure 617 includes logical rows 617a-617h, each of which includes augmentation data for a particular ID. In this example, the data structure 617 includes a logical column of ID data (not shown). By combining the ID data in data structure 620 with the ID data in data structure 617, collection unit 610 creates an augmented record for a particular ID, e.g., a record that includes data received from a real-time data stream and is then augmented with the augmented data.
[0087] The collection unit 610 also filters and correlates the data records 601-608 received in the data stream 619. In one transformation, the collection unit 610 collects records received from a batch search, e.g., from a data store. In this example, the collection unit 610 filters records containing non-Boston values in the "location" field. That is, the collection unit 610 filters the data records 601-608 to include only data records with a value of "Boston" in the location field. The collection unit 610 also correlates the records 602, 603, and 606 remaining after filtering. For example, the collection unit 610 correlates record 602 with record 603 (because both are associated with the same ID). The collection unit 610 associates records 602 and 603 with each other by including logical rows 631 and 632 in the data structure 630 that represent these records next to each other. In one transformation, the collection unit 610 merges the correlated records 602 and 603 into a single record. In yet another example, the collection unit 610 performs the correlation by creating correlated aggregations.
[0088] In correlated aggregation, the aggregated value is not the same as the returned value. Rather, the collection unit 610 performs the correlation using various fields in the records (separate in time). In this example, the correlation is particularly complex because the collection unit 610 combines records that are separated in time based on, for example, specific values of fields in those records. Below is an example of the correlation performed by the collection unit 610: [For each client, for each data record of type trade (where trade action = "buy"), calculate the symbol with the maximum trading volume seen in the last 20 minutes.] In the above example, the underlined portions represent the portions of the aggregate definition that are parameters that may be specified. Below is a list of possible parameters for correlated aggregates. In this example, these parameters may be specified as part of the collection application:
[0089] [Table 1]
[0090] In this example, the collection unit 610 uses a "field" or "expression" along with a "filtering expression" and a "selection function" to select appropriate data records from a "time window." We then calculate a value from the data records and use that value as the aggregate value. In one example, a data record has two fields, a quantity field and a symbol field. For the collection unit 610 to identify the symbol corresponding to the data record with the greatest quantity (within a given period), in the collection application, the "selection function" is set to "latest max," the "field" or "expression" to "quantity," and the "expression to be calculated" to "symbol."
[0091] In yet another transformation, the collection unit 610 correlates data records that are separated in time, for example, by collecting particular data records that are related by or contain a particular key and then waiting a specified amount of time for another data record that contains or is associated with the same key. When the collection unit 610 collects correlated data records (e.g., data records that are related by the same key), the collection unit 610 packetizes these correlated data records, for example, by merging these data records into a single record and creating a data packet that includes this single merged record. In one example, because the data records are separated in time, the collection unit 610 stores data indicative of the collected data records in memory and waits a specified period of time to see if another data record is received that has a key that matches the key of the previously stored (e.g., temporary in memory) data record. Because of the volume of data records received by the collection unit 610, the collection unit 610 provides an in-memory grid (or another in-memory data structure) to track and store keyed data representing received data records and (optionally) timestamps representing the times at which these records were received. When a new record is received, the collection unit 610 identifies the key of the newly received record and searches the in-memory data grid for a matching key. If a matching key is present, the collection unit 610 correlates the data records for the matching key, for example, by merging the records into a single record. The collection unit 610 can also update the in-memory grid with data indicating that the next data record for that key has been received. The collection unit 610 is also configured to purge or otherwise remove entries in the in-memory grid after a specified amount of time has elapsed since the time indicated in the timestamp.
[0092] In this example, data structure 630 represents the filtered and correlated records and includes logical rows 631-633 and logical columns 630a-630e. Data structure 630 is not a static data structure, but rather a dynamic data structure that is updated intermittently, for example, as new records are received. An entry (e.g., a logical row) is removed from data structure 630 after dynamic logic (described in more detail below) is applied to the particular entry. Following application of the dynamic logic to a particular entry, data structure 630 no longer retains or continues to store that entry. In this example, logical column 630e represents augmented data, for example, augmented data for a particular ID. Logical row 631 represents a version of record 602 that has been augmented by including augmented data associated with ID "384343."
[0093] The collection unit 610 forwards the data structure 630 to the detection module 634. In this example, the detection module 634 executes the executable logic 644 as described herein. The detection module 634 also executes the dynamic logic (e.g., performing dynamic segmentation), for example, by executing logic contained in a dynamic logic data structure 640. In one example, the dynamic logic includes executable logic configured to dynamically process input data records on the fly as they arrive. In this example, the data structure 640 includes logical rows 646, 648 and logical columns 640a-640d that specify rules. In this example, the detection module 634 is configured to execute the data structure 640 against the data structure 630 to determine records represented in the data structure 630 that satisfy the logic defined by the data structure 640. After detecting records that satisfy this logic, the detection module 634 executes one or more portions of the executable logic 644 that are associated with portions of the logic contained in the data structure 640 (e.g., segments defined by the data structure 640). That is, executable logic 644 has different portions (e.g., rules) associated with different portions of this dynamic logic. Not all portions of executable logic 644 are executable for all portions of this dynamic logic. Detection module 634 executes a portion of executable logic 644 for detected records that satisfy a portion of this dynamic logic. For example, record 603 represented by logical row 632 satisfies the dynamic logic contained in logical row 646 in data structure 640. In this example, a portion of executable logic 644 is defined as being executable for the logic defined by logical row 646 in data structure 640. Thus, detection module 634 executes that portion of executable logic 644 for record 603 (or for the data contained in logical row 632). Based on the execution of that portion of executable logic, detection module 634 creates instructions 642 to send a targeted message (e.g., an offer to replenish cell phone airtime) to a client device associated with the user represented by the ID contained in logical row 632 of logic column 630a.
[0094] In one example, dynamic logic included in data structure 640 specifies various segments, e.g., population segments. For example, logic included in logic row 646 specifies a particular segment. In this example, the segmentation performed by detection module 634 is dynamic because it is performed “on the fly” as data records are received and processed in real time by system 600. Traditionally, customer records are stored on disk (e.g., in a data repository), and then customers are “segmented,” for example, by applying various segmentation rules to the customer records. This is an example of static segmentation because an unchanging collection of customer records is segmented. In contrast, here, there is no unchanging collection of records. Rather, records are continuously and / or intermittently received by system 600. As records are received, system 600 dynamically segments them, for example, by using a continuous flow of on-the-fly processing of the records, part of which includes segmentation. Because the records are dynamically segmented, system 600 can detect when a user or client device enters a particular geographic location, e.g., a mall, and send a targeted message to the client device at that time while the user is still in the mall. In one example, the dynamic segmentation rules can be specified by a parameterized application, e.g., one setting of a parameterized detection application.
[0095] In one example, the collection unit 610 or another component of the CDA system generates a target group (TG) and a control group (CG) against which, for example, dynamic logic is executed. Typically, the control group includes a set of users (e.g., subscribers) who meet specified criteria of the dynamic logic and are therefore candidates for inclusion in a particular group (e.g., for inclusion in a campaign), but who are specifically outside the group so that their behavior can be compared with other users who are actually within the group (e.g., and therefore users who are being offered the campaign, i.e., target group users). In this example, the collection unit dynamically determines the target group and control group on the fly in real time as data records are received. That is, rather than calculating the control group and target group from a static data set, the collection unit 610 determines the target group and control group from a dynamically changing data set (e.g., intermittently updated and changing as new records are received).
[0096] In one example, a control group is defined as follows: if data records are associated with a particular logic (e.g., present in a given campaign), they are assigned to a TG or CG by the collection unit 610. In one example, a data record is associated with a particular logic if the key of that data record is designated to be associated with that logic. To ensure that TGs and CGs have similar characteristics and attributes, the collection unit 610 creates match panels that divide the data records and stores them in memory. The match panel (also called a match panel cube) contains a multidimensional grid (e.g., including the entire subscriber base) that represents the keys of the received records, with each key (e.g., subscriber) assigned to one "cube" (cell). These dimensions describe various aspects associated with the key (e.g., data fields indicating: average revenue per user (ARPU), geography (rural or urban), age-on-network, etc.). Keys with the same values of these dimensions appear in the same cube of the patch panel.
[0097] Next, as described below, for example, the collection unit 610 can determine which keys (or data records) to place in the TG or CG for each cube in the match panel using specified logic. Target vs. control group membership determination is also made in a specified chart (e.g., created from the collection application). Late-arriving records are handled by the chart logic if they affect the (campaign) logic. For example, if the logic relies on data record A arriving before data record B, but due to operational delays, data record B may arrive first, the chart can be created to accommodate this possibility. If there is a possibility of time distortion between the systems transmitting these data records, the time chart logic also accommodates this time distortion. In one example, logic to assign keys (or records) to either the TG or the CG (or to neither—for data records where both the TG and CG are saturated) is executed by the collection unit 610 as follows: If the logic is configured not to require a CG, the collection unit 610 assigns the key to the TG.
[0098] The logic is as follows: for each key, determine (e.g., in a table and / or memory) whether the key is pre-designated as belonging to a CG. If so, the collection unit 610 assigns the key to the CG. For each cube in the match panel, the collection unit 610 modifies the entry in the CG as follows: if less than a designated amount (e.g., a designated percentage) of a predetermined group of keys (e.g., representing a campaign population) belonging to that cube are present in the CG, a shortage condition exists. If a shortage condition exists, move the TG keys in that cube to the CG for that cube. If more than a designated amount of a predetermined group of keys belonging to that cube are present in the CG, an excess condition exists. If an excess condition exists, the collection unit 610 marks these keys as not to be placed in either the TG or the CG. In one example, the collection unit 610 calculates a measure of the CG's suitability (scored as blue, yellow, or red). This measure is a dynamic measure that changes intermittently because the CG changes intermittently.
[0099] Referring to FIG. 5A, an exemplary networked environment 700 for performing real-time CDA functions is shown. In this example, the networked environment 700 includes a CDA system 702, a client system 704, an external data source 706, a network data source 708, and an external system 710. The CDA system 702 includes an ingestion / integration module 712 that performs the functions of, for example, the ingestion module described above. In this example, the ingestion module 712 stores data in a data warehouse 732 (indicated by arrow 756). The ingestion module 712 stores data both in near real-time (e.g., as data is received) and in batches. The ingestion module 712 also retrieves data from the data warehouse 732 (indicated by arrow 758), for example, to enrich data records and create long-form records. In this example, the ingestion module 712 includes an in-memory archive 715 that stores data retrieved from the data warehouse 732. In this example, the collection module 712 may also send some of the collected data back to the external data source 706, as indicated by arrow 717, for example, to facilitate data feedback and form a data feedback loop. The CDA system 702 also includes an application 714 that includes a detection module 716 and an action module 730 that perform the tasks and functions described above.
[0100] In this example, the detection module 716 includes an in-memory data store 718 that stores profile data (e.g., of users of the CDA system 702) as well as state data (e.g., indicating the execution state of applications), data flow charts, campaigns (e.g., including a series of applications and / or data flow charts), and the like. As shown in this example, profiles are stored in memory rather than on disk, e.g., to reduce and / or eliminate delays in data retrieval and enable real-time maintenance of state. In this example, the detection module 716 performs enrichment, e.g., by enriching data records 723 with profile data stored in the in-memory data store 718. In this example, the detection module 716 performs enrichment for a parameterized detection application. The enrichment performed by the detection module 716 involves retrieving profile data associated with a key (e.g., profile data for a particular user) from the memory 718 and writing the retrieved profile data to one or more fields of the data record 723 (for that key) according to instructions specified by the parameterized detection application specification.
[0101] The CDA system 702 also continuously or intermittently updates a user's profile, for example, by updating the profile data in the in-memory data store 718. For example, if the operations module 730 sends an offer to a user, the CDA system 702 updates the user's profile data with data indicating that an offer was sent and which offer it was. For example, each offer includes or is associated with a key or unique identifier. The profile data also includes or is associated with a key or other unique identifier. After the operations module 730 sends an offer, the CDA system 702 identifies the key associated with the offer and updates the profile data associated with that same key with data indicating that the offer was sent.
[0102] Based on this profile data, the CDA system 702 also generates KPIs. In one example, KPIs are measurable values that indicate the extent to which a predetermined objective is being effectively achieved. In one example, the KPIs are generated by and stored in the detection module 716. In this example, the KPIs include data indicating the time an offer was sent to a customer, the customer to whom the offer was sent, whether the customer responded to the offer, etc. In another example, the KPIs represent other metrics, such as the number of times a user deposits more than $10 into an ATM, or other predetermined metrics or events. In this example, the CDA system 702 receives input data records. The CDA system 702 is configured to specify which metrics are KPIs. Thus, the KPIs (and / or their definitions) are incorporated as part of the CDA system 702's initial configuration. In yet another example, the CDA system 702 updates the profile data and / or KPIs (for a particular key) with other data indicating each event represented by a particular record received for that particular key. This profile data therefore tracks and includes all received events (for a particular user) and / or data representative of a given event or type of event.
[0103] In another example, the discovery module 716 (or parameterized discovery application) creates one or more KPIs for a particular value of a key. In this example, the KPI specifies one or more values of a data item associated with the particular value of the key. The CDA system 702 receives data for the particular value of the key. The received data indicates one or more actions initiated by the action module 730 and / or feedback regarding the received data, including input data records. In this example, the discovery module 716 updates the KPI for the particular value of the key with the feedback data and stores the KPI in the in-memory data store 718.
[0104] The CDA system 702 may forward the KPIs to other systems, for example, to enable those systems to track and manage offer effectiveness and / or customer journey. As previously mentioned, the profile data along with the KPIs may be maintained in memory, for example, with the profile data stored in memory and the KPIs stored as part of the profile data.
[0105] The detection module 716 is also configured to add external data to the in-memory profile (which is stored in the in-memory data store 718). This external data is retrieved by the CDA system 702 from one or more external data sources 706. In this example, the detection module 716 adds the external data to the in-memory profile, for example, to enrich the detection of data records and to further reduce delays in building records containing profile data. In this example, the detection module 716 performs in-memory key-based processing (e.g., of the collected data records). In this example, this processing is key-based processing because the detection module 716 processes data records associated with a particular key (or identifier) and processes these keyed data records (e.g., data records associated with a particular key) according to the in-memory state (for an application or set of applications) of that key, as described in U.S. Patent Application No. 62 / 270,257. In this example, the collection module 712 sends records 723 to a queue 721 for retrieval by the detection module 716. In this example, the collection module 712 sends data and / or data records to the detection module 716 both in batches and in real time (e.g., as data records are received by the collection module 712). The collection module 712 transfers data in batches to the detection module 716, for example, by sending (as part of records 723) batch data received from one or more data repositories.
[0106] In this example, the detection module 716 executes various rules (e.g., rules specified by various applications and / or flow charts). Based on the execution of these rules, the detection module 716 determines whether a state (e.g., for a particular key) needs to be updated. For example, based on the execution of a rule, the state of an application (for a particular key) transitions from one state to another. In this example, the detection module 716 updates the state for that key accordingly. In one example, the detection module 716 stores the state as shared variables, for example, via a persistent in-memory keyed data store that is isolated from a particular running dataflow graph or application. This reduces the latency required to determine the state because, for example, each application and / or module can retrieve the in-memory value of its shared variables.
[0107] The detection module 716 also includes applications 720, 722, 724, and 726, each of which performs various types of data processing. In this example, the collection module 712 creates a data record 723 from one or more collected data records, for example, by including the collected data records as subrecords in the data record 723 and augmenting the data record 723 with profile data and / or other stored data. The detection module 716 sends the data record 723 to each of the applications 720, 722, 724, and 726, each of which is configured to perform data processing and apply rules to perform data record detection. In conventional methods of data record detection, separate data records are created and transferred to each application (e.g., in a format appropriate for each application), resulting in increased system resources, increased memory storage, and increased system latency compared to the system resource consumption, memory consumption, and resulting latency of transferring a single large record of data records (e.g., record 723) to each of applications 720, 722, 724, and 726, as described, for example, in U.S. patent application Ser. No. 62 / 270,257. However, when each application is already integrated into a particular system or module, the applications share a common format, and thus a single record can be transferred to each of the applications.
[0108] Based on the execution of one or more of the applications 720, 722, 724, 726, the detection module 716 detects one or more predetermined data records. For each detected data record, the detection module 716 adds an action trigger (e.g., an instruction to perform or cause the execution of one or more actions) to a queue 728, which forwards the action trigger to an action module 730. Based on the action trigger content, the action module 730 causes the execution of one or more actions (e.g., sending an email, a text message, an SMS message, etc.). In some examples, based on the content of the action trigger, the action module 730 creates a message or content and customizes the message / content for the user to whom the message / content is addressed. The action module 730 forwards the customized message to one or more external systems 710 (as indicated by arrow 766). The external system sends the message / content or performs further actions based on the customized message.
[0109] In this example, data warehouse 732 includes a query engine 734, a data warehouse 738, and an analysis engine 736. Query engine 734 queries data (e.g., data indicating detected data records, data indicating processed data records, etc.) from discovery module 716, data warehouse 738, or other data sources and forwards the queried data to analysis engine 736 for performing data analysis (as indicated by arrow 760). In this example, analysis engine 736 stores data analysis results in data warehouse 738 (as indicated by arrow 762). In addition to performing data record discovery, discovery module 716 is also configured to process warehouse data (e.g., data stored in data warehouse 738 or another data store) to perform various analyses. In this example, analysis engine 736 uses various personalization rules 713 that specify which rules to apply to data records associated with which keys. In this example, the personalization rules also include segmentation rules that specify how to personalize or target instructions for records associated with various predetermined segments. The analysis engine 736, for example, sends the personalization rules 713 to the application 714 for execution of the personalization rules (as indicated by arrow 764).
[0110] In this example, the operational module 730 sends data and / or messages to the network data source 708 (as indicated by arrow 740), which feeds data back to the CDA system 702 (as indicated by arrow 752). In one example, the operational module 730 sends customized messages to the network data source 708, which re-forwards these customized messages to the client system 704 (as indicated by arrow 744). The client system sends a confirmation message (not shown) to the network data source 708 (as indicated by arrow 746), which forwards the confirmation message to the CDA system 702 as feedback (as indicated by arrow 752). Thus, the network environment 700 implements a feedback loop (via one or more paths indicated by arrows 740, 744, 746, 752) that allows the CDA system 702 to audit or track receipt of messages.
[0111] In this example, the operational module 730 also sends data (e.g., customized messages or other customized data) to the external data sources 706 (as indicated by arrow 742) to implement a data feedback loop. In response, the one or more external data sources 706 send this data (received from the operational module 730) to the client system 704 (as indicated by arrow 748), which in turn can send a confirmation back to the one or more external data sources 706 (as indicated by arrow 750), which in turn sends the confirmation back to the CDA system 702 as feedback (as indicated by arrow 754). Thus, the network environment 700 implements another feedback loop (via one or more paths indicated by arrows 742, 748, 750, 754). In this example, the analytics engine 736 or another module of the CDA system 702 tracks the end-to-end customer journey by tracking, for example, whether the offer was delivered as planned (e.g., offer execution), the user's response to the offer, offer fulfillment, etc. In this example, the analytics engine 736 tracks the customer journey (e.g., offer delivery and offer fulfillment) by one of the feedback loops described above and / or by obtaining data from multiple systems, for example, from external data source 706, external system 710, or other external system that tracks offer fulfillment and delivery.
[0112] In some examples, the CDA system 702 can evaluate the received data as part of one of the feedback loops described above. For example, this received data can indicate the effectiveness of an offer or an action output by an action module. Based on this received data, the CDA system 702 evaluates whether a parameter value (e.g., a user input) led to a desired system response, e.g., a desired action being performed. Using this feedback, the CDA system 702 or a user can change or adjust one or more parameter values of a collection, detection, or action application. For example, a detection application can be configured with parameter values that specify the detection of an event (or data record) indicating that a user's data plan is below a threshold of remaining data. In response to the detection of this event, the action application can be configured with one or more parameterized values that specify a particular action or output to be produced. In this example, this output can be a message informing the user of a special promotional opportunity to replenish or purchase more data. Using one of the feedback loops described above, the CDA system 702 tracks the effectiveness of the output, for example, by tracking the fulfillment of an offer. In one example, the output from the action module includes a key (or other identifier) that uniquely identifies the output or the user to whom the output is addressed or destined. Actions on the output (e.g., clicking on a link or other selectable portion in the output) are tracked (e.g., by an external system), for example, by a cookie containing the key or identifier (or another identifier associated with the output's key), by associating the action with a digital signature that includes the output's key or identifier, etc.
[0113] In this example, based on data received from the feedback loop, the CDA system 702 can identify that a particular output is not particularly effective. For example, the output may not result in at least a threshold at which the user would purchase additional data. Based on the feedback data, the CDA system 702 or the user can adjust one or more parameter values of a detection application (e.g., specifying detection of various or diverse events) and / or an action application (e.g., specifying various or diverse outputs in response to detected events). In this example, the CDA system 702 determines that an action taken was incorrect (e.g., did not achieve a desired result) or was not performed correctly, and then executes a correction loop that adjusts the parameter values by changing the values to obtain the correct action to be performed. In some examples, the correction loop is heuristic-based and includes a set of rules that specify changes to one or more applications to be changed and / or to obtain a specified result.
[0114] In another example, the detection module 716 sends feedback directly to the data warehouse 732, as indicated by arrow 701. For example, the detection module 716 can send data indicative of the success of a campaign to the data warehouse 732. The detection module 716 can also send data (e.g., feedback data) to the data warehouse 732 indicative of the paths taken (in the detection application), the decisions made by the detection application, and the branching logic followed by the detection application. In this example, the CDA system 702 can run one or more machine learning programs, heuristic networks, or neural networks on the feedback data received by the data warehouse 732 from the detection module 714. Based on this execution, the detection module 716 can update the values of one or more parameters in the detection application or the types of parameters included in the detection application. For example, if a particular area of the branching logic is underutilized (as indicated by the feedback data), the neural network can use the feedback data indicative of which logic is accessed and, based on the feedback data, adjust the parameters of the underutilized branching logic to expedite traversal of that logic.
[0115] The above-described feedback and correction loops may also be implemented in the following manner to ensure proper operation of the underlying system, e.g., to ensure that one or more actions are performed correctly, i.e., as desired. The one or more actions may involve data processing tasks or network communications, such as one or more of: sending a text message to an external device, sending an email to an external system, opening a work order ticket in a case management system, immediately disconnecting a mobile phone connection, providing a web service to a target device, and transmitting a data packet of one or more converted data items with a notification, and running a data processing application hosted on one or more external computers on the one or more converted data items. In some examples, the one or more actions include offering a user, e.g., 10 additional voice minutes or additional data. In this example, the system described herein is configured to communicate with a provisioning system (e.g., a system that maintains or manages minutes and / or data usage, such as a telephone network) to indicate to the provisioning system that the user has additional voice minutes and / or data. In another example, the one or more actions include a system that provides a benefit to a user, such as providing money. In this example, the system is configured to provide money through a financial transaction with the user's account maintained at a financial institution. In another example, the one or more actions include downgrading the user's service when the user is roaming to prevent the user from incurring roaming charges.
[0116] The underlying system and / or user may receive a feedback message in a feedback loop indicating whether the one or more operations (i) completed successfully or (ii) failed. The one or more operations are considered to have failed if any portion of the one or more operations did not complete. The feedback message may also optionally indicate which portion of the one or more failed operations did not complete. For example, the feedback message may indicate whether a data processing application hosted on the one or more external computers performed correctly or not on the one or more converted data items, e.g., certain data processing tasks of the application were not performed on certain ones of the data items. The feedback message may indicate which data processing tasks (e.g., which portions of program code) were not performed. For example, an operation to immediately disconnect a mobile telephone connection was not performed correctly because the mobile telephone connection was only disconnected after a certain delay that exceeds a predetermined delay considered acceptable. The feedback message may indicate this certain delay. The feedback message may also indicate result data (eg, a certain delay or result experienced by the data processing application) of the one or more operations that have been successfully completed and / or failed.
[0117] The result data may be compared with predetermined data associated with successful completion of execution of the one or more operations (e.g., a predetermined delay considered acceptable or a desired result to be produced by the data processing application). It may then be determined based on the comparison whether execution of the one or more operations was completed successfully or whether execution of the one or more operations failed. Execution of the one or more operations may be considered successful if the result data deviates from the predetermined data by less than a predetermined amount (e.g., in the form of an absolute value or a percentage), and execution of the one or more operations is determined to have failed if the result data deviates from the desired data by at least the predetermined amount.
[0118] That is, the correct execution of one or more operations can be represented by predetermined data and / or predetermined quantities. Considering the above possible operations to be performed, such deviations of the result data from the predetermined data can occur with respect to the receipt of confirmation data, the sending time or number of characters of a text / email message, the ticket number of a ticket, the time taken to disconnect a mobile phone connection, the network bandwidth available for web services, the amount of data sent for that data packet, a data processing task performed in a data processing application on an external device, or similar parameters or properties characterizing that operation.
[0119] Subsequently, the one or more specified values of one or more parameters of the first, second, and / or third parameterized applications may be modified (by the system 702 or by the user) based on the result data, and the first, second, and / or third parameterized applications may be re-executed by the system under the modified one or more specified values. To give the user a further opportunity to evaluate and / or initiate correct operation of the underlying system, one or more of the following interactions may be provided during operation of one or more user interfaces: displaying one or more user interface elements indicative of the predetermined data and the predetermined quantities; outputting via the displayed one or more user interface elements whether the one or more operations (i) completed successfully or (ii) failed; and outputting the result data via the displayed one or more user interface elements. The user interface may receive the predetermined data and the predetermined quantities from the user (e.g., via icons as graphic elements).
[0120] Alternatively, one or both of the predetermined data and the predetermined amount are stored in a memory accessible and retrievable by the system. The user interface can graphically output (e.g., via icons as user interface elements) whether the one or more operations (i) completed successfully or (ii) failed. The user interface can graphically output result data (e.g., via icons as user interface elements) for user inspection or automatic use by the system. Via one or more displayed (graphical) user interface elements, one or more user-specified modified specified values of one or more parameters of the first, second, and / or third parameterized applications can be received (based on the result data), and the first, second, and / or third parameterized applications can be automatically re-executed by a system (e.g., system 702) under one or more of the modified specified values. The sending of the one or more instructions causing the (re-)execution of the one or more operations is performed automatically by using the output of the third parameterized application specifying the execution of the one or more operations as input. The one or more instructions may be sent over a network connection to cause (re)execution of the one or more actions on an external device.
[0121] The user interface can provide a graphical element that, when activated by a user, automatically initiates the re-execution with the modified values. For example, a user can be provided with the results data on a user interface and recognize or be provided with information that the results data is not as desired for correct execution of the one or more operations. The user or a system (e.g., system 702) can then initiate the re-execution of the operations with the modified values in a correction loop. In this case, the modified values are calculated by the system (or user input) based on the results data to ensure correct / desired execution of the operations. This can help, for example, to ensure proper operation of an underlying system by ensuring that the one or more operations are executed correctly (i.e., as desired).
[0122] As described herein, the CDA system 702 provides end-to-end operational robustness (e.g., the CDA system 702 can be utilized to process data for operational systems as well as archival / analysis systems) and almost unlimited scalability in terms of required hardware, due to, for example, application reuse, in-memory state and profiles, and large records of data records that are generated and then multi-published to various applications, including, for example, applications 720, 722, 724, and 726.
[0123] Referring to Figure 5B, network environment 751 is a transformation of network environment 700 described in Figure 5A. In this example, network environment 751 includes client systems 792, networked systems 768, external input systems 769 (streaming or inputting data), external output systems 770 (receiving output data), and CDA system 753. The CDA system includes collection applications 780, applications 755 (which include detection applications 775 and action applications 776), profile data structures 771, and data analysis applications 782. In this application, data analysis applications 782 include a data repository 765 (e.g., a data lake) that stores data (e.g., in native formats), a data warehouse 784, analytics applications 786 (e.g., an analytics engine that implements machine learning and data correlation), and visualization data 763 (e.g., data for generating data analysis visualizations, profile data, etc.). The data analysis application 782 also includes a query engine 767 that queries the data repository 765, the data warehouse 784, the profile data structure 771, etc., e.g., to query data for processing or analysis by the analysis application 786 and / or for inclusion in the visualization data 763. In one example, the profile data structure 771 is stored in volatile memory (e.g., to reduce memory storage requirements and to reduce latency in retrieving the profile data compared to, e.g., the latency in retrieving the profile data if the profile data were stored on disk) and / or in non-volatile memory (e.g., in the data repository 765 and / or the data warehouse 784). The CDA system 753 also includes a data layer 788 (e.g., a service layer) that sends portions of the near real-time profile data from the profile data structure 771 (e.g., in batch or real-time) to the external output system 770 and / or the external input system 769.
[0124] In operation, one or more client systems 792 send data (e.g., data packets) including (or indicating) an event to the networked system 768 and the external input system 769. In one example, one of the client systems 792 is a smartphone. One of the external input systems 769 is a telephone company's telecommunications system. In this example, the smartphone sends a text message. When sending the text message, the smartphone also sends data (indicating that the text message was sent) to the telecommunications system. In this example, the data indicating that the text message was sent is an event. In response, the telecommunications system sends the event to the collection application 780. Typically, the collection application 780 is configured to collect data records (e.g., in batch and / or real time) from data sources and data streams. In this example, the collection application 780 sends the received data records to the application 755. In this example, the collection application 780 sends the received data records in batch and real time by adding the collected data records to a queue 759.
[0125] In this example, the detection application 775 is configured to execute various rule sets 790a, 790b, 790c...790n. In this example, each rule set executes a particular campaign, loyalty program, fraud detection program, etc. The detection application 775 executes the various rule sets on data records received from the collection application 780. In this example, the detection application 775 includes a profile repository 773 that dynamically updates user profiles (e.g., in memory) with data contained in received data records. In this example, the profile data configuration 771 intermittently transfers profile data (e.g., associated with particular keys and / or all keys) to the profile repository 773. The profile repository 773 stores the profile data in memory to reduce the time it takes for the detection application 775 to retrieve the profile data when it is needed, for example, to execute the rules and / or determine whether one or more of the rules have been satisfied. When the detection application 775 receives a new data record, the detection module 775 updates the profile repository 773 with the profile data for the appropriate profile (e.g., by matching the key associated with the received data record with the key associated with the profile data stored in the profile repository 773). In one example, specific fields of the data record contain the profile data. In this example, the profile repository 773 is updated with the contents of those fields (e.g., for the appropriate keys). The profile repository 773 intermittently transfers the updated profile data to the profile data structure 771, updating the profile data structure 771.
[0126] Based on the execution of one or more rules, the detection application 775 identifies one or more actions to be performed. In this example, the detection application 775 adds instructions to a queue 761 that cause the execution of these one or more actions. The instructions in the queue 761 are sent to an action application 776 that includes a profile repository 772, which stores profile data in memory. In this example, the profile repository 772 intermittently retrieves profile data from a profile data structure 771 and stores the retrieved profile data in memory. The profile repository 772 does this to enable the retrieval of profile data by the action application 776 with reduced latency compared to the latency of searching a profile data disk. For each action specified in the instructions, the action application 776 performs or causes the execution of the action (e.g., by adding an execution instruction to a queue 774 and sending these execution instructions in the queue 774 to the external output system 770). In either example, the action application 776 uses the profile data in the profile repository 772 to customize the action to be performed (with profile data specific to the recipient of the action) or customize the execution instructions (e.g., by adding profile data specific to the recipient of the action). By storing the profile data in the profile repository 772 in memory rather than on disk, the action application 776 retrieves the profile data with reduced latency compared to the latency of retrieving the profile data from disk. Based on this reduced latency, the action application 776 can add profile data to the instructions and / or the action in near real time. In some examples, the action application 776 updates the profile repository 772 with new profile data (e.g., data indicating that an offer has been sent to a particular user).The profile repository 772 intermittently transfers profile data and / or updates to copies of that profile data stored in the profile repository 772 to the profile data structure 771. Based on the updates and / or profile data received from the profile repositories 773, 772, the profile data structure 771 maintains near real-time customer profiles. In this example, the profile data structure 771 also receives profile data from a data warehouse 784, which further enables the profile data structure 771 to maintain versions of the profile data associated with specific keys and / or all keys.
[0127] In this example, CDA system 753 is configured to transfer (e.g., in batch and / or real time) the profile data in profile data structure 771 to collection application 780, which in turn transfers the profile data to external input system 769 and / or networked system 768. CDA system 753 is also configured to transfer (e.g., in batch and / or real time) the profile data from profile data structure 771 to data layer 788 for transfer to external output system 770, networked system 768 shown in data flow 757 (e.g., to enable work scheduling, monitoring, audit trails, and performance tracking based on data included in the profile data, such as data indicating actions performed and when they were performed), and to external input system 769 shown in data flow 791 (e.g., to enable data profiling based on metadata and reference data included in the profile data, and data governance and data management, such as quality and data lineage). In this example, the profile data stored in profile data structure 771 includes data representing transformations and other operations performed on one or more portions of the data series data and the profile data (e.g., associated with a particular key).
[0128] In this example, collection application 780 and data analysis application 782 are configured to send data to each other in batches. In doing so, collection application 780 forwards collected data records and other data received from systems 768, 769 to data analysis application 782 for storage in data repository 765 and / or data warehouse 784. Similarly, data analysis application 782 sends analyzed data and / or other stored data to collection application 780 for subsequent forwarding to one or more of systems 768, 769, for example, to facilitate data integration between various systems.
[0129] Referring to Figure 6, a graphic user interface 800 shows a versatile wizard for configuring the detection module, for example by defining the initial settings of the program for performing the detection, for example by setting parameters such as the duration of the campaign or program, data records, data sources, and the definition of the target population. In this example, the graphic user interface 800 includes a portion 802 for specifying one or more characteristics of the program (e.g., a program for running a campaign), a portion 804 for specifying one or more characteristics and / or parameter values that define the event to be detected, a portion 806 for specifying one or more characteristics and / or parameter values for performing the detection, and a portion 808 for specifying state transitions (e.g., that are part of the detection), and one or more characteristics and / or parameter values for viewing the results.
[0130] In this example, portion 804 includes sub-portions 804a, 804b, and 804c that specify values for various parameters (e.g., contained in parameterized logic or parameterized applications) for performing the detection. For example, sub-portion 804a allows the user to specify which types of data records (e.g., data records or data records with specific characteristics) should be detected. Sub-portion 804b allows the user to specify keys that should not be detected, e.g., to filter our records associated with specific keys. Sub-portion 804c allows the user to specify which records should be used as test records, e.g., when testing the processing of data records by an application in real time. Portion 806 allows the user to configure which rules and / or logic should be executed by the detection module.
[0131] Referring now to Figure 7, the transformation of Figure 6 is shown. In this example, portion 804 includes a selectable portion 812 (e.g., a link), the selection of which causes the display of overlay 810. In this example, overlay 810 includes one or more selectable portions (e.g., checkboxes) that specify which data records to process (e.g., by applying rules) against those data records. In this example, data entered into overlay 810 functionally specifies parameter values in the same manner as described in connection with Figure 2C, e.g., as interface 250 functioned when setting parameter values for the underlying graph and / or parameterized application (or parameterized logic) and / or application.
[0132] Referring to FIG. 8, a graphical user interface 900 provides configuration of the collection module and illustrates the correspondence between portions 902, 904, 906, and 908 of the configuration interface 901 and portions 912, 914, 916, and 918 of a dataflow graph 910. The values of parameters in the dataflow graph 910 are set through the configuration interface 901, e.g., by performing the functions described above in connection with FIGS. 2A-2C. In this example, the graphical user interface 900 illustrates a generic application that may be configured by a non-developer. In this example, the dataflow graph 910 includes an application that collects, transforms, and enriches data records received from multiple sources. In this example, the dataflow graph 910 includes an application with various parameters, the values of which are set via the configuration interface 901. After values for the parameters in the dataflow graph 910 are specified, the CDA system creates an instance of the dataflow graph, e.g., an instance of the dataflow graph 910, in which the values of the parameters in the dataflow graph are set to the values specified by the configuration interface.
[0133] In this example, portion 902 of configuration interface 901 specifies source files to be collected by the configuration module. Input to portion 902 specifies values for parameters included in portion 912 of dataflow graph 910. Based on the selection of one or more selectable portions in portion 902, the user specifies the source of data and collection of data records. Portion 904 specifies transformations, e.g., reformatting. As described herein, CDA systems handle arbitrarily large data volumes, low latency, and the complexity of multiple data formats. In this example, portion 904 enables handling of multiple data formats. Input to portion 904 specifies values for parameters included in portion 914 of dataflow graph 910. Input to portion 906 specifies data transformations, e.g., augmenting and adding profile data to long records. In this example, portion 906 corresponds to portion 916 of dataflow graph 910 and specifies values for one or more parameters included in portion 916 of dataflow graph 910. In this example, portion 908 includes one or more selectable portions that specify values for one or more parameters included in portion 918 of dataflow graph 910. Portion 908 specifies values for parameters related to the output of data, e.g., data to be output and instructions for transfer to various devices.
[0134] Referring to FIG. 9, a graphical user interface 1000 displays controls for configuring the detection module. The graphical user interface 1000 includes a portion 1002 that displays inputs available in the graphical user interface 1000, e.g., inputs to a defined logic. In this example, the portion 1002 defines a palette as described in U.S. Patent Application No. 62 / 270,257. The palette includes all data records contained in a long record (e.g., a record generated from a collection of records received by a CDA system). In this example, the palette includes inputs 1026, 1054 (e.g., data records) and augmented inputs 1052 (e.g., augmented data records). Typically, augmented inputs include, for example, inputs based on data retrieved or collected from a data store or memory storage rather than data received externally from the system. In this example, the inputs include, for example, pre-computed aggregates as described in U.S. Patent Application No. 62 / 270,257. For input 1026, there are various types of inputs, including inputs 1028, 1030, 1032, 1034, 1036, 1038, 1040, 1042, 1044, 1046, 1048, and 1050. Each of these types of inputs is contained in a long record. Each of the inputs displayed in portion 1002 is selectable (e.g., by dragging and dropping) as an input to a cell in portion 1004 for use in defining logic, for example.
[0135] The graphical user interface 1000 also includes a portion 1004 for generating logic (e.g., rules) to be executed by the detection module. In one example, the portion 1004 includes a business rules editor. In this example, the portion 1004 displays visualizations of one or more parameterized applications, e.g., detection applications and action applications. The portion 1004 also displays editable cells for specifying values for parameters in the parameterized applications. In one example, the business rules editor and / or the portion 1004 are used in defining the specification, for example, after inputting values for parameters of the parameterized applications. That is, the inputs to the portion 1004 and the rules specified by the portion 1004 together form the specification.
[0136] In this example, portion 1004 includes logic portion 1006, which specifies one or more triggers or conditions that, when met, cause the execution of one or more actions (e.g., by causing the execution of an output). Logic portion 1008 specifies various outputs to be initiated and / or performed when the conditions or triggers are met. In this example, logic portion 1006 includes state portion 1006a, which specifies one or more state triggers, e.g., a condition that specifies that if the state of a particular key corresponds to a particular value or type of state, the detection module should perform the corresponding action of that condition. Logic portion 1008 also includes state portion 1008a, which specifies a new state to which the logic transitions when the corresponding trigger is met, e.g.,
[0137] In this example, logical portion 1006 includes logical subportions 1010, 1012, 1014, and 1016, each of which specifies one or more conditions. For example, logical subportion 1010 specifies that the condition for the user associated with the keyed data to be processed is that the user is a new user. Logical subportion 1012 specifies that the condition for the state of the keyed data to be processed is that the state has a value of "CampaignState.Eligible" for "is_null." Logical subportion 1014 specifies that the condition for the keyed data to be processed is any data record type. Logical subportion 1016 specifies that the condition for SMS usage is greater than 500. In this example, the combination of logical subportions 1010, 1012, 1014, and 1016 together form a trigger, the satisfaction of which causes the execution and / or realization of a corresponding output by the detection module. In this example, logical subportions 1010, 1012, 1014, and 1016 correspond to logical subportions 1018, 1020, 1022, 1024, and 1025, which together specify a particular output. The detection module creates instructions indicating the particular output and sends these instructions to an action module, for example, to cause the particular output to be realized and / or to customize the output and send data indicating the customized output (or the customized output itself) to an external system for execution. In this example, logical subportion 1018 specifies that the state of the execution logic is updated to a value of "NewCampaignState.SendOffer." Logical subportion 1020 specifies the action that an SMS message is sent. Logical subportion 1022 specifies the content of the message. Logical subportion 1024 specifies an expiration date or value for the duration the message is active. Logical subportion 1025 specifies a fulfillment plan code. In this example, the combination of logic subportions 1018, 1020, 1022, 1024, and 1025 together form an output that is realized after satisfying the conditions specified in logic subportions 1010, 1012, 1014, and 1016, for example.
[0138] In one example, the palette displayed in portion 1002 is used in generating the logic contained in logical subportions 1012, 1016, 1018, 1020, 1022, and 1024. For example, the logic of logical subportion 1016 is generated by dragging and dropping data record 1040 into logical subportion 1016 (which, in this example, contains editable cells). The user then further edits logical subportion 1016 by entering the text “>500” into the editable cell (which is logical subportion 1016). In this example, data record 1040 corresponds to a parameter in the associated data flow graph or application. The user editing logical subportion 1016 then enters the value of that parameter, i.e., the value “>500.” Using the logic displayed in portion 1004, the CDA system creates an application or data flow diagram that implements that logic, for example, by setting the values of parameters in the parameterized logic to the values displayed in portion 1004. In some cases, this executable logic may be specified as a flowchart rather than being specified in a table.
[0139] Referring to FIG. 10 , specification 1100 includes a flowchart 1102 that includes nodes 1102a-1102g. Typically, the chart includes an application for processing data records, e.g., a parameterized application. In this example, a specification of values (e.g., inputs) for parameters of the parameterized application creates the specification. As described above, the specification represents executable logic based on states reached from executing the executable logic on preceding data items and specifies various states of the executable logic. Typically, executable logic includes source code and other computer instructions. Each node in the chart represents one or more portions of the executable logic. For example, a node includes one or more logical expressions (hereinafter referred to as “logic”) that generate the executable logic. In another example, a node corresponds to one or more particular portions of the executable logic when the executable logic is in a particular state. In this example, the executable logic in the chart is generated using, for example, a palette (described above) to select various inputs for inclusion in one or more nodes of the chart.
[0140] The application (represented by a chart) includes graphical units of logic for responding to input data records and creating output data records, e.g., data records generated based on logic contained in the specification. Generally, the graphical units of logic include logic that is generated at least in part graphically, e.g., by dragging and dropping various nodes from an application (not shown) into a window for building a chart. In one example, a node includes logic (not shown) that specifies how to process input data records, how to set values for variables used by executable logic, what output data records to generate, etc., when conditions specified by the logic are met. In one example, a node is programmable by a user entering values for parameters and / or variables used within the node's logic.
[0141] The chart itself is executable because the logic within the nodes is compiled into executable logic, with each node corresponding to one or more portions of that executable logic. For example, the system transforms the specification (and / or the chart within the specification) by compiling the logic within the nodes into executable logic. Because the chart itself is executable, the chart itself can process data records and can be stopped, started, and paused. The system also maintains the state of the flowchart 1102, for example, by tracking which of nodes 1102a-1102g is currently executing. The state of the flowchart 1102 corresponds to the state of the executable file represented by the flowchart 1102. For example, each node in the flowchart 1102 represents a particular state of the executable logic (one or more portions of the executable logic are executable in that state). When the flowchart 1102 is executed for different values of the key, the system maintains the state of the flowchart 1102 for each value of the key, for example, by maintaining per-instance state, as described in more detail below. In this example, flowchart 1102 includes a state transition diagram in which each input data record drives transitions between nodes, and data records are evaluated based on states reached by processing past data records. Links between nodes in flowchart 1102 represent the flow of logic over time.
[0142] Node 1102a represents the start of executable logic. After completion of node 1102a, the state of flowchart 1102 transitions to node 1102b, which represents one or more other portions of executable logic. Node 1102b includes a wait node (hereinafter, wait node 1102b). Wait node 1102b represents a wait state in which a portion of executable logic (corresponding to wait node 1102b) waits for input data records that satisfy one or more conditions. In one example, the wait state may be part of another state of flowchart 1102, such as a state in which the system executes the wait node (to implement the wait state) and then executes one or more other nodes. After completion of the portion of executable logic represented by wait node 1102b, the system exits the wait state and executes node 1102c, which represents executable logic for making a decision. In this example, node 1102c includes a decision node. Generally, a decision node includes a node that includes logic for making a decision (e.g., logic that evaluates for a Boolean value).
[0143] Based on the outcome of the decision, the state of flowchart 1102 transitions to node 1102g (which transitions state back to node 1102a) or to another wait node, node 1102d. After completion of the portion of executable logic represented by wait node 1102d, the state of flowchart 1102 transitions to node 1102e, which comprises a send node. Generally, a send node comprises a node representing executable logic for causing transmission of data to another system. After completing execution of the portion of executable logic represented by node 1102e, the state of flowchart 1102 transitions to node 1102f, which comprises an end node. Generally, an end node represents completion of execution of executable logic.
[0144] In one example, a wait node represents a transition between states that originates at the wait node, e.g., a transition from one state to another. In this example, flowchart 1102 differs from a state transition diagram because not all nodes in flowchart 1102 represent wait nodes that represent state transitions. Rather, some nodes represent actions that are performed, for example, when flowchart 1102 is already in a particular state. In some examples, a system processes flowchart 1102 to generate a state machine diagram or state machine instructions.
[0145] In this example, flowchart 1102 includes two states: a first state represented by nodes 1102b, 1102c, and 1102g, and a second state represented by nodes 1102d, 1102e, and 1102f. In this first state, the system waits for a particular data record (as represented by node 1102b) and then executes node 1102c, which causes a transition (of specification 1100 and / or chart 1102) to a second state (the beginning of which is represented by node 1102d) or the execution of node 1102g. Once in the second state, the system again waits for a particular data record (as represented by node 1102d) and then executes nodes 1102e and 1102f. By including nodes other than wait nodes, flowchart 1102 includes a logical graph of the temporal processing of data records. In this example, chart 1102 includes link 2i, which represents the transition of chart 1102 from a first state to a second state, and also represents the flow of data from node 1102c to node 1102d.
[0146] Chart 1102 also includes link 1102j between node 1102a and node 1102b and link 1102k between node 1102b and node 1102c to represent a user-specified execution order for the portions of executable logic in the first state corresponding to nodes 1102a, 1102b, and 1102c. In this example, the portions of executable logic in the first state (hereinafter the "executable logic of the first state") include statements (e.g., logical statements, instructions, etc. (collectively referred to herein without limitation as "statements").) Generally, the execution order includes the order in which executable logic and / or statements are executed. Each of nodes 1102a, 1102b, and 1102c corresponds to one or more of those statements (e.g., to one or more portions of the executable logic of the first state). Thus, link 1102j represents the execution order of the executable logic of the first state by indicating that statements in the executable logic of the first state represented by node 1102a are executed by the system before other statements in the executable logic of the first state represented by node 1102b are executed. Link 1102k also represents the execution order of the executable logic of the first state by indicating that statements in the executable logic of the first state represented by node 1102b are executed by the system before other statements in the executable logic of the first state represented by node 1102c are executed.
[0147] The specification 1100 also includes a key 1102h, which identifies that the flowchart 1102 processes a data record that contains or is associated with the key 1102h. In this example, a custom identifier (ID) is used as the key. The key 1102h may correspond to one of the fields of the data record (i.e., a data record field), such as a subscriber_ID field, a customer_ID field, a session_ID field, etc. In this example, the customer_ID field is the key field. For a particular data record, the system determines the value of the key for that data record by identifying the value of the key field for that data record.
[0148] In this example, flowchart 1102 accepts data records of a specified type (e.g., specified during configuration of flowchart 1102). In this example, flowchart 1102 accepts a data record that includes key 1102h. In this example, flowchart 1102 and the data record share a key. Generally, a flowchart accepts these data record types by including logic for processing data records that include the flowchart's key. When processing of a data record begins, the system starts a new flowchart instance for each new value of the key for that flowchart, e.g., by maintaining the state of executable logic (represented in the flowchart) for each new value of the key. The system processes data records by configuring the flowchart instance (and thus the underlying executable logic) to respond to data records for a particular key value. In one example, a flowchart accepts a customer's short message service (SMS) data record. A flowchart instance for a particular customer ID manages data records for that customer. There can be as many flowchart instances as there are customer IDs encountered in input data records. In some examples, the systems described herein provide a user interface for configuration of flowchart 1102, e.g., to allow a user to easily enter values into the various components of the flowchart.
[0149] 11, diagram 1107 includes flowchart instances 1103, 1104, 1105 and data records 1106a, 1106b, 1106c generated by the system from, for example, flowchart 1102 (FIG. 10). That is, a new copy or instance of flowchart 1102 is created for each new key found in data records 1106a, 1106b, and 1106c.
[0150] Each of the flowchart instances 1103, 1104, and 1105 is associated with a "customer_id" key. Flowchart instance 1103 processes data records that contain a value of "VBN3419" in its "customer_id" field, which is a key field in this example. Flowchart instance 1104 processes data records that contain a value of "CND8954" in its "customer_id" field. Flowchart instance 1105 processes data records that contain a value of "MGY6203" in its "customer_id" field. In this example, the system does not re-execute executable logic for each flowchart instance. Rather, the system implements the flowchart instance by executing executable logic and then maintaining state for each value of the key. Thus, one example of "a flowchart instance processes data records" is the system executing executable logic (represented by a flowchart), maintaining state for each value of the key, and processing data records associated with a particular value of the key (based on the state of the state machine for that particular value of the key).
[0151] In this example, flowchart instance 1103 includes nodes 1103a-1103g that correspond to nodes 1102a-1102g, respectively, in Figure 10. Flowchart instance 1104 includes nodes 1104a-1104g that correspond to nodes 1102a-1102g, respectively, in Figure 10. Flowchart instance 1105 includes nodes 1105a-1105g that correspond to nodes 1102a-1102g, respectively, in Figure 10.
[0152] A flowchart instance is itself executable. After a system receives an input data record associated with a particular value of a key, the flowchart instance for that particular value of key processes the input data record, for example, by the system executing one or more portions of executable logic corresponding to the flowchart instance (or one or more nodes of the flowchart instance). The flowchart instance continues to process the input data record until the input data record reaches an end node or a wait node. In this example, the input data record continues to be processed by the system, which continues to process the input data record until the flowchart instance reaches a portion of executable logic corresponding to, for example, an end node or a wait node. When the input data record reaches a wait node, the flowchart instance pauses until a certain amount of time has passed or until a suitable new input data record arrives. Generally, suitable data records include data records that satisfy one or more specified conditions or criteria (e.g., contained within the logic of the node). When the input data record reaches an end node, execution of the flowchart instance is completed.
[0153] A flowchart instance has its own life cycle. When a data record arrives, it changes the current state or status of the flowchart instance; the data record triggers a decision, or returns to the beginning of the flowchart instance, or sends a message to the customer. When a data record reaches an end node, the flowchart instance for the customer ends.
[0154] In this example, the system starts flowchart instance 1103 for a value of "VBN3419" in a key field (e.g., customer_id=VBN3419). Flowchart instance 1103 processes a subset of data records 1106a, 1106b, and 1106c that contain a customer_id of VBN3419. In this example, flowchart instance 1103 processes data record 1106a, which has a value of "VBN3419" in the customer_id key field. Nodes 1103a, 1103b, and 1103c of flowchart instance 1103 process data record 1106a. The current state of flowchart instance 1103 is waiting for a data record, as represented by the dashed line at node 1103d. Upon reaching node 1103d, flowchart instance 1103 waits for another data record with customer ID=VBN3419 to process through nodes 1103d, 1103e, and 1103f of flowchart instance 1103.
[0155] The system starts flowchart instance 1104 for the "CND8954" value of the key (e.g., customer_id=CND8954). Flowchart instance 1104 processes the subset of data records 1106a, 1106b, 1106c that contain a customer_id of CND8954. In this example, flowchart instance 1104 includes wait nodes 1104b and 1104d. Each data record can satisfy only one wait node per flowchart instance. Thus, flowchart instance 1104 processes data record 1106b with customer_id=CND8954 from node 1104b to node 1104d, then waits for a second data record with the same key before proceeding to node 1104f. The system starts flowchart instance 1105 for the "MGY6203" value of the key (e.g., customer_id=MGY6203). Flowchart instance 1105 processes the subset of data records 1106a, 1106b, 1106c that contain a customer_id of MGY6203. In this example, flowchart instance 1105 processes data record 1106c with customer_id=MGY6203 through nodes 1105b-1105d, then waits for a second data record with the same key before proceeding to node 1105e where a message is sent. In this example, the system does not stop at non-wait nodes, and therefore does not stop at node 1105e, but rather sends a message and continues to node 1105f.
[0156] In the modified version of Figure 11, the system generates multiple flowchart instances for a single key value. For example, there may be several flowchart instances for the same customer with different start and end dates, or several flowchart instances for the same customer for different marketing campaigns.
[0157] In this example, the system maintains the state of the instances by storing state data in a data repository or in-memory data grid, e.g., data indicating which node for each instance is currently executing. Generally, state data includes data that indicates a state. In this example, an instance is associated with a key value. The data repository or in-memory data grid stores the key values. The system maintains the state of the instances by storing state data for each key value. Upon completing processing of a data record for a particular value of the key, the system updates the state data in the data repository to specify that the next node (in flowchart 1102) represents the current state for that key value. When another data record then arrives, the system looks up the current state for that key value in the data repository and executes the portion of the executable logic that corresponds to the node that represents the current state of the executable logic for that key value.
[0158] Referring to FIG. 12, logic flow 1112 is based on the execution of logic (e.g., rule application logic generated using the palette described above) that uses long records created using the techniques described above in its execution. In this example, logic flow 1112 specifies various data record triggers and actions based on the data records contained in one or more particular subscriber data records. Logic flow 1112 includes various decision points (e.g., "Has the subscriber consumed 50 SMS messages?"). For a particular key, the detection module determines which branch of logic flow 1112 to follow based on the data records (or lack thereof) contained in that subscriber's data record and the state of the executable logic for that key. Typically, a state refers to a particular component (e.g., a particular data record trigger or a particular action) to which the logic transitioned during its execution. For example, a state indicates which data record trigger or action in the logic is currently occurring for a particular key. In some examples, the detection module waits a specified time before selecting a branch in logic flow 1112. By waiting for this specified period of time, the detection module analyzes the new data records that have been inserted into the data records.
[0159] In this example, logic flow 1112 includes a data record trigger 1119 that specifies that after activation of a particular subscriber's service, the action module performs a start operation 1120 that monitors the total number of SMS messages consumed by that particular subscriber over the following two days. In this example, data record 1119 is a condition precedent to the rule executed by logic flow 1112. After satisfying data record 1119, the action module performs start operation 1120. A detection module determines when a particular subscriber has satisfied data record trigger 1119 by detecting an activation data record in the long record, and determines the subscriber associated with the activation data record (by subscriber ID).
[0160] In this example, if the subscriber has consumed at least 50 SMS messages in the last two days (e.g., as indicated by the SMS usage data record aggregate in the data record), data record trigger 1113 is executed. Data record trigger 1113 executes suggested replenishment action 1114, which prompts the action module to allow this particular subscriber to replenish. When the subscriber replenishes, an entry in the data record for that particular subscriber is updated with a data record representing the replenishment. This data record update causes logic flow 1112 to execute data record trigger 1115, which in turn specifies the execution of action 1116, which sends a packet offer SMS to the subscriber after a successful replenishment. Typically, the packet offer is an offer to purchase a package or bundled service.
[0161] When the user sends a response to the package offer SMS, the data record is updated with a data record representing the response and indicating that the response is received in less than three hours. The detection module detects the data record update and causes execution of data record trigger 1117. Data record trigger 1117 specifies execution of operation 1118 to end the campaign (for that particular subscriber) when the subscriber completes the package purchase if the response is received in less than ten hours. If the particular subscriber's input in the data record indicates that the particular subscriber did not send a response to operation 1116, logic flow 1112 also specifies operation 1125 to end the campaign for that particular subscriber.
[0162] In one example, an entry for a particular subscriber in a data record indicates that the subscriber did not refill, for example, due to the absence of a refill data record or due to a derived data record indicating the absence of a refill. In this example, logic flow 1112 specifies data record trigger 1123 that, for example, waits three hours to monitor whether the user refills within the next three hours. After three hours, data record trigger 1123 causes execution of reminder action 1124, which sends a refill reminder SMS to the subscriber. If the subscriber does not respond to the reminder SMS, logic flow 1112 specifies action 1126, which ends the campaign for that particular subscriber.
[0163] In response to operation 1120, the input for that particular subscriber may indicate that the subscriber did not consume at least 50 SMS within the last two days. The input may indicate this by a derived data record indicating a lack of consumption of 50 SMS, or by an aggregate SMS usage data record indicating that the consumption was less than 50 SMS. In this example, logic flow 1112 includes a data record trigger 1121 that waits five days and then executes operation 1122 to send a reminder SMS. After sending the reminder, if the subscriber still does not consume 50 SMS within the next five days (e.g., as indicated by the subscriber's data record in the data records), logic flow 1112 specifies a data record trigger 1127 that executes operation 1128 to send an alert to a customer recovery team (e.g., to notify the team that the consumer is not using the service) and to end the campaign for that particular subscriber.
[0164] Referring to Figure 13, a system (e.g., a CDA system described herein, e.g., system 100) performs process 1200 in processing data items in multiple distinct data streams. In operation, the system accesses 1202 first, second, and third parameterized applications (e.g., parameterized applications). In this example, each of the parameterized applications includes one or more parameters that define one or more characteristics of the parameterized application.
[0165] The system (or collection module 114 of the system described above) executes 1204 the first parameterized application with one or more specified values for one or more parameters of the first parameterized application to execute a collection module (e.g., collection module 114 of the system described above) that processes the data records. In one example, the first parameterized application includes a dataflow graph 910 ( FIG. 9 ). In this example, values for one or more parameters of the dataflow graph 910 are specified by inputting values into one or more of the portions 902, 904, 906, and 908. In this example, the portions 902, 904, 906, and 908 provide for input of one or more values into one or more selectable and / or editable areas of the portions 902, 904, 906, and 908. These selectable and / or editable areas are mapped to or otherwise associated with parameters of the dataflow graph 910 to, for example, enable setting of values for these parameters.
[0166] The system (or the collection module 114 of the system described above) processes data items (e.g., data records) by collecting 1206 data items from multiple data streams, each separate (e.g., distinct). In this example, a first portion of the data items collected from one stream has a different format than a second portion of the data items collected from a different stream. Furthermore, the data items are associated with key values and may thus be referred to as keyed data. The system (or the collection module 114 of the system described above) also transforms 1208 the first portion of the data items and the second portion of the data items according to one or more specifications of a first parameterized application. In one example, the system transforms the data items (or portions thereof) by formatting the data into a format suitable for a collection module, a detection module, and / or an action module (e.g., the collection module 114, the detection module 116, and / or the action module 118 described above). By doing so, the system solves the problem of data integration and data record management provisioning without the need for additional technology, as described above. This functionality simplifies and accelerates end-to-end integration, as described above. In particular, received data records require formatting and one-time validation (e.g., by collection module 114) and can then be processed by the collection, detection, and action modules (e.g., collection module 114, detection module 116, and / or action module 118 described above). In traditional approaches, separate systems perform collection, detection, and action. Thus, data must be formatted for the collection phase, then reformatted into a format appropriate for the detection phase, and then reformatted into a format appropriate for the action phase. This repeated formatting and reformatting introduces significant delays in the collection, detection, and action execution and real-time (or near-real-time) action recording as data records are received. This repeated formatting and reformatting also consumes significant bandwidth and system memory resources.Thus, in contrast to conventional methods, the system described herein provides reduced bandwidth and memory consumption as well as reduced latency in processing to perform real-time operations.
[0167] The collection module also transforms the received data records by parsing, verifying, and augmenting (if necessary) them with more slowly changing data (e.g., profile data from a data memory). As previously mentioned, this augmentation includes generating all data records (or a subset thereof) in the data record palette to enable real-time execution of logic on the received data records without the delay of performing a database lookup to retrieve the data necessary for logic execution. The system stores these verified and augmented data records in memory, for example, to enable real-time detection. The system also stores these verified and augmented data records on disk to ensure archival and successful recovery. The collection module is configured to handle multiple data sources (and / or almost any type of data source), thus addressing the complexities of large data volumes, low latency, and multiple data formats at its disposal, while enabling rapid and autonomous integration into the CDA system. Following transformation of the collected data, the collection module adds 1210 an input representing the transformed data item to a queue. The queue then forwards the data item to a second parameterization application, for example, the detection module.
[0168] In this example, the system executes 1212 a second parameterized application, for example, to execute a detection module (e.g., detection module 116 described above). In particular, the system executes the second parameterized application with one or more specified values for one or more parameters of the second parameterized application to process the transformed data in the queue. In this example, the second parameterized application represents a specification that includes rules and respective conditions for the rules. Additionally, state of the specification is maintained for each value of the key.
[0169] In one example, processing transformed data items in the queue includes: For one or more transformed data items associated with a particular value of the key, the system (and / or its detection module) detects (1214) that at least one of the one or more transformed data items satisfies one or more conditions of at least one of the rules of the specification. In response to the detection, the detection module generates an output that specifies the performance of one or more actions and adds the output to the queue. For example, the output is a message or an offer to the user. In one example, the output, along with one or more user actions (or lack thereof) related to the output, is tracked by the CDA system. Tracking these user actions provides feedback to the CDA system. In the process, the CDA system tracks the output and subsequently receives data (e.g., from an external system or from the CDA system itself) indicating the delivery of the output, one or more user interactions with the output, one or more actions related to the output, etc. The detection module also causes (1216) a transition of the specification from a current state to a subsequent state for the particular value of the key. For example, portion 1004 of Figure 9 displays state portion 1006a, which indicates the current state of the executable logic. Portion 1004 also displays state portion 1008a, which indicates the next state or destination state of the executable logic.
[0170] In this example, the system also executes 1218 a third parameterized application with one or more specified values for one or more parameters of the third parameterized application to perform tasks including sending another command to cause the performance of one or more actions. In this example, the third parameterized application executes an action module. As previously described, this action module customizes an output and sends it to a user and / or forwards the customized output to an external system.
[0171] The techniques described above can be implemented using software for execution on a computer. For example, the software forms procedures by one or more computer programs executing on one or more programmed or programmable computer systems (which can be of various architectures, such as distributed, client / server, grid, etc.), each including at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. The software may form one or more modules of a larger program that provides other services related to, for example, the design and construction of charts and flowcharts. The nodes, links, and elements of a chart can be implemented as data structures stored in a computer-readable medium or other organized data conforming to a data model stored in a data repository.
[0172] The techniques described herein can be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or combinations thereof. An apparatus can also be implemented by a computer program product tangibly embodied in or stored in a machine-readable storage device (e.g., a non-transitory machine-readable storage device, a machine-readable hardware storage device, etc.) for execution by a programmable processor, and the actions of a method can be performed by the programmable processor executing a program of instructions to perform functions by operating on input data and generating output. The embodiments described herein and other claimed embodiments, as well as the techniques described herein, can advantageously be implemented by one or more computer programs executable on a programmable system including at least one programmable processor coupled to exchange data and instructions with a data storage system, at least one input device, and at least one output device. Each computer program can be implemented in a high-level procedural or object-oriented programming language, or in assembly or machine language if desired; in any case, the language can be a compiled or interpreted language.
[0173] Processors suitable for executing a computer program include, by way of example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operatively coupled to receive data therefrom, transfer data thereto, or both. Computer-readable media for embodying computer program instructions and data include all types of non-volatile memory, including, by way of example, semiconductor memory devices, such as EPROMs, EEPROMs, flash memory devices, magnetic disks, such as internal hard disks and removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory can be supplemented by, or incorporated in, special-purpose logic circuitry. Any of the foregoing may be supplemented by, or incorporated in, ASICs (application-specific integrated circuits).
[0174] To enable user interaction, embodiments can be implemented on a computer that has a display device, such as an LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices can also be used to enable user interaction, for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, including acoustic input, voice input, or tactile input.
[0175] Embodiments may be implemented by a computing system including a back-end component, e.g., a data server; by a computing system including a middleware component, e.g., an application server; by a computing system including a front-end component, e.g., a client computer having a graphical user interface or web browser through which a user can interact with an implementation of an embodiment; or by a computing system including any combination of such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, e.g., a communications network. Examples of communications networks include local area networks (LANs) and wide area networks (WANs), e.g., the Internet.
[0176] The systems and methods, or portions thereof, may use the "World Wide Web" (Web or WWW), a collection of servers on the Internet that utilize the Hypertext Transfer Protocol (HTTP). HTTP is a well-known application protocol that provides users with access to resources, which may be information in various formats, such as text, graphics, images, audio, video, Hypertext Markup Language (HTML), programs, etc. When a link is specified by a user, the client computer makes a TCP / IP request to the Web server and receives information, which may be another Web page formatted according to HTML. The user may also access other pages on the same or other servers by following on-screen prompts, entering specific data, or clicking selected icons. It should also be noted that embodiments that use Web pages may use any type of selection device known to those skilled in the art, such as check boxes or drop-down boxes, to allow users to select options for a given component. The servers run on a variety of platforms, including UNIX machines, although other platforms, such as Windows 2000 / 2003, Windows NT, Sun, Linux, and Macintosh, may also be used. A computer user can view information available on a network or server on the Web using browsing software such as Firefox, Netscape Navigator, Microsoft Internet Explorer, or Mosaic browser. A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0177] Other embodiments are within the scope and spirit of this specification and claims. For example, due to the nature of software, the functionality described above may be implemented using software, hardware, firmware, hardwiring, or any combination thereof. The features implementing the functionality may also be physically located in various locations, including being distributed, such that portions of the functionality are implemented in various physical locations. The use of the term "a" throughout this specification and application is not intended to be limiting and thus is not intended to exclude a plural or "one or more" meaning of the term "a." Additionally, to the extent priority is claimed to a provisional patent application, it should be understood that the provisional patent application includes examples of how the techniques described herein may be implemented, but is not limiting.
[0178] Having described several embodiments of the present invention, those skilled in the art will nonetheless understand that various modifications can be made without departing from the spirit and scope of the claims and techniques described herein.
Claims
1. 1. A method implemented by a data processing system for finding aggregate values for keyed data items, wherein each keyed data item associates a particular data item with a particular key value and, for each of the one or more key values, causes the execution of one or more operations; The one or more actions include: receiving keyed data items collected from a plurality of data sources, the keyed data items including data items associated with particular data values for a key value; maintaining, for each key value, a state of the application, the state identifying one or more portions of the application executing in that state, the state of the application for a key value being based on a state previously reached from executing one or more portions of the application on one or more data items associated with the key value; storing, for each of a plurality of key values, one or more aggregated values for that key value in a first data store, the aggregated values for that key value based on (i) one or more data values for that key value retrieved from a second data source, and (ii) one or more data values of one or more respective data items for that key value included in the keyed data items collected from the data source; displaying one or more interface elements specifying one or more properties or one or more parameter values for processing the keyed data item; executing the application having the one or more properties or the one or more parameter values specified by the one or more interface elements; For a specific key value in the input, a keyed data item, identifying the current state of the application for that particular key value; accessing the data for that particular key value, executing one or more portions of the application designated in its current state for execution against the accessed data for that particular key value in accordance with the one or more properties or the one or more parameter values specified by the one or more interface elements; detecting a particular aggregated value for a particular key value in the accessed data, the particular aggregated value for the particular key value being based on (i) one or more data values for the particular key value retrieved from a second data source, and (ii) one or more data values for one or more respective data items for that key value included in the keyed data items collected from the data source; upon detecting a particular aggregated value for a particular key value in the accessed data, outputting an instruction specifying detection of the particular aggregated value for the particular key value; and transitioning the application from a current state to a subsequent state for a particular key value; executing the application by instructing one or more communication channels to perform one or more operations for a particular key value in response to receiving the command for the particular key value; A method comprising:
2. 10. The method of claim 1, the data items include data records; The method further includes reformatting the data records according to a particular format.
3. 10. The method of claim 1, the data item is a data record, The method further includes augmenting the data record with data from a profile of a user associated with the data record; The method wherein the enrichment is in accordance with instructions specified by the application to retrieve profile data about a user and add the profile data to one or more fields of the data record.
4. 10. The method of claim 1, The method further comprising the step of performing a feedback loop to one or more third party systems to seek confirmation of performance of the one or more actions.
5. 10. The method of claim 1, receiving data for a particular key value, the received data indicating feedback regarding at least one of the one or more actions; updating a KPI for a particular key value based on the received data indicative of the feedback by aggregating the received data with one or more portions of data contained in or associated with the KPI; The method further comprises:
6. 10. The method of claim 1, The one or more actions include: Sending a text message to an external device; Sending emails to external systems; Opening a work order ticket in the case management system; Disconnecting the mobile phone connection; providing a web service to a target device; sending a data packet of one or more transformed data items together with the notification; and executing a data processing application hosted on one or more external computers on one or more transformed data items; The method includes one or more of the following:
7. 10. The method of claim 1, one or more instructions are sent over the network connection to cause execution of at least one of the one or more actions on the external device; The method comprises: The method further comprising receiving a feedback message indicating whether at least one of the one or more operations (i) completed successfully or (ii) failed.
8. 10. The method of claim 1, The method further comprising determining a range of values for processing the keyed data item based on one or more characteristics, and executing the application according to the range of values.
9. 10. The method of claim 1, The method, wherein the one or more parameter values are ranges of values.
10. one or more non-transitory machine-readable hardware storage devices for finding aggregate values for keyed data items, each keyed data item associating a particular data item with a particular key value and causing the performance of one or more operations for each of the one or more key values, the one or more non-transitory machine-readable hardware storage devices storing instructions executable by one or more processing devices to perform the operations; The operation is receiving keyed data items collected from a plurality of data sources, the keyed data items including data items associated with particular data values for a key value; maintaining, for each key value, a state of the application, the state identifying one or more portions of the application executing in that state, the state of the application for a key value being based on a state previously reached from executing one or more portions of the application on one or more data items associated with the key value; storing, for each of a plurality of key values, one or more aggregated values for that key value in a first data store, the aggregated values for that key value based on (i) one or more data values for that key value retrieved from a second data source, and (ii) one or more data values of one or more respective data items for that key value included in the keyed data items collected from the data source; displaying one or more interface elements specifying one or more properties or one or more parameter values for processing the keyed data item; executing the application having the one or more properties or the one or more parameter values specified by the one or more interface elements; For a specific key value in the input, a keyed data item, identifying the current state of the application for that particular key value; accessing the data for that particular key value, executing one or more portions of the application designated in its current state for execution against the accessed data for that particular key value in accordance with the one or more properties or the one or more parameter values specified by the one or more interface elements; detecting a particular aggregated value for a particular key value in the accessed data, the particular aggregated value for the particular key value being based on (i) one or more data values for the particular key value retrieved from a second data source, and (ii) one or more data values for one or more respective data items for that key value included in the keyed data items collected from the data source; upon detecting a particular aggregated value for a particular key value in the accessed data, outputting an instruction specifying detection of the particular aggregated value for the particular key value; and transitioning the application from a current state to a subsequent state for a particular key value; executing the application by instructing one or more communication channels to perform one or more operations for a particular key value in response to receiving the command for the particular key value; one or more non-transitory machine-readable hardware storage devices, including:
11. 11. The one or more non-transitory machine-readable hardware storage devices of claim 10, the data item is a data record, The operations further include augmenting the data record with data from a profile of a user associated with the data record; The augmentation comprises one or more non-transitory machine-readable hardware storage devices that follow instructions specified by the application to retrieve profile data about a user and add the profile data to one or more fields of the data record.
12. 11. The one or more non-transitory machine-readable hardware storage devices of claim 10, The operation is receiving data for a particular key value, the received data indicating feedback regarding at least one of the one or more actions; updating a KPI for a particular key value based on the received data indicative of the feedback by aggregating the received data with one or more portions of data contained in or associated with the KPI; one or more non-transitory machine-readable hardware storage devices, including:
13. 11. The one or more non-transitory machine-readable hardware storage devices of claim 10, The one or more actions include: Sending a text message to an external device; Sending emails to external systems; Opening a work order ticket in the case management system; Disconnecting the mobile phone connection; providing a web service to a target device; sending a data packet of one or more transformed data items together with the notification; and executing a data processing application hosted on one or more external computers on one or more transformed data items; one or more non-transitory machine-readable hardware storage devices, including one or more of:
14. 11. The one or more non-transitory machine-readable hardware storage devices of claim 10, one or more non-transitory machine-readable hardware storage devices, wherein the one or more parameter values are ranges of values.
15. 1. A method executed by a data processing system for executing a computer program for processing one or more keyed data items associated with one or more key values, and for selecting a particular rule to be applied from one or more rules for a particular key value when the computer program is in a particular state with respect to the particular key value, comprising: identifying a computer program for processing one or more data items, said one or more data items being associated with said one or more key values, at least one of said one or more data items comprising a data record to be processed by said data processing system; maintaining, for each key value, a state of the application, the state identifying one or more portions of the application executing in that state, the state of the application for a key value being based on a state previously reached from executing one or more portions of the application on one or more data items associated with the key value; augmenting the data record by combining data from a profile associated with the data record, the augmenting being in accordance with instructions based on the computer program including instructions for retrieving profile data from a data store associated with the computer program and adding the profile data to one or more fields of the data record; executing the computer program to process the one or more data items, wherein one or more states of the computer program are maintained for the one or more key values; The execution identifying a state of the computer program for one or more data items associated with the particular key value, the state being associated with the particular key value; identifying one or more rules based on the identified state for the particular key value, a rule specifying one or more attributes and further specifying one or more actions to be performed when a rule is selected, the rule being selected when the computer program is at least in the identified state for the particular key value; selecting, according to the computer program in the identified state for the particular key value, from one or more rules, a rule that specifies one or more attributes of the one or more data items associated with the particular key value; performing one or more actions specified by the identified one or more rules; transitioning the computer program from a first state to a second state for the particular key value; A method comprising:
16. 16. The method of claim 15, The method, wherein the particular key value includes a user identifier associated with a particular user.
17. 16. The method of claim 15, The one or more attributes identify a population segment associated with the particular key value.
18. 16. The method of claim 15, The method, wherein the one or more rules are configured based on a parameterized application.
19. 16. The method of claim 15, The method further includes displaying one or more user interface elements for specifying one or more values of the one or more attributes.
20. 16. The method of claim 15, The one or more actions include: Sending a text message to an external device; Sending emails to external systems; Opening a work order ticket in the case management system; Disconnecting the mobile phone connection; providing a web service to a target device; sending a data packet of one or more transformed data items together with the notification; and executing a data processing application hosted on one or more external computers on one or more transformed data items; The method includes one or more of the following:
21. 1. A data processing system for finding aggregated values for keyed data items, each keyed data item associating a particular data item with a particular key value and causing the execution of one or more operations for each of the one or more key values; The data processing system includes: at least one processing device; at least one memory in communication with the at least one processing unit, the at least one memory storing instructions that, when executed by the at least one processing unit, cause the at least one processing unit to perform operations; The operation is receiving keyed data items collected from a plurality of data sources, the keyed data items including data items associated with particular data values for a key value; maintaining, for each key value, a state of the application, the state identifying one or more portions of the application executing in that state, the state of the application for a key value being based on a state previously reached from executing one or more portions of the application on one or more data items associated with the key value; storing, for each of a plurality of key values, one or more aggregated values for that key value in a first data store, the aggregated values for that key value based on (i) one or more data values for that key value retrieved from a second data source, and (ii) one or more data values of one or more respective data items for that key value included in the keyed data items collected from the data source; displaying one or more interface elements specifying one or more properties or one or more parameter values for processing the keyed data item; executing the application having the one or more properties or the one or more parameter values specified by the one or more interface elements; For a specific key value in the input, a keyed data item, identifying the current state of the application for that particular key value; accessing the data for that particular key value, executing one or more portions of the application designated in its current state for execution against the accessed data for that particular key value in accordance with the one or more properties or the one or more parameter values specified by the one or more interface elements; detecting a particular aggregated value for a particular key value in the accessed data, the particular aggregated value for the particular key value being based on (i) one or more data values for the particular key value retrieved from a second data source, and (ii) one or more data values for one or more respective data items for that key value included in the keyed data items collected from the data source; upon detecting a particular aggregated value for a particular key value in the accessed data, outputting an instruction specifying detection of the particular aggregated value for the particular key value; and transitioning the application from a current state to a subsequent state for a particular key value; executing the application by instructing one or more communication channels to perform one or more operations for a particular key value in response to receiving the command for the particular key value; a data processing system including:
22. 22. The data processing system of claim 21, the data items include data records; The operations further include reformatting the data records according to a particular format.
23. 22. The data processing system of claim 21, the data item is a data record, The operations further include augmenting the data record with data from a profile of a user associated with the data record; The enrichment is in accordance with instructions specified by the application to retrieve profile data about a user and add the profile data to one or more fields of the data record.
24. 22. The data processing system of claim 21, The data processing system, wherein the actions further include performing a feedback loop to one or more third party systems to seek confirmation of performance of the one or more actions.
25. 22. The data processing system of claim 21, The operation is receiving data for a particular key value, the received data indicating feedback regarding at least one of the one or more actions; updating a KPI for a particular key value based on the received data indicative of the feedback by aggregating the received data with one or more portions of data contained in or associated with the KPI; 20. The data processing system according to claim 19, further comprising:
26. 22. The data processing system of claim 21, The one or more actions include: Sending a text message to an external device; Sending emails to external systems; Opening a work order ticket in the case management system; Disconnecting the mobile phone connection; providing a web service to a target device; sending a data packet of one or more transformed data items together with the notification; and executing a data processing application hosted on one or more external computers on one or more transformed data items; 1. A data processing system comprising one or more of:
27. 1. A data processing system for processing one or more keyed data items associated with one or more key values, and for selecting a particular rule to be applied from one or more rules for a particular key value when a computer program is in a particular state with respect to the particular key value, comprising: The data processing system includes: at least one processing device; at least one memory in communication with the at least one processing unit, the at least one memory storing instructions that, when executed by the at least one processing unit, cause the at least one processing unit to perform operations; The operation is identifying a computer program for processing one or more data items, said one or more data items being associated with said one or more key values, at least one of said one or more data items comprising a data record to be processed by said data processing system; maintaining, for each key value, a state of the application, the state identifying one or more portions of the application executing in that state, the state of the application for a key value being based on a state previously reached from executing one or more portions of the application on one or more data items associated with the key value; augmenting the data record by combining data from a profile associated with the data record, the augmenting being in accordance with instructions based on the computer program including instructions for retrieving profile data from a data store associated with the computer program and adding the profile data to one or more fields of the data record; executing the computer program to process the one or more data items, wherein one or more states of the computer program are maintained for the one or more key values; The execution identifying a state of the computer program for one or more data items associated with the particular key value, the state being associated with the particular key value; identifying one or more rules based on the identified state for the particular key value, a rule specifying one or more attributes and further specifying one or more actions to be performed when a rule is selected, the rule being selected when the computer program is at least in the identified state for the particular key value; selecting, according to the computer program in the identified state for the particular key value, from one or more rules, a rule that specifies one or more attributes of the one or more data items associated with the particular key value; performing one or more actions specified by the identified one or more rules; transitioning the computer program from a first state to a second state for the particular key value; a data processing system including:
28. 28. The data processing system of claim 27, The particular key value includes a user identifier associated with a particular user.
29. 28. The data processing system of claim 27, The one or more attributes identify a population segment associated with the particular key value.
30. 28. The data processing system of claim 27, The one or more rules are configured based on a parameterized application.
31. 28. The data processing system of claim 27, The one or more actions include: Sending a text message to an external device; Sending emails to external systems; Opening a work order ticket in the case management system; Disconnecting the mobile phone connection; providing a web service to a target device; sending a data packet of one or more transformed data items together with the notification; and executing a data processing application hosted on one or more external computers on one or more transformed data items; 1. A data processing system comprising one or more of:
32. one or more non-transitory machine-readable hardware storage devices for storing a computer program for processing one or more keyed data items associated with one or more key values and for selecting a particular rule from one or more rules to be applied for a particular key value when the computer program is in a particular state with respect to the particular key value, the one or more non-transitory machine-readable hardware storage devices storing instructions executable by one or more processing devices to perform operations; The operation is identifying a computer program that processes one or more data items, the one or more data items being associated with the one or more key values, at least one of the one or more data items comprising a data record to be processed by a data processing system; maintaining, for each key value, a state of the application, the state identifying one or more portions of the application executing in that state, the state of the application for a key value being based on a state previously reached from executing one or more portions of the application on one or more data items associated with the key value; augmenting the data record by combining data from a profile associated with the data record, the augmenting being in accordance with instructions based on the computer program including instructions for retrieving profile data from a data store associated with the computer program and adding the profile data to one or more fields of the data record; executing the computer program to process the one or more data items, wherein one or more states of the computer program are maintained for the one or more key values; The execution identifying a state of the computer program for one or more data items associated with the particular key value, the state being associated with the particular key value; identifying one or more rules based on the identified state for the particular key value, a rule specifying one or more attributes and further specifying one or more actions to be performed when a rule is selected, the rule being selected when the computer program is at least in the identified state for the particular key value; selecting, according to the computer program in the identified state for the particular key value, from one or more rules, a rule that specifies one or more attributes of the one or more data items associated with the particular key value; performing one or more actions specified by the identified one or more rules; transitioning the computer program from a first state to a second state for the particular key value; one or more non-transitory machine-readable hardware storage devices, including:
33. 33. The one or more non-transitory machine-readable hardware storage devices of claim 32, One or more non-transitory machine-readable hardware storage devices, wherein the particular key value contains a user identifier associated with a particular user.
34. 33. The one or more non-transitory machine-readable hardware storage devices of claim 32, One or more non-transitory machine-readable hardware storage devices, wherein the one or more attributes identify a population segment associated with the particular key value.
35. 33. The one or more non-transitory machine-readable hardware storage devices of claim 32, One or more non-transitory machine-readable hardware storage devices, wherein the one or more rules are configured based on a parameterized application.
36. 33. The one or more non-transitory machine-readable hardware storage devices of claim 32, The one or more actions include: Sending a text message to an external device; Sending emails to external systems; Opening a work order ticket in the case management system; Disconnecting the mobile phone connection; providing a web service to a target device; sending a data packet of one or more transformed data items together with the notification; and executing a data processing application hosted on one or more external computers on one or more transformed data items; one or more non-transitory machine-readable hardware storage devices, including one or more of:
Citation Information
Patent Citations
Specifying user interface elements
JP2013513864A
Dynamic service integration system and method
JP2015505387A
Data aggregation in mediation systems
JP2015528967A
Techniques For Specifying And Collecting Data Aggregations
US20100185618A1